Mixed NAL unit type-based image encoding / decoding method and device, and method for transmitting bitstream

JP2025166165APending Publication Date: 2025-11-05LG ELECTRONICS INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025134978
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-03-05
Filing Date
2025-08-14
Publication Date
2025-11-05

Smart Images

  • Figure 2025166165000001_ABST
    Figure 2025166165000001_ABST
Patent Text Reader

Abstract

To provide a mixed NAL unit type-based image encoding / decoding method.SOLUTION: A method comprises: obtaining, from a bitstream, video coding layer (VCL) network abstraction layer (NAL) unit type information of a current picture and first flag information specifying whether a subpicture included in the current picture is treated as one picture; determining a NAL unit type of each of a plurality of slices included in the current picture, based on the VCL NAL unit type information; and decoding the plurality of slices based on the determined NAL unit type and the first flag information. The current picture may comprise two or more subpictures based on at least some of the plurality of slices having different NAL unit types.SELECTED DRAWING: Figure 16
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an image encoding / decoding method and apparatus, and more particularly to an image encoding / decoding method and apparatus based on hybrid NAL unit types, and a method for transmitting a bitstream generated by the image encoding method / apparatus of the present disclosure. [Background technology]

[0002] Recently, demand for high-resolution, high-quality images, such as HD (High Definition) images and UHD (Ultra High Definition) images, has been increasing in various fields. As image data becomes higher in resolution and quality, the amount of information or bits to be transmitted increases relatively compared to conventional image data. The increase in the amount of information or bits to be transmitted results in an increase in transmission costs and storage costs.

[0003] This requires highly efficient image compression techniques for effectively transmitting, storing, and reproducing high-resolution, high-quality image information. Summary of the Invention [Problem to be solved by the invention]

[0004] An object of the present disclosure is to provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.

[0005] Another object of the present disclosure is to provide an image encoding / decoding method and apparatus based on hybrid NAL unit types.

[0006] Another object of the present disclosure is to provide an image encoding / decoding method and apparatus based on two or more sub-pictures having different NAL unit types.

[0007] Another object of the present disclosure is to provide a computer-readable recording medium storing a bitstream generated by the image encoding method or apparatus according to the present disclosure.

[0008] Another object of the present disclosure is to provide a computer-readable recording medium storing a bitstream that is received by an image decoding device according to the present disclosure, decoded, and used to restore an image.

[0009] Another object of the present disclosure is to provide a method for transmitting a bitstream generated by the image encoding method or apparatus according to the present disclosure.

[0010] The technical problems to be solved by the present disclosure are not limited to the above-mentioned technical problems, and other technical problems not mentioned above will be clearly understood by a person having ordinary skill in the technical field to which the present disclosure pertains from the following description. [Means for solving the problem]

[0011] An image decoding method according to one aspect of the present disclosure includes the steps of: acquiring, from a bitstream, video coding layer (VCL) NAL (network abstraction layer) unit type information of a current picture and first flag information indicating whether sub-pictures included in the current picture are treated as one picture; determining a NAL unit type for each of a plurality of slices in the current picture based on the acquired VCL NAL unit type information; and decoding the plurality of slices based on the determined NAL unit type and the first flag information, wherein the current picture includes two or more sub-pictures based on at least some of the plurality of slices having different NAL unit types, and the first flag information may have a predetermined value indicating that each of the two or more sub-pictures is treated as one picture based on at least some of the plurality of slices having different NAL unit types.

[0012]

[0013] According to another aspect of the present disclosure, an image decoding device includes a memory and at least one processor, wherein the at least one processor acquires, from a bitstream, video coding layer (VCL) network abstraction layer (NAL) unit type information of a current picture and first flag information indicating whether sub-pictures included in the current picture are treated as one picture; determines a NAL unit type of each of a plurality of slices in the current picture based on the acquired VCL NAL unit type information; and decodes the plurality of slices based on the determined NAL unit type and the first flag information, wherein the current picture includes two or more sub-pictures based on at least some of the plurality of slices having different NAL unit types; and the first flag information may have a predetermined value indicating that each of the two or more sub-pictures is treated as one picture based on at least some of the plurality of slices having different NAL unit types.

[0013] An image encoding method according to another aspect of the present disclosure includes the steps of dividing a current picture into one or more sub-pictures, determining a NAL (network abstraction layer) unit type for each of a plurality of slices included in the one or more sub-pictures, and encoding the plurality of slices based on the determined NAL unit types, wherein the current picture is divided into two or more sub-pictures based on at least some of the plurality of slices having different NAL unit types, and each of the two or more sub-pictures can be treated as a single picture based on at least some of the plurality of slices having different NAL unit types.

[0014] A computer-readable recording medium according to another aspect of the present disclosure can store a bitstream generated by the image encoding method or image encoding device of the present disclosure.

[0015] A transmission method according to another aspect of the present disclosure can transmit a bitstream generated by the image encoding device or image encoding method of the present disclosure.

[0016] The features described above in this brief summary of the present disclosure are merely exemplary embodiments of the detailed description of the present disclosure that follows and are not intended to limit the scope of the present disclosure. [Effects of the Invention]

[0017] According to the present disclosure, an image encoding / decoding method and apparatus with improved encoding / decoding efficiency can be provided.

[0018] Furthermore, according to the present disclosure, an image encoding / decoding method and apparatus based on a hybrid NAL unit type can be provided.

[0019] Furthermore, the present disclosure can provide an image encoding / decoding method and apparatus based on two or more sub-pictures having different NAL unit types.

[0020] The present disclosure also provides a method for transmitting a bitstream generated by the image encoding method or apparatus according to the present disclosure.

[0021] Furthermore, according to the present disclosure, a recording medium storing a bitstream generated by the image encoding method or apparatus according to the present disclosure can be provided.

[0022] The present disclosure may also provide a recording medium storing a bitstream that is received by an image decoding device according to the present disclosure, decoded, and used to restore an image.

[0023] The effects obtained by the present disclosure are not limited to the effects described above, and other effects not described above will be clearly understood by those having ordinary skill in the art to which the present disclosure pertains from the following description. [Brief explanation of the drawings]

[0024] [Figure 1] 1 is a diagram illustrating a video coding system to which embodiments of the present disclosure can be applied; [Figure 2] 1 is a diagram schematically illustrating an image encoding device to which an embodiment of the present disclosure can be applied. [Figure 3] FIG. 1 is a diagram schematically illustrating an image decoding device to which an embodiment of the present disclosure can be applied. [Figure 4] 1 is a flowchart illustrating an outline of an image decoding procedure to which an embodiment of the present disclosure can be applied. [Figure 5] 1 is a flowchart illustrating an outline of an image encoding procedure to which an embodiment of the present disclosure can be applied. [Figure 6] FIG. 1 shows an example of a hierarchical structure for coded images / video. [Figure 7] FIG. 1 is a diagram illustrating a picture parameter set (PPS) according to one embodiment of the present disclosure. [Figure 8] FIG. 10 illustrates a slice header according to one embodiment of the present disclosure. [Figure 9] FIG. 10 is a diagram showing an example of a sub-picture. [Figure 10] A diagram showing an example of a picture having mixed NAL unit types. [Figure 11] FIG. 1 is a diagram illustrating an example of a picture parameter set (PPS) according to an embodiment of the present disclosure. [Figure 12] FIG. 1 is a diagram illustrating an example of a sequence parameter set (SPS) according to an embodiment of the present disclosure. [Figure 13] 10A and 10B are diagrams illustrating the decoding order and output order for each picture type. [Figure 14] 10 is a flowchart illustrating a method for determining a NAL unit type of a current picture according to one embodiment of the present disclosure. [Figure 15] 1 is a flowchart illustrating an image encoding method according to an embodiment of the present disclosure. [Figure 16]1 is a flowchart illustrating an image decoding method according to an embodiment of the present disclosure. [Figure 17] 1 is a diagram illustrating an exemplary content streaming system to which an embodiment of the present disclosure can be applied; DETAILED DESCRIPTION OF THE INVENTION

[0025] The present disclosure will be described in detail below with reference to the accompanying drawings, so that those skilled in the art can easily implement the present disclosure. However, the present disclosure may be embodied in various different forms and is not limited to the embodiments described herein.

[0026] In describing the embodiments of the present disclosure, if it is determined that a detailed description of a known configuration or function may obscure the gist of the present disclosure, the detailed description thereof will be omitted. In addition, in the drawings, parts that are not related to the description of the present disclosure will be omitted, and similar parts will be designated by similar reference numerals.

[0027] In this disclosure, when a component is referred to as being "coupled," "coupled," or "connected" to another component, this includes not only a direct connection, but also an indirect connection where another component exists between them. Furthermore, when a component is referred to as "including" or "having" another component, this does not mean that the other component is excluded, but that the component can further include the other component, unless otherwise specified.

[0028] In this disclosure, terms such as "first" and "second" are used only to distinguish one component from another component, and do not limit the order or importance of the components unless otherwise specified. Therefore, within the scope of this disclosure, a first component in one embodiment may be referred to as a second component in another embodiment, and similarly, a second component in one embodiment may be referred to as a first component in another embodiment.

[0029] In this disclosure, components that are distinguished from one another are used to clearly describe the characteristics of each component and do not necessarily mean that the components are separate. In other words, multiple components may be integrated into a single hardware or software unit, or a single component may be distributed into multiple hardware or software units. Therefore, even if not otherwise specified, such integrated or distributed embodiments are also included within the scope of this disclosure.

[0030] In this disclosure, the components described in various embodiments are not necessarily essential components, and some may be optional components. Therefore, an embodiment consisting of a subset of the components described in one embodiment is also within the scope of this disclosure. Furthermore, an embodiment including other components in addition to the components described in various embodiments is also within the scope of this disclosure.

[0031] The present disclosure relates to image encoding and decoding, and terms used in this disclosure may have their ordinary meaning in the technical field to which the present disclosure belongs unless they are newly defined in this disclosure.

[0032] In this disclosure, a "picture" generally refers to a unit representing any one image in a specific time period, and a slice / tile is a coding unit constituting a part of a picture, and one picture may be composed of one or more slices / tiles. Furthermore, a slice / tile may include one or more coding tree units (CTUs).

[0033] In this disclosure, "pixel" or "pel" may refer to the smallest unit constituting one picture (or image). Also, "sample" may be used as a term corresponding to pixel. A sample may generally indicate a pixel or a pixel value, may indicate only a pixel / pixel value of a luma component, or may indicate only a pixel / pixel value of a chroma component.

[0034] In this disclosure, the term "unit" may refer to a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to that region. The term "unit" may be used interchangeably with terms such as "sample array," "block," or "area," depending on the situation. In general, an M×N block may include a set (or array) of samples or transform coefficients consisting of M columns and N rows.

[0035] In the present disclosure, a "current block" may refer to any one of a "current coding block," a "current coding unit," a "block to be coded," a "block to be decoded," or a "block to be processed." When prediction is performed, a "current block" may refer to a "current predicted block" or a "block to be predicted." When transformation (inverse transformation) / quantization (inverse quantization) is performed, a "current block" may refer to a "current transformed block" or a "block to be transformed." When filtering is performed, a "current block" may refer to a "block to be filtered."

[0036] Furthermore, in this disclosure, unless explicitly stated as a chroma block, the term "current block" may refer to a block including both a luma component block and a chroma component block, or to the "luma block of the current block." The luma component block of the current block may be expressed by explicitly including the term "luma block" or "current luma block." The chroma component block of the current block may be expressed by explicitly including the term "chroma block" or "current chroma block."

[0037] In the present disclosure, " / " and "," can be interpreted as "and / or." For example, "A / B" and "A, B" can be interpreted as "A and / or B." Also, "A / B / C" and "A, B, C" can mean "at least one of A, B, and / or C."

[0038] In this disclosure, "or" can be interpreted as "and / or." For example, "A or B" can mean 1) only "A," 2) only "B," or 3) "A and B." Alternatively, in this disclosure, "or" can mean "additionally or alternatively."

[0039] Video Coding System Overview

[0040] FIG. 1 is a diagram illustrating a video coding system to which embodiments of the present disclosure can be applied.

[0041] A video coding system according to one embodiment may include an encoding device 10 and a decoding device 20. The encoding device 10 may transmit encoded video and / or image information or data to the decoding device 20 in a file or streaming format via a digital storage medium or a network.

[0042] An encoding device 10 according to an embodiment may include a video source generation unit 11, an encoding unit 12, and a transmission unit 13. A decoding device 20 according to an embodiment may include a reception unit 21, a decoding unit 22, and a rendering unit 23. The encoding unit 12 may be referred to as a video / image encoding unit, and the decoding unit 22 may be referred to as a video / image decoding unit. The transmission unit 13 may be included in the encoding unit 12. The reception unit 21 may be included in the decoding unit 22. The rendering unit 23 may include a display unit, which may be configured as a separate device or an external component.

[0043] The video source generation unit 11 can acquire video / images through a video / image capture, synthesis, or generation process. The video source generation unit 11 can include a video / image capture device and / or a video / image generation device. The video / image capture device can include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. The video / image generation device can include, for example, a computer, a tablet, a smartphone, etc., and can (electronically) generate video / images. For example, virtual video / images can be generated via a computer, etc., in which case the video / image capture process can be replaced with a process in which related data is generated.

[0044] The encoder 12 may encode the input video / image. The encoder 12 may perform a series of steps such as prediction, transformation, and quantization for compression and coding efficiency. The encoder 12 may output the encoded data (encoded video / image information) in a bitstream format.

[0045] The transmitter 13 may transmit the encoded video / image information or data output in a bitstream format to the receiver 21 of the decoding device 20 in a file or streaming format via a digital storage medium or a network. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray®, HDD, and SSD. The transmitter 13 may include elements for generating a media file in a predetermined file format and elements for transmitting via a broadcasting / communication network. The receiver 21 may extract / receive the bitstream from the storage medium or network and transmit it to the decoder 22.

[0046] The decoding unit 22 can decode the video / image by performing a series of steps such as inverse quantization, inverse transformation, and prediction corresponding to the operations of the encoding unit 12.

[0047] The rendering unit 23 can render the decoded video / images, and the rendered video / images can be displayed via the display unit.

[0048] Overview of the image encoding device

[0049] FIG. 2 is a diagram schematically illustrating an image encoding device to which an embodiment of the present disclosure can be applied.

[0050] 2, the image encoding device 100 may include an image division unit 110, a subtraction unit 115, a transform unit 120, a quantization unit 130, an inverse quantization unit 140, an inverse transform unit 150, an addition unit 155, a filtering unit 160, a memory 170, an inter prediction unit 180, an intra prediction unit 185, and an entropy encoding unit 190. The inter prediction unit 180 and the intra prediction unit 185 may be collectively referred to as a "prediction unit." The transform unit 120, the quantization unit 130, the inverse quantization unit 140, and the inverse transform unit 150 may be included in a residual processing unit. The residual processing unit may further include a subtraction unit 115.

[0051] Depending on the embodiment, all or at least some of the components constituting the image encoding device 100 may be realized by a single hardware component (e.g., an encoder or a processor). Also, the memory 170 may include a decoded picture buffer (DPB) and may be realized by a digital storage medium.

[0052] The image division unit 110 may divide an input image (or picture, frame) input to the image encoding device 100 into one or more processing units. As an example, the processing units may be called coding units (CUs). The coding units may be obtained by recursively dividing a coding tree unit (CTU) or a largest coding unit (LCU) using a QT / BT / TT (quad-tree / binary-tree / ternary-tree) structure. For example, one coding unit may be divided into multiple coding units at deeper depths based on a quad-tree structure, a binary-tree structure, and / or a ternary-tree structure. To divide the coding units, the quad-tree structure may be applied first, and then the binary-tree structure and / or the ternary-tree structure may be applied later. The coding procedure according to the present disclosure may be performed based on the final coding unit that is not further divided. The maximum coding unit may be used as the final coding unit, or a lower-depth coding unit obtained by dividing the maximum coding unit may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and / or reconstruction, which will be described later. As another example, a processing unit of the coding procedure may be a prediction unit (PU) or a transform unit (TU). The prediction unit and the transform unit may be divided or partitioned from the final coding unit, respectively. The prediction unit may be a unit of sample prediction, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.

[0053] The prediction unit (inter prediction unit 180 or intra prediction unit 185) may perform prediction on a current block (current block) to generate a predicted block including prediction samples for the current block. The prediction unit may determine whether intra prediction or inter prediction is applied to the current block or CU. The prediction unit may generate various information related to prediction of the current block and transmit it to the entropy coding unit 190. The prediction information may be coded by the entropy coding unit 190 and output in a bitstream format.

[0054] The intra prediction unit 185 may predict the current block by referring to samples in the current picture. The referenced samples may be located in the neighborhood of the current block or may be located far away from the current block according to the intra prediction mode and / or intra prediction technique. The intra prediction modes may include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes may include, for example, a DC mode and a planar mode. The directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the granularity of the prediction direction. However, this is merely an example, and more or less directional prediction modes may be used depending on the settings. The intra prediction unit 185 may also determine the prediction mode to be applied to the current block using the prediction modes applied to neighboring blocks.

[0055] The inter prediction unit 180 may derive a predicted block for a current block based on a reference block (reference sample array) identified by a motion vector on a reference picture. To reduce the amount of motion information transmitted in inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation between the motion information of neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on the inter prediction direction (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block may be called a collocated reference block, a collocated CU (colCU), etc. The reference picture including the temporal neighboring block may be called a collocated picture (colPic). For example, the inter predictor 180 may construct a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to derive a motion vector and / or a reference picture index for the current block. Inter prediction may be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the inter predictor 180 may use motion information of neighboring blocks as motion information for the current block. In the case of skip mode, unlike in merge mode, a residual signal may not be transmitted.In the case of a motion vector prediction (MVP) mode, the motion vector of a neighboring block is used as a motion vector predictor, and the motion vector of the current block can be signaled by encoding a motion vector difference and an indicator for the motion vector predictor. The motion vector difference may mean the difference between the motion vector of the current block and the motion vector predictor.

[0056] The predictor may generate a prediction signal based on various prediction methods and / or prediction techniques, which will be described later. For example, the predictor may apply intra prediction or inter prediction to predict the current block, or may simultaneously apply intra prediction and inter prediction. A prediction method that simultaneously applies intra prediction and inter prediction to predict the current block may be referred to as combined inter and intra prediction (CIIP). The predictor may also perform intra block copy (IBC) to predict the current block. Intra block copy can be used for content image / video coding, such as screen content coding (SCC), for games. IBC is a method of predicting a current block using an already reconstructed reference block in a current picture that is located a predetermined distance away from the current block. When IBC is applied, the position of the reference block in the current picture may be coded as a vector (block vector) corresponding to the predetermined distance. IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that a reference block is derived within the current picture. That is, the IBC may use at least one of the inter prediction techniques described in this disclosure.

[0057] The prediction signal generated by the prediction unit may be used to generate a restored signal or a residual signal. The subtraction unit 115 may subtract the prediction signal (predicted block, predicted sample array) output from the prediction unit from the input image signal (original block, original sample array) to generate a residual signal (residual signal, residual block, residual sample array). The generated residual signal may be transmitted to the conversion unit 120.

[0058] The transform unit 120 may generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loeve transform (KLT), a graph-based transform (GBT), or a conditionally non-linear transform (CNT). Here, the GBT refers to a transform obtained from a graph representing inter-pixel relationship information. The CNT refers to a transform obtained based on a predicted signal generated using all previously reconstructed pixels. The transform process may be applied to pixel blocks having the same square size or to non-square blocks of variable size.

[0059] The quantization unit 130 may quantize the transform coefficients and transmit the quantized transform coefficients to the entropy coding unit 190. The entropy coding unit 190 may encode the quantized signal (information about the quantized transform coefficients) and output the encoded signal in a bitstream format. The information about the quantized transform coefficients may be referred to as residual information. The quantization unit 130 may rearrange the quantized transform coefficients in a block format into a one-dimensional vector format based on a coefficient scan order, and may generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector format.

[0060] The entropy coding unit 190 may perform various coding methods, such as exponential Golomb, context-adaptive variable length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy coding unit 190 may also code information necessary for video / image restoration (e.g., values ​​of syntax elements) together with or separately from the quantized transform coefficients. The coded information (e.g., coded video / image information) may be transmitted or stored in a bitstream format in network abstraction layer (NAL) units. The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / image information may also include general constraint information. The signaling information, transmitted information and / or syntax elements mentioned in this disclosure may be encoded through the above-described encoding procedure and included in the bitstream.

[0061] The bitstream may be transmitted via a network or stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as a USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitting unit (not shown) that transmits and / or a storing unit (not shown) that stores the signal output from the entropy encoding unit 190 may be provided as an internal / external element of the image encoding device 100, or the transmitting unit may be provided as a component of the entropy encoding unit 190.

[0062] The quantized transform coefficients output from the quantization unit 130 can be used to generate a residual signal. For example, the residual signal (residual block or residual sample) can be reconstructed by applying inverse quantization and inverse transform to the quantized transform coefficients via the inverse quantization unit 140 and the inverse transform unit 150.

[0063] The adder 155 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter prediction unit 180 or the intra prediction unit 185. When there is no residual for the current block to be processed, such as when a skip mode is applied, the predicted block may be used as the reconstructed block. The adder 155 may be referred to as a reconstruction unit or a reconstructed block generation unit. The generated reconstructed signal may be used for intra prediction of the next current block to be processed in the current picture, and may also be used for inter prediction of the next picture after filtering, as will be described later.

[0064] The filtering unit 160 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 160 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture and store the modified reconstructed picture in the memory 170, specifically, in the DPB of the memory 170. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, etc. The filtering unit 160 may generate various information related to filtering and transmit it to the entropy coding unit 190, as will be described later in connection with each filtering method. The filtering information may be coded by the entropy coding unit 190 and output in a bitstream format.

[0065] The modified reconstructed picture transmitted to the memory 170 can be used as a reference picture in the inter prediction unit 180. When inter prediction is applied through this, the image encoding device 100 can avoid a prediction mismatch between the image encoding device 100 and the image decoding device, and can also improve encoding efficiency.

[0066] The DPB in the memory 170 may store modified reconstructed pictures for use as reference pictures in the inter predictor 180. The memory 170 may store motion information of blocks from which motion information in the current picture is derived (or coded) and / or motion information of already reconstructed intra-picture blocks. The stored motion information may be transmitted to the inter predictor 180 to be used as motion information of spatially surrounding blocks or temporally surrounding blocks. The memory 170 may store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra predictor 185.

[0067] Overview of the image decoding device

[0068] FIG. 3 is a diagram schematically illustrating an image decoding device to which an embodiment of the present disclosure can be applied.

[0069] 3, the image decoding apparatus 200 may include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an adder 235, a filtering unit 240, a memory 250, an inter prediction unit 260, and an intra prediction unit 265. The inter prediction unit 260 and the intra prediction unit 265 may be collectively referred to as a "prediction unit." The inverse quantization unit 220 and the inverse transform unit 230 may be included in a residual processing unit.

[0070] Depending on the embodiment, all or at least some of the components constituting the image decoding device 200 may be realized by a single hardware component (e.g., a decoder or a processor). Also, the memory 170 may include a DPB and may be realized by a digital storage medium.

[0071] The image decoding device 200, which receives a bitstream including video / image information, can reconstruct an image by performing a process corresponding to the process performed by the image encoding device 100 of FIG. 2. For example, the image decoding device 200 can perform decoding using a processing unit applied in the image encoding device. Therefore, the decoding processing unit can be, for example, a coding unit. The coding unit can be obtained by dividing a coding tree unit or a maximum coding unit. The reconstructed image signal decoded and output by the image decoding device 200 can be reproduced by a reproduction device (not shown).

[0072] The image decoding apparatus 200 may receive a signal output from the image encoding apparatus of FIG. 2 in a bitstream format. The received signal may be decoded via an entropy decoding unit 210. For example, the entropy decoding unit 210 may parse the bitstream to derive information (e.g., video / image information) necessary for image reconstruction (or picture reconstruction). The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / image information may also include general constraint information. The image decoding apparatus may further use the information on the parameter sets and / or the general constraint information to decode an image. The signaling information, received information, and / or syntax elements referred to in the present disclosure may be obtained from the bitstream by being decoded via the decoding procedure. For example, the entropy decoding unit 210 may decode information in a bitstream based on a coding method such as Exponential-Golomb coding, CAVLC, or CABAC, and output values ​​of syntax elements required for image restoration and quantized values ​​of transform coefficients related to residuals. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element from the bitstream, determines a context model using information on the syntax element to be decoded and decoded information on neighboring blocks and the block to be decoded, or information on symbols / bins decoded in a previous step, predicts the occurrence probability of the bins based on the determined context model, and performs arithmetic decoding of the bins to generate symbols corresponding to the values ​​of each syntax element. After determining the context model, the CABAC entropy decoding method may update the context model using information on the decoded symbol / bin for the context model of the next symbol / bin.Among the information decoded by the entropy decoding unit 210, information related to prediction is provided to the prediction units (inter prediction unit 260 and intra prediction unit 265), and residual values ​​entropy decoded by the entropy decoding unit 210, i.e., quantized transform coefficients and related parameter information, may be input to the inverse quantization unit 220. Also, among the information decoded by the entropy decoding unit 210, information related to filtering may be provided to the filtering unit 240. Meanwhile, a receiving unit (not shown) for receiving a signal output from the image encoding device may be further provided as an internal / external element of the image decoding device 200, or the receiving unit may be provided as a component of the entropy decoding unit 210.

[0073] Meanwhile, the image decoding apparatus according to the present disclosure may be referred to as a video / image / picture decoding apparatus. The image decoding apparatus may include an information decoder (video / image / picture information decoder) and / or a sample decoder (video / image / picture sample decoder). The information decoder may include an entropy decoding unit 210, and the sample decoder may include at least one of an inverse quantization unit 220, an inverse transform unit 230, an adder 235, a filtering unit 240, a memory 250, an inter prediction unit 260, and an intra prediction unit 265.

[0074] The inverse quantization unit 220 may inverse quantize the quantized transform coefficients and output the transform coefficients. The inverse quantization unit 220 may rearrange the quantized transform coefficients in a two-dimensional block format. In this case, the rearrangement may be performed based on the coefficient scanning order performed in the image encoding device. The inverse quantization unit 220 may perform inverse quantization on the quantized transform coefficients using a quantization parameter (e.g., quantization step size information) to obtain transform coefficients.

[0075] The inverse transform unit 230 can inversely transform the transform coefficients to obtain a residual signal (residual block, residual sample array).

[0076] The prediction unit may perform prediction on a current block and generate a predicted block including prediction samples for the current block. The prediction unit may determine whether intra prediction or inter prediction is applied to the current block based on information about the prediction output from the entropy decoding unit 210, and may determine a specific intra / inter prediction mode (prediction technique).

[0077] The prediction unit can generate a prediction signal based on various prediction methods (techniques) described below, as described in the description of the prediction unit of the image encoding device 100.

[0078] The intra predictor 265 may predict the current block by referring to samples in the current picture. The description of the intra predictor 185 may also be applied to the intra predictor 265.

[0079] The inter prediction unit 260 may derive a predicted block for a current block based on a reference block (reference sample array) identified by a motion vector on a reference picture. To reduce the amount of motion information transmitted in inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on correlations between motion information of neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on an inter prediction direction (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. For example, the inter prediction unit 260 may construct a motion information candidate list based on the neighboring blocks and derive a motion vector and / or a reference picture index for the current block based on received candidate selection information. Inter prediction may be performed based on various prediction modes (techniques), and the prediction information may include information indicating the inter prediction mode (technique) for the current block.

[0080] The adder 235 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to a prediction signal (predicted block, predicted sample array) output from a prediction unit (including the inter prediction unit 260 and / or intra prediction unit 265). When there is no residual for the current block, such as when a skip mode is applied, the predicted block can be used as the reconstructed block. The description of the adder 155 also applies to the adder 235. The adder 235 may also be referred to as a reconstruction unit or a reconstructed block generator. The generated reconstructed signal may be used for intra prediction of the next current block in the current picture, and may also be used for inter prediction of the next picture via filtering, as described below.

[0081] The filtering unit 240 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 240 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture, and may store the modified reconstructed picture in the memory 250, specifically, in a DPB of the memory 250. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, etc.

[0082] The (modified) reconstructed picture stored in the DPB of the memory 250 can be used as a reference picture in the inter predictor 260. The memory 250 can store motion information of a block from which motion information in the current picture is derived (or decoded) and / or motion information of a block in an already reconstructed picture. The stored motion information can be transmitted to the inter predictor 260 to be used as motion information of a spatially surrounding block or a temporally surrounding block. The memory 250 can store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra predictor 265.

[0083] In this specification, the embodiments described for the filtering unit 160, inter prediction unit 180 and intra prediction unit 185 of the image encoding device 100 can also be applied in a similar or corresponding manner to the filtering unit 240, inter prediction unit 260 and intra prediction unit 265 of the image decoding device 200, respectively.

[0084] General image / video coding procedures

[0085] In image / video coding, pictures constituting an image / video can be coded / decoded according to a sequence of decoding orders. A picture order corresponding to an output order of decoded pictures can be set to be different from the decoding order. Based on this, not only forward prediction but also backward prediction can be performed during inter prediction.

[0086] FIG. 4 is a flow chart diagram illustrating an outline of an image decoding procedure to which the embodiments of the present disclosure can be applied.

[0087] Each procedure shown in Fig. 4 may be performed by the image encoding apparatus of Fig. 3. For example, in Fig. 4, step S410 may be performed by the entropy decoding unit 210, step S420 may be performed by a prediction unit including an intra prediction unit 265 and an inter prediction unit 260, step S430 may be performed by a residual processing unit including an inverse quantization unit 220 and an inverse transform unit 230, step S440 may be performed by the addition unit 235, and step S450 may be performed by the filtering unit 240. Step S410 may include the information decoding procedure described in this disclosure, step S420 may include the inter / intra prediction procedure described in this disclosure, step S430 may include the residual processing procedure described in this disclosure, step S440 may include the block / picture reconstruction procedure described in this disclosure, and step S450 may include the in-loop filtering procedure described in this disclosure.

[0088] Referring to Figure 4, as shown in the description of Figure 3, the picture decoding procedure may generally include an image / video information acquisition procedure (S410) from a bitstream (through decoding), a picture reconstruction procedure (S420-S440), and an in-loop filtering procedure for the reconstructed picture (S450). The picture reconstruction procedure may be performed based on prediction samples and residual samples obtained through the inter / intra prediction (S420) and residual processing (S430, inverse quantization and inverse transform of quantized transform coefficients) processes described in this disclosure. A modified reconstructed picture may be generated through an in-loop filtering procedure for the reconstructed picture generated by the picture reconstruction procedure. The modified reconstructed picture may be output as a decoded picture or may be stored in a decoded picture buffer or memory 250 of the decoding device and used as a reference picture in the inter prediction procedure when decoding a subsequent picture. In some cases, the in-loop filtering procedure may be omitted. In this case, the reconstructed picture may be output as a decoded picture or may be stored in a decoded picture buffer or memory 250 of the decoding device and used as a reference picture in an inter-prediction procedure when decoding a subsequent picture. As described above, the in-loop filtering procedure (S450) may include a deblocking filtering procedure, a sample adaptive offset (SAO) procedure, an adaptive loop filter (ALF) procedure, and / or a bilateral filter procedure, some or all of which may be omitted. Furthermore, one or some of the deblocking filtering procedure, the sample adaptive offset (SAO) procedure, the adaptive loop filter (ALF) procedure, and the bilateral filter procedure may be applied sequentially, or all of them may be applied sequentially. For example, the SAO procedure may be performed after the deblocking filtering procedure is applied to the reconstructed picture.Or, for example, the ALF procedure can be performed after a deblocking filtering procedure is applied to the reconstructed picture, which can also be performed in the encoding device.

[0089] FIG. 5 is a flowchart that schematically illustrates an image encoding procedure to which the embodiments of the present disclosure can be applied.

[0090] Each procedure shown in Fig. 5 may be performed by the image encoding apparatus of Fig. 2. For example, step S510 may be performed by a prediction unit including an intra prediction unit 185 or an inter prediction unit 180, step S520 may be performed by a residual processing unit including a transform unit 120 and / or a quantization unit 130, and step S530 may be performed by an entropy encoding unit 190. Step S510 may include an inter / intra prediction procedure described in this disclosure, step S520 may include a residual processing procedure described in this disclosure, and step S530 may include an information encoding procedure described in this disclosure.

[0091] Referring to FIG. 5, the picture encoding procedure, as described with reference to FIG. 2, may include not only a procedure of encoding information for picture reconstruction (e.g., prediction information, residual information, partitioning information, etc.) and outputting it in a bitstream format, but also a procedure of generating a reconstructed picture for a current picture and an optional procedure of applying in-loop filtering to the reconstructed picture. The encoding apparatus may derive (modified) residual samples from quantized transform coefficients via the inverse quantization unit 140 and the inverse transform unit 150, and may generate a reconstructed picture based on the prediction samples output in step S510 and the (modified) residual samples. The reconstructed picture generated in this manner may be the same as the reconstructed picture generated by the decoding apparatus described above. A modified reconstructed picture may be generated through an in-loop filtering procedure on the reconstructed picture, which may be stored in the decoded picture buffer or memory 170 and used as a reference picture in the inter prediction procedure when encoding a subsequent picture, as in the decoding apparatus. As described above, in some cases, some or all of the in-loop filtering procedure may be omitted. When the in-loop filtering procedure is performed, (in-loop) filtering-related information (parameters) can be coded by the entropy coding unit 190 and output in bitstream format, and the decoding device can perform the in-loop filtering procedure in the same manner as the coding device based on the filtering-related information.

[0092] This in-loop filtering procedure can reduce noise that occurs during image / video coding, such as blocking artifacts and ringing artifacts, and improve subjective / objective visual quality. Also, by performing the in-loop filtering procedure in both the encoding device and the decoding device, the encoding device and the decoding device can derive the same prediction result, thereby improving the reliability of picture coding and reducing the amount of data to be transmitted for picture coding.

[0093] As described above, a picture reconstruction procedure may be performed not only in a decoding device but also in an encoding device. Reconstructed blocks may be generated based on intra prediction / inter prediction for each block, and a reconstructed picture including the reconstructed blocks may be generated. If a current picture / slice / tile group is an I picture / slice / tile group, blocks included in the current picture / slice / tile group may be reconstructed based only on intra prediction. On the other hand, if the current picture / slice / tile group is a P or B picture / slice / tile group, blocks included in the current picture / slice / tile group may be reconstructed based on intra prediction or inter prediction. In this case, inter prediction may be applied to some blocks in the current picture / slice / tile group, and intra prediction may be applied to the remaining blocks. Color components of a picture may include luma components and chroma components, and unless explicitly limited in this disclosure, methods and embodiments proposed in this disclosure may be applied to luma components and chroma components.

[0094] Example of coding hierarchy and structure

[0095] Video / images coded according to this disclosure may be processed, for example, according to the coding hierarchy and structure described below.

[0096] FIG. 6 shows an example of a hierarchical structure for coded images / video.

[0097] The coded image / video can be divided into the VCL (video coding layer), which handles the image / video decoding process and itself, the lower system, which transmits and stores the coded information, and the NAL (network abstraction layer), which exists between the VCL and the lower system and is responsible for network adaptation functions.

[0098] The VCL can generate VCL data containing compressed image data (slice data), or it can generate parameter sets containing information such as a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), and a Video Parameter Set (VPS), or an SEI (Supplemental Enhancement Information) message that is additionally required for image decoding processing.

[0099] In NAL, NAL units can be generated by adding header information (NAL unit header) to RBSP (Raw Byte Sequence Payload) generated by VCL. RBSP refers to slice data, parameter sets, SEI messages, etc. generated by VCL. The NAL unit header can include NAL unit type information identified by the RBSP data included in the corresponding NAL unit.

[0100] As shown in Figure 6, NAL units can be classified into VCL NAL units and non-VCL NAL units according to the type of RBSP generated in the VCL. A VCL NAL unit can refer to a NAL unit containing information about an image (slice data), and a non-VCL NAL unit can refer to a NAL unit containing information necessary for decoding an image (parameter set or SEI message).

[0101] The VCL NAL unit and non-VCL NAL unit described above can be transmitted over a network with header information attached according to the data standard of the lower system. For example, the NAL unit can be transformed into a data format of a predetermined standard such as the H.266 / VVC file format, the Real-time Transport Protocol (RTP), or the Transport Stream (TS) and then transmitted over various networks.

[0102] As described above, the NAL unit type of an NAL unit can be identified according to the RBSP data structure included in the NAL unit, and information about the NAL unit type can be stored and signaled in the NAL unit header. For example, NAL units can be broadly classified into VCL NAL unit types and non-VCL NAL unit types depending on whether the NAL unit includes information about an image (slice data). VCL NAL unit types can be classified according to the nature and type of pictures included in the VCL NAL unit, and non-VCL NAL unit types can be classified according to the type of parameter set.

[0103] Below is a list of examples of NAL unit types identified by the type of parameter set / information included in the non-VCL NAL unit type.

[0104] -DCI (Decoding capability information) NAL unit type (NUT): Type for NAL units including DCI

[0105] -VPS (Video Parameter Set) NUT: Type for NAL units containing VPS

[0106] -SPS (Sequence Parameter Set) NUT: Type for NAL units containing SPS

[0107] -PPS (Picture Parameter Set) NUT: Type for NAL units containing PPS

[0108] -APS (Adaptation Parameter Set) NUT: Type for NAL units containing APS

[0109] -PH(Picture header) NUT: Type for NUL units containing picture headers

[0110] The above-mentioned NAL unit type has syntax information for the NAL unit type, and the syntax information can be stored in a NAL unit header and signaled. For example, the syntax information is nal_unit_type, and the NAL unit type can be identified using the value of nal_unit_type.

[0111] Meanwhile, one picture may include multiple slices, and one slice may include a slice header and slice data. In this case, one picture header may be added to multiple slices (slice header and slice data set) in one picture. The picture header (picture header syntax) may include information / parameters commonly applicable to the picture. The slice header (slice header syntax) may include information / parameters commonly applicable to the slices. The APS (APS syntax) or PPS (PPS syntax) may include information / parameters commonly applicable to one or more slices or pictures. The SPS (SPS syntax) may include information / parameters commonly applicable to one or more sequences. The VPS (VPS syntax) may include information / parameters commonly applicable to multiple layers. The DCI may include information / parameters related to decoding capability.

[0112] In the present disclosure, the high level syntax (HLS) may include at least one of the APS syntax, PPS syntax, SPS syntax, VPS syntax, DCI syntax, picture header syntax, and slice header syntax. Also, in the present disclosure, the low level syntax (LLS) may include, for example, slice data syntax, CTU syntax, coding unit syntax, transform unit syntax, etc.

[0113] Meanwhile, in the present disclosure, image / video information encoded from an encoding device to a decoding device and signaled in a bitstream format may include not only intra-picture partitioning-related information, intra / inter prediction information, residual information, in-loop filtering information, etc., but also the slice header information, the picture header information, the APS information, the PPS information, the SPS information, the VPS information, and / or the DCI information. Also, the image / video information may further include general constraint information and / or NAL unit header information.

[0114] Entry points signaling overview

[0115] As described above, a VCL NAL unit can include slice data as Raw Byte Sequence Payload (RBSP). The slice data is byte-aligned within the VCL NAL unit and can include one or more subsets. At least one entry point for random access (RA) can be defined for the subsets, and parallel processing can be performed based on the entry point.

[0116] The VVC standard supports wavefront parallel processing (WPP), one of several parallel processing techniques, which allows multiple slices in a picture to be coded and decoded in parallel.

[0117] To activate parallel processing capability, entry point information can be signaled. Based on the entry point information, the image decoding device can directly access the start point of a data segment included in an NAL unit. Here, the start point of a data segment can mean the start point of a tile in a slice or the start point of a CTU row in a slice.

[0118] The entry point information can be signaled in higher level syntax, for example, a picture parameter set (PPS) and / or a slice header.

[0119] FIG. 7 is a diagram illustrating a picture parameter set (PPS) according to one embodiment of the present disclosure, and FIG. 8 is a diagram illustrating a slice header according to one embodiment of the present disclosure.

[0120] First, referring to FIG. 7, a picture parameter set (PPS) may include an entry_point_offsets_present_flag as a syntax element indicating whether or not entry point information is signaled.

[0121] The entry_point_offsets_present_flag may indicate whether or not signaling of entry point information is present in a slice header that references a picture parameter set (PPS). For example, an entry_point_offsets_present_flag having a first value (e.g., 0) may indicate that signaling of entry point information for a tile or specific CTU rows within a tile is not present in the slice header. In contrast, an entry_point_offsets_present_flag having a second value (e.g., 1) may indicate that signaling of entry point information for a tile or specific CTU rows within a tile is present in the slice header.

[0122] 7 illustrates a case where the entry_point_offsets_present_flag is included in a picture parameter set (PPS), but this is merely an example and the embodiments of the present disclosure are not limited thereto. For example, the entry_point_offsets_present_flag may be included in a sequence parameter set (SPS).

[0123] Next, referring to FIG. 8, the slice header may include offset_len_minus1 and entry_point_offset_minus1[i] as syntax elements for identifying the entry point.

[0124] offset_len_minus1 may indicate a value obtained by subtracting 1 from the bit length of entry_point_offset_minus1[i]. The value of offset_len_minus1 may range from 0 to 31. In one example, offset_len_minus1 may be signaled based on a variable NumEntryPoints representing the total number of entry points. For example, offset_len_minus1 may be signaled only when NumEntryPoints is greater than 0. Also, offset_len_minus1 may be signaled based on entry_point_offsets_present_flag described above with reference to FIG. 7. For example, offset_len_minus1 may be signaled only when entry_point_offsets_present_flag has a second value (e.g., 1) (i.e., when entry point information is signaled in the slice header).

[0125] entry_point_offset_minus1[i] can represent the i-th entry point offset in bytes, and can be expressed by adding 1 bit to offset_len_minus1. The slice data in an NAL unit can include the same number of subsets as the value obtained by adding 1 to NumEntryPoints, and the index value indicating each of the subsets can range from 0 to NumEntryPoints. The first byte of the slice data in an NAL unit can be represented by byte 0.

[0126] If entry_point_offset_minus1[i] is signaled, emulation prevention bytes included in slice data within a NAL unit can be counted as part of the slice data for subset identification. Subset 0, the first subset of slice data, can have a structure from byte 0 to entry_point_offset_minus1[0]. Similarly, subset k, the kth subset of slice data, can have a structure from firstByte[k] to lastByte[k]. Here, firstByte[k] can be derived as shown in Equation 1 below, and lastByte[k] can be derived as shown in Equation 2 below.

[0127]

number

[0128]

number

[0129] In Equation 1 and Equation 2, k is equal to or greater than 1 and can have a value in the range of NumEntryPoints minus 1.

[0130] The last subset of the slice data (i.e., the NumEntryPoints-th subset) may consist of the remaining bytes of the slice data.

[0131] On the other hand, if a predetermined synchronization process for context variables is not performed (e.g., sps_entropy_coding_sync_enabled_flag==0) before decoding the CTU containing the first CTB in the CTB row in each tile, and the slice contains one or more complete tiles, each subset of slice data can consist of all coded bits for all CTUs in the same tile. In this case, the total number of subsets of slice data can be the same as the total number of tiles in the slice.

[0132] Alternatively, if the predetermined synchronization process is not performed and the slice contains one subset for a row of CTUs in a single tile, NumEntryPoints may be 0. In this case, one subset of slice data may consist of all coded bits for all CTUs in the slice.

[0133] Alternatively, if the predetermined synchronization process is performed (e.g., sps_entropy_coding_sync_enabled_flag==1), each subset may consist of all coding bits for all CTUs in one CTU row within one tile. In this case, the total number of subsets of slice data may be the same as the total number of CTU rows for each tile within the slice.

[0134] Overview of mixed NAL unit types

[0135] Generally, one NAL unit type can be set for one picture. As described above, syntax information indicating the NAL unit type can be stored and signaled in a NAL unit header within the NAL unit. For example, the syntax information is nal_unit_type, and the NAL unit type can be identified using the nal_unit_type value.

[0136] An example of an NAL unit type to which the embodiment of the present disclosure can be applied is shown in Table 1 below.

[0137] [Table 1-1] [Table 1-2]

[0138] Referring to Table 1, VCL NAL unit types can be classified into NAL unit types 0 to 12 depending on the nature and type of picture, etc. Non-VCL NAL unit types can be classified into NAL unit types 13 to 31 depending on the type of parameter set, etc.

[0139] Examples of VCL NAL unit types are:

[0140] -IRAP (Intra Random Access Point) NAL unit type (NUT): Type for the NAL unit of an IRAP picture, set in the range from IDR_W_RADL to CRA_NUT.

[0141] IDR (Instantaneous Decoding Refresh) NUT: Type for the NAL unit of an IDR picture, set to IDR_W_RADL or IDR_N_LP.

[0142] CRA (Clean Random Access) NUT: Type for NAL units of CRA pictures, set to CRA_NUT.

[0143] RADL (Random Access Decodable Leading) NUT: Type for the NAL unit of a RADL picture, set to RADL_NUT.

[0144] RASL (Random Access Skipped Leading) NUT: Type for the NAL unit of a RASL picture, set to RASL_NUT.

[0145] - Trailing NUT: Type for the NAL unit of the trailing picture, set to TRAIL_NUT.

[0146] GDR (Gradual Decoding Refresh) NUT: Type for NAL units of GDR pictures, set to GDR_NUT.

[0147] STSA (Step-wise Temporal Sublayer Access) NUT: Type for the NAL unit of an STSA picture, set to STSA_NUT.

[0148] Meanwhile, the VVC standard allows a single picture to include multiple slices with different NAL unit types. For example, a single picture may include at least one first slice with a first NAL unit type and at least one second slice with a second NAL unit type different from the first NAL unit type. In this case, the NAL unit type of the picture may be referred to as a mixed NAL unit type. As such, the VVC standard's support for mixed NAL unit types allows multiple pictures to be more easily reconstructed / combined during content composition, encoding / decoding, and other processes.

[0149] However, the existing scheme for hybrid NAL unit types has a limitation that only hybridization between two NAL unit types is allowed. Furthermore, if a picture has a hybrid NAL unit type, one or more VCL NAL units of the picture must have a NAL unit type ranging from IDR_W_RADL to CRA_NUT, and the remaining VCL NAL units of the picture must have a NAL unit type ranging from TRAIL_NUT to RSV_VCL_6. As a result, although hybrid NAL unit types are useful in image processing, they cannot be generally utilized.

[0150] To solve this problem, according to an embodiment of the present disclosure, mixing between two or more NAL unit types is permitted, and more diverse types of mixed NAL unit types can be provided.

[0151] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings.

[0152] If a picture contains two or more slices, each of which has a different NAL unit type, the picture may be restricted to contain two or more sub-pictures. That is, if a picture has a mixed NAL unit type, the picture may contain two or more sub-pictures.

[0153] A sub-picture may include one or more slices and may form a rectangular area within a picture. The sizes of the sub-pictures within a picture may be set to be different from each other. Alternatively, the size and position of a particular individual sub-picture may be set to be the same for all pictures in a sequence.

[0154] FIG. 9 is a diagram showing an example of a sub-picture.

[0155] Referring to Figure 9, one picture can be divided into 18 tiles. 12 tiles can be arranged on the left side of the picture, and each tile can include one slice consisting of a 4x4 CTU. 6 tiles can be arranged on the right side of the picture, and each tile can include two slices, each consisting of a 2x2 CTU, stacked vertically. As a result, the picture includes 24 sub-pictures and 24 slices, and each sub-picture can include one slice.

[0156] In one embodiment, each sub-picture in a picture can be treated as a single picture to support mixed NAL unit types. When a sub-picture is treated as a single picture, the sub-picture can be independently coded / decoded regardless of the coding / decoding results of other sub-pictures. Here, independent coding / decoding may mean that the block division structure (e.g., single tree structure, dual tree structure, etc.), prediction mode type (e.g., intra prediction, inter prediction, etc.), decoding procedure, etc. of the sub-picture can be set differently from other sub-pictures. For example, if a first sub-picture is coded / decoded based on an intra prediction mode, a second sub-picture adjacent to the first sub-picture and treated as a single picture can be coded / decoded based on an inter prediction mode.

[0157] Thus, if a picture contains two or more independent sub-pictures, each of which has a different NAL unit type, the picture can have a mixed NAL unit type.

[0158] FIG. 10 is a diagram showing an example of a picture having mixed NAL unit types.

[0159] 10, a picture 1000 can include first to third sub-pictures 1010 to 1030. The first and third sub-pictures 1010 and 1030 can each include two slices, while the second sub-picture 1020 can include four slices.

[0160] When the first to third subpictures 1010 to 1030 are treated as one picture, they can be coded independently to form different bitstreams. For example, coded slice data of the first subpicture 1010 can be encapsulated into one or more NAL units having a NAL unit type such as RASL_NUT to form a first bitstream (Bitstream 1). Coded slice data of the second subpicture 1020 can be encapsulated into one or more NAL units having a NAL unit type such as RADL_NUT to form a second bitstream (Bitstream 2). Coded slice data of the third subpicture 1030 can be encapsulated into one or more NAL units having a NAL unit type such as RASL_NUT to form a third bitstream (Bitstream 3). As a result, one picture 1000 can have a mixed NAL unit type, which is a combination of RASL_NUT and RADL_NUT.

[0161] In one embodiment, all slices included in each sub-picture within a picture can be restricted to have the same NAL unit type. For example, two slices included in the first sub-picture 1010 can each have a NAL unit type such as RASL_NUT. Four slices included in the second sub-picture 1020 can each have a NAL unit type such as RADL_NUT. Two slices included in the third sub-picture 1030 can each have a NAL unit type such as RASL_NUT.

[0162] Information about subpictures can be signaled in higher level syntax, such as a picture parameter set (PPS) and a sequence parameter set (SPS), and information about the application of hybrid NAL unit types can be signaled in higher level syntax, such as a picture parameter set (PPS).

[0163] FIG. 11 is a diagram illustrating an example of a picture parameter set (PPS) according to an embodiment of the present disclosure, and FIG. 12 is a diagram illustrating an example of a sequence parameter set (SPS) according to an embodiment of the present disclosure.

[0164] First, referring to FIG. 11, a picture parameter set (PPS) may include pps_mixed_nalu_types_in_pic_flag (or mixed_nalu_types_in_pic_flag) as a syntax element regarding whether or not mixed NAL unit types are applied.

[0165] The pps_mixed_nalu_types_in_pic_flag may indicate whether the current picture has mixed NAL unit types. For example, a pps_mixed_nalu_types_in_pic_flag with a first value (e.g., 0) may indicate that the current picture does not have mixed NAL unit types. In this case, the current picture may have the same NAL unit type for all VCL NAL units, for example, the same NAL unit type as the coded slice NAL unit. In contrast, a pps_mixed_nalu_types_in_pic_flag with a second value (e.g., 1) may indicate that the current picture has mixed NAL unit types.

[0166] In one embodiment, if the current picture has mixed NAL unit types (eg, pps_mixed_nalu_types_in_pic_flag==1), the VCL NAL units of the current picture can be restricted to not have NAL unit types such as GDR_NUT.

[0167] In one embodiment, if the current picture has mixed NAL unit types (e.g., pps_mixed_nalu_types_in_pic_flag==1), if any VCL NAL unit of the current picture has a NAL unit type such as IDR_W_RADL, IDR_N_LP, or CRA_NUT (NAL unit type A), all other VCL NAL units of the current picture can be restricted to have a NAL unit type such as NAL unit type A or a NAL unit type such as TRAIL_NUT. For example, if any VCL NAL unit of the current picture has a NAL unit type such as IDR_W_RADL, all other NAL units of the current picture can have a NAL unit type such as IDR_W_RADL or TRAIL_NUT.

[0168] If the current picture has a mixed NAL unit type (e.g., pps_mixed_nalu_types_in_pic_flag==1), each sub-picture in the current picture can have any of the VCL NAL unit types described above with reference to Table 1. For example, if a sub-picture in the current picture is an IDR sub-picture, the sub-picture can have a NAL unit type such as IDR_W_RADL or IDR_N_LP. Alternatively, if a sub-picture in the current picture is a trailing sub-picture, the sub-picture can have a NAL unit type such as TRAIL_NUT.

[0169] Therefore, a pps_mixed_nalu_types_in_pic_flag having a second value (e.g., 1) may indicate that a picture referencing a picture parameter set (PPS) may contain slices with different NAL unit types. Here, the picture may originate from a sub-picture bitstream merging operation in which the encoder must ensure matching of bitstream structures and alignment between parameters of the original bitstreams. As an example of the alignment, if a reference picture list (RPL) syntax element for a slice with an NAL unit type such as IDR_W_RADL or IDR_N_LP is not present in the slice header (e.g., sps_idr_rpl_present_flag==0) and the current picture including the slice has mixed NAL unit types (e.g., pps_mixed_nalu_types_in_pic_flag==1), the current picture may be restricted not to contain slices with NAL unit types such as IDR_W_RADL or IDR_N_LP.

[0170] On the other hand, if mixed NAL unit types are restricted from being applied to all pictures in an Output Layer Set (OLS) (e.g., gci_no_mixed_nalu_types_in_pic_constraint_flag==1), pps_mixed_nalu_types_in_pic_flag may have the first value (e.g., 0).

[0171] Furthermore, the picture parameter set (PPS) may include pps_no_pic_partition_flag (or no_pic_partion_flag) as a syntax element indicating whether or not a picture is partitioned.

[0172] The pps_no_pic_partition_flag may indicate whether picture partitioning is applicable to the current picture. For example, a pps_no_pic_partition_flag having a first value (e.g., 0) may indicate that the current picture cannot be partitioned. In contrast, a pps_no_pic_partition_flag having a second value (e.g., 1) may indicate that the current picture can be partitioned into two or more tiles or slices. If the current picture has mixed NAL unit types (e.g., pps_mixed_nalu_types_in_pic_flag==1), the current picture may be restricted to being partitioned into two or more tiles or slices (e.g., pps_no_pic_partition_flag=1).

[0173] Furthermore, the picture parameter set (PPS) may include pps_num_subpics_minus1 (or num_subpics_minus1) as a syntax element that indicates the number of subpictures.

[0174] pps_num_subpics_minus1 may represent the number of subpictures included in the current picture minus 1. pps_num_subpics_minus1 can be signaled only if picture partitioning is applicable to the current picture (e.g., pps_no_pic_partition_flag==1). If pps_num_subpics_minus1 is not signaled, the value of pps_num_subpics_minus1 can be inferred as 0. On the other hand, a syntax element representing the number of subpictures can also be signaled in a higher-level syntax different from the picture parameter set (PPS), for example, a sequence parameter set (SPS).

[0175] In one embodiment, if the current picture contains only one subpicture (e.g., pps_num_subpics_minus1==0), the current picture can be restricted to not have mixed NAL unit types (e.g., pps_mixed_nalu_types_in_pic_flag=0), i.e., if the current picture has mixed NAL unit types (e.g., pps_mixed_nalu_types_in_pic_flag==1), the current picture can be restricted to contain two or more subpictures (e.g., pps_num_subpics_minus1>0).

[0176] Next, referring to FIG. 12, the sequence parameter set (SPS) may include sps_subpic_treated_as_pic_flag[i] (or subpic_treated_as_pic_flag[i]) as a syntax element related to the handling of subpictures during encoding / decoding.

[0177] sps_subpic_treated_as_pic_flag[i] may indicate whether each subpicture in the current picture is treated as one picture. For example, a first value (e.g., 0) of sps_subpic_treated_as_pic_flag[i] may indicate that the i-th subpicture in the current picture is not treated as one picture. In contrast, a second value (e.g., 1) of sps_subpic_treated_as_pic_flag[i] may indicate that the i-th subpicture in the current picture is treated as one picture in the encoding / decoding process except for the in-loop filtering operation. If sps_subpic_treated_as_pic_flag[i] is not signaled, sps_subpic_treated_as_pic_flag[i] may be inferred to have the second value (e.g., 1).

[0178] In one embodiment, if the current picture contains two or more subpictures (e.g., pps_num_subpics_minus1>0) and at least one of the subpictures is not treated as one picture (e.g., sps_subpic_treated_as_pic_flag[i]==0), the current picture can be restricted to not have mixed NAL unit types (e.g., pps_mixed_nalu_types_in_pic_flag=0). That is, if the current picture has mixed NAL unit types (e.g., pps_mixed_nalu_types_in_pic_flag==1), any subpictures in the current picture can be restricted to be treated as one picture (e.g., sps_subpic_treated_as_pic_flag[i]=1).

[0179] Hereinafter, NAL unit types according to an embodiment of the present disclosure will be described in detail for each picture type.

[0180] (1) IRAP (Intra Random Access Point) picture

[0181] An IRAP picture is a randomly accessible picture and may have an NAL unit type such as IDR_W_RADL, IDR_N_LP, or CRA_NUT, as described above with reference to Table 1. An IRAP picture may not refer to any other picture other than the IRAP picture for inter prediction during decoding. An IRAP picture may include an instantaneous decoding refresh (IDR) picture and a clean random access (CRA) picture.

[0182] The first picture in a bitstream in decoding order may be restricted to be an IRAP picture or a Gradual Decoding Refresh (GDR) picture. For a single-layer bitstream, if a mandatory parameter set to refer to is available, the IRAP picture and all non-RASL pictures following the IRAP picture in decoding order can be correctly decoded, even if none of the pictures preceding the IRAP picture in decoding order are decoded.

[0183] In one embodiment, an IRAP picture may not have mixed NAL unit types. That is, the pps_mixed_nalu_types_in_pic_flag described above for an IRAP picture may have a first value (e.g., 0), and all slices in an IRAP picture may have the same NAL unit type in the range of IDR_W_RADL to CRA_NUT. As a result, if the first decoded slice in a picture has an NAL unit type in the range of IDR_W_RADL to CRA_NUT, the picture can be determined to be an IRAP picture.

[0184] (2) CRA (Clean Random Access) Pictures

[0185] A CRA picture is one of the IRAP pictures and may have a NAL unit type such as CRA_NUT as described above with reference to Table 1. A CRA picture may not need to refer to any other picture other than the CRA picture for inter prediction during decoding.

[0186] A CRA picture may be the first picture in the bitstream in decoding order, or any picture after the first. A CRA picture can be associated with a RADL or RASL picture.

[0187] When NoIncorrectPicOutputFlag has a second value (e.g., 1) for a CRA picture, the RASL picture associated with the CRA picture cannot be decoded because it references a picture that does not exist in the bitstream, and as a result, may not be output by the image decoding device. Here, NoIncorrectPicOutputFlag may indicate whether a picture preceding a recovery point picture in decoding order can be output before the recovery point picture. For example, NoIncorrectPicOutputFlag having a first value (e.g., 0) may indicate that a picture preceding a recovery point picture in decoding order can be output before the recovery point picture. In this case, the CRA picture may not be the first picture in the bitstream or the first picture following an End Of Sequence (EOS) NAL unit in decoding order. This may indicate that random access does not occur. Alternatively, NoIncorrectPicOutputFlag having a second value (e.g., 1) may indicate that a picture preceding the recovery point picture in decoding order cannot be output before the recovery point picture. In this case, the CRA picture may be the first picture in the bitstream or the first picture following the EOS NAL unit in decoding order. This may indicate the case where random access occurs. Meanwhile, NoIncorrectPicOutputFlag may also be referred to as NoOutputBeforeRecoveryFlag depending on the embodiment.

[0188] For all picture units (PUs) following the current picture in decoding order in a coded layer video sequence (CLVS), reference picture list 0 (e.g., RefPicList[0]) and reference picture list 1 (e.g., RefPicList[1]) for one slice included in a CRA sub-picture belonging to the picture units (PUs) may be restricted to not include any picture in an active entry that precedes the picture including the CRA sub-picture in decoding order. Here, a picture unit (PU) may refer to a set of NAL units that are related to one another according to a predetermined classification rule for one coded picture and include multiple NAL units that are consecutive in decoding order.

[0189] (3) IDR (Instantaneous Decoding Refresh) Picture

[0190] An IDR picture is one of the IRAP pictures and may have an NAL unit type such as IDR_W_RADL or IDR_N_LP, as described above with reference to Table 1. An IDR picture may not need to refer to any other picture other than the IDR picture for inter prediction during decoding.

[0191] An IDR picture may be the first picture in a bitstream in decoding order, or any picture after the first. Each IDR picture may be the first picture in a Coded Video Sequence (CVS) in decoding order.

[0192] If an IDR picture has a NAL unit type such as IDR_W_RADL for each NAL unit, the IDR picture may have an associated RADL picture. Alternatively, if an IDR picture has a NAL unit type such as IDR_N_LP for each NAL unit, the IDR picture may not have an associated leading picture. On the other hand, the IDR picture may not be associated with a RADL picture.

[0193] For all picture units (PUs) that follow the current picture in decoding order in a Coded Layer Video Sequence (CLVS), reference picture list 0 (e.g., RefPicList[0]) and reference picture list 1 (e.g., RefPicList[1]) for one slice included in an IDR sub-picture belonging to the picture units (PUs) can be restricted to not include any pictures in active entries that precede the picture containing the IDR sub-picture in decoding order.

[0194] (4) RADL (Random Access Decodable Leading) Picture

[0195] A RADL picture is one of the leading pictures and can have a NAL unit type such as RADL_NUT as described above with reference to Table 1.

[0196] An RADL picture may not be used as a reference picture during decoding of a trailing picture associated with the same IRAP picture as the RADL picture. When field_seq_flag has a first value (e.g., 0) for an RADL picture, the RADL picture may precede all non-leading pictures associated with the same IRAP picture in decoding order. Here, field_seq_flag may indicate whether a Coded Layer Video Sequence (CLVS) conveys pictures representing fields or frames. For example, field_seq_flag having a first value (e.g., 0) may indicate that the CLVS conveys pictures representing frames. In contrast, field_seq_flag having a second value (e.g., 1) may indicate that the CLVS conveys pictures representing fields.

[0197] (5) RASL (Random Access Skipped Leading) Picture

[0198] A RASL picture is one of the leading pictures and can have a NAL unit type such as RASL_NUT as described above with reference to Table 1.

[0199] In one example, all RASL pictures may be leading pictures of the associated CRA picture. If NoIncorrectPicOutputFlag has a second value (e.g., 1) for the CRA picture, the RASL picture cannot be decoded because it references a picture that does not exist in the bitstream, and as a result, it may not be output by the image decoding device.

[0200] An RASL picture does not need to be used as a reference picture during the decoding of a non-RASL picture, but if there is an RADL picture that belongs to the same layer as the RASL picture and is associated with the same CRA picture, the RASL picture can be used as a collocated reference picture for inter-prediction of an RADL sub-picture included in the RADL picture.

[0201] If field_seq_flag has the first value (eg, 0) for an RASL picture, the RASL picture may precede in decoding order all non-leading pictures of the CRA picture associated with the RASL picture.

[0202] (6) Trailing Picture

[0203] A trailing picture is a non-IRAP picture that follows the associated IRAP or GDR picture in output order, but may not be an STSA picture. A trailing picture can also follow the associated IRAP picture in decoding order. That is, a trailing picture that follows the associated IRAP picture in output order but precedes it in decoding order is not allowed.

[0204] (7) GDR (Gradual Decoding Refresh) Picture

[0205] A GDR picture is a randomly accessible picture and can have a NAL unit type such as GDR_NUT, as described above with reference to Table 1.

[0206] (8) STSA (Step-wise Temporal Sublayer Access) Picture

[0207] An STSA picture is a randomly accessible picture and can have a NAL unit type such as STSA_NUT as described above with reference to Table 1.

[0208] An STSA picture may not reference a picture with the same TemporalId as the STSA picture for inter prediction. Here, TemporalId may be an identifier indicating a temporal hierarchy, such as a temporal sublayer, in scalable video coding. As an example, an STSA picture may be restricted to have a TemporalId greater than 0.

[0209] A picture having the same Temporal ID as an STSA picture and following the STSA picture in decoding order may not refer to a picture having the same Temporal ID as the STSA picture and preceding the STSA picture in decoding order for inter prediction. An STSA picture may activate up-switching from an immediately lower sublayer of the current sublayer to which the STSA picture belongs to the current sublayer.

[0210] FIG. 13 is a diagram for explaining the decoding order and output order for each picture type.

[0211] Pictures can be classified as I pictures, P pictures, or B pictures depending on the prediction method. An I picture refers to a picture to which only intra prediction can be applied and can be decoded without reference to other pictures. An I picture can also be called an intra picture and can include the above-mentioned IRAP picture. A P picture refers to a picture to which intra prediction and unidirectional inter prediction can be applied and can be decoded with reference to one other picture. A B picture refers to a picture to which intra prediction and bidirectional / unidirectional inter prediction can be applied and can be decoded with reference to one or two other pictures. P pictures and B pictures are also called inter pictures and can include the above-mentioned RADL pictures, RASL pictures, and trailing pictures.

[0212] Inter-pictures can be further classified into leading pictures (LP) and non-leading pictures (NLP) according to the decoding order and output order. A leading picture refers to a picture that follows an IRAP picture in decoding order and precedes the IRAP picture in output order, and can include the above-mentioned RADL picture and RASL picture. A non-leading picture refers to a picture that follows an IRAP picture in decoding order and output order, and can include the above-mentioned trailing picture.

[0213] In Figure 13, each picture name may indicate the picture type described above. For example, I5 picture may be an I picture, B0, B2, B3, B4, and B6 pictures may be B pictures, and P1 and P7 pictures may be P pictures. Furthermore, in Figure 13, each arrow may indicate a reference direction between pictures. For example, B0 picture can be decoded by referring to P1 picture.

[0214] 13, the I5 picture may be an IRAP picture, for example, a CRA picture. When random access to the I5 picture is performed, the I5 picture may be the first picture in decoding order.

[0215] The B0 and P1 pictures precede the I5 picture in decoding order and may form a separate video sequence from the I5 picture, while the B2, B3, B4, B6, and P7 pictures follow the I5 picture in decoding order and may form a video sequence together with the I5 picture.

[0216] The B2, B3, and B4 pictures can be classified as leading pictures because they follow the I5 picture in decoding order and precede the I5 picture in output order. The B2 picture can be decoded by referring to the P1 picture preceding the I5 picture in decoding order. Therefore, when random access to the I5 picture occurs, the B2 picture cannot be correctly decoded because it references the P1 picture, which does not exist in the bitstream. A picture type such as the B2 picture can be called a RASL picture. In contrast, the B3 picture can be decoded by referring to the I5 picture preceding the B3 picture in decoding order. Therefore, when random access to the I5 picture occurs, the B3 picture can be correctly decoded by referring to the already decoded I5 picture. Furthermore, the B4 picture can be decoded by referring to the I5 picture and the B3 picture preceding the B4 picture in decoding order. Therefore, when random access to the I5 picture occurs, the B4 picture can be correctly decoded by referring to the already decoded I5 picture and the B3 picture. Picture types such as the B3 and B4 pictures can be called RADL pictures.

[0217] On the other hand, pictures B6 and P7 follow picture I5 ​​in decoding order and output order, so they can be classified as non-leading pictures. Picture B6 can be decoded by referring to pictures I5 and P7, which precede picture I5 ​​in decoding order. Therefore, when random access to picture I5 ​​occurs, picture B6 can be correctly decoded by referring to pictures I5 and P7, which have already been decoded.

[0218] Thus, within one video sequence, the decoding process and the output process may be performed in different orders depending on the picture type. For example, if one video sequence includes IRAP pictures, leading pictures, and non-leading pictures, the decoding process may be performed in the order of IRAP pictures, leading pictures, and non-leading pictures, but the output process may be performed in the order of leading pictures, IRAP pictures, and non-leading pictures.

[0219] FIG. 14 is a flowchart illustrating a method for determining a NAL unit type of a current picture according to one embodiment of the present disclosure.

[0220] Referring to FIG. 14, the image coding apparatus may determine whether the current picture includes two or more subpictures (S1410). Partition information for the current picture may be signaled using one or more syntax elements in higher-level syntax. For example, no_pic_partition_flag, which indicates whether the current picture is partitioned, and pps_num_subpics_minus1, which indicates the number of subpictures included in the current picture, may be signaled via the picture parameter set (PPS) described above with reference to FIG. 7. If the current picture includes two or more subpictures, no_pic_partition_flag may have a first value (e.g., 0), and pps_num_subpics_minus may have a value greater than 0.

[0221] If the current picture does not include two or more sub-pictures ("NO" in S1410), the image coding apparatus may determine that the current picture has a single NAL unit type (S1450). That is, all slices in the current picture can have any of the multiple NAL unit types described above with reference to Table 1.

[0222] Alternatively, if the current picture includes two or more sub-pictures ("YES" in S1410), the image coding apparatus may determine whether each sub-picture in the current picture is treated as one picture (S1420). Information regarding whether a sub-picture is treated as one picture may be signaled using a predetermined syntax element in a higher-level syntax. For example, sps_subpic_treated_as_pic_flag[i], which indicates whether a sub-picture is treated as one picture, may be signaled via the sequence parameter set (SPS) described above with reference to FIG. 12. If each sub-picture in the current picture is treated as one picture, sps_subpic_treated_as_pic_flag[i] may have a second value (e.g., 1).

[0223] If the sub-pictures in the current picture are not treated as one picture ("NO" in S1420), the image coding apparatus may determine that the current picture has a single NAL unit type (S1450).

[0224] Alternatively, if each sub-picture in the current picture is treated as one picture ("YES" in S1420), the image coding device can determine whether at least some of the slices in the current picture have different NAL unit types (S1430).

[0225] If at least some of the slices in the current picture have different NAL unit types ("YES" in S1430), the image coding apparatus may determine that the current picture has mixed NAL unit types (S1440).

[0226] Alternatively, if all of the slices in the current picture have the same NAL unit type (NO in S1430), the image coding apparatus may determine that the current picture has a single NAL unit type (S1450).

[0227] If the current picture has a hybrid NAL unit type, the image coding apparatus can generate a sub-bitstream for each sub-picture included in the current picture, and the multiple sub-bitstreams generated for the current picture can form a single bitstream.

[0228] Information regarding whether the current picture has a mixed NAL unit type may be signaled using a predetermined syntax element in the higher-level syntax. For example, pps_mixed_nalu_types_in_pic_flag, which indicates whether the current picture has a mixed NAL unit type, may be signaled via the picture parameter set (PPS) described above with reference to FIG. 11. In this case, the image decoding apparatus may determine whether the current picture has a mixed NAL unit type based on pps_mixed_nalu_types_in_pic_flag. For example, if pps_mixed_nalu_types_in_pic_flag has a first value (e.g., 0), the image decoding apparatus may determine that the current picture has a single NAL unit type. In contrast, if pps_mixed_nalu_types_in_pic_flag has a second value (e.g., 1), the image decoding apparatus may determine that the current picture has a mixed NAL unit type.

[0229] Meanwhile, all slices in each sub-picture included in the current picture may have the same NAL unit type. That is, each sub-picture may have only a single NAL unit type. Also, if the current picture has mixed NAL unit types, at least some of the sub-pictures included in the current picture may have different NAL unit types. For example, a first sub-picture included in the current picture may have a first NAL unit type, and a second sub-picture included in the current picture may have a second NAL unit type different from the first NAL unit type.

[0230] The type of each sub-picture included in the current picture may be determined based on the NAL unit type of the sub-picture. For example, as described above with reference to Table 1, the type of a sub-picture having a NAL unit type such as TRAIL_NUT may be determined as a trailing sub-picture. Also, the type of a sub-picture having a NAL unit type such as STSA_NUT may be determined as an STSA sub-picture, and the type of a sub-picture having a NAL unit type such as GDR_NUT may be determined as a GDR sub-picture. Also, the type of a sub-picture having a NAL unit type such as RADL_NUT may be determined as an RADL sub-picture, and the type of a sub-picture having a NAL unit type such as RASL_NUT may be determined as an RASL sub-picture. Also, a sub-picture type having a NAL unit type such as IDR_W_RADL or IDR_N_LP can be determined as an IDR sub-picture, and a sub-picture type having a NAL unit type such as CRA_NUT can be determined as a CRA sub-picture.

[0231] As described above, according to the embodiments of the present disclosure, when a current picture includes two or more sub-pictures, each of which is treated as a single picture, the current picture can have a mixed NAL unit type, which can further improve image encoding / decoding efficiency in various applications and use cases.

[0232] Hereinafter, an image encoding / decoding method according to an embodiment of the present disclosure will be described in detail with reference to FIGS.

[0233] FIG. 15 is a flowchart illustrating an image encoding method according to an embodiment of the present disclosure.

[0234] The image coding method of Fig. 15 can be performed by the image coding apparatus of Fig. 2. For example, step S1510 can be performed by the image dividing unit 110, and steps S1520 and S1530 can be performed by the entropy coding unit 190.

[0235] Referring to FIG. 15, the image coding apparatus can divide a current picture into one or more sub-pictures (S1510).

[0236] As an example, if the current picture has mixed NAL units, the current picture can be restricted to include two or more sub-pictures, i.e., the current picture can be divided into two or more sub-pictures based on the fact that at least some of the slices included in the current picture have different NAL unit types.

[0237] Partition information of the current picture may be coded / signaled using a predetermined syntax element in the higher level syntax. For example, pps_no_pic_partition_flag indicating whether the current picture is partitioned and flag information pps_num_subpics_minus1 indicating the number of subpictures included in the current picture may be coded / signaled via the picture parameter set (PPS) described above with reference to Figure 11. In this case, if the current picture is partitioned into two or more subpictures, pps_no_pic_partition_flag may have a first value (e.g., 0), and pps_num_subpics_minus may have a value greater than 0.

[0238] In one embodiment, when a current picture has a hybrid NAL unit type, each sub-picture included in the current picture can be treated as one picture. For example, a current picture having a hybrid NAL unit type can include a first sub-picture and a second sub-picture, and the first sub-picture and the second sub-picture can be coded / decoded independently except for in-loop filtering operations.

[0239] Information regarding whether subpictures included in the current picture are treated as one picture may be coded / signaled using a predetermined syntax element in higher-level syntax. For example, flag information sps_subpic_treated_as_pic_flag[i] indicating whether the i-th subpicture included in the current picture is treated as one picture may be coded / signaled via the sequence parameter set (SPS) described above with reference to FIG. 12. In this case, if the i-th subpicture included in the current picture is not treated as one picture, sps_subpic_treated_as_pic_flag[i] may have a first value (e.g., 0). On the other hand, if the i-th subpicture included in the current picture is treated as one picture, sps_subpic_treated_as_pic_flag[i] may have a second value (e.g., 1).

[0240] The image coding apparatus can determine the NAL unit type of each of the slices included in the current picture (S1520).

[0241] In one embodiment, all slices included in each sub-picture in the current picture may have the same NAL unit type. For example, all slices included in a first sub-picture of the current picture may have a first NAL unit type, and all slices included in a second sub-picture of the current picture may have a second NAL unit type. In this case, if the current picture has mixed NAL unit types, the second NAL unit type may be different from the first NAL unit type. Alternatively, if the current picture has a single NAL unit type, the second NAL unit type may be the same as the first NAL unit type.

[0242] In one embodiment, whether at least some of the slices included in the current picture have different NAL unit types (i.e., whether the current picture has mixed NAL unit types) may be coded / signaled using a predetermined syntax element in higher-level syntax. For example, flag information pps_mixed_nalu_types_in_pic_flag (or mixed_nalu_types_in_pic_flag), which indicates whether the current picture has mixed NAL unit types, may be coded / signaled via the Picture Parameter Set (PPS) described above with reference to FIG. 11. In this case, if the current picture has a single NAL unit type, pps_mixed_nalu_types_in_pic_flag may have a first value (e.g., 0). On the other hand, if the current picture has mixed NAL unit types, pps_mixed_nalu_types_in_pic_flag may have a second value (e.g., 1).

[0243] In one embodiment, the value of pps_mixed_nalu_types_in_pic_flag may be determined based on whether hybrid NAL unit types are applicable to all pictures in an Output Layer Set (OLS). For example, if hybrid NAL unit types are not applicable to all pictures in an Output Layer Set (OLS), pps_mixed_nalu_types_in_pic_flag may be restricted to have a first value (e.g., 0). Alternatively, if hybrid NAL unit types are applicable to all pictures in an Output Layer Set (OLS), pps_mixed_nalu_types_in_pic_flag may have a first value (e.g., 0) or a second value (e.g., 1).

[0244] Meanwhile, whether a hybrid NAL unit type is applicable to all pictures in an output layer set (OLS) may be coded / signaled using a predetermined syntax element in a higher level syntax. For example, flag information gci_no_mixed_nalu_types_in_pic_constraint_flag (or no_mixed_nalu_types_in_pic_constraint_flag) indicating whether a hybrid NAL unit type is restricted may be coded / signaled via general_constraints_info including general constraint information. In this case, if a hybrid NAL unit type is applicable to all pictures in an output layer set (OLS), gci_no_mixed_nalu_types_in_pic_constraint_flag may have a first value (e.g., 0). Alternatively, if mixed NAL unit types are not applicable to all pictures in an output layer set (OLS), gci_no_mixed_nalu_types_in_pic_constraint_flag may have a second value (eg, 1).

[0245] In one embodiment, the value of pps_mixed_nalu_types_in_pic_flag may be determined based on whether the current picture includes only one subpicture. For example, if the current picture includes only one subpicture, pps_mixed_nalu_types_in_pic_flag may be restricted to have a first value (e.g., 0). Alternatively, if the current picture includes two or more subpictures, pps_mixed_nalu_types_in_pic_flag may have a first value (e.g., 0) or a second value (e.g., 1).

[0246] In one embodiment, the value of pps_mixed_nalu_types_in_pic_flag may be determined based on whether all sub-pictures included in the current picture are treated as a single picture. For example, if at least one of the sub-pictures included in the current picture is not treated as a single picture, pps_mixed_nalu_types_in_pic_flag may be restricted to have a first value (e.g., 0). Alternatively, if all sub-pictures included in the current picture are treated as a single picture, pps_mixed_nalu_types_in_pic_flag may have a first value (e.g., 0) or a second value (e.g., 1).

[0247] The image coding apparatus may encode the slices included in the current picture based on the NAL unit type determined in step S1520 (S1530). As described above, the encoding process for each slice may be performed in units of coding units (CUs) based on a predetermined prediction mode. Meanwhile, each sub-picture may be encoded independently to form different sub-bitstreams. For example, a first sub-bitstream containing coding information for a first sub-picture may be formed, and a second sub-bitstream containing coding information for a second sub-picture may be formed.

[0248] FIG. 16 is a flowchart illustrating an image decoding method according to an embodiment of the present disclosure.

[0249] The image decoding method of Fig. 16 may be performed by the image decoding apparatus of Fig. 3. For example, steps S1610 and S1620 may be performed by the entropy decoding unit 210, and step S1630 may be performed by any of the inverse quantization unit 220 to the intra prediction unit 265.

[0250] Referring to FIG. 16, the image decoding device can obtain, from the bitstream, VCL NAL unit type information of the current picture and first flag information indicating whether the sub-pictures included in the current picture are treated as one picture (S1610).

[0251] The VCL NAL unit type information of the current picture may include a NAL unit type value of a VCL NAL unit that contains coded image data (e.g., slice data) of the current picture, which may be obtained by parsing a syntax element nal_unit_type included in the NAL unit header of the VCL NAL unit.

[0252] The first flag information can be obtained by parsing sps_subpic_treated_as_pic_flag[i] included in a higher level syntax, for example, a sequence parameter set (SPS). In this case, if the i-th subpicture included in the current picture is not treated as one picture, sps_subpic_treated_as_pic_flag[i] may have a first value (e.g., 0). In contrast, if the i-th subpicture included in the current picture is treated as one picture, sps_subpic_treated_as_pic_flag[i] may have a second value (e.g., 1).

[0253] The image decoding apparatus can determine the NAL unit type of each of the slices in the current picture based on the VCL NAL unit type information of the current picture (S1620).

[0254] In one embodiment, if at least some of the slices in the current picture have different NAL unit types (i.e., if the current picture has mixed NAL unit types), the current picture may contain two or more sub-pictures.

[0255] In one embodiment, if at least some of the slices in the current picture have different NAL unit types, all sub-pictures included in the current picture may be treated as a single picture. In this case, the first flag information (e.g., sps_subpic_treated_as_pic_fla[1]) may have a second value (e.g., 1) for all sub-pictures included in the current picture.

[0256] In one embodiment, all slices included in each sub-picture in the current picture may have the same NAL unit type. For example, all slices included in a first sub-picture of the current picture may have a first NAL unit type, and all slices included in a second sub-picture of the current picture may have a second NAL unit type. In this case, if the current picture has mixed NAL unit types, the second NAL unit type may be different from the first NAL unit type. Alternatively, if the current picture has a single NAL unit type, the second NAL unit type may be the same as the first NAL unit type.

[0257] In one embodiment, whether at least some of the slices in the current picture have different NAL unit types may be determined based on second flag information (e.g., pps_mixed_nalu_types_in_pic_flag) obtained from higher-level syntax, e.g., a picture parameter set (PPS). For example, if the second flag information has a first value (e.g., 0), all slices in the current picture may have the same NAL unit type. In contrast, if the second flag information has a second value (e.g., 1), at least some of the slices in the current picture may have different NAL unit types.

[0258] In one embodiment, the second flag information (e.g., pps_mixed_nalu_types_in_pic_flag) may have a predetermined value based on whether a hybrid NAL unit type is applicable to all pictures in an Output Layer Set (OLL). For example, if a hybrid NAL unit type is not applicable to all pictures in an Output Layer Set (OLS), the second flag information may have only a first value (e.g., 0). In this case, the image decoding apparatus may infer that the value of the second flag information is the first value (e.g., 0) without separately parsing the second flag information. In contrast, if a hybrid NAL unit type is applicable to all pictures in an Output Layer Set (OLS), the second flag information may have a first value (e.g., 0) or a second value (e.g., 1).

[0259] In one embodiment, the second flag information (e.g., pps_mixed_nalu_types_in_pic_flag) may have a predetermined value based on whether the current picture includes only one sub-picture. For example, if the current picture includes only one sub-picture, the second flag information may have only a first value (e.g., 0). In this case, the image decoding apparatus may infer that the value of the second flag information is the first value (e.g., 0) without separately parsing the second flag information. In contrast, if the current picture includes two or more sub-pictures, the second flag information may have a first value (e.g., 0) or a second value (e.g., 1).

[0260] In one embodiment, the second flag information (e.g., pps_mixed_nalu_types_in_pic_flag) may have a predetermined value based on whether all sub-pictures included in the current picture are treated as one picture. For example, if at least one of the sub-pictures included in the current picture is not treated as one picture, the second flag information may have only a first value (e.g., 0). In this case, the image decoding apparatus may infer that the value of the second flag information is the first value (e.g., 0) without separately parsing the second flag information. In contrast, if all sub-pictures included in the current picture are treated as one picture, the second flag information may have a first value (e.g., 0) or a second value (e.g., 1).

[0261] Meanwhile, the type of each sub-picture included in the current picture may be determined based on the NAL unit type of the sub-picture. For example, as described above with reference to Table 1, the type of a sub-picture having a NAL unit type such as TRAIL_NUT may be determined as a trailing sub-picture. Also, the type of a sub-picture having a NAL unit type such as STSA_NUT may be determined as an STSA sub-picture, and the type of a sub-picture having a NAL unit type such as GDR_NUT may be determined as a GDR sub-picture. Also, the type of a sub-picture having a NAL unit type such as RADL_NUT may be determined as an RADL sub-picture, and the type of a sub-picture having a NAL unit type such as RASL_NUT may be determined as an RASL sub-picture. In addition, a sub-picture type having a NAL unit type such as IDR_W_RADL or IDR_N_LP can be determined as an IDR sub-picture, and a sub-picture type having a NAL unit type such as CRA_NUT can be determined as a CRA sub-picture.

[0262] The image decoding apparatus may encode a plurality of slices in the current picture based on the first flag information acquired in step S1610 and the NAL unit type determined in step S1620 (S1630). At this time, the decoding process for each slice may be performed in units of CU (coding yunit) based on a predetermined prediction mode, as described above.

[0263] As described above, according to one embodiment of the present disclosure, the current picture can have two or more NAL unit types based on the sub-picture structure.

[0264] The names of syntax elements described in this disclosure may include information about the location where the syntax element is signaled. For example, a syntax element beginning with "sps_" may mean that the syntax element is signaled in a sequence parameter set (SPS). Furthermore, a syntax element beginning with "pps_", "ph_", "sh_", etc. may mean that the syntax element is signaled in a picture parameter set (PPS), picture header, slice header, etc., respectively.

[0265] Although the exemplary method of the present disclosure is expressed as a series of operations for clarity of explanation, this is not intended to limit the order in which the steps are performed, and the steps may be performed simultaneously or in a different order if necessary. To achieve the method according to the present disclosure, the steps illustrated may include other steps, or some steps may be omitted and the remaining steps may be included, or some steps may be omitted and additional other steps may be included.

[0266] In the present disclosure, an image encoding device or an image decoding device that performs a predetermined operation (step) can perform the operation (step) to check the execution conditions and circumstances of the operation (step). For example, if it is described that a predetermined operation is performed when a predetermined condition is satisfied, the image encoding device or the image decoding device can perform the predetermined operation after performing an operation to check whether the predetermined condition is satisfied.

[0267] The various embodiments of the present disclosure are not intended to enumerate all possible combinations, but are intended to describe representative aspects of the present disclosure, and the matters described in the various embodiments may be applied independently or in combination of two or more.

[0268] Additionally, various embodiments of the present disclosure may be implemented using hardware, firmware, software, or a combination thereof, etc. In the case of a hardware implementation, the implementation may be using one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), general processors, controllers, microcontrollers, microprocessors, etc.

[0269] In addition, an image decoding apparatus and an image encoding apparatus to which an embodiment of the present disclosure is applied may be included in a multimedia broadcast transmitting / receiving apparatus, a mobile communication terminal, a home cinema video apparatus, a digital cinema video apparatus, a surveillance camera, a video conversation apparatus, a real-time communication apparatus such as video communication, a mobile streaming apparatus, a storage medium, a camcorder, a video on demand (VoD) service providing apparatus, an over-the-top (OTT) video apparatus, an internet streaming service providing apparatus, a three-dimensional (3D) video apparatus, an image telephone video apparatus, a medical video apparatus, etc., and may be used to process a video signal or a data signal. For example, an over-the-top (OTT) video apparatus may include a game console, a Blu-ray player, an internet-connected TV, a home theater system, a smartphone, a tablet PC, a digital video recorder (DVR), etc.

[0270] FIG. 17 is a diagram illustrating a content streaming system to which an embodiment of the present disclosure can be applied.

[0271] As shown in FIG. 17, a content streaming system to which an embodiment of the present disclosure is applied can broadly include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.

[0272] The encoding server compresses content input from a multimedia input device such as a smartphone, camera, or camcorder into digital data to generate a bitstream and transmits the bitstream to the streaming server. As another example, if a multimedia input device such as a smartphone, camera, or video camera directly generates a bitstream, the encoding server can be omitted.

[0273] The bitstream can be generated by an image encoding method and / or image encoding device to which an embodiment of the present disclosure is applied, and the streaming server can temporarily store the bitstream during the process of transmitting or receiving the bitstream.

[0274] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server serves as an intermediary for informing the user of available services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, which then transmits the multimedia data to the user. In this case, the content streaming system may include a separate control server, which may control commands and responses between devices in the content streaming system.

[0275] The streaming server may receive content from a media storage and / or an encoding server. For example, when receiving content from the encoding server, the content may be received in real time. In this case, the streaming server may store the bitstream for a certain period of time to provide a smooth streaming service.

[0276] Examples of the user device include a mobile phone, a smartphone, a laptop computer, a digital broadcasting terminal, a personal digital assistant (PDA), a portable multimedia player (PMP), a navigation system, a slate PC, a tablet PC, an ultrabook, a wearable device such as a smartwatch, smart glass, a head mounted display (HMD), a digital TV, a desktop computer, and digital signage.

[0277] Each server in the content streaming system can be operated as a distributed server, in which case data received from each server can be processed in a distributed manner.

[0278] The scope of the present disclosure includes software or machine-executable commands (e.g., operating systems, applications, firmware, programs, etc.) that cause operations according to the methods of various embodiments to be performed on a device or computer, and non-transitory computer-readable medium on which such software or commands can be stored and executed on a device or computer. [Industrial Applicability]

[0279] Embodiments according to the present disclosure can be used to encode / decode images.

Claims

1. An image decoding method performed by an image decoding device, comprising: obtaining, from a bitstream, video coding layer (VCL) network abstraction layer (NAL) unit type information of a current picture and first flag information indicating whether a sub-picture included in the current picture is treated as one picture; determining a NAL unit type of each of a plurality of slices included in the current picture based on the obtained VCL NAL unit type information; decoding the slices based on the determined NAL unit type and the first flag information; the current picture includes two or more sub-pictures based on at least some of the slices having different NAL unit types; the first flag information has a predetermined value indicating that each of the two or more sub-pictures is treated as one picture based on the fact that at least some of the plurality of slices have different NAL unit types; all slices in each sub-picture included in the current picture have the same NAL unit type; The image decoding method, wherein the first flag information is obtained from a sequence parameter set (SPS) of the bitstream based on whether the current picture includes two or more sub-pictures.

2. The image decoding method according to claim 1 , wherein whether at least some of the slices have different NAL unit types from each other is determined based on second flag information obtained from a picture parameter set (PPS).

3. 3. The image decoding method of claim 2, wherein the second flag information has a predetermined value indicating that all of the plurality of slices have the same NAL unit based on the fact that a mixed NAL unit type is not applicable to all pictures in an output layer set (OLS).

4. 3. The image decoding method of claim 2, wherein, based on the current picture including only one sub-picture, the second flag information has a predetermined value indicating that all of the plurality of slices have the same NAL unit type.

5. 3. The image decoding method of claim 2, wherein the second flag information has a predetermined value indicating that all of the plurality of slices have the same NAL unit type based on the fact that at least one of the sub-pictures included in the current picture is not treated as one picture.

6. The image decoding method of claim 1 , wherein the sub-picture type of each sub-picture included in the current picture is determined based on the NAL unit type of each sub-picture.

7. An image coding method performed by an image coding device, comprising: Dividing a current picture into one or more sub-pictures; determining a network abstraction layer (NAL) unit type for each of a plurality of slices included in the one or more sub-pictures; encoding the plurality of slices based on the determined NAL unit types; the current picture is divided into two or more sub-pictures based on at least some of the slices having different NAL unit types; Each of the two or more sub-pictures is treated as one picture based on the fact that at least some of the slices have different NAL unit types; all slices in each sub-picture included in the current picture have the same NAL unit type; An image coding method, wherein first flag information indicating whether the sub-pictures included in the current picture are treated as one picture based on the fact that the current picture includes two or more sub-pictures is coded into a sequence parameter set (SPS) of a bitstream.

8. The image encoding method according to claim 7 , wherein whether at least some of the plurality of slices have different NAL unit types is encoded based on predetermined flag information included in a picture parameter set (PPS).

9. 9. The image coding method of claim 8, wherein the flag information has a predetermined value indicating that all of the plurality of slices have the same NAL unit type based on the fact that a mixed NAL unit type is not applicable to all pictures in an output layer set (OLS).

10. The image encoding method according to claim 8 , wherein, based on the fact that the current picture includes only one sub-picture, the flag information has a predetermined value indicating that all of the plurality of slices have the same NAL unit type.

11. 9. The image encoding method of claim 8, wherein the flag information has a predetermined value indicating that all of the plurality of slices have the same NAL unit type based on the fact that at least one of the sub-pictures included in the current picture is not treated as one picture.

12. 1. A method for transmitting a bitstream generated by an image coding method, comprising: The image encoding method includes: Dividing a current picture into one or more sub-pictures; determining a network abstraction layer (NAL) unit type for each of a plurality of slices included in the one or more sub-pictures; encoding the plurality of slices based on the determined NAL unit types; the current picture is divided into two or more sub-pictures based on at least some of the slices having different NAL unit types; Each of the two or more sub-pictures is treated as one picture based on the fact that at least some of the slices have different NAL unit types; all slices in each sub-picture included in the current picture have the same NAL unit type; A method in which first flag information indicating whether the sub-pictures included in the current picture are treated as one picture based on the current picture including two or more sub-pictures is coded in a sequence parameter set (SPS) of the bitstream.

Citation Information

Patent Citations

  • METHOD AND APPARATUS FOR IMAGE ENCODING / DECODING BASED ON HYBRID NAL UNIT TYPES AND METHOD FOR TRANSMITTING BITSTREAM - Patent application

    JP7415030B2

  • Video codec allowing sub-picture or region wise random access and concept for video composition using the same

    WO2020157287A1

  • Pictures with mixed NAL unit types

    WO2020185922A1

  • Decoder, encoder and methods for mixing NAL units of different NAL unit types in video streams

    WO2021122817A1