Image encoding / decoding method and apparatus for sub-layers based on required number of sub-layers and bitstream transmission method

By determining the number of layers and sublayers in the output layer set, and based on the maximum number of sublayers required by the direct reference layer, the problem of low encoding/decoding efficiency for high-resolution and high-quality images is solved, and efficient image information transmission and storage are achieved.

CN115769579BActive Publication Date: 2025-12-16LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202180041964.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-04-15
Filing Date
2021-04-15
Publication Date
2025-12-16
Estimated Expiration
2041-04-15

AI Technical Summary

Technical Problem

Existing technologies suffer from low encoding/decoding efficiency in the transmission of high-resolution and high-quality images, leading to increased transmission and storage costs.

Method used

Encoding/decoding efficiency is improved and bitstreams are generated by determining the number of layers and sublayers in the Output Layer Set (OLS) based on the maximum number of sublayers required by the direct reference layer.

Benefits of technology

It improves the efficiency of image encoding/decoding, reduces transmission and storage costs, and achieves efficient image information transmission and storage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115769579B_ABST
    Figure CN115769579B_ABST
Patent Text Reader

Abstract

Image encoding / decoding methods and apparatuses are provided. A method for image decoding apparatus according to the present disclosure includes the steps of determining at least one layer belonging to an output layer set (OLS), and determining a number of sub-layers of the at least one layer.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to an image encoding / decoding method and apparatus, and more particularly, to an image encoding and decoding method and apparatus that determine sub-layers based on the number of required sub-layers, and a method of transmitting a bitstream generated by the image encoding method / apparatus of the present disclosure. BACKGROUND

[0002] Recently, the demand for high-resolution and high-quality images, such as high-definition (HD) images and ultra-high-definition (UHD) images, is increasing in various fields. As the resolution and quality of image data are improved, the amount of information or bits to be transmitted is relatively increased compared to existing image data. The increase in the amount of information or bits to be transmitted results in an increase in transmission and storage costs.

[0003] Therefore, an efficient image compression technique is required to effectively transmit, store, and reproduce information about high-resolution and high-quality images. SUMMARY

[0004] TECHNICAL PROBLEM

[0005] An object of the present disclosure is to provide an image encoding / decoding method and apparatus having improved encoding / decoding efficiency.

[0006] Another object of the present disclosure is to provide an image encoding / decoding method and apparatus that improve encoding / decoding efficiency by determining sub-layers based on the number of required sub-layers.

[0007] Another object of the present disclosure is to provide a method of transmitting a bitstream generated by the image encoding method or apparatus according to the present disclosure.

[0008] Another object of the present disclosure is to provide a recording medium that stores a bitstream generated by the image encoding method or apparatus according to the present disclosure.

[0009] Another object of the present disclosure is to provide a recording medium that stores a bitstream generated by the image encoding method or apparatus according to the present disclosure.

[0010] The technical problems solved by the present disclosure are not limited to the above-mentioned technical problems, and other technical problems not described herein will become apparent to those skilled in the art from the following description.

[0011] TECHNICAL SOLUTION

[0012] An image decoding method performed by an image decoding apparatus according to an aspect of the present disclosure can include the steps of determining at least one layer in an output layer set (OLS), and determining a number of sub-layers of the at least one layer.

[0013] An image decoding apparatus according to an aspect of the present disclosure can include a memory and at least one processor. The at least one processor can be configured to determine at least one layer in an output layer set (OLS), and determine a number of sub-layers of the at least one layer. The number of sub-layers of a current layer in the OLS can be determined based on a maximum number of required sub-layers of a direct reference layer, the direct reference layer being a layer capable of directly referring to the current layer.

[0014] An image encoding method performed by an image encoding apparatus according to an aspect of the present disclosure can include the steps of determining at least one layer in an output layer set (OLS), and determining a number of sub-layers of the at least one layer. The number of sub-layers of a current layer in the OLS can be determined based on a maximum number of required sub-layers of a direct reference layer, the direct reference layer being a layer capable of directly referring to the current layer.

[0015] Further, a transmitting method according to another aspect of the present disclosure can transmit a bitstream generated by an image encoding apparatus or method according to the present disclosure.

[0016] Further, a computer-readable recording medium according to another aspect of the present disclosure can store a bitstream generated by an image encoding method or apparatus according to the present disclosure.

[0017] Further, a computer-readable recording medium according to another aspect of the present disclosure can store a bitstream for enabling a decoding apparatus to perform an image decoding method according to the present disclosure.

[0018] The features briefly described above in relation to the present disclosure are merely exemplary aspects of the following detailed description of the present disclosure, and do not limit the scope of the present disclosure.

[0019] Advantageous Effects

[0020] According to the present disclosure, it is possible to provide an image encoding / decoding method and apparatus having improved encoding / decoding efficiency.

[0021] Further, according to the present disclosure, it is possible to provide an image encoding / decoding method and apparatus that improves encoding / decoding efficiency by determining sub-layers based on a number of required sub-layers.

[0022] Further, according to the present disclosure, it is possible to provide a method of transmitting a bitstream generated by an image encoding method or apparatus according to the present disclosure.

[0023] Further, according to the disclosure, a recording medium storing a bitstream generated by the image encoding method or apparatus according to the disclosure can be provided.

[0024] Further, according to the disclosure, a recording medium storing a bitstream received, decoded, and used for reconstructing an image by the image decoding apparatus according to the disclosure can be provided.

[0025] Those skilled in the art will understand that the effects realized by the disclosure are not limited to what has been particularly described hereinabove and other advantages of the disclosure will be more clearly understood from the following detailed description. BRIEF DESCRIPTION OF DRAWINGS

[0026] FIG. 1 is a view schematically showing a video encoding system to which embodiments of the disclosure are applicable.

[0027] FIG. 2 is a view schematically showing an image encoding apparatus to which embodiments of the disclosure are applicable.

[0028] FIG. 3 is a view schematically showing an image decoding apparatus to which embodiments of the disclosure are applicable.

[0029] FIG. 4 and FIG. 5 is a view showing an example of a picture decoding and encoding process according to an embodiment.

[0030] FIG. 6 is a view showing a layer structure of an image for encoding according to an embodiment.

[0031] FIG. 7 to FIG. 8 is a view exemplifying multi-layer based encoding and decoding.

[0032] FIG. 9 is a view exemplifying a syntax structure of a VPS according to an embodiment of the disclosure.

[0033] FIG. 10 to FIG. 14 is a view exemplifying a pseudo code for deriving VPS related variables according to an embodiment of the disclosure.

[0034] FIG. 15 is a view exemplifying a syntax structure of a VPS according to another embodiment of the disclosure.

[0035] FIG. 16 to FIG. 17 is a view exemplifying a pseudo code for deriving VPS related variables according to another embodiment of the disclosure.

[0036] FIG. 18 is a view exemplifying a method of deriving a sub-bitstream according to an embodiment of the disclosure.

[0037] FIG. 19 is a view illustrating an encoding and / or decoding method according to another embodiment of the disclosure.

[0038] FIG. 20 is a view showing a content streaming system to which an embodiment of the disclosure is applicable. DETAILED DESCRIPTION

[0039] Hereinafter, embodiments of the disclosure will be described in detail with reference to the accompanying drawings so as to be easily carried out by one of ordinary skill in the art. However, the disclosure can be implemented in various different forms and is not limited to the embodiments described herein.

[0040] In describing the disclosure, if it is determined that a detailed description of the relevant known function or configuration makes the scope of the disclosure unnecessarily obscure, a detailed description thereof will be omitted. In the drawings, parts irrelevant to the description of the disclosure are omitted, and like reference numerals are assigned to like parts.

[0041] In the disclosure, when one component is "connected", "coupled", or "linked" to another component, it can include not only a direct connection relationship but also an indirect connection relationship in which a middle component exists. In addition, when one component "includes" or "has" another component, it means that it can further include the other component unless otherwise specified, rather than excluding the other component.

[0042] In the disclosure, the terms first, second, and the like can be used only for the purpose of distinguishing one component from other components, and do not limit the order or importance of the components unless otherwise specified. Accordingly, within the scope of the disclosure, a first component in one embodiment can be referred to as a second component in another embodiment, and similarly, a second component in one embodiment can be referred to as a first component in another embodiment.

[0043] In the disclosure, components distinguished from each other are intended to clearly describe each feature, and do not mean that the components must be separated. That is, a plurality of components can be integrated and implemented in one hardware or software unit, or one component can be distributed and implemented in a plurality of hardware or software units. Therefore, even if not otherwise specified, embodiments in which components are integrated or components are distributed are included in the scope of the disclosure.

[0044] In the disclosure, components described in each embodiment are not necessarily essential components, and some components can be optional components. Therefore, embodiments consisting of a subset of the components described in the embodiments are also included in the scope of the disclosure. In addition, embodiments including other components in addition to the components described in the various embodiments are included in the scope of the disclosure.

[0045] The present disclosure relates to encoding and decoding of images, and unless redefined in the present disclosure, the terms used in the present disclosure can have the general meanings commonly used in the technical field to which the present disclosure belongs.

[0046] The methods / embodiments disclosed in the present disclosure are applicable to the methods disclosed in the Versatile Video Coding (VVC) standard. In addition, the methods / embodiments disclosed in the present disclosure are applicable to the methods disclosed in the Essential Video Coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the second generation Audio Video Coding standard (AVS2), or the next generation video / image coding standard (e.g., H.267 or H.268).

[0047] In the present disclosure, various embodiments of video / image coding are provided, and the embodiments not described in the present disclosure can be performed in combination.

[0048] In the present disclosure, "video" can mean a set of images over time. "Picture" generally refers to a unit representing one image at a specific time, and a slice / tile is a coding unit that constitutes a part of a picture in coding. A slice / tile can include one or more coding tree units (CTUs). A CTU can be partitioned into one or more CUs.

[0049] One picture can be composed of one or more slices / tiles. A tile is a rectangular region within a specific tile row and a specific tile column in a picture, and can be composed of a plurality of CTUs. A tile column can be defined as a rectangular region of CTUs, and can have a height equal to the height of the picture and a width specified by syntax elements signaled from a bitstream portion such as a picture parameter set. A tile row can be defined as a rectangular region of CTUs, and can have a width equal to the width of the picture and a height specified by syntax elements signaled from a bitstream portion such as a picture parameter set.

[0050] Tile scanning is a specific sequential ordering of partitioning CTUs of a picture. Here, the CTUs are sequentially ordered in CTU raster scan in a tile, and the tiles in a picture are sequentially ordered in raster scan of tiles of the picture. A slice includes an integer number of consecutive complete CTU rows or an integer number of complete tiles within a tile of a picture. A slice can be exclusively included in a single NAL unit.

[0051] One picture can be partitioned into two or more sub-pictures. A sub-picture can be a rectangular region of one or more slices in a picture.

[0052] A picture can include one or more tile groups. A tile group can include one or more tiles. A birck can denote a rectangular region of CTU rows within a tile in a picture. A tile can include one or more bircks. A birck can denote a rectangular region of CTU rows in a tile. A tile can be divided into multiple bircks, and each birck can include one or more CTU rows belonging to the tile. A tile that is not divided into multiple bircks can also be regarded as a birck.

[0053] A "pixel" or a "pel" can mean a minimum unit constituting a picture (or an image). Also, a "sample" can be used as a term corresponding to a pixel. A sample can generally denote a pixel or a value of a pixel, and can denote only a pixel / value of a luma component or only a pixel / value of a chroma component.

[0054] In the disclosure, a "unit" can denote a basic unit of image processing. A unit can include at least one of a specific region of a picture and information related to the region. A unit can include one luma block and two chroma blocks (e.g., Cb and Cr). In some cases, a unit can be used interchangeably with terms such as "sample array", "block", or "region". In general, an MxN block can include a set (or an array) of M columns and N rows of samples (or sample array) or transform coefficients.

[0055] In the disclosure, a "current block" can mean one of a "current coding block", a "current coding unit", a "coding target block", a "decoding target block", or a "processing target block". When performing prediction, a "current block" can mean a "current prediction block" or a "prediction target block". When performing transform (inverse transform) / quantization (dequantization), a "current block" can mean a "current transform block" or a "transform target block". When performing filtering, a "current block" can mean a "filtering target block".

[0056] Also, in the disclosure, unless explicitly stated as a chroma block, a "current block" can mean a "luma block of a current block". A "chroma block of a current block" can be expressed by including an explicit description of a chroma block such as "chroma block" or "current chroma block".

[0057] In the disclosure, a slash " / " or a comma "," should be interpreted to indicate "and / or". For example, expressions "A / B" and "A, B" can mean "A and / or B". Also, "A / B / C" and "A / B / C" can mean "at least one of A, B, and / or C".

[0058] In the disclosure, the term "or" is to be interpreted as indicating "and / or". For example, the expression "A or B" can include 1) only "A", 2) only "B", and / or 3) both of "A and B". In other words, in the disclosure, the term "or" is to be interpreted as indicating "additionally or alternatively".

[0059] In the disclosure, "at least one of A and B" can mean "only A", "only B", or "both of A and B". Also, in the disclosure, "at least one of A or B" or "at least one of A and / or B" can be interpreted as the same as "at least one of A and B".

[0060] Also, in the disclosure, "at least one of A, B, and C" can mean "only A", "only B", "only C", or "any combination of A, B, and C". Also, in the disclosure, "at least one of A, B, or C" or "at least one of A, B, and / or C" can be interpreted as the same as "at least one of A, B, and C".

[0061] Also, the parentheses used in the disclosure can mean "for example". Specifically, when describing "prediction (intra prediction)", "intra prediction" can be proposed as an example of "prediction". In other words, the "prediction" of the disclosure is not limited to "intra prediction", and "intra prediction" can be proposed as an example of "prediction". Also, even when describing "prediction (i.e., intra prediction)", "intra prediction" can be proposed as an example of "prediction".

[0062] In the disclosure, technical features described separately in one drawing can be implemented individually or simultaneously.

[0063] Overview of a video coding system

[0064] FIG. 1 is a view showing a video coding system according to the disclosure.

[0065] The video coding system according to the embodiment can include a source device 10 and a receiving device 20. The source device 10 can deliver encoded video and / or image information or data in the form of a file or a stream to the receiving device 20 via a digital storage medium or a network.

[0066] The source device 10 according to the embodiments can include a video source generator 11, an encoding device 12, and a transmitter 13. The reception device 20 according to the embodiments can include a receiver 21, a decoding device 22, and a Tenderer 23. The encoding device 12 can be referred to as a video / image encoding device, and the decoding device 22 can be referred to as a video / image decoding device. The transmitter 13 can be included in the encoding device 12. The receiver 21 can be included in the decoding device 22. The Tenderer 23 can include a display, and the display can be configured as a separate device or an external component.

[0067] The video source generator 11 can acquire a video / image through a process of capturing, synthesizing, or generating a video / image. The video source generator 11 can include a video / image capturing device and / or a video / image generating device. The video / image capturing device can include, for example, one or more cameras, a video / image archive including previously captured videos / images, or the like. The video / image generating device can include, for example, a computer, a tablet, and a smart phone, and can generate a video / image (electronically). For example, a virtual video / image can be generated through a computer or the like. In this case, the video / image capturing process can be replaced by a process of generating related data.

[0068] The encoding device 12 can encode an input video / image. For compression and coding efficiency, the encoding device 12 can perform a series of processes such as prediction, transform, and quantization. The encoding device 12 can output encoded data (encoded video / image information) in the form of a bitstream.

[0069] The transmitter 13 can transmit the encoded video / image information or data output in the form of a bitstream to the receiver 21 of the reception device 20 in the form of a file or a stream through a digital storage medium or a network. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, or the like. The transmitter 13 can include an element for generating a media file through a predetermined file format and can include an element for transmission through a broadcasting / communication network. The receiver 21 can extract / receive a bitstream from a storage medium or a network and transmit the bitstream to the decoding device 22.

[0070] The decoding device 22 can decode a video / image by performing a series of processes such as dequantization, inverse transform, and prediction, which correspond to the operations of the encoding device 12.

[0071] The Tenderer 23 can render the decoded video / image. The rendered video / image can be displayed through a display.

[0072] Overview of an image encoding device

[0073] FIG. 2FIG. 1 is a view schematically illustrating an image encoding apparatus to which embodiments of the present disclosure are applicable.

[0074] As shown in FIG. 1, the image source apparatus 100 can include an image partitioner 110, a subtractor 115, a transformer 120, a quantizer 130, a dequantizer 140, an inverse transformer 150, an adder 155, a filter 160, a memory 170, an inter-predictor 180, an intra-predictor 185, and an entropy encoder 190. The inter-predictor 180 and the intra-predictor 185 can be collectively referred to as a "predictor." The transformer 120, the quantizer 130, the dequantizer 140, and the inverse transformer 150 can be included in a residual processor. The residual processor can further include the subtractor 115. FIG. 2

[0075] In some embodiments, all or at least some of the plurality of components configuring the image source apparatus 100 can be configured by one hardware component (e.g., an encoder or a processor). In addition, the memory 170 can include a decoded picture buffer (DPB) and can be configured by a digital storage medium.

[0076] The image partitioner 110 can partition an input image (or picture or frame) input to the image source apparatus 100 into one or more processing units. For example, the processing units can be referred to as coding units (CUs). The coding units can be obtained by recursively partitioning a coding tree unit (CTU) or a largest coding unit (LCU) according to a quadtree binary tree ternary (QT / BT / TT) structure. For example, one coding unit can be partitioned into a plurality of coding units of a deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. For the partitioning of the coding units, the quadtree structure can be applied first, and the binary tree structure and / or the ternary tree structure can be applied later. The encoding process according to the present disclosure can be performed based on a final coding unit that is no longer partitioned. The largest coding unit can be used as the final coding unit, and a coding unit of a deeper depth obtained by partitioning the largest coding unit can be used as the final coding unit. Here, the encoding process can include a process of prediction, transform, and reconstruction that will be described later. As another example, the processing unit of the encoding process can be a prediction unit (PU) or a transform unit (TU). The prediction unit and the transform unit can be divided or partitioned from the final coding unit. The prediction unit can be a sample prediction unit, and the transform unit can be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.

[0077] ​The predictor (inter-predictor 180 or intra-predictor 185) can perform prediction on a block (current block) to be processed and generate a prediction block including predicted samples of the current block. The predictor can determine whether to apply intra-prediction or inter-prediction on a basis of the current block or CU. The predictor can generate various information related to prediction of the current block and transmit the generated information to the entropy encoder 190. The information about prediction can be encoded in the entropy encoder 190 and output in the form of a bitstream.

[0078] The intra-predictor 185 can predict the current block by referring to samples in the current picture. The reference samples can be located in the neighbors of the current block or can be placed separately according to the intra-prediction mode and / or the intra-prediction technique. The intra-prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes can include, for example, a DC mode and a planar mode. The directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes according to the level of detail of the prediction direction. However, this is merely an example, and more or less directional prediction modes can be used according to settings. The intra-predictor 185 can determine a prediction mode applied to the current block by using a prediction mode applied to a neighboring block.

[0079] The inter predictor 180 can derive a prediction block of a current block based on a reference block (a reference sample array) specified by a motion vector on a reference picture. In this case, to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of a block, a sub-block, or a sample based on the correlation of the motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the neighboring blocks can include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block can be the same or different. The temporal neighboring block can be referred to as a collocated reference block, a collocated CU (colCU), etc. The reference picture including the temporal neighboring block can be referred to as a collocated picture (colPic). For example, the inter predictor 180 can configure a motion information candidate list based on the neighboring blocks and generate information indicating which candidate to use to derive the motion vector and / or the reference picture index of the current block. The inter prediction can be performed based on various prediction modes. For example, in the case of a skip mode and a merge mode, the inter predictor 180 can use the motion information of the neighboring blocks as the motion information of the current block. In the case of the skip mode, unlike the merge mode, a residual signal can not be transmitted. In the case of a motion vector prediction (MVP) mode, the motion vector of the neighboring block can be used as a motion vector predictor, and the motion vector of the current block can be signaled by encoding a motion vector difference and an indicator of the motion vector predictor. The motion vector difference can mean a difference between the motion vector of the current block and the motion vector predictor.

[0080] The predictor can generate a prediction signal based on various prediction methods and prediction techniques described below. For example, the predictor can not only apply intra prediction or inter prediction, but also simultaneously apply both intra prediction and inter prediction to predict the current block. The prediction method of simultaneously applying both intra prediction and inter prediction to predict the current block can be referred to as combined inter and intra prediction (CIIP). In addition, the predictor can perform intra block copy (IBC) to predict the current block. Intra block copy can be used for content image / video encoding of games, etc., for example, screen content coding (SCC). IBC is a method of predicting a current picture using a reference block previously reconstructed in the current picture at a position apart by a predetermined distance. When IBC is applied, the position of the reference block in the current picture can be encoded as a vector (block vector) corresponding to the predetermined distance. IBC basically performs prediction in the current picture, but can be performed similarly to inter prediction, in that the reference block is derived within the current picture. That is, IBC can use at least one of the inter prediction techniques described in the disclosure.

[0081] The prediction signal generated by the predictor can be used to generate a reconstructed signal or to generate a residual signal. The subtractor 115 can generate a residual signal (a residual block or a residual sample array) by subtracting the prediction signal (a prediction block or a prediction sample array) output from the predictor from the input image signal (an original block or an original sample array). The generated residual signal can be transmitted to the transformer 120.

[0082] The transformer 120 can generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique can include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a karhunen-loève transform (KLT), a graph-based transform (GBT), or a conditional non-linear transform (CNT). Here, the GBT refers to a transform obtained from a graph when relationship information between pixels is represented by a graph. The CNT refers to a transform obtained based on a prediction signal generated using all previously reconstructed pixels. In addition, the transform process can be applied to a square pixel block having the same size or can be applied to a block having a variable size other than a square.

[0083] The quantizer 130 can quantize the transform coefficients and transmit them to the entropy encoder 190. The entropy encoder 190 can encode the quantized signal (information about the quantized transform coefficients) and output a bitstream. The information about the quantized transform coefficients can be referred to as residual information. The quantizer 130 can rearrange the quantized transform coefficients of the block type into a one-dimensional vector form based on a coefficient scan order, and generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form.

[0084] The entropy encoder 190 can perform various encoding methods, for example, such as exponential golomb, context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), etc. The entropy encoder 190 can encode information (for example, values of syntax elements, etc.) required for video / image reconstruction other than the quantized transform coefficients together or individually. The encoded information (for example, encoded video / image information) can be transmitted or stored in the form of a bitstream in units of a network abstraction layer (NAL). The video / image information can further include information about various parameter sets, such as an adaptive parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information can further include general constraint information. The information signaled, transmitted, and / or syntax elements described in the disclosure can be encoded through the above-described encoding process and included in the bitstream.

[0085] The bitstream can be transmitted through a network or can be stored in a digital storage medium. The network can include a broadcasting network and / or a communication network, and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) that transmits a signal output from the entropy encoder 190 and / or a storage unit (not shown) that stores the signal can be included as an internal / external element of the image source device 100. Alternatively, the transmitter can be provided as a component of the entropy encoder 190.

[0086] The quantized transform coefficients output from the quantizer 130 can be used to generate a residual signal. For example, the residual signal (residual block or residual sample) can be reconstructed by applying dequantization and inverse transform to the quantized transform coefficients through the dequantizer 140 and the inverse transformer 150.

[0087] The adder 155 adds the reconstructed residual signal to a prediction signal output from the inter-predictor 180 or the intra-predictor 185 to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). If there is no residual for a block to be processed, such as the case where a skip mode is applied, the prediction block can be used as the reconstructed block. The adder 155 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-prediction of a next block to be processed in the current picture, and can be used for inter-prediction of a next picture by filtering as described below.

[0088] The filter 160 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 160 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and store the modified reconstructed picture in the memory 170, specifically, in the DPB of the memory 170. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filtering, bilateral filtering, etc. The filter 160 can generate various information related to filtering and transmit the generated information to the entropy encoder 190, as described later in the description of each filtering method. The information related to filtering can be encoded by the entropy encoder 190 and output in the form of a bitstream.

[0089] The modified reconstructed picture transmitted to the memory 170 can be used as a reference picture in the inter-predictor 180. When inter-prediction is applied through the image source device 100, prediction mismatch between the image source device 100 and an image decoding apparatus can be avoided and coding efficiency can be improved.

[0090] The DPB of the memory 170 can store the modified reconstructed picture to be used as a reference picture in the inter prediction 180. The memory 170 can store motion information of a block for deriving (or encoding) motion information in the current picture and / or motion information of a block already reconstructed in the picture. The stored motion information can be transmitted to the inter prediction 180 and used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory 170 can store reconstructed samples of a reconstructed block in the current picture and can transfer the reconstructed samples to the intra prediction 185.

[0091] Overview of an image decoding device

[0092] FIG. 3 FIG. 1 is a view schematically showing an image encoding apparatus to which embodiments of the present disclosure are applicable.

[0093] As shown in FIG. 3 FIG. 2, the image receiving apparatus 200 can include an entropy decoder 210, a dequantizer 220, an inverse transformer 230, an adder 235, a filter 240, a memory 250, an inter prediction 260, and an intra prediction 265. The inter prediction 260 and the intra prediction 265 can be collectively referred to as a "predictor". The dequantizer 220 and the inverse transformer 230 can be included in a residual processor.

[0094] According to embodiments, all or at least some of the plurality of components configured to the image receiving apparatus 200 can be configured by hardware components (e.g., a decoder or a processor). In addition, the memory 250 can include a decoded picture buffer (DPB) or can be configured by a digital storage medium.

[0095] The image receiving apparatus 200 having received a bitstream including video / image information can reconstruct an image by performing a process corresponding to a process performed by the image source apparatus 100. FIG. 2 The image receiving apparatus 200 can perform decoding using a processing unit applied in the image encoding apparatus. Accordingly, the processing unit for decoding can be, for example, a coding unit. The coding unit can be acquired by partitioning a coding tree unit or a largest coding unit. The reconstructed image signal decoded and output by the image receiving apparatus 200 can be reproduced through a reproduction apparatus (not shown).

[0096] The image receiving apparatus 200 can receive a bitstream in the form of a stream from the image source apparatus 100. FIG. 2The received signal can be decoded by the entropy decoder 210. For example, the entropy decoder 210 can parse a bitstream to derive information (e.g., video / image information) required for image reconstruction (or picture reconstruction). The video / image information can further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information can further include general constraint information. The image decoding device can further decode a picture based on the information on the parameter sets and / or the general constraint information. The information and / or the syntax elements described in the disclosure to be signaled / received can be decoded through a decoding process and obtained from the bitstream. For example, the entropy decoder 210 decodes information in the bitstream based on an encoding method such as exponential Golomb coding, CAVLC, or CABAC, and outputs values of syntax elements required for image reconstruction and quantized values of transform coefficients of a residual. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element in the bitstream, determine a context model using decoded information of a target syntax element, neighboring blocks, and a decoding target block, or information of a previously decoded symbol / bin, perform arithmetic decoding on the bins by predicting a probability of occurrence of the bins according to the determined context model, and generate a symbol corresponding to a value of each syntax element. In this case, the CABAC entropy decoding method can update the context model by using the decoded information of the symbol / bin for the context model of the next symbol / bin after determining the context model. Information related to prediction among the information decoded by the entropy decoder 210 can be provided to the predictors (inter-predictor 260 and intra-predictor 265), and residual values (that is, quantized transform coefficients and related parameter information) on which entropy decoding is performed in the entropy decoder 210 can be input to the dequantizer 220. In addition, information on filtering among the information decoded by the entropy decoder 210 can be provided to the filter 240. Further, a receiver (not shown) for receiving a signal output from the image encoding device can be further configured as an internal / external element of the image receiving apparatus 200, or the receiver can be a component of the entropy decoder 210.

[0097] Further, the image decoding device according to the disclosure can be referred to as a video / image / picture decoding device. The image decoding device can be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder can include the entropy decoder 210. The sample decoder can include at least one of the dequantizer 220, the inverse transformer 230, the adder 235, the filter 240, the memory 250, the inter-predictor 260, or the intra-predictor 265.

[0098] The dequantizer 220 can dequantize the quantized transform coefficients and output the transform coefficients. The dequantizer 220 can rearrange the quantized transform coefficients in the form of a two-dimensional block. In this case, the rearrangement can be performed based on a coefficient scan order performed in the image encoding apparatus. The dequantizer 220 can perform dequantization on the quantized transform coefficients by using a quantization parameter (e.g., quantization step length information) and obtain the transform coefficients.

[0099] The inverse transformer 230 can inverse-transform the transform coefficients to obtain a residual signal (a residual block, a residual sample array).

[0100] The predictor can perform prediction on the current block and generate a prediction block including prediction samples of the current block. The predictor can determine whether to apply intra prediction or inter prediction to the current block based on information about prediction output from the entropy decoder 210, and can determine a specific intra / inter prediction mode (prediction technique).

[0101] The same as described in the predictor of the image source apparatus 100, the predictor can generate a prediction signal based on various prediction methods (techniques) which will be described later.

[0102] The intra predictor 265 can predict the current block by referring to samples in the current picture. The description of the intra predictor 185 is equally applicable to the intra predictor 265.

[0103] The inter predictor 260 can derive a prediction block of the current block based on a reference block (a reference sample array) specified by a motion vector on a reference picture. In this case, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of a block, a sub-block, or a sample based on the correlation of the motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the neighboring blocks can include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. For example, the inter predictor 260 can configure a motion information candidate list based on the neighboring blocks, and derive a motion vector and / or a reference picture index of the current block based on received candidate selection information. The inter prediction can be performed based on various prediction modes, and the information about prediction can include information indicating an inter prediction mode of the current block.

[0104] The adder 235 can generate a reconstructed signal (a reconstructed picture, a reconstructed block, a reconstructed sample array) by adding the obtained residual signal to a prediction signal (a prediction block, a prediction sample array) output from the predictor (including the inter-predictor 260 and / or the intra-predictor 265). If a block to be processed has no residual, such as when a skip mode is applied, the prediction block can be used as the reconstructed block. The description of the adder 155 is equally applicable to the adder 235. The adder 235 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-prediction of a next block to be processed in the current picture, and can be used for inter-prediction of a next picture by filtering as described below.

[0105] The filter 240 can improve subjective / objective picture quality by applying filtering to the reconstructed signal. For example, the filter 240 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and store the modified reconstructed picture in the memory 250, specifically, in the DPB of the memory 250. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filtering, bilateral filtering, etc.

[0106] The (modified) reconstructed picture stored in the DPB of the memory 250 can be used as a reference picture in the inter-predictor 260. The memory 250 can store motion information for a block for which motion information is derived (or decoded) in the current picture and / or motion information for a block that has been reconstructed in the picture. The stored motion information can be transmitted to the inter-predictor 260 to be used as motion information of a spatially neighboring block or motion information of a temporally neighboring block. The memory 250 can store and deliver a reconstructed sample of a reconstructed block in the current picture to the intra-predictor 265.

[0107] In the present disclosure, the embodiments described in the filter 160, the inter-predictor 180, and the intra-predictor 185 of the image source device 100 can be equally or correspondingly applied to the filter 240, the inter-predictor 260, and the intra-predictor 265 of the image receiving device 200.

[0108] General image / video encoding process

[0109] In image / video encoding, pictures constituting an image / video can be encoded / decoded according to a decoding order. A picture order corresponding to an output order of decoded pictures can be set to be different from the decoding order, and based thereon, not only forward prediction but also backward prediction can be performed during inter-prediction.

[0110] FIG. 4 An example of a schematic picture decoding process to which embodiments of the present disclosure are applicable is shown. In FIG. 4In this process, S410 can be executed in the entropy decoder 210 of the decoding device, S420 can be executed in the predictor including the intra-frame predictor 265 and the inter-frame predictor 260, S430 can be executed in the residual processor including the dequantizer 220 and the inverse transformer 230, S440 can be executed in the adder 235, and S450 can be executed in the filter 240. S410 can include the information decoding process described in this disclosure, S420 can include the inter-frame / intra-frame prediction process described in this disclosure, S430 can include the residual processing process described in this disclosure, S440 can include the block / frame reconstruction process described in this disclosure, and S450 can include the in-loop filtering process described in this disclosure.

[0111] refer to FIG. 4 The image decoding process can schematically include a process of obtaining image / video information from the bitstream (through decoding) (S410), an image reconstruction process (S420 to S440), and an in-loop filtering process for the reconstructed image (S450). The image reconstruction process can be performed based on prediction samples and residual samples obtained through inter-frame / intra-frame prediction (S420) and residual processing (S430) (dequantization and inverse transform of quantization transform coefficients) as described in this disclosure. For the reconstructed image generated by the image reconstruction process, a modified reconstructed image can be generated through the in-loop filtering process. The modified reconstructed image can be used as the decoded image output, stored in the decoded image buffer or memory 250 of the decoding device, and used as a reference image in the inter-frame prediction process when decoding images later. In some cases, the in-loop filtering process can be omitted. In this case, the reconstructed image can be used as the decoded image output, stored in the decoded image buffer or memory 250 of the decoding device, and used as a reference image in the inter-frame prediction process when decoding images later. The in-loop filtering process (S450) may include a deblocking filtering process, a sample adaptive offset (SAO) process, an adaptive loop filter (ALF) process, and / or a bilateral filter process, some or all of which may be omitted as described above. Furthermore, one or more of the deblocking filtering process, the sample adaptive offset (SAO) process, the adaptive loop filter (ALF) process, and / or the bilateral filter process may be applied sequentially, or all of them may be applied sequentially. For example, the SAO process may be performed after the deblocking filtering process has been applied to the reconstructed frame. Alternatively, for example, the ALF process may be performed after the deblocking filtering process has been applied to the reconstructed frame. This can be performed similarly even in an encoding device.

[0112] FIG. 5 An example of an illustrative screen encoding process to which embodiments of this disclosure apply is shown. FIG. 5 In the above reference, S510 can be found... FIG. 2The encoding device described can perform in a predictor including the intra predictor 185 or the inter predictor 180, S520 can perform in a residual processor including the transformer 120 and / or the quantizer 130, and S530 can perform in the entropy encoder 190. S510 can include the inter / intra prediction process described in the present disclosure, S520 can include the residual processing process described in the present disclosure, and S530 can include the information encoding process described in the present disclosure.

[0113] With reference to FIG. 5 , the picture encoding process can not only illustratively include a process for encoding information (e.g., prediction information, residual information, partition information, etc.) used for picture reconstruction and outputting the information in the form of a bitstream, but also can include a process for generating a reconstructed picture of the current picture and an (optional) process for applying in-loop filtering to the reconstructed picture, as described with respect to FIG. 2 The encoding device can derive (modified) residual samples from the quantized transform coefficients through the dequantizer 140 and the inverse transformer 150, and generate a reconstructed picture based on the prediction samples as the output of S510 and the (modified) residual samples. The reconstructed picture generated in this way can be equal to the reconstructed picture generated in the decoding device. The modified reconstructed picture can be generated through the in-loop filtering process for the reconstructed picture, can be stored in the decoded picture buffer or the memory 170, and can be used as a reference picture in the inter prediction process when encoding a picture later, similarly to the case in the decoding device. As described above, in some cases, some or all of the in-loop filtering process can be omitted. When the in-loop filtering process is performed, (in-loop) filtering-related information (parameters) can be encoded in the entropy encoder 190 and output in the form of a bitstream, and the decoding device can perform the in-loop filtering process based on the filtering-related information using the same method as the encoding device.

[0114] Through such an in-loop filtering process, noise (such as blocking artifacts and ringing artifacts) generated during image / video encoding can be reduced, and subjective / objective visual quality can be improved. In addition, by performing the in-loop filtering process in both the encoding device and the decoding device, the encoding device and the decoding device can derive the same prediction results, can increase the reliability of picture encoding, and can reduce the amount of data transmitted for picture encoding.

[0115] As described above, the picture reconstruction process can be performed not only in a decoding apparatus but also in an encoding apparatus. The reconstructed blocks can be generated based on the intra prediction / inter prediction in a unit of block, and the reconstructed picture including the reconstructed blocks can be generated. When the current picture / slice / tile group is an I picture / slice / tile group, the blocks included in the current picture / slice / tile group can be reconstructed based on only the intra prediction. Also, when the current picture / slice / tile group is a P or B picture / slice / tile group, the blocks included in the current picture / slice / tile group can be reconstructed based on the intra prediction or the inter prediction. In this case, the inter prediction can be applied to some blocks in the current picture / slice / tile group, and the intra prediction can be applied to the remaining blocks. The color components of a picture can include a luma component and a chroma component, and the methods and embodiments of the disclosure can be applied to the luma component and the chroma component unless the disclosure is explicitly limited.

[0116] Examples of coding layers and structures

[0117] An encoded video / image according to the disclosure can be processed, for example, according to the encoding layer and structure to be described below.

[0118] FIG. 6 is a view showing a layer structure of an encoded image. The encoded image can be classified into a video coding layer (VCL) for a process of an image decoding process and itself, a lower system for transmitting and storing encoded information, and a network abstraction layer (NAL) existing between the VCL and the lower system and responsible for a network adaptation function.

[0119] In the VCL, VCL data including compressed image data (slice data) can be generated, or supplemental enhancement information (SEI) messages additionally required for a decoding process of an image such as a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS) can be generated.

[0120] In the NAL, header information (a NAL unit header) can be added to the raw byte sequence payload (RBSP) generated in the VCL to generate a NAL unit. In this case, the RBSP refers to the slice data, the parameter set, and the SEI message generated in the VCL. The NAL unit header can include NAL unit type information designated according to the RBSP data included in the corresponding NAL unit.

[0121] As illustrated, the NAL unit can be classified into a VCL NAL unit and a non-VCL NAL unit according to the RBSP generated in the VCL. The VCL NAL unit can refer to a NAL unit including information about an image (slice data), and the non-VCL NAL unit can refer to a NAL unit including information required for decoding an image (a parameter set or an SEI message).

[0122] The VCL NAL unit and the non-VCL NAL unit can be attached with header information and transmitted through a network according to a data standard of a lower system. For example, the NAL unit can be modified into a data format of a predetermined standard such as an H.266 / VVC file format, an RTP (Real-time Transport Protocol), or a TS (Transport Stream), and transmitted through various networks.

[0123] As described above, in the NAL unit, the NAL unit type can be designated according to the RBSP data structure included in the corresponding NAL unit, and information about the NAL unit type can be stored in the NAL unit header and signaled.

[0124] For example, according to whether the NAL unit includes information (slice data) about an image, this can be primarily classified into a VCL NAL unit type and a non-VCL NAL unit type. The VCL NAL unit type can be classified according to the characteristics and type of a picture included in the VCL NAL unit, and the non-VCL NAL unit type can be classified according to the type of a parameter set.

[0125] Examples of the NAL unit type designated according to the type of the parameter set / information included in the non-VCL NAL unit type will be listed below.

[0126] - DCI (Decoding Capability Information) NAL unit: type of NAL unit including DCI

[0127] - VPS (Video Parameter Set) NAL unit: type of NAL unit including VPS

[0128] - SPS (Sequence Parameter Set) NAL unit: type of NAL unit including SPS

[0129] - PPS (Picture Parameter Set) NAL unit: type of NAL unit including PPS

[0130] - APS (Adaptive Parameter Set) NAL unit: type of NAL unit including APS

[0131] - PH (Picture Header) NAL unit: type of NAL unit including PH

[0132] The above-described NAL unit type can have syntax information for the NAL unit type, and the syntax information can be stored in the NAL unit header and signaled. For example, the syntax information can be nal_unit_type, and the NAL unit type can be designated as a nal_unit_type value.

[0133] Further, as described above, one picture can include a plurality of slices, and one slice can include a slice header and slice data. In this case, one picture header can be further added to the plurality of slices (slice header and slice data sets) in one picture. The picture header (picture header syntax) can include information / parameters that are generally applicable to a picture.

[0134] The slice header (slice header syntax) can include information / parameters that are generally applicable to a slice. The APS (APS syntax) or the PPS (PPS syntax) can include information / parameters that are generally applicable to one or more slices or pictures. The SPS (SPS syntax) can include information / parameters that are generally applicable to one or more sequences. The VPS (VPS syntax) can include information / parameters that are generally applicable to a plurality of layers. The DCI (DCI syntax) can include information / parameters that are generally applicable to an entire video. The DCI can include information / parameters related to decoding capability. In the disclosure, a high-level syntax (HLS) can include at least one of the APS syntax, the PPS syntax, the SPS syntax, the VPS syntax, the DCI syntax, the picture header syntax, or the slice header syntax. Further, in the disclosure, a low-level syntax (LLS) can include, for example, a slice data syntax, a CTU syntax, a coding unit syntax, a transform unit syntax, and the like.

[0135] In the disclosure, image / video information encoded in an encoding device and signaled to a decoding device in the form of a bitstream can include not only picture-in- partitioning-related information, intra / inter prediction information, residual information, in-loop filtering information, but also information about a slice header, information about a picture header, information about an APS, information about a PPS, information about an SPS, information about a VPS, and / or information about a DCI. In addition, the image / video information can further include general constraint information and / or information about a NAL unit header.

[0136] Picture information signaling NAL units

[0137] Picture information can be signaled in units of NAL units. For example, the picture information can be signaled as in the following description. A sublayer is a temporal scalable layer of a temporal scalable bitstream, which consists of VCL NAL units with a certain value of the variable Temporalld and associated non-VCL NAL units. The variable Temporalld can be derived as follows:

[0138] [Equation 1]

[0139] Temporalld = nuh_temporal_id_plus1 - 1

[0140] The syntax element nuh temporal id plusl for signaling the value of the variable Temporalld can be signaled by the NAL unit header of a NAL unit. When the value of nal unit type in the NAL unit header is in the range of IDR W RADL to RSV IRAP 12, inclusive, the value of Temporalld shall be equal to 0. When the value of nal unit type is equal to STSA NUT and the value of vps independent layer flag[GeneralLayerldx[nuh layer id]] is equal to 1, the value of Temporalld shall not be equal to 0. The value of Temporalld shall be the same for all VCL NAL units of an AU. The value of Temporalld for a coded picture, PU, or AU is the value of Temporalld for the VCL NAL units of the coded picture, PU, or AU. The value of Temporalld for a sublayer representation is the maximum value of Temporalld for all VCL NAL units in the sublayer representation.

[0141] The value of Temporalld for a non-VCL NAL unit can be constrained as follows:

[0142] If nal unit type is equal to DCI NUT, VPS NUT, or SPS NUT, Temporalld shall be equal to 0 and the value of Temporalld for the AU containing the NAL unit shall be equal to 0.

[0143] Otherwise, if nal unit type is equal to PH NUT, the value of Temporalld shall be equal to the Temporalld of the PU containing the NAL unit.

[0144] Otherwise, if nal unit type is equal to EOS NUT or EOB NUT, Temporalld shall be equal to 0.

[0145] Otherwise, if nal unit type is equal to AUD NUT, FD NUT, PREFIX SEI NUT, or SUFFIX SEI NUT, Temporalld shall be equal to the Temporalld of the AU containing the NAL unit.

[0146] Otherwise, when nal unit type is equal to PPS NUT, PREFIX APS NUT, or SUFFIX APS NUT, the value of Temporalld shall be greater than or equal to the value of Temporalld of the PU containing the NAL unit.

[0147] For example, when the NAL unit is a non-VCL NAL unit, the value of Temporalld can be equal to the minimum of the Temporalld values of all AUs to which the non-VCL NAL unit applies. When the value of nal unit type is equal to PPS NUT, PREFIX APS NUT, or SUFFIX APS NUT, the value of Temporalld can be greater than or equal to the value of Temporalld of the AU containing it, as all PPS and APS can be included in the beginning of the bitstream (e.g., when they are transmitted out-of-band and the receiver places them at the beginning of the bitstream). Here, the first coded picture can have a Temporalld equal to 0.

[0148] In embodiments, the coded pictures obtained from the bitstream according to the NAL unit information can be signaled by the encoding device and identified by the decoding device as follows. However, this is an example and other methods can be used to identify the pictures.

[0149] An intra random access point (IRAP) picture is a coded picture for which all VCL NAL units have the same value of nal unit type in the range of IDR W RADL to CRA NUT, inclusive. In embodiments, an IRAP picture does not reference any picture other than itself for inter prediction in its decoding process. An IRAP picture can be a CRA picture or an IDR picture. The first picture in a bitstream in decoding order must be an IRAP or GDR picture. An IRAP picture and all subsequent non-RASL pictures in CVS in decoding order can be decoded correctly without performing the decoding process of any picture that precedes the IRAP picture in decoding order if the necessary parameter sets are available when needed.

[0150] A clean random access (CRA) picture is an IRAP picture with each VCL NAL unit having a nal unit type equal to CRA NUT. For example, a CRA picture does not reference any picture other than itself for inter prediction in its decoding process, and can be the first picture in a bitstream in decoding order, or can appear later in the bitstream. A CRA picture can have associated RADL or RASL pictures. When a CRA picture has NoIncorrectPicOutputFlag equal to 1, the associated RASL pictures can not be output by the decoder, as they can not be decodable, as they can contain references to pictures that are not present in the bitstream. In implementations, when incomplete pictures are not output during picture decoding, a CRA picture can have NoIncorrectPicOutputFlag equal to 1 if the CRA picture is an incomplete picture.

[0151] An instantaneous decoding refresh (IDR) picture is an IRAP picture with each VCL NAL unit having a nal unit type equal to IDR W RADL or IDR N LP. For example, an IDR picture does not reference any picture other than itself for inter prediction in its decoding process. In addition, an IDR picture can be the first picture in a bitstream in decoding order, or can appear later in the bitstream. Each IDR picture is the first picture of a CVS in decoding order. When the IDR picture has each VCL NAL unit having a nal unit type equal to IDR W RADL, the IDR picture can have associated RADL pictures. When the IDR picture has each VCL NAL unit having a nal unit type equal to IDR N LP, the IDR picture has no associated leading pictures. An IDR picture has no associated RASL pictures.

[0152] A random access decodable leading (RADL) picture is a coded picture with each VCL NAL unit having a nal unit type equal to RADL NUT. In implementations, all RADL pictures are leading pictures. A RADL picture is not used as a reference picture in the decoding process of trailing pictures of the same associated IRAP picture. When a syntax element field seq flag, obtained from the bitstream, is equal to 0, all RADL pictures, when present, precede, in decoding order, all non-leading pictures of the same associated IRAP picture.

[0153] A random access skipped leading (RASL) picture can be a coded picture with each VCL NAL unit having nal unit type equal to RASL NUT. In an implementation, all RASL pictures are leading pictures of an associated CRA picture. When the associated CRA picture has NoIncorrectPicOutputFlag equal to 1, the RASL picture is not output and cannot be correctly decoded because the RASL picture can contain references to pictures that are not present in the bitstream. A RASL picture is not used as a reference picture for the decoding process of non-RASL pictures. When field_seq_flag is equal to 0, all RASL pictures, when present, precede, in decoding order, all non-leading pictures of the same associated CRA picture.

[0154] A trailing picture is a non-IRAP picture that follows, in output order, an associated IRAP picture and is not an STSA picture. A trailing picture associated with an IRAP picture also follows, in decoding order, the IRAP picture. A picture that follows, in output order, an associated IRAP picture and precedes, in decoding order, the associated IRAP picture is not allowed.

[0155] A gradual decoding refresh (GDR) picture is a picture with each VCL NAL unit having nal unit type equal to GDR NUT.

[0156] A stepwise temporal sub-layer access (STSA) picture is a coded picture with each VCL NAL unit having nal unit type equal to STSA NUT. An STSA picture can not use pictures with the same Temporalld as the STSA picture for inter prediction reference. A picture with the same Temporalld as the STSA picture that follows the STSA picture in decoding order can not use pictures with the same Temporalld as the STSA picture that precede the STSA picture in decoding order for inter prediction reference.

[0157] An STSA picture enables up-switching from the immediately lower sub-layer to the sub-layer containing the STSA picture at the STSA picture. An STSA picture must have a Temporalld greater than 0.

[0158] In an implementation, for a single-layer or multi-layer bitstream, one or more of the following constraints can apply:

[0159] - Each picture, except the first picture in the bitstream in decoding order, can be considered to be associated with the preceding IRAP picture in decoding order.

[0160] - When a picture is a leading picture of an IRAP picture, it will be a RADL or RASL picture.

[0161] - When the picture is a trailing picture of an IRAP picture, it shall not be a RADL or RASL picture.

[0162] - There shall not be a RASL picture associated with an IDR picture in the bitstream.

[0163] - There shall not be a RADL picture associated with an IDR picture with nal unit type equal to IDR N LP in the bitstream. (For example, random access can be performed at the location of an IRAP PU by discarding all PUs before the IRAP PU and correctly decoding the IRAP picture and all subsequent non-RASL pictures in decoding order, provided that each parameter set is available in the bitstream when referred to or is obtainable by an external device.)

[0164] - Any picture that precedes the IRAP picture in decoding order shall precede the IRAP picture in output order, and shall precede any RADL picture associated with the IRAP picture in output order.

[0165] - Any RASL picture associated with a CRA picture shall precede any RADL picture associated with the CRA picture in output order.

[0166] - Any RASL picture associated with a CRA picture shall be located in output order after any IRAP picture that precedes the CRA picture in decoding order.

[0167] - If field_seq_flag is equal to 0 and the current picture is a leading picture associated with an IRAP picture, it precedes all non-leading pictures associated with the same IRAP picture in decoding order. Otherwise, let picA and picB be the first leading picture and the last leading picture associated with an IRAP picture in decoding order, respectively, there shall be at most one non-leading picture preceding picA in decoding order, and there shall be no non-leading picture between picA and picB in decoding order.

[0168] Coding based on multiple layers

[0169] The image / video encoding according to the disclosure can include multi-layer based image / video encoding. The multi-layer based image / video encoding can include scalable encoding. In the multi-layer based encoding or the scalable encoding, an input signal can be processed for each layer. According to the layer, the input signal (input image / video) can have different values in terms of at least one of resolution, frame rate, bit depth, color format, aspect ratio, or view angle. In this case, redundant information transmission / processing can be reduced and compression efficiency can be increased by performing inter-layer prediction using a difference between layers (e.g., based on scalability).

[0170] FIG. 7 is a schematic block diagram of a multi-layer encoding apparatus 700 to which embodiments of the disclosure are applied and which performs encoding of a multi-layer video / image signal.

[0171] FIG. 7 The multi-layer encoding apparatus 700 of FIG. 2 may include an encoding apparatus. In comparison with FIG. 2 , the image partitioner 110 and the adder 155 are not shown in the multi-layer encoding apparatus 700 of FIG. 7 , and the multi-layer encoding apparatus 700 can include the image partitioner 110 and the adder 155. In an embodiment, the image partitioner 110 and the adder 155 can be included in a unit of layer. Hereinafter, the prediction based on the multi-layer will be focused on the description of FIG. 7 . For example, in addition to the following description, the multi-layer encoding apparatus 700 can include the technical ideas of the encoding apparatus described above with reference to FIG. 2 .

[0172] For convenience of description, a multi-layer structure composed of two layers is shown in FIG. 7 . However, embodiments of the disclosure are not limited to two layers, and the multi-layer structure to which embodiments of the disclosure are applied can include two or more layers.

[0173] Referring to FIG. 7 , the encoding apparatus 700 includes an encoder 700-1 of layer 1 and an encoder 700-0 of layer 0. Layer 0 can be a base layer, a reference layer, or a lower layer, and layer 1 can be an enhancement layer, a current layer, or a higher layer.

[0174] The encoder 700-1 of layer 1 can include a predictor 720-1, a residual processor 730-1, a filter 760-1, a memory 770-1, an entropy encoder 740-1, and a multiplexer (MUX) 770. In an embodiment, the MUX can be included as an external component.

[0175] The encoder 700-0 of layer 0 can include a predictor 720-0, a residue processor 730-0, a filter 760-0, a memory 770-0, and an entropy encoder 740-0.

[0176] The predictors 720-0 and 720-1 can perform prediction for input pictures based on various prediction schemes as described above. For example, the predictors 720-0 and 720-1 can perform inter prediction and intra prediction. The predictors 720-0 and 720-1 can perform prediction in a predetermined processing unit. The processing unit can be a coding unit (CU) or a transform unit (TU). A prediction block (including prediction samples) can be generated according to a result of prediction, and based on this, the residue processor can derive a residual block (including residual samples).

[0177] Through inter prediction, prediction can be performed based on information about at least one of a previous preceding picture and / or a next picture of a current picture, thereby generating a prediction block. Through intra prediction, prediction can be performed based on neighboring samples in the current picture, thereby generating a prediction block.

[0178] As an inter prediction mode or method, various prediction modes or methods described above can be used. In inter prediction, a reference picture can be selected for a current block to be predicted, and a reference block corresponding to the current block can be selected from the reference picture. The predictors 720-0 and 720-1 can generate a prediction block based on the reference block.

[0179] In addition, the predictor 720-1 can perform prediction for layer 1 using information about layer 0. In the disclosure, a method of predicting information about a current layer using information about another layer for convenience of description is referred to as inter-layer prediction.

[0180] Information about a current layer predicted using information about another layer (e.g., predicted through inter-layer prediction) can be at least one of texture, motion information, unit information, or a predetermined parameter (e.g., a filtering parameter, etc.).

[0181] In addition, information about another layer used for prediction of a current layer (e.g., for inter-layer prediction) can be at least one of texture, motion information, unit information, or a predetermined parameter (e.g., a filtering parameter, etc.).

[0182] For inter-layer prediction, a current block can be a block in a current picture in a current layer (e.g., layer 1), and can be a block to be encoded. A reference block is a block in a picture (a reference picture) on a layer (a reference layer, e.g., layer 0) to which a current block refers, which belongs to the same access unit (AU) as a picture (a current picture) to which the current block belongs, and can be a block corresponding to the current block.

[0183] As an example of inter-layer prediction, there is inter-layer motion prediction that predicts motion information of a current layer using motion information of a reference layer. According to inter-layer motion prediction, motion information of a current block can be predicted using motion information of a reference block. That is, when deriving motion information according to an inter prediction mode to be described below, motion information candidates can be derived based on motion information of an inter-layer reference block rather than a temporal neighboring block.

[0184] When inter-layer motion prediction is applied, the predictor 720-1 can scale and use reference block (that is, inter-layer reference block) motion information of a reference layer.

[0185] As another example of inter-layer prediction, inter-layer texture prediction can use texture of a reconstructed reference block as a prediction value of a current block. In this case, the predictor 720-1 can scale the texture of the reference block by up-scaling. Inter-layer texture prediction can be referred to as inter-layer (reconstructed) sample prediction or simply inter-layer prediction.

[0186] In inter-layer parameter prediction, which is another example of inter-layer prediction, a derived parameter of a reference layer can be reused in a current layer, or a parameter for the current layer can be derived based on a parameter used in the reference layer.

[0187] In inter-layer residual prediction, which is another example of inter-layer prediction, residual information of another layer can be used to predict residual information of a current layer, and based on this, prediction of a current block can be performed.

[0188] In inter-layer difference prediction, which is another example of inter-layer prediction, prediction of a current block can be performed using a difference between images obtained by up-sampling or down-sampling a reconstructed picture of a current layer and a reconstructed picture of a reference layer.

[0189] In inter-layer syntax prediction, which is another example of inter-layer prediction, syntax information of a reference layer can be used to predict or generate texture of a current block. In this case, the syntax information of the reference layer referred to can include information on an intra prediction mode and motion information.

[0190] When predicting a specific block, a plurality of prediction methods using the above-described inter-layer prediction can be utilized.

[0191] Here, as examples of inter-layer prediction, although inter-layer texture prediction, inter-layer motion prediction, inter-layer unit information prediction, inter-layer parameter prediction, inter-layer residual prediction, inter-layer difference prediction, inter-layer syntax prediction, and the like are described, inter-layer prediction applicable to the present disclosure is not limited thereto.

[0192] For example, inter-layer prediction can be applied as an extension of inter prediction for a current layer. That is, inter prediction can be performed for a current block by including a reference picture derived from a reference layer in reference pictures that can be referred to for inter prediction of the current block.

[0193] In this case, the inter-layer reference picture can be included in a reference picture list for the current block. The predictor 720-1 can perform inter prediction for the current block using the inter-layer reference picture.

[0194] Here, the inter-layer reference picture can be a reference picture constructed by sampling a reconstructed picture of a reference layer to correspond to a current layer. Thus, when the reconstructed picture of the reference layer corresponds to a picture of the current layer, the reconstructed picture of the reference layer can be used as the inter-layer reference picture without sampling. For example, when a width and a height of a sample are the same in the reconstructed picture of the reference layer and the reconstructed picture of the current layer and an offset between an upper left end, an upper right end, a lower left end, and a lower right end in the picture of the reference layer and an upper left end, an upper right end, a lower left end, and a lower right end in the picture of the current layer is 0, the reconstructed picture of the reference layer can be used as the inter-layer reference picture for the current layer without being sampled again.

[0195] In addition, the reconstructed picture of the reference layer from which the inter-layer reference picture is derived can be a picture belonging to the same AU as a current picture to be encoded.

[0196] When inter prediction for a current block is performed by including an inter-layer reference picture in a reference picture list, a position of the inter-layer reference picture in the reference picture list can be different between a reference picture list L0 and L1. For example, in the reference picture list L0, the inter-layer reference picture can be located after short-term reference pictures before the current picture, and in the reference picture list L1, the inter-layer reference picture can be located at the end of the reference picture list.

[0197] Here, the reference picture list L0 is a reference picture list for inter prediction of a P slice or a reference picture list used as a first reference picture list in inter prediction of a B slice. The reference picture list L1 can be a second reference picture list for inter prediction of a B slice.

[0198] Thus, the reference picture list L0 can consist of short-term reference pictures before the current picture, the inter-layer reference picture, short-term reference pictures after the current picture, and long-term reference pictures in the following order. The reference picture list L1 can consist of short-term reference pictures after the current picture, short-term reference pictures before the current picture, long-term reference pictures, and the inter-layer reference picture in the following order.

[0199] In this case, a P (Prediction) slice is a slice for which inter prediction is performed using at most one motion vector and a reference picture index per prediction block or intra prediction is performed. A B (Bi-prediction) slice is a slice for which prediction is performed using at most two motion vectors and a reference picture index per prediction block or intra prediction is performed. In this regard, an I (Intra) slice is a slice for which only intra prediction is applied.

[0200] In addition, when inter prediction of a current block is performed based on a reference picture list including an inter-layer reference picture, the reference picture list can include a plurality of inter-layer reference pictures derived from a plurality of layers.

[0201] When a plurality of inter-layer reference pictures are included, the inter-layer reference pictures can be alternately arranged in the reference picture lists L0 and L1. For example, assume that two inter-layer reference pictures, such as an inter-layer reference picture ILRPi and an inter-layer reference picture ILRPj, are included in a reference picture list for inter prediction of a current block. In this case, in the reference picture list L0, ILRPi can be located after a short-term reference picture preceding the current picture and ILRPj can be located at the end of the list. In addition, in the reference picture list L1, ILRPi can be located at the end of the list and ILRPj can be located after a short-term reference picture following the current picture.

[0202] In this case, the reference picture list L0 can consist of a short-term reference picture preceding the current picture, the inter-layer reference picture ILRPi, a short-term reference picture following the current picture, a long-term reference picture, and the inter-layer reference picture ILRPj in the following order. The reference picture list L1 can consist of a short-term reference picture following the current picture, the inter-layer reference picture ILRPj, a short-term reference picture preceding the current picture, a long-term reference picture, and the inter-layer reference picture ILRPi in the following order.

[0203] In addition, one of the two inter-layer reference pictures can be an inter-layer reference picture derived from a scalable layer for resolution and the other can be an inter-layer reference picture derived from a layer for providing another view. In this case, for example, if ILRPi is an inter-layer reference picture derived from a layer for providing a different resolution and ILRPj is an inter-layer reference picture derived from a layer for providing a different view, in the case of scalable video coding supporting scalability excluding view scalability only, the reference picture list L0 can consist of a short-term reference picture preceding the current picture, the inter-layer reference picture ILRPi, a short-term reference picture following the current picture, and a long-term reference picture in the following order, and the reference picture list L1 can consist of a short-term reference picture following the current picture, a short-term reference picture preceding the current picture, a long-term reference picture, and the inter-layer reference picture ILRPi in the following order.

[0204] Further, in inter-layer prediction, as information about the inter-layer reference picture, only sample values can be used, only motion information (motion vector) can be used, or both sample values and motion information can be used. When the reference picture index indicates the inter-layer reference picture, according to information received from the encoding apparatus, the predictor 720-1 can use only sample values of the inter-layer reference picture, can use only motion information (motion vector) of the inter-layer reference picture, or can use both sample values and motion information of the inter-layer reference picture.

[0205] When only sample values of the inter-layer reference picture are used, the predictor 720-1 can derive samples of a block specified by a motion vector from the inter-layer reference picture as prediction samples of the current block. In the case of scalable video encoding without considering a view angle, a motion vector in inter-layer prediction (inter-layer prediction) using the inter-layer reference picture can be set to a fixed value (for example, 0).

[0206] When only motion information of the inter-layer reference picture is used, the predictor 720-1 can use a motion vector specified by the inter-layer reference picture as a motion vector predictor for deriving a motion vector of the current block. In addition, the predictor 720-1 can use the motion vector specified by the inter-layer reference picture as a motion vector of the current block.

[0207] When both sample values and motion information of the inter-layer reference picture are used, the predictor 720-1 can use samples of a region of the inter-layer reference picture corresponding to the current block and motion information (motion vector) specified in the inter-layer reference picture for prediction of the current block.

[0208] When inter-layer prediction is applied, the encoding apparatus can transmit a reference index indicating an inter-layer reference picture in a reference picture list to the decoding apparatus, and can transmit information for specifying which information (sample information, motion information, or sample information and motion information) to use from the inter-layer reference picture (that is, information for specifying a dependency type of dependency of inter-layer prediction between two layers) to the decoding apparatus.

[0209] FIG. 8 is a schematic block diagram of a decoding apparatus to which embodiments of the present disclosure are applicable and which performs decoding of a multi-layer video / image signal. FIG. 8 The decoding apparatus of FIG. 3 The decoding apparatus of FIG. 8 The realigner shown in can be omitted or included in the dequantizer. In the description of the diagram, focus will be placed on prediction based on multiple layers. In addition, the description of the decoding apparatus of FIG. 3 The decoding apparatus of

[0210] In FIG. 8In the example of FIG. 8, a multi-layer structure composed of two layers will be described for convenience of description. However, it should be noted that embodiments of the present disclosure are not limited thereto, and embodiments of the present disclosure apply to a multi-layer structure composed of two or more layers.

[0211] Referring to FIG. 8, FIG. 8 The decoding apparatus 800 can include a decoder 800-1 of layer 1 and a decoder 800-0 of layer 0. The decoder 800-1 of layer 1 can include an entropy decoder 810-1, a residue processor 820-1, a predictor 830-1, an adder 840-1, a filter 850-1, and a memory 860-1. The decoder 800-0 of layer 0 can include an entropy decoder 810-0, a residue processor 820-0, a predictor 830-0, an adder 840-0, a filter 850-0, and a memory 860-0.

[0212] When a bitstream including image information is received from the encoding apparatus, the demultiplexer 805 can demultiplex information of each layer and transmit the information to the decoding apparatus for each layer.

[0213] The entropy decoders 810-1 and 810-0 can perform decoding corresponding to the encoding method used in the encoding apparatus. For example, when CABAC is used in the encoding apparatus, the entropy decoders 810-1 and 810-0 can perform entropy decoding using CABAC.

[0214] When the prediction mode of the current block is an intra prediction mode, the predictors 830-1 and 830-0 can perform intra prediction on the current block based on neighboring reconstructed samples in the current picture.

[0215] When the prediction mode of the current block is an inter prediction mode, the predictors 830-1 and 830-0 can perform inter prediction on the current block based on information included in at least one of a picture before or after the current picture. Some or all of the motion information necessary for inter prediction can be derived by checking information received from the encoding apparatus.

[0216] When a skip mode is applied as the inter prediction mode, no residual is transmitted from the encoding apparatus, and the predicted block can be the reconstructed block.

[0217] In addition, the predictor 830-1 of layer 1 can perform inter prediction or intra prediction using only information about layer 1, and perform inter-layer prediction using information about another layer (layer 0).

[0218] As information about the current layer predicted using information about another layer (e.g., predicted by inter-layer prediction), there can be at least one of texture, motion information, unit information, predetermined parameters (e.g., filtering parameters, etc.).

[0219] As information about another layer for prediction of the current layer (e.g., for inter-layer prediction), there can be at least one of a texture, motion information, a unit information, a predetermined parameter (e.g., a filtering parameter, etc.).

[0220] In inter-layer prediction, a current block can be a block in a current picture in a current layer (e.g., layer 1), and can be a block to be decoded. A reference block can be a block in a picture (reference picture) on a layer (reference layer, e.g., layer 0) to which a current block refers for prediction, and which belongs to the same access unit (AU) as a picture (current picture) to which the current block belongs, and can be a block corresponding to the current block.

[0221] The multi-layer decoding device 800 can perform inter-layer prediction as described in the multi-layer encoding device 700. For example, the multi-layer decoding device 800 can perform inter-layer texture prediction, inter-layer motion prediction, inter-layer unit information prediction, inter-layer parameter prediction, inter-layer residual prediction, inter-layer difference prediction, inter-layer syntax prediction, etc. as described in the multi-layer encoding device 700, and inter-layer prediction applicable in the present disclosure is not limited thereto.

[0222] When the reference picture index received from the encoding device or the reference picture index derived from the neighboring block indicates an inter-layer reference picture in the reference picture list, the predictor 830-1 can perform inter-layer prediction using the inter-layer reference picture. For example, when the reference picture index indicates the inter-layer reference picture, the predictor 830-1 can derive sample values of a region in the inter-layer reference picture specified by the motion vector as a prediction block of the current block.

[0223] In this case, the inter-layer reference picture can be included in the reference picture list for the current block. The predictor 830-1 can perform inter prediction on the current block using the inter-layer reference picture.

[0224] As described above in the multi-layer encoding device 700, in the operation of the multi-layer decoding device 800, the inter-layer reference picture can be a reference picture constructed by sampling a reconstructed picture of a reference layer to correspond to a current layer. The processing of the case where the reconstructed picture of the reference layer corresponds to the picture of the current layer can be performed in the same manner as the encoding process.

[0225] In addition, as described above in the multi-layer encoding device 700, in the operation of the multi-layer decoding device 800, the reconstructed picture of the reference layer from which the inter-layer reference picture is derived can be a picture belonging to the same AU as a current picture to be encoded.

[0226] Also, as described above in the multi-layer encoding apparatus 700, in the operation of the multi-layer decoding apparatus 800, when inter-layer reference pictures in the reference picture list are included to perform inter prediction of the current block, the positions of the inter-layer reference pictures in the reference picture list can be different between the reference picture lists L0 and L1.

[0227] Also, as described above in the multi-layer encoding apparatus 700, in the operation of the multi-layer decoding apparatus 800, when inter-layer reference pictures in the reference picture list are included to perform inter prediction of the current block, the positions of the inter-layer reference pictures in the reference picture list can be different between the reference picture lists L0 and L1.

[0228] Also, as described above in the multi-layer encoding apparatus 700, in the operation of the multi-layer decoding apparatus 800, as information about the inter-layer reference picture, only sample values can be used, only motion information (motion vector) can be used, or both sample values and motion information can be used.

[0229] The multi-layer decoding apparatus 800 can receive a reference index indicating an inter-layer reference picture in a reference picture list from the multi-layer encoding apparatus 700, and perform inter-layer prediction based on the reference index. Also, the multi-layer decoding apparatus 800 can receive information for specifying which information (sample information, motion information, or both sample information and motion information) to use from an inter-layer reference picture from the multi-layer encoding apparatus 700 (that is, information for specifying a dependency type of dependency between inter-layer predictions of two layers).

[0230] High-level syntax (HLS) signaling and semantics

[0231] As described above, the HLS can be encoded and / or signaled for video and / or image encoding. As described above, the video / image information of the present disclosure can be included in the HLS. Also, an image / video encoding method can be performed based on such image / video information.

[0232] Video parameter set signaling

[0233] A video parameter set (VPS) is a parameter set that is used for the carriage of layer information. Layer information can include, for example, information about output layer sets (OLSs), information about profile tier level, information about the relationship between an OLS and a hypothetical reference decoder, and information about the relationship between an OLS and a decoded picture buffer (DPB). A VPS can not be necessary for the decoding of a bitstream. A VPS raw byte sequence payload (RBSP) is included in at least one access unit (AU) with Temporalld equal to 0 before it is referenced or provided through external means, such that it shall be available to the decoding process. All VPS NAL units with a particular value of vps_video_parameter_set_id in a coded video sequence (CVS) shall have the same content.

[0234] FIG. 9 is a view illustrating a syntax structure of a VPS according to an embodiment of the disclosure. Hereinafter, syntax elements of FIG. 9 will be described.

[0235] vps_video_parameter_set_id provides an identifier for a VPS. Other syntax elements can refer to the VPS using vps_video_parameter_set_id. The value of vps_video_parameter_set_id shall be greater than 0.

[0236] vps_max_layers_minus1 can specify the maximum allowed number of layers in each CVS referring to the VPS. For example, vps_max_layers_minus1 plus 1 can specify the maximum allowed number of layers in each CVS referring to the VPS.

[0237] vps_max_sublayer_minus1 plus 1 can specify the maximum number of temporal sub-layers that can be present in layers in each CVS referring to the VPS.

[0238] vps_all_layers_same_num_sublayers_flag equal to 1 can specify that the number of temporal sub-layers is the same for all layers in each CVS referring to the VPS. vps_all_layers_same_num_sublayers_flag equal to 0 can specify that the number of temporal sub-layers can be the same or different for layers in each CVS referring to the VPS. When the value of vps_all_layers_same_num_sublayers_flag is not provided in the bitstream, the value of vps_all_layers_same_num_sublayers_flag can be inferred to be equal to 1.

[0239] vps_all_independent_layers_flag equal to 1 can specify that all layers in the CVS are coded independently without using inter-layer prediction. vps_all_independent_layers_flag equal to 0 can specify that one or more of the layers in the CVS can be coded using inter-layer prediction.

[0240] vps_layer_id[ i ] can specify the nuh layer id value of the i-th layer. For any two non-negative integer values of m and n, the value of vps_layer_id[ m ] shall be less than vps_layer_id[ n ] when m is less than n. Here, nuh layer id is a syntax element signaled in the NAL unit header and can specify an identifier of the NAL unit.

[0241] vps_independent_layer_flag[ i ] equal to 1 can specify that the layer with index i does not use inter-layer prediction. vps_independent_layer_flag[ i ] equal to 0 can specify that the layer with index i can use inter-layer prediction and the syntax element vps_direct_ref_layer_flag[ i ][ j ] can be obtained from the VPS. Here, j can be in the range of 0 to i - 1, inclusive. When the value of vps_independent_layer_flag[ i ] is not present in the bitstream, the value of vps_independent_layer_flag[ i ] can be inferred to be equal to 1.

[0242] vps_direct_ref_layer_flag[ i ][ j ] equal to 0 can specify that the layer with index j is not a direct reference layer of the layer with index i. vps_direct_ref_layer_flag[ i ][ j ] equal to 1 can specify that the layer with index j is a direct reference layer of the layer with index i. When the value of vps_direct_ref_layer_flag[ i ][ j ] is not obtained from the bitstream for i and j in the range of 0 to vps_max_layers_minus1, inclusive, the value of vps_direct_ref_layer_flag[ i ][ j ] can be inferred to be equal to 0. When vps_independent_layer_flag[ i ] is equal to 0, there shall exist at least one value of j in the range of 0 to i - 1, inclusive, such that the value of vps_direct_ref_layer_flag[ i ][ j ] is equal to 1.

[0243] In an embodiment, the pseudo code of FIG. 10 derives the variables NumDirectRefLayers[i], DirectRefLayerIdx[i][d], NumRefLayers[i], RefLayerIdx[i][r], and LayerUsedAsRefLayerFlag[j].

[0244] A variable GeneralLayerIdx[i] that specifies the layer index of the layer with nuh layer id equal to vps layer id[i] can be derived as follows.

[0245] [Equation 2]

[0246] for(i = 0; i <= vps max layers minusl; i++)

[0247] GeneralLayerIdx[vps layer id[i]] = i

[0248] max_tid_ref_present_flag[i] equal to 1 can specify that the syntax element max_tid_il_ref_pics_plusl[i] is provided from the bitstream. max_tid_ref_present_flag[i] equal to 0 can specify that the syntax element max_tid_il_ref_pics_plusl[i] is not provided from the bitstream.

[0249] max_tid_il_ref_pics_plusl[i] equal to 0 can specify that inter-layer prediction is not used by non-IRAP pictures of the i-th layer. max_tid_il_ref_pics_plusl[i] greater than 0 can specify that there are no pictures with Temporalld greater than max_tid_il_ref_pics_plusl[i] - 1 that are used as ILRPs (inter-layer reference pictures) for decoded pictures of the i-th layer. When the value of max_tid_il_ref_pics_plusl[i] is not obtained from the bitstream, the value of max_tid_il_ref_pics_plusl[i] can be inferred to be equal to 7.

[0250] A syntax element each_layer_is_an_ols_flag equal to 1 can specify that each OLS contains only one layer and that each layer itself in the CVS of the VPS is an OLS with the single included layer being the only output layer. A each_layer_is_an_ols_flag equal to 0 can specify that an OLS can contain more than one layer. In implementations, if vps_max_layers_minus1 is equal to 0, the value of each_layer_is_an_ols_flag can be inferred to be equal to 1. Otherwise, when vps_all_independent_layers_flag is equal to 0, the value of each_layer_is_an_ols_flag can be inferred to be equal to 0.

[0251] An ols_mode_idc equal to 0 can specify that the total number of OLSs specified by the VPS is equal to vps_max_layers_minus1 + 1. The i-th OLS can include layers with layer indices from 0 to i, inclusive, and for each OLS, only the highest layer in the OLS can be output.

[0252] An ols_mode_idc equal to 1 can specify that the total number of OLSs specified by the VPS is equal to vps_max_layers_minus1 + 1. The i-th OLS can include layers with layer indices from 0 to i, inclusive, and for each OLS, all layers in the OLS can be output.

[0253] An ols_mode_idc equal to 2 can specify that the total number of OLSs specified by the VPS is explicitly signaled and for each OLS, the output layers are explicitly signaled and the other layers are layers that are directly or indirectly referenced layers of the output layers of the OLS.

[0254] The value of ols_mode_idc shall be in the range of 0 to 2, inclusive. The value of 3 for ols_mode_idc is reserved for future use. When vps_all_independent_layers_flag is equal to 1 and each_layer_is_an_ols_flag is equal to 0, the value of ols_mode_idc can be inferred to be equal to 2.

[0255] num_output_layer_sets_minus1 plus 1 can specify the total number of OLSs specified by the VPS when ols_mode_idc is equal to a predetermined value (e.g., 2)

[0256] As FIG. 11As shown, the variable TotalNumOlss is derived that specifies the total number of OLSs specified by the VPS.

[0257] An ols_output_layer_flag[ i ][ j ] equal to 1 can specify that the layer with nuh layer id equal to vps layer id[ j ] is an output layer of the i-th OLS when ols mode idc is equal to 2. An ols_output_layer_flag[ i ][ j ] equal to 0 can specify that the layer with nuh layer id equal to vps layer id[ j ] is not an output layer of the i-th OLS when ols mode idc is equal to 2.

[0258] The variable NumOutputLayersInOls[ i ], which specifies the number of output layers in the i-th OLS, the variable NumSubLayersInLayerInOLS[ i ][ j ], which specifies the number of sub-layers in the j-th layer in the i-th OLS, the variable OutputLayerIdInOls[ i ][ j ], which specifies the nuh layer id value of the j-th output layer in the i-th OLS, and the variable LayerUsedAsOutputLayerFlag[ k ], which specifies whether the k-th layer is used as an output layer in at least one OLS, can be derived similarly to the pseudo code of FIG. 12

[0259] For each value of i in the range of 0 to vps max layers minus 1, inclusive, the values of LayerUsedAsRefLayerFlag[ i ] and LayerUsedAsOutputLayerFlag[ i ] shall not both be equal to 0. In other words, there shall not exist a layer that is neither an output layer of at least one OLS nor a direct reference layer of any other layer.

[0260] For each OLS, there shall exist at least one layer that is an output layer. In other words, for any value of i in the range of 0 to TotalNumOlss - 1, inclusive, the value of NumOutputLayersInOls[ i ] shall be greater than or equal to 1.

[0261] The variable NumLayersInOls[ i ], which specifies the number of layers in the i-th OLS, and the variable LayerIdInOls[ i ][ j ], which specifies the nuh layer id value of the j-th layer in the i-th OLS, can be derived as shown in FIG. 13 ​​

[0262] In an implementation, the 0thOLS contains only the lowest layer. The lowest layer can represent a layer with nuh layer id equal to vps layer id[0]. In addition, for the 0thOLS, only the layer included can be output.

[0263] The variable OlsLayerIdx[i][j] that specifies the OLS layer index of a layer with nuh layer id equal to LayerIdInOls[i][j] can be derived as shown in FIG. 14

[0264] The lowest layer in each OLS will be an independent layer. For example, for each i in the range of 0 to TotalNumOlss - 1, inclusive, the value of vps_independent_layer_flag[GeneralLayerIdx[LayerIdInOls[i][0]]] shall be equal to 1.

[0265] Each layer will be included in at least one OLS specified by the VPS.

[0266] Restriction on signaling of max_tid_il_ref_pics_plusl[i]

[0267] The signaling related to the above syntax element max_tid_il_ref_pics_plus1[i] has issues in terms of its semantics and functionality. For example, the semantics of max_tid_il_ref_pics_plus1[i] indicates that when its value is greater than 0, it means that a picture in the ith layer uses at most max_tid_il_ref_ref_pics_plus1[i] - 1 sub-layers from the reference pictures when performing its inter-layer prediction.

[0268] This means that a picture in the ith layer will use at most max_tid_il_ref_ref_pics_plus1[i] - 1 sub-layers from the reference pictures for its inter-layer prediction.

[0269] The syntax element max_tid_il_ref_pics_plus1[i] is designed to allow the derivation of the variable NumSubLayersInLayerInOLS[i][j] that can specify the number of sub-layers in the jth layer in the ith OLS. Here, OLS is an abbreviation for output layer set and can mean a set of at least one layer specified by an output layer.

[0270] ​The variable NumSubLayersInLayerInOLS [i] [j] can be used in the bitstream extraction process to remove pictures in sub-layers within a layer that are not output layers in the extracted output layer set.

[0271] When a layer uses more than 1 reference layer for inter-layer prediction and the number of sub-layers used from each of the reference layers is not the same, the mechanism is not optimal.

[0272] For example, assume layer 2 references layer 0 and layer 1 for inter-layer prediction. For inter-layer prediction from layer 0, only 2 sub-layers are used, while inter-layer prediction from layer 1 uses 3 sub-layers. In this example, then according to the signaling mechanism described above, layer 2 uses 3 sub-layers for inter-layer prediction, and in this approach, sub-layer 3 or above cannot be removed from layer 0.

[0273] Improvements

[0274] The following embodiments provide a solution to the above problem. The embodiments can be applied individually or in combination.

[0275] Improvement 1. For the signaling of the maximum number of sub-layers for inter-layer prediction of a layer, instead of signaling one value for all reference layers, for each reference layer of a layer, the maximum number of sub-layers used from that reference layer can be signaled.

[0276] Improvement 2: The maximum number of sub-layers used from reference layer j for layer i can only be present when layer j is a direct reference layer of layer i. Improvement 3: The flag specifying whether the signaling of the maximum number of sub-layers for inter-layer prediction by a layer is provided can be signaled one for all layers provided in the VPS.

[0277] The syntax element max_tid_ref_present_flag [i] can be changed to max_tid_ref_present_flag.

[0278] When all layers are independent layers, max_tid_ref_present_flag can not be present, and when max_tid_ref_present_flag is not present, max_tid_ref_present_flag can be inferred to be equal to 0.

[0279] Improvement 4. The signaling of max_tid_il_ref_pics_plusl[i][j] can be used to derive the values of the array NumSubLayersInLayerInOLS[m][n], where m is in the range of 0 to TotalNumOlss - 1, inclusive, and n is in the range of 0 to vps_max_layers_minusl, inclusive.

[0280] Improvement 5. The derivation of NumSubLayersInLayerInOLS for both direct and indirect reference layers is addressed, especially for OLS mode 0 and OLS mode 2.

[0281] a) When the j-th layer in the bitstream is an output layer in the i-th OLS, the value of NumSubLayersInLayerInOLS[i][j] can be set equal to the maximum allowed number of sub-layers.

[0282] b) For non-output layers in an OLS, a process for deriving the value of NumSubLayersInLayerInOLS[i][j] can be performed (i.e., in the example where the j-th layer in the bitstream is not an output layer in the i-th OLS), and the process can be invoked starting from the highest non-output layer in the OLS towards the lowest layer in the OLS. When the value of NumSubLayersInLayerInOLS[i][j] is greater than the smaller value between max_tid_il_ref_pics_plusl[k][j] and NumSubLayersInLayerInOLS[i][k], the value of NumSubLayersInLayerInOLS[i][j] can be updated, where the k-th layer in the bitstream is also included in the i-th OLS, and the k-th layer can reference the j-th layer.

[0283] Implementation 1

[0284] In an embodiment, the improvements 1, 2, 4 and 5 can be implemented according to the changed VPS syntax shown in Table 1. FIG. 15 In the following, the syntax elements changed from the above VPS syntax will be described. FIG. 15

[0285] ​max_tid_ref_present_flag[ i ] equal to 1 can specify that the syntax element max_tid_il_ref_pics_plus1[ i ][ j ] is present in the bitstream. max_tid_ref_present_flag[ i ] equal to 0 can specify that the syntax element max_tid_il_ref_pics_plus1[ i ][ j ] is not present in the bitstream.

[0286] max_tid_il_ref_pics_plus1[ i ][ j ] equal to 0 can specify that the j-th layer is not used as a reference layer for inter-layer prediction from non-IRAP pictures of the i-th layer. max_tid_il_ref_pics_plus1[ i ][ j ] greater than 0 can specify that for the decoding of pictures of the i-th layer, no picture from the j-th layer with Temporalld greater than max_tid_il_ref_pics_plus1[ i ][ j ] - 1 is used as ILRP. When the value of max_tid_il_ref_pics_plus1[ i ][ j ] is not obtained from the bitstream, the value of max_tid_il_ref_pics_plus1[ i ][ j ] can be inferred to be equal to 7.

[0287] In addition, using the syntax elements and variables changed according to the above description, the variables NumOutputLayersInOls[ i ], NumSubLayersInLayerInOLS[ i ][ j ], OutputLayerIdInOls[ i ][ j ], and LayerUsedAsOutputLayerFlag[ k ] can be determined similarly to the pseudo code of FIG. 16 to FIG. 17 Furthermore, the variable NumSubLayersInLayerInOLS[ i ][ j ] specifies the number of sub-layers of the j-th layer in the i-th OLS, as described above. However, according to the derivation of FIG. 16 and FIG. 17 the definitions can be changed and used so that the variable NumSubLayersInLayerInOLS[ i ][ j ] specifies the number of sub-layers of the j-th layer in the bitstream.

[0288] Similar to the method described above, for each reference layer of a layer, a credit can be used to notify the maximum number of sub-layers for the reference layer. In addition, the maximum number of sub-layers used by the reference layer j for layer i can be provided only when layer j is a direct reference layer of layer i. In addition, max_tid_il_ref_pics_plusl[i][j] can be used to derive the values of the array NumSubLayersInLayerInOLS[m][n]. In addition, when the OLS mode is 0 and 2, NumSubLayersInLayerInOLS can be derived for direct and indirect reference layers.

[0289] Implementation 2

[0290] In embodiments, a sub-bitstream can be derived based on the number of sub-layers derived according to Embodiment 1 according to the following sub-bitstream extraction process. The sub-bitstream extraction process can be a specified process by which NAL units in a bitstream that do not belong to a target set determined by a target OLS index and a target highest Temporalld are removed from the bitstream, where the output sub-bitstream contains NAL units in the bitstream that belong to the target set.

[0291] More specifically, the sub-bitstream extraction process can receive as inputs a variable inBitstream specifying an input bitstream, a target OLS index targetOlsIdx, and a target Temporalld tIdTarget. A variable OutBitstream specifying an output sub-bitstream can be derived as follows.

[0292] First, the sub-bitstream can be determined by the value of the input bitstream (S1810). For example, a variable outBitstream specifying the bitstream can be set to be the same as a variable inBitstream specifying the bitstream.

[0293] Next, NAL units having a predetermined Temporalld can be removed from the sub-bitstream (S1820). For example, all NAL units having a Temporalld greater than tIdTarget can be removed from outBitstream.

[0294] Next, NAL units having a predetermined nal_unit_type can be removed from the sub bitstream (S1830). For example, all NAL units having a nal_unit_type not equal to any one of VPS NUT (VPS NAL unit type), DCI NUT (Decoding Capability Information NAL unit type), and EOB NUT (End of bitstream NAL unit type) and having a nuh layer id not included in LayerldlnOls[targetOlsIdx] can be removed from outBitstream.

[0295] Next, predetermined NAL units can be removed from the sub bitstream based on NumSubLayersInLayerInOLS[][] (S1840). For example, all NAL units for which all of the following conditions are true can be removed from outBitstream. Here, nal_unit_type, nuh_layer_id, and Temporalld are values of the NAL unit to be determined for removal.

[0296] (Condition 1) nal_unit_type is not equal to IDR W RADL (Instantaneous Decoding Refresh with Random Access Decodable Leading), IDR N LP (Instantaneous Decoding Refresh without Leading Picture), or CRA NUT (Clean Random Access NAL unit type).

[0297] (Condition 2) nuh_layer_id is equal to LayerldlnOls[targetOlsIdx][j], where j is in the range of 0 to NumLayerslnOls[targetOlsIdx] - 1, inclusive.

[0298] (Condition 3) Temporalld has a value greater than or equal to NumSubLayersInLayerInOLS[targetOlsIdx][j].

[0299] Further, Condition 3 can be replaced with the following Condition 4 based on GeneralLayerIdx[nuh_layer_id] specifying the VPS layer index corresponding to the nuh layer id, as described in Equation 2 above.

[0300] (Condition 4) Temporalld has a value greater than or equal to NumSubLayersInLayerInOLS[targetOlsIdx][GeneralLayerIdx[nuh_layer_id]].

[0301] Next, predetermined SEI NAL units can be removed from the sub-bitstream (S1850). For example, all SEI NAL units containing scalable nesting SEI messages can be removed from outBitstream. Here, the scalable nesting messages can be messages for which nesting_ols_flag is equal to 1 and for which there is no value of i in the range of 0 to nesting_num_olss_minus1, inclusive, such that NestingOlsIdx[i] is equal to targetOlsIdx.

[0302] Here, nesting_ols_flag equal to 1 can specify that the scalable nesting SEI message applies to a particular OLS. nesting_ols_flag equal to 0 can specify that the scalable nesting SEI message applies to a particular layer. The variable NestingOlsIdx[i] can specify the OLS index of the i-th OLS for which the scalable nesting SEI message applies when sn_ols_flag is equal to 1. nesting_num_olss_minus1 plus 1 can specify the number of OLSs for which the scalable nesting SEI message applies.

[0303] Encoding and decoding method

[0304] Hereinafter, an image encoding method and an image decoding method performed by an image encoding apparatus and an image decoding apparatus according to embodiments will be described. FIG. 19 is a view illustrating a method of determining a number of sub-layers of a current layer so that an image encoding apparatus encodes an image and / or an image decoding apparatus decodes an image according to an embodiment.

[0305] An image decoding apparatus according to an embodiment includes a memory and a processor, and the decoding apparatus can perform decoding according to the embodiments described below by operation of the processor. An image encoding apparatus according to an embodiment includes a memory and a processor, and the encoding apparatus can perform encoding by operation of the processor in a manner corresponding to decoding of the decoding apparatus according to the embodiments described below. Hereinafter, for convenience of description, the operation of the decoding apparatus will be described, but the following description applies to the encoding apparatus as well.

[0306] The decoding apparatus according to an embodiment can determine at least one layer in an output layer set (OLS) (S1910). Next, the decoding apparatus can determine a number of sub-layers of the at least one layer (S1920).

[0307] For example, the number of sub-layers of a current layer in an OLS can be determined based on a maximum number of required sub-layers of a direct reference layer that can directly reference the current layer. Here, the number of sub-layers of the current layer in the OLS can be NumSubLayersInLayerInOLS[i][k] as described above. Whether a layer with index / can directly reference the current layer when the index of the current layer is k can be identified by the pseudo code shown in the following equation as FIG. 17

[0308] [Equation 3]

[0309] for (1 = k + 1; 1 <= vps max layers minus 1; 1++)

[0310] if (vps direct ref layer flag[1][k])

[0311] The process will be executed when the condition is true;

[0312] In addition, the number of sub-layers of a current layer is determined based on an OLS mode, and the OLS mode can be determined based on OLS mode index information (e.g., ols mode idc) obtained from a bitstream.

[0313] In addition, the number of sub-layers of a current layer can be determined based on whether the current layer is an output layer. For example, as shown in FIG. 17 FIG. 17 whether the current layer with index k is an output layer can be determined based on the pseudo code according to the following equation.

[0314] [Equation 4]

[0315] if (!ols output layer flag[i][k])

[0316] The process will be executed when the condition is true;

[0317] For example, based on the current layer being an output layer, the number of sub-layers of the current layer can be set to be equal to the maximum allowed number of sub-layers (e.g., vps max sub layers minus 1 + 1). Alternatively, based on the current layer not being an output layer, the number of sub-layers of the current layer can be set to be the maximum value of the maximum number of required sub-layers (e.g., maxSublayerNeeded) determined for at least one direct reference layer that can directly reference the current layer. Here, the layer index (e.g., index 1) of the direct reference layer is greater than the layer index (e.g., index k) of the current layer. FIG. 17 FIG. 17 ​​​and the network abstraction layer (NAL) unit identifier of the directly referenced layer should be greater than the NAL unit identifier of the current layer.

[0318] Further, the maximum number of required sub-layers determined for the directly referenced layer can be determined based on any one of the first value and the second value, the number of sub-layers of the directly referenced layer (e.g., FIG. 17 The first value can be determined based on NumSubLayersInLayerInOLS[i][l] of the directly referenced layer, the maximum temporal identifier among temporal identifiers specified for at least one picture of the current layer that can be referenced by the directly referenced layer (e.g., FIG. 17 The second value can be determined based on max_tid_il_ref_pics_plus1[l][k] of the directly referenced layer. For example, the maximum number of required sub-layers can be set to the smaller one of the first value and the second value. For example, the pseudo code of maxSublayerNeeded = min(NumSublayerlnLayerlnOLS[i][l], max_tid_il_ref_pics_plus1[l][k]) can be executed. Application implementation The pseudo code of maxSublayerNeeded = min(NumSublayerlnLayerlnOLS[i][l], max_tid_il_ref_pics_plus1[l][k]) can be executed.

[0319] In addition, the number of sub-layers can be determined for all non-output layers in the OLS, and the number of sub-layers can be sequentially determined in the order of non-output layers having lower layer indices, starting from the non-output layer having the highest layer index. For example, the pseudo code according to the following equation can be executed.

[0320] [Equation 5]

[0321] for (k = highestIncludedLayer - 1; k >= 0; k--)

[0322] if (layerIncludedInOlsFlag[i][k])

[0323] if (!ols_output_layer_flag[i][k])

[0324] a process for determining the number of sub-layers of a layer corresponding to index k;

[0325] In addition, a step of obtaining a sub-bitstream from a bitstream based on the number of sub-layers of the current layer can also be included. For example, the sub-bitstream can be obtained by removing predetermined network abstraction layer (NAL) units from the bitstream, and the predetermined NAL units can be determined based on a comparison between temporal identifiers of the predetermined NAL units and the number of sub-layers determined for a layer corresponding to the predetermined NAL units.

[0326] Here, the predetermined NAL unit is determined based on whether a value of a temporal identifier of the predetermined NAL unit (e.g., Temporalld in the description of S1840) is equal to or greater than a value of a number of sub-layers determined for a layer in which the predetermined NAL unit corresponds to a predetermined OLS (e.g., NumSubLayersInLayerInOLS[targetOlsIdx][GeneralLayerIdx[nuh_layer_id]] in the description of S1840), and the OLS can be any one of OLSs identified by information signaled by a video parameter set (VPS).

[0327] For example, in an embodiment, the image decoding method can include the steps of obtaining, from a bitstream, output layer set (OLS) mode index information indicating an OLS mode of a current video parameter set (VPS), sub-layer index information indicating a maximum available number of temporal sub-layers in a layer of the current VPS, and temporal identifier information indicating a maximum temporal identifier of a reference layer of a layer to be decoded; determining an OLS mode of the current VPS based on the OLS mode index information; determining at least one layer in at least one OLS of the current VPS; and determining a number of required sub-layers of the at least one layer.

[0328] Based on the current layer being an output layer in the current OLS, a number of sub-layers of the current layer can be set to the maximum available number of temporal sub-layers in the layer of the current VPS. Based on the current layer not being an output layer, a number of required sub-layers of the current layer can be set to a maximum value of a maximum number of sub-layers of at least one direct reference layer in the current OLS.

[0329] A layer index of each direct reference layer can be greater than a layer index of the current layer, and a network abstraction layer (NAL) unit identifier of each direct reference layer can be greater than a NAL unit identifier of the current layer.

[0330] A maximum number of required sub-layers of the direct reference layer can be set to a smaller one of a first number and a second number, the first number being a number of sub-layers of the direct reference layer, and the second number can be a maximum temporal identifier of the current layer to be referred to by the direct reference layer.

[0331] In addition, a number of sub-layers can be determined for all non-output layers in the at least one OLS, and the number of sub-layers can be sequentially determined starting from a non-output layer having a highest layer index to a lowest layer index.

[0332] In addition, the decoding method can further include the step of obtaining a sub-bitstream from the bitstream based on the number of sub-layers of the at least one layer. Here, the sub-bitstream can be obtained by removing a target network abstraction layer (NAL) from the bitstream, and the target NAL unit can have a temporal identifier (e.g., Temporalld in the description of S1840) having a value equal to or greater than the number of sub-layers (e.g., NumSubLayersInLayerInOLS[targetOlsIdx][GeneralLayerIdx[nuh_layer_id]] in the description of S1840) of a layer having an index (e.g., GeneralLayerIdx[nuh_layer_id] in the description of S1840) identified based on a layer identifier (e.g., nuh_layer_id in the description of S1840) of the target NAL unit.

[0333] In an embodiment, the image encoding method can include the steps of determining an output layer set (OLS) mode of a current video parameter set (VPS), a maximum available number of temporal sub-layers in a layer of the current VPS, a maximum temporal identifier of a reference layer of a layer to be decoded, and a number of sub-layers of at least one layer in at least one OLS of the current VPS, and generating a bitstream including OLS mode index information indicating the OLS mode of the current VPS, sub-layer index information indicating the maximum available number of temporal sub-layers in the layer of the current VPS, and temporal identifier information indicating the maximum temporal identifier of the reference layer of the layer to be decoded.

[0334] Based on the current layer being an output layer in the current OLS, the number of sub-layers of the current layer can be set to the maximum available number of temporal sub-layers in the layer of the current VPS.

[0335] Based on the current layer not being an output layer, the number of sub-layers of the current layer can be set to a maximum value of a maximum number of required sub-layers of a direct reference layer in the current OLS.

[0336] Here, the current layer can be one of the direct reference layers, a layer index of each of the direct reference layers can be greater than a layer index of the current layer, and a network abstraction layer (NAL) unit identifier of each of the direct reference layers can be greater than a NAL unit identifier of the current layer.

[0337] The maximum number of required sub-layers of the direct reference layer can be set to a smaller one of a first number and a second number, the first number can be a number of sub-layers of the direct reference layer, and the second number can be a maximum temporal identifier of the current layer to be referred to by the direct reference layer.

[0338] Also, in the encoding method, a sub-bitstream can be obtained by removing target network abstraction layer (NAL) units from a bitstream based on the number of sub-layers of the at least one layer. Here, the target NAL units can include a temporal identifier (e.g., Temporalld in the description of S1840) having a value equal to or greater than the number of sub-layers of the layer identified based on the layer identifier of the target NAL unit (e.g., NumSubLayersInLayerInOLS[targetOlsIdx][GeneralLayerIdx[nuh_layer_id]] in the description of S1840).

[0339] FIG. 20

[0340] Although the above-described exemplary methods of the disclosure are represented as a series of operations for clarity of description, the order of performing the steps is not intended to be limited, and the steps can be performed simultaneously or in a different order if necessary. To implement the methods according to the disclosure, the described steps can further include other steps, can include the remaining steps other than some steps, or can include other additional steps other than some steps.

[0341] In the disclosure, an image encoding apparatus or an image decoding apparatus that performs a predetermined operation (step) can perform an operation (step) that confirms an execution condition or a situation corresponding to the operation (step). For example, if it is described that a predetermined operation is performed when a predetermined condition is satisfied, the image encoding apparatus or the image decoding apparatus can perform the predetermined operation after determining whether the predetermined condition is satisfied.

[0342] The various embodiments of the disclosure are not a list of all possible combinations and are intended to describe representative aspects of the disclosure, and matters described in the various embodiments can be applied independently or in combination of two or more.

[0343] The various embodiments of the disclosure can be implemented in hardware, firmware, software, or a combination thereof. In the case of implementing the disclosure through hardware, the disclosure can be implemented by an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a general purpose processor, a controller, a microcontroller, a microprocessor, etc.

[0344] In addition, the image decoding apparatus and the image encoding apparatus to which embodiments of the disclosure are applied can be included in a multimedia broadcast transmitting and receiving device, a mobile communication terminal, a home theater video device, a digital theater video device, a surveillance camera, a video chat device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camcorder, a video on demand (VoD) service providing device, an over the top video (OTT video) device, an Internet streaming service providing device, a three-dimensional (3D) video device, a video phone video device, a medical video device, etc., and can be used to process a video signal or a data signal. For example, the OTT video device can include a game console, a Blu-ray player, an Internet access television, a home theater system, a smart phone, a tablet PC, a digital video recorder (DVR), etc.

[0345] FIG. 20 is a view showing a content streaming system to which embodiments of the disclosure can be applied.

[0346] As ​ indicated, the content streaming system to which embodiments of the disclosure are applied can mainly include an encoding server, a streaming server, a web server, a media storage device, a user device, and a multimedia input device.

[0347] The encoding server compresses content input from a multimedia input device such as a smart phone, a camera, a camcorder, etc., into digital data to generate a bitstream and transmit the bitstream to the streaming server. As another example, when a multimedia input device such as a smart phone, a camera, a camcorder, etc., directly generates a bitstream, the encoding server can be omitted.

[0348] The bitstream can be generated by an image encoding method or an image encoding apparatus to which embodiments of the disclosure are applied, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.

[0349] The streaming server transmits multimedia data to a user device based on a request of the user through the web server, and the web server serves as a medium to inform a service to the user. When the user requests a desired service to the web server, the web server can deliver it to the streaming server, and the streaming server can transmit multimedia data to the user. In this case, the content streaming system can include a separate control server. In this case, the control server is used to control commands / responses between devices in the content streaming system.

[0350] The streaming server can receive content from the media storage device and / or the encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store a bitstream for a predetermined time.

[0351] Examples of the user device can include a mobile phone, a smart phone, a laptop computer, a digital broadcasting terminal, a personal digital assistant (PDA), a portable multimedia player (PMP), a navigation device, a slate PC, a tablet PC, an ultrabook, a wearable device (e.g., a smart watch, smart glasses, a head-mounted display), a digital TV, a desktop computer, a digital signage, etc.

[0352] The respective servers in the content streaming system can operate as distributed servers, in which case the data received from the respective servers can be distributed.

[0353] The scope of the present disclosure includes software or machine-executable commands (e.g., operating systems, applications, firmware, programs, etc.) for enabling operations of the methods according to various embodiments to be executed on a device or computer, a non-transitory computer-readable medium having such software or commands stored thereon and executable on a device or computer.

[0354] Industrial applicability

[0355] Embodiments of the present disclosure can be used to encode or decode an image.

Claims

1. An image decoding method performed by an image decoding apparatus, the image decoding method comprising the steps of: determining at least one layer in an output layer set (OLS); and determining a number of sub-layers for the at least one layer, wherein the number of sub-layers for a current layer in the OLS is determined based on a minimum value between a first value and a second value, wherein the first value is determined based on a number of sub-layers for a first layer, and wherein the second value is determined based on a maximum temporal identifier among temporal identifiers specified for at least one picture of the current layer to be referenced by the first layer.

2. The image decoding method according to claim 1, wherein the number of sub-layers for the current layer is determined based on an OLS mode, and wherein the OLS mode is determined based on OLS mode index information obtained from a bitstream.

3. The image decoding method according to claim 1, wherein, a layer index of the first layer is larger than a layer index of the current layer, and wherein a network abstraction layer (NAL) unit identifier of the first layer is larger than a NAL unit identifier of the current layer.

4. The image decoding method according to claim 1, further comprising the steps of: obtaining a sub-bitstream from a bitstream based on the number of sub-layers for the current layer.

5. The image decoding method according to claim 4, wherein the sub-bitstream is obtained by removing predetermined network abstraction layer (NAL) units from the bitstream, and wherein the predetermined NAL units are determined based on a comparison between a temporal identifier of the predetermined NAL units and the number of sub-layers determined for a second layer corresponding to the predetermined NAL units.

6. The image decoding method according to claim 5, wherein the predetermined NAL units are determined based on whether a value of the temporal identifier of the predetermined NAL units is equal to or larger than a value of the number of sub-layers determined for the second layer among layers of a predetermined OLS corresponding to the predetermined NAL units, and wherein the OLS is identified based on information signaled by a video parameter set (VPS).

7. An image encoding method performed by an image encoding apparatus, the image encoding method comprising the steps of: determining at least one layer in an output layer set (OLS); and determining a number of sub-layers for the at least one layer, wherein the number of sub-layers for a current layer in the OLS is determined based on a minimum value between a first value and a second value, wherein the first value is determined based on a number of sub-layers for a first layer, and wherein the second value is determined based on a maximum temporal identifier among temporal identifiers specified for at least one picture of the current layer to be referenced by the first layer.

8. A method for transmitting a bitstream generated by an image encoding method, the image encoding method comprising the steps of: determining at least one layer in an output layer set (OLS); and determining a number of sub-layers for the at least one layer, wherein the number of sub-layers for a current layer in the OLS is determined based on a minimum value between a first value and a second value, ​ ​ wherein the first value is determined based on a number of sub-layers for the first layer, and wherein the second value is determined based on a maximum temporal identifier among temporal identifiers specified for at least one picture of the current layer to be referenced by the first layer.