Image encoding / decoding method and apparatus for determining sub-layer based on inter-layer reference and method for transmitting bitstream

The image encoding/decoding method improves efficiency by determining sub-layers based on inter-layer references, addressing the increased costs associated with high-resolution images through optimized bitstream transmission.

JP2026020407APending Publication Date: 2026-02-06LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025211375
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-04-08
Filing Date
2025-12-01
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

The increasing demand for high-resolution, high-quality images leads to a significant increase in transmission and storage costs due to the higher amount of information required, necessitating highly efficient image compression techniques.

Method used

An image encoding/decoding method that determines a sub-layer based on whether inter-layer reference is used, along with a bitstream transmission method, to improve encoding/decoding efficiency.

Benefits of technology

The method enhances encoding/decoding efficiency by optimizing the use of inter-layer references, allowing for more efficient transmission and storage of high-resolution, high-quality images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026020407000001_ABST
    Figure 2026020407000001_ABST
Patent Text Reader

Abstract

Provided are an image encoding / decoding method and apparatus for determining a sub-layer based on inter-layer reference.SOLUTION: An image decoding method performed by an image decoding apparatus includes determining whether there is inter-layer direct reference, and determining the number of sub-layers of a current layer based on whether there is inter-layer direct reference. Based on the current layer not being the output layer, the number of sub-layers of the current layer may be determined based on the maximum identifier information obtained based on the inter-layer direct reference. The maximum identifier information indicates that a picture having a temporal identifier greater than a value identified by the maximum identifier information among a plurality of pictures of a first layer is not used as an inter-layer reference picture to decode a target picture of a second layer.SELECTED DRAWING: Figure 19
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an image encoding / decoding method and apparatus, and more particularly to an image encoding / decoding method and apparatus that determine a sublayer based on whether or not there is inter-layer reference, and a method for transmitting a bitstream generated by the image encoding method / apparatus of the present disclosure. [Background technology]

[0002] Recently, demand for high-resolution, high-quality images, such as HD (High Definition) images and UHD (Ultra High Definition) images, has been increasing in various fields. As image data becomes higher in resolution and quality, the amount of information or bits to be transmitted increases relatively compared to conventional image data. The increase in the amount of information or bits to be transmitted results in an increase in transmission costs and storage costs.

[0003] This requires highly efficient image compression techniques for effectively transmitting, storing, and reproducing high-resolution, high-quality image information. Summary of the Invention [Problem to be solved by the invention]

[0004] An object of the present disclosure is to provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.

[0005] Another object of the present disclosure is to provide an image encoding / decoding method and apparatus that improves encoding / decoding efficiency by determining a sub-layer based on whether inter-layer reference is used.

[0006] Another object of the present disclosure is to provide a method for transmitting a bitstream generated by the image encoding method or apparatus according to the present disclosure.

[0007] Another object of the present disclosure is to provide a recording medium storing a bitstream generated by the image encoding method or apparatus according to the present disclosure.

[0008] Another object of the present disclosure is to provide a recording medium storing a bitstream that is received by an image decoding device according to the present disclosure, decoded, and used to restore an image. For example, the recording medium may store a bitstream that causes a decoding device according to the present disclosure to perform an image decoding method according to the present disclosure.

[0009] The technical problems to be solved by the present disclosure are not limited to the above-mentioned technical problems, and other technical problems not mentioned above will be clearly understood by a person having ordinary skill in the technical field to which the present disclosure pertains from the following description. [Means for solving the problem]

[0010] An image decoding method performed by an image decoding device according to one aspect of the present disclosure may include a step of determining whether inter-layer direct reference is performed, and a step of determining the number of sub-layers of a current layer based on whether inter-layer direct reference is performed.

[0011] Furthermore, an image decoding device according to one aspect of the present disclosure is an image decoding device including a memory and at least one processor, wherein the at least one processor can determine whether inter-layer direct reference is performed and determine the number of sub-layers of a current layer based on whether inter-layer direct reference is performed.

[0012] In addition, an image encoding method performed by an image encoding device according to one aspect of the present disclosure can determine whether inter-layer direct reference is performed, and determine the number of sub-layers of the current layer based on whether inter-layer direct reference is performed.

[0013] Furthermore, a transmission method according to another aspect of the present disclosure can transmit a bitstream generated by the image encoding device or image encoding method of the present disclosure.

[0014] Furthermore, a computer-readable recording medium according to another aspect of the present disclosure can store a bitstream generated by the image encoding method or image encoding device of the present disclosure.

[0015] Furthermore, a computer-readable recording medium according to another aspect of the present disclosure can store a bitstream that causes a decoding device to perform the image decoding method of the present disclosure.

[0016] Furthermore, the features described above in the brief summary of the present disclosure are merely exemplary embodiments of the detailed description of the present disclosure that follows, and are not intended to limit the scope of the present disclosure. [Effects of the Invention]

[0017] According to the present disclosure, an image encoding / decoding method and apparatus with improved encoding / decoding efficiency can be provided.

[0018] Furthermore, according to the present disclosure, an image encoding / decoding method and apparatus can be provided that can improve the efficiency of encoding / decoding by determining a sub-layer based on whether or not inter-layer reference is used.

[0019] The present disclosure also provides a method for transmitting a bitstream generated by the image encoding method or apparatus according to the present disclosure.

[0020] Furthermore, according to the present disclosure, a recording medium storing a bitstream generated by the image encoding method or apparatus according to the present disclosure can be provided.

[0021] Furthermore, according to the present disclosure, it is possible to provide a recording medium that stores a bitstream that is received by the image decoding device according to the present disclosure, decoded, and used to restore an image.

[0022] The effects obtained by the present disclosure are not limited to the effects described above, and other effects not described above will be clearly understood by those having ordinary skill in the art to which the present disclosure pertains from the following description. [Brief explanation of the drawings]

[0023] [Figure 1] 1 is a diagram illustrating a video coding system to which embodiments of the present disclosure can be applied; [Figure 2] 1 is a diagram schematically illustrating an image encoding device to which an embodiment of the present disclosure can be applied. [Figure 3] FIG. 1 is a diagram schematically illustrating an image decoding device to which an embodiment of the present disclosure can be applied. [Figure 4] FIG. 2 illustrates an example of a picture decoding and encoding procedure according to an embodiment. [Figure 5] FIG. 2 illustrates an example of a picture decoding and encoding procedure according to an embodiment. [Figure 6] FIG. 2 illustrates a hierarchical structure for coded images according to one embodiment. [Figure 7] FIG. 1 is a diagram illustrating multi-layer-based encoding and decoding. [Figure 8] FIG. 1 is a diagram illustrating multi-layer-based encoding and decoding. [Figure 9] FIG. 2 is a diagram illustrating a syntax structure of a VPS according to one embodiment of the present disclosure. [Figure 10] FIG. 10 illustrates pseudocode for deriving VPS-related variables according to one embodiment of the present disclosure. [Figure 11] FIG. 10 illustrates pseudocode for deriving VPS-related variables according to one embodiment of the present disclosure. [Figure 12] FIG. 10 illustrates pseudocode for deriving VPS-related variables according to one embodiment of the present disclosure. [Figure 13] FIG. 10 illustrates pseudocode for deriving VPS-related variables according to one embodiment of the present disclosure. [Figure 14] FIG. 10 illustrates pseudocode for deriving VPS-related variables according to one embodiment of the present disclosure. [Figure 15] FIG. 10 is a diagram illustrating a syntax structure of a VPS according to another embodiment of the present disclosure. [Figure 16] FIG. 10 illustrates pseudocode for deriving VPS-related variables according to another embodiment of the present disclosure. [Figure 17] FIG. 10 is a diagram illustrating a syntax structure of a VPS according to another embodiment of the present disclosure. [Figure 18] FIG. 10 illustrates pseudocode for deriving VPS-related variables according to another embodiment of the present disclosure. [Figure 19] FIG. 10 illustrates an encoding and / or decoding method according to another embodiment of the present disclosure. [Figure 20] FIG. 1 illustrates a content streaming system to which an embodiment of the present disclosure can be applied. DETAILED DESCRIPTION OF THE INVENTION

[0024] The present disclosure will be described in detail below with reference to the accompanying drawings, so that those skilled in the art can easily implement the present disclosure. However, the present disclosure may be embodied in various different forms and is not limited to the embodiments described herein.

[0025] In describing the embodiments of the present disclosure, if it is determined that a detailed description of a known configuration or function may obscure the gist of the present disclosure, the detailed description thereof will be omitted. In addition, in the drawings, parts that are not related to the description of the present disclosure will be omitted, and similar parts will be designated by similar reference numerals.

[0026] In this disclosure, when a component is referred to as being "coupled," "coupled," or "connected" to another component, this includes not only a direct connection, but also an indirect connection where another component exists between them. Furthermore, when a component is referred to as "including" or "having" another component, this does not mean that the other component is excluded, but that the component can further include the other component, unless otherwise specified.

[0027] In this disclosure, terms such as "first" and "second" are used only to distinguish one component from another component, and do not limit the order or importance of the components unless otherwise specified. Therefore, within the scope of this disclosure, a first component in one embodiment may be called a second component in another embodiment, and similarly, a second component in one embodiment may be called a first component in another embodiment.

[0028] In this disclosure, components that are distinguished from one another are used to clearly describe the characteristics of each component and do not necessarily mean that the components are separate. In other words, multiple components may be integrated into a single hardware or software unit, or a single component may be distributed into multiple hardware or software units. Therefore, even if not otherwise specified, such integrated or distributed embodiments are also included within the scope of this disclosure.

[0029] In this disclosure, the components described in various embodiments are not necessarily essential components, and some may be optional components. Therefore, an embodiment consisting of a subset of the components described in one embodiment is also within the scope of this disclosure. Furthermore, an embodiment including other components in addition to the components described in various embodiments is also within the scope of this disclosure.

[0030] The present disclosure relates to image encoding and decoding, and terms used in this disclosure may have their ordinary meaning in the technical field to which the present disclosure belongs unless they are newly defined in this disclosure.

[0031] The methods / embodiments disclosed in this disclosure may be applied to methods disclosed in the versatile video coding (VVC) standard, the essential video coding (EVC) standard, the AOV1 (Audio Media Video 1) standard, the second generation audio video coding (AVS2) standard, or next-generation video / image coding standards (e.g., H.267 or H.268).

[0032] This disclosure presents various embodiments related to video / image coding, and unless otherwise stated, the embodiments in this disclosure may be performed in combination with each other.

[0033] In this disclosure, "video" may refer to a collection of a series of images over time. A "picture" generally refers to a unit representing any one image in a specific time period, and a slice / tile is a coding unit that constitutes a part of a picture during encoding. A slice / tile may include one or more coding tree units (CTUs). The CUT may be divided into one or more CUs.

[0034] A picture may consist of one or more slices / tiles. A tile is a rectangular region located in a specific tile row and a specific tile column within a picture, and may consist of multiple CTUs. A tile column may be defined as a rectangular region of CTUs, have the same height as the picture, and may have a width specified by a syntax element signaled from a bitstream portion such as a picture parameter set. A tile row may be defined as a rectangular region of CTUs, have the same width as the picture, and may have a height specified by a syntax element signaled from a bitstream portion such as a picture parameter set.

[0035] A tile scan is a predetermined sequential ordering method of CTUs that divides a picture. Here, CTUs may be sequentially ordered within a tile according to a CTU raster scan, and tiles within a picture may be sequentially ordered according to the raster scan order of tiles in the picture. A slice can contain an integer number of complete tiles or an integer number of consecutive complete CTU rows within a tile of a picture. A slice can be contained exclusively in one single NAL unit.

[0036] A picture can be partitioned into two or more sub-pictures, which can be rectangular regions of one or more slices within the picture.

[0037] A picture can be made up of one or more tile groups. A tile group can contain one or more tiles. A brick can represent a rectangular area of ​​a CTU row within a tile in a picture. A tile can contain one or more bricks. A brick can represent a rectangular area of ​​a CTU row within a tile. A tile can be divided into multiple bricks, and each brick can contain one or more CTU rows belonging to the tile. A tile that is not divided into multiple bricks can also be treated as a brick.

[0038] In this disclosure, "pixel" or "pel" may refer to the smallest unit constituting one picture (or image). Also, "sample" may be used as a term corresponding to pixel. A sample may generally indicate a pixel or a pixel value, may indicate only a pixel / pixel value of a luma component, or may indicate only a pixel / pixel value of a chroma component.

[0039] In this disclosure, the term "unit" may refer to a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to that region. One unit may include one luma block and two chroma (e.g., Cb, Cr) blocks. The term "unit" may be used interchangeably with terms such as "sample array," "block," or "area," depending on the situation. In general, an M×N block may include a set (or array) of samples or transform coefficients consisting of M columns and N rows.

[0040] In the present disclosure, a "current block" may refer to any one of a "current coding block," a "current coding unit," a "block to be coded," a "block to be decoded," or a "block to be processed." When prediction is performed, a "current block" may refer to a "current predicted block" or a "block to be predicted." When transformation (inverse transformation) / quantization (inverse quantization) is performed, a "current block" may refer to a "current transformed block" or a "block to be transformed." When filtering is performed, a "current block" may refer to a "block to be filtered."

[0041] In this disclosure, "current block" may mean "luma block of the current block" unless explicitly stated as a chroma block. "Chroma block of the current block" may be expressed explicitly as "chroma block" or "current chroma block" including the explicit statement of a chroma block.

[0042] In the present disclosure, " / " and "," can be interpreted as "and / or." For example, "A / B" and "A, B" can be interpreted as "A and / or B." Also, "A / B / C" and "A, B, C" can mean "at least one of A, B, and / or C."

[0043] In this disclosure, "or" can be interpreted as "and / or." For example, "A or B" can mean 1) only "A," 2) only "B," or 3) "A and B." Alternatively, in this disclosure, "or" can mean "additionally or alternatively."

[0044] In the present disclosure, "at least one of A and B" can mean "A only," "B only," or "both A and B." Also, in the present disclosure, the expressions "at least one of A or B" and "at least one of A and / or B" can be interpreted as being the same as "at least one of A and B."

[0045] Additionally, in this disclosure, "at least one of A, B, and C" can mean "A only," "B only," "C only," or "any combination of A, B, and C." Additionally, "at least one of A, B, or C" or "at least one of A, B, and / or C" can mean "at least one of A, B, and C."

[0046] Furthermore, parentheses used in the present disclosure may mean "for example." Specifically, when "prediction (intra prediction)" is displayed, "intra prediction" may be proposed as an example of "prediction." In other words, "prediction" in the present disclosure is not limited to "intra prediction," and "intra prediction" may be proposed as an example of "prediction." Furthermore, when "prediction (i.e., intra prediction)" is displayed, "intra prediction" may be proposed as an example of "prediction."

[0047] In the present disclosure, technical features described separately in one drawing may be realized separately or simultaneously.

[0048] Video Coding System Overview

[0049] FIG. 1 illustrates a video coding system according to this disclosure.

[0050] A video coding system according to one embodiment may include a source device 10 and a receiving device 20. The source device 10 may transmit encoded video and / or image information or data to the receiving device 20 in file or streaming format via a digital storage medium or a network.

[0051] According to one embodiment, the source device 10 may include a video source generation unit 11, an encoding unit 12, and a transmission unit 13. According to one embodiment, the receiving device 20 may include a receiving unit 21, a decoding unit 22, and a rendering unit 23. The encoding unit 12 may be referred to as a video / image encoding unit, and the decoding unit 22 may be referred to as a video / image decoding unit. The transmission unit 13 may be included in the encoding unit 12. The receiving unit 21 may be included in the decoding unit 22. The rendering unit 23 may include a display unit, which may be configured as a separate device or an external component.

[0052] The video source generation unit 11 can acquire video / images through a video / image capture, synthesis, or generation process. The video source generation unit 11 can include a video / image capture device and / or a video / image generation device. The video / image capture device can include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. The video / image generation device can include, for example, a computer, a tablet, a smartphone, etc., and can (electronically) generate video / images. For example, virtual video / images can be generated via a computer, etc., in which case the video / image capture process can be replaced with a process in which related data is generated.

[0053] The encoding device 12 can encode the input video / image. The encoding device 12 can perform a series of steps such as prediction, transformation, quantization, etc. for compression and encoding efficiency. The encoding device 12 can output the encoded data (encoded video / image information) in a bitstream format.

[0054] The transmitting unit 13 can transmit the encoded video / image information or data output in a bitstream format to the receiving unit 21 of the receiving device 20 in a file or streaming format via a digital storage medium or a network. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray®, HDD, and SSD. The transmitting unit 13 can include elements for generating a media file in a predetermined file format and elements for transmitting via a broadcasting / communication network. The receiving unit 21 can extract / receive the bitstream from the storage medium or network and transmit it to the decoding device 22.

[0055] The decoder 22 can decode the video / image by performing a series of steps such as inverse quantization, inverse transformation, and prediction corresponding to the operations of the encoder 12 .

[0056] The rendering unit 23 can render the decoded video / images, and the rendered video / images can be displayed via the display unit.

[0057] Overview of the image encoding device

[0058] FIG. 2 is a diagram schematically illustrating an image encoding device to which an embodiment of the present disclosure can be applied.

[0059] 2, the image encoding device 100 may include an image division unit 110, a subtraction unit 115, a transform unit 120, a quantization unit 130, an inverse quantization unit 140, an inverse transform unit 150, an addition unit 155, a filtering unit 160, a memory 170, an inter prediction unit 180, an intra prediction unit 185, and an entropy encoding unit 190. The inter prediction unit 180 and the intra prediction unit 185 may be collectively referred to as a "prediction unit." The transform unit 120, the quantization unit 130, the inverse quantization unit 140, and the inverse transform unit 150 may be included in a residual processing unit. The residual processing unit may further include a subtraction unit 115.

[0060] Depending on the embodiment, all or at least some of the components constituting the image encoding device 100 may be realized by a single hardware component (e.g., an encoder or a processor). Also, the memory 170 may include a decoded picture buffer (DPB) and may be realized by a digital storage medium.

[0061] The image division unit 110 may divide an input image (or picture, frame) input to the image encoding device 100 into one or more processing units. As an example, the processing units may be called coding units (CUs). The coding units may be obtained by recursively dividing a coding tree unit (CTU) or a largest coding unit (LCU) using a QT / BT / TT (quad-tree / binary-tree / ternary-tree) structure. For example, one coding unit may be divided into multiple coding units at deeper depths based on a quad-tree structure, a binary-tree structure, and / or a ternary-tree structure. To divide the coding units, the quad-tree structure may be applied first, and then the binary-tree structure and / or the ternary-tree structure may be applied later. The coding procedure according to the present disclosure may be performed based on the final coding unit that is not further divided. The maximum coding unit may be used as the final coding unit, or a lower-depth coding unit obtained by dividing the maximum coding unit may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and / or reconstruction, which will be described later. As another example, a processing unit of the coding procedure may be a prediction unit (PU) or a transform unit (TU). The prediction unit and the transform unit may be divided or partitioned from the final coding unit, respectively. The prediction unit may be a unit of sample prediction, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.

[0062] The prediction unit (inter prediction unit 180 or intra prediction unit 185) may perform prediction on a current block (current block) to generate a predicted block including prediction samples for the current block. The prediction unit may determine whether intra prediction or inter prediction is applied to the current block or CU. The prediction unit may generate various information related to the prediction of the current block and transmit it to the entropy coding unit 190. The prediction information may be coded by the entropy coding unit 190 and output in a bitstream format.

[0063] The intra prediction unit 185 may predict the current block by referring to samples in the current picture. The referenced samples may be located in the neighborhood of the current block or may be located far away from the current block according to the intra prediction mode and / or intra prediction technique. The intra prediction modes may include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes may include, for example, a DC mode and a planar mode. The directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the granularity of the prediction direction. However, this is merely an example, and more or less directional prediction modes may be used depending on the settings. The intra prediction unit 185 may also determine the prediction mode to be applied to the current block using the prediction modes applied to neighboring blocks.

[0064] The inter prediction unit 180 may derive a predicted block for a current block based on a reference block (reference sample array) identified by a motion vector on a reference picture. To reduce the amount of motion information transmitted in inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation between the motion information of neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on the inter prediction direction (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block may be called a collocated reference block, a collocated CU (colCU), etc. The reference picture including the temporal neighboring block may be called a collocated picture (colPic). For example, the inter predictor 180 may construct a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to derive a motion vector and / or a reference picture index for the current block. Inter prediction may be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the inter predictor 180 may use motion information of neighboring blocks as motion information for the current block. In the case of skip mode, unlike in merge mode, a residual signal may not be transmitted.In the case of a motion vector prediction (MVP) mode, the motion vector of a neighboring block is used as a motion vector predictor, and the motion vector of the current block can be signaled by encoding a motion vector difference and an indicator for the motion vector predictor. The motion vector difference may mean the difference between the motion vector of the current block and the motion vector predictor.

[0065] The predictor may generate a prediction signal based on various prediction methods and / or prediction techniques, which will be described later. For example, the predictor may apply intra prediction or inter prediction to predict the current block, or may simultaneously apply intra prediction and inter prediction. A prediction method that simultaneously applies intra prediction and inter prediction to predict the current block may be referred to as combined inter and intra prediction (CIIP). The predictor may also perform intra block copy (IBC) to predict the current block. Intra block copy can be used for content image / video coding, such as screen content coding (SCC), for games. IBC is a method of predicting a current block using an already reconstructed reference block in a current picture that is located a predetermined distance away from the current block. When IBC is applied, the position of the reference block in the current picture may be coded as a vector (block vector) corresponding to the predetermined distance. IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that a reference block is derived within the current picture. That is, the IBC may use at least one of the inter prediction techniques described in this disclosure.

[0066] The prediction signal generated by the prediction unit may be used to generate a restored signal or a residual signal. The subtraction unit 115 may subtract the prediction signal (predicted block, predicted sample array) output from the prediction unit from the input image signal (original block, original sample array) to generate a residual signal (residual signal, residual block, residual sample array). The generated residual signal may be transmitted to the conversion unit 120.

[0067] The transform unit 120 may generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loeve transform (KLT), a graph-based transform (GBT), or a conditionally non-linear transform (CNT). Here, the GBT refers to a transform obtained from a graph representing inter-pixel relationship information. The CNT refers to a transform obtained based on a predicted signal generated using all previously reconstructed pixels. The transform process may be applied to pixel blocks having the same square size or to non-square blocks of variable size.

[0068] The quantization unit 130 may quantize the transform coefficients and transmit the quantized transform coefficients to the entropy coding unit 190. The entropy coding unit 190 may encode the quantized signal (information about the quantized transform coefficients) and output the encoded signal in a bitstream format. The information about the quantized transform coefficients may be referred to as residual information. The quantization unit 130 may rearrange the quantized transform coefficients in a block format into a one-dimensional vector format based on a coefficient scan order, and may generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector format.

[0069] The entropy coding unit 190 may perform various coding methods, such as exponential Golomb, context-adaptive variable length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy coding unit 190 may also code information necessary for video / image restoration (e.g., values ​​of syntax elements) together with or separately from the quantized transform coefficients. The coded information (e.g., coded video / image information) may be transmitted or stored in a bitstream format in network abstraction layer (NAL) units. The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / image information may also include general constraint information. The signaling information, transmitted information and / or syntax elements mentioned in this disclosure may be encoded through the above-described encoding procedure and included in the bitstream.

[0070] The bitstream may be transmitted via a network or stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as a USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitting unit (not shown) that transmits and / or a storing unit (not shown) that stores the signal output from the entropy encoding unit 190 may be provided as an internal / external element of the image encoding device 100, or the transmitting unit may be provided as a component of the entropy encoding unit 190.

[0071] The quantized transform coefficients output from the quantization unit 130 can be used to generate a residual signal. For example, the residual signal (residual block or residual sample) can be reconstructed by applying inverse quantization and inverse transform to the quantized transform coefficients via the inverse quantization unit 140 and the inverse transform unit 150.

[0072] The adder 155 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter prediction unit 180 or the intra prediction unit 185. When there is no residual for the current block to be processed, such as when a skip mode is applied, the predicted block may be used as the reconstructed block. The adder 155 may be referred to as a reconstruction unit or a reconstructed block generation unit. The generated reconstructed signal may be used for intra prediction of the next current block to be processed in the current picture, and may also be used for inter prediction of the next picture after filtering, as will be described later.

[0073] The filtering unit 160 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 160 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture and store the modified reconstructed picture in the memory 170, specifically, in the DPB of the memory 170. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, etc. The filtering unit 160 may generate various information related to filtering and transmit it to the entropy coding unit 190, as will be described later in connection with each filtering method. The filtering information may be coded by the entropy coding unit 190 and output in a bitstream format.

[0074] The modified reconstructed picture transmitted to the memory 170 can be used as a reference picture in the inter prediction unit 180. When inter prediction is applied through this, the image encoding device 100 can avoid a prediction mismatch between the image encoding device 100 and the image decoding device, and can also improve encoding efficiency.

[0075] The DPB in the memory 170 may store modified reconstructed pictures for use as reference pictures in the inter predictor 180. The memory 170 may store motion information of blocks from which motion information in the current picture is derived (or coded) and / or motion information of already reconstructed intra-picture blocks. The stored motion information may be transmitted to the inter predictor 180 to be used as motion information of spatially surrounding blocks or temporally surrounding blocks. The memory 170 may store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra predictor 185.

[0076] Overview of the image decoding device

[0077] FIG. 3 is a diagram schematically illustrating an image decoding device to which an embodiment of the present disclosure can be applied.

[0078] 3, the image decoding apparatus 200 may include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an adder 235, a filtering unit 240, a memory 250, an inter prediction unit 260, and an intra prediction unit 265. The inter prediction unit 260 and the intra prediction unit 265 may be collectively referred to as a "prediction unit." The inverse quantization unit 220 and the inverse transform unit 230 may be included in a residual processing unit.

[0079] Depending on the embodiment, all or at least some of the components constituting the image decoding device 200 may be realized by a single hardware component (e.g., a decoder or a processor). Also, the memory 170 may include a DPB and may be realized by a digital storage medium.

[0080] The image decoding device 200, which receives a bitstream including video / image information, can reconstruct an image by performing a process corresponding to the process performed by the image encoding device 100 of FIG. 2. For example, the image decoding device 200 can perform decoding using a processing unit applied in the image encoding device. Therefore, the decoding processing unit can be, for example, a coding unit. The coding unit can be obtained by dividing a coding tree unit or a maximum coding unit. The reconstructed image signal decoded and output by the image decoding device 200 can be reproduced by a reproduction device (not shown).

[0081] The image decoding apparatus 200 may receive a signal output from the image encoding apparatus of FIG. 2 in a bitstream format. The received signal may be decoded via an entropy decoding unit 210. For example, the entropy decoding unit 210 may parse the bitstream to derive information (e.g., video / image information) necessary for image reconstruction (or picture reconstruction). The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / image information may also include general constraint information. The image decoding apparatus may further use the information on the parameter sets and / or the general constraint information to decode an image. The signaling information, received information, and / or syntax elements referred to in the present disclosure may be obtained from the bitstream by being decoded via the decoding procedure. For example, the entropy decoding unit 210 may decode information in a bitstream based on a coding method such as Exponential-Golomb coding, CAVLC, or CABAC, and output values ​​of syntax elements required for image restoration and quantized values ​​of transform coefficients related to residuals. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element from the bitstream, determines a context model using information on the syntax element to be decoded and decoded information on neighboring blocks and the block to be decoded, or information on symbols / bins decoded in a previous step, predicts the occurrence probability of the bins based on the determined context model, and performs arithmetic decoding of the bins to generate symbols corresponding to the values ​​of each syntax element. After determining the context model, the CABAC entropy decoding method may update the context model using information on the decoded symbol / bin for the context model of the next symbol / bin.Among the information decoded by the entropy decoding unit 210, information related to prediction is provided to the prediction units (inter prediction unit 260 and intra prediction unit 265), and residual values ​​entropy decoded by the entropy decoding unit 210, i.e., quantized transform coefficients and related parameter information, may be input to the inverse quantization unit 220. Also, among the information decoded by the entropy decoding unit 210, information related to filtering may be provided to the filtering unit 240. Meanwhile, a receiving unit (not shown) for receiving a signal output from the image encoding device may be further provided as an internal / external element of the image decoding device 200, or the receiving unit may be provided as a component of the entropy decoding unit 210.

[0082] Meanwhile, the image decoding apparatus according to the present disclosure may be referred to as a video / image / picture decoding apparatus. The image decoding apparatus may include an information decoder (video / image / picture information decoder) and / or a sample decoder (video / image / picture sample decoder). The information decoder may include an entropy decoding unit 210, and the sample decoder may include at least one of an inverse quantization unit 220, an inverse transform unit 230, an adder 235, a filtering unit 240, a memory 250, an inter prediction unit 260, and an intra prediction unit 265.

[0083] The inverse quantization unit 220 may inverse quantize the quantized transform coefficients and output the transform coefficients. The inverse quantization unit 220 may rearrange the quantized transform coefficients in a two-dimensional block format. In this case, the rearrangement may be performed based on the coefficient scanning order performed in the image encoding device. The inverse quantization unit 220 may perform inverse quantization on the quantized transform coefficients using a quantization parameter (e.g., quantization step size information) to obtain transform coefficients.

[0084] The inverse transform unit 230 can inversely transform the transform coefficients to obtain a residual signal (residual block, residual sample array).

[0085] The prediction unit may perform prediction on a current block and generate a predicted block including prediction samples for the current block. The prediction unit may determine whether intra prediction or inter prediction is applied to the current block based on information about the prediction output from the entropy decoding unit 210, and may determine a specific intra / inter prediction mode (prediction technique).

[0086] The prediction unit can generate a prediction signal based on various prediction methods (techniques) described below, as described in the description of the prediction unit of the image encoding device 100.

[0087] The intra predictor 265 may predict the current block by referring to samples in the current picture. The description of the intra predictor 185 may also be applied to the intra predictor 265.

[0088] The inter prediction unit 260 may derive a predicted block for a current block based on a reference block (reference sample array) identified by a motion vector on a reference picture. To reduce the amount of motion information transmitted in inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on correlations between motion information of neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on an inter prediction direction (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. For example, the inter prediction unit 260 may construct a motion information candidate list based on the neighboring blocks and derive a motion vector and / or a reference picture index for the current block based on received candidate selection information. Inter prediction may be performed based on various prediction modes (techniques), and the prediction information may include information indicating the inter prediction mode (technique) for the current block.

[0089] The adder 235 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to a prediction signal (predicted block, predicted sample array) output from a prediction unit (including the inter prediction unit 260 and / or intra prediction unit 265). When there is no residual for the current block, such as when a skip mode is applied, the predicted block can be used as the reconstructed block. The description of the adder 155 also applies to the adder 235. The adder 235 may also be referred to as a reconstruction unit or a reconstructed block generator. The generated reconstructed signal may be used for intra prediction of the next current block in the current picture, and may also be used for inter prediction of the next picture via filtering, as described below.

[0090] The filtering unit 240 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 240 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture, and may store the modified reconstructed picture in the memory 250, specifically, in a DPB of the memory 250. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, etc.

[0091] The (modified) reconstructed picture stored in the DPB of the memory 250 can be used as a reference picture in the inter predictor 260. The memory 250 can store motion information of a block from which motion information in the current picture is derived (or decoded) and / or motion information of a block in an already reconstructed picture. The stored motion information can be transmitted to the inter predictor 260 to be used as motion information of a spatially surrounding block or a temporally surrounding block. The memory 250 can store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra predictor 265.

[0092] In this specification, the embodiments described for the filtering unit 160, inter prediction unit 180 and intra prediction unit 185 of the image encoding device 100 can also be applied in a similar or corresponding manner to the filtering unit 240, inter prediction unit 260 and intra prediction unit 265 of the image decoding device 200, respectively.

[0093] General image / video coding procedures

[0094] In image / video coding, pictures constituting an image / video can be coded / decoded according to a sequence of decoding orders. A picture order corresponding to an output order of decoded pictures can be set to be different from the decoding order. Based on this, not only forward prediction but also backward prediction can be performed during inter prediction.

[0095] 4 shows an example of a schematic picture decoding procedure to which embodiments of the present disclosure can be applied. In FIG. 4, S410 may be performed by the entropy decoding unit 210 of the decoding device described above in FIG. 3, S420 may be performed by a prediction unit including an intra prediction unit 265 and an inter prediction unit 260, S430 may be performed by a residual processing unit including an inverse quantization unit 220 and an inverse transform unit 230, S440 may be performed by the adder 235, and S450 may be performed by the filtering unit 240. S410 may include the information decoding procedure described in this disclosure, S420 may include the inter / intra prediction procedure described in this disclosure, S430 may include the residual processing procedure described in this disclosure, S440 may include the block / picture reconstruction procedure described in this disclosure, and S450 may include the in-loop filtering procedure described in this disclosure.

[0096] Referring to Figure 4, as shown in the description of Figure 3, the picture decoding procedure may generally include an image / video information acquisition procedure (S410) from a bitstream (through decoding), a picture reconstruction procedure (S420-S440), and an in-loop filtering procedure for the reconstructed picture (S450). The picture reconstruction procedure may be performed based on prediction samples and residual samples obtained through the inter / intra prediction (S420) and residual processing (S430, inverse quantization and inverse transform of quantized transform coefficients) processes described in this disclosure. A modified reconstructed picture may be generated through an in-loop filtering procedure for the reconstructed picture generated by the picture reconstruction procedure. The modified reconstructed picture may be output as a decoded picture or may be stored in a decoded picture buffer or memory 250 of the decoding device and used as a reference picture in the inter prediction procedure when decoding a subsequent picture. In some cases, the in-loop filtering procedure may be omitted. In this case, the reconstructed picture may be output as a decoded picture or may be stored in a decoded picture buffer or memory 250 of the decoding device and used as a reference picture in an inter-prediction procedure when decoding a subsequent picture. As described above, the in-loop filtering procedure (S450) may include a deblocking filtering procedure, a sample adaptive offset (SAO) procedure, an adaptive loop filter (ALF) procedure, and / or a bilateral filter procedure, some or all of which may be omitted. Furthermore, one or some of the deblocking filtering procedure, the sample adaptive offset (SAO) procedure, the adaptive loop filter (ALF) procedure, and the bilateral filter procedure may be applied sequentially, or all of them may be applied sequentially. For example, the SAO procedure may be performed after the deblocking filtering procedure is applied to the reconstructed picture.Or, for example, the ALF procedure can be performed after a deblocking filtering procedure is applied to the reconstructed picture, which can also be performed in the encoding device.

[0097] 5 shows an example of a schematic picture encoding procedure to which the embodiments of the present disclosure can be applied. In FIG. 5, S510 may be performed in a prediction unit including the intra prediction unit 185 or the inter prediction unit 180 of the encoding device described above in FIG. 2, S520 may be performed in a residual processing unit including the transform unit 120 and / or the quantization unit 130, and S530 may be performed in the entropy encoding unit 190. S510 may include the inter / intra prediction procedure described in the present disclosure, S520 may include the residual processing procedure described in the present disclosure, and S530 may include the information encoding procedure described in the present disclosure.

[0098] Referring to FIG. 5, the picture encoding procedure, as described with reference to FIG. 2, may include not only a procedure of encoding information for picture reconstruction (e.g., prediction information, residual information, partitioning information, etc.) and outputting it in a bitstream format, but also a procedure of generating a reconstructed picture for a current picture and an optional procedure of applying in-loop filtering to the reconstructed picture. The encoding apparatus may derive (modified) residual samples from quantized transform coefficients via the inverse quantization unit 140 and the inverse transform unit 150, and may generate a reconstructed picture based on the prediction samples output by S510 and the (modified) residual samples. The reconstructed picture generated in this manner may be the same as the reconstructed picture generated by the decoding apparatus described above. A modified reconstructed picture may be generated through an in-loop filtering procedure on the reconstructed picture, which may be stored in the decoded picture buffer or memory 170 and used as a reference picture in the inter prediction procedure when encoding a subsequent picture, as in the decoding apparatus. As described above, in some cases, some or all of the in-loop filtering procedure may be omitted. When the in-loop filtering procedure is performed, (in-loop) filtering-related information (parameters) can be coded by the entropy coding unit 190 and output in bitstream format, and the decoding device can perform the in-loop filtering procedure in the same manner as the coding device based on the filtering-related information.

[0099] This in-loop filtering procedure can reduce noise that occurs during image / video coding, such as blocking artifacts and ringing artifacts, and improve subjective / objective visual quality. Also, by performing the in-loop filtering procedure in both the encoding device and the decoding device, the encoding device and the decoding device can derive the same prediction result, thereby improving the reliability of picture coding and reducing the amount of data to be transmitted for picture coding.

[0100] As described above, a picture reconstruction procedure may be performed not only in a decoding device but also in an encoding device. Reconstructed blocks may be generated based on intra prediction / inter prediction for each block, and a reconstructed picture including the reconstructed blocks may be generated. If a current picture / slice / tile group is an I picture / slice / tile group, blocks included in the current picture / slice / tile group may be reconstructed based only on intra prediction. On the other hand, if the current picture / slice / tile group is a P or B picture / slice / tile group, blocks included in the current picture / slice / tile group may be reconstructed based on intra prediction or inter prediction. In this case, inter prediction may be applied to some blocks in the current picture / slice / tile group, and intra prediction may be applied to the remaining blocks. Color components of a picture may include luma components and chroma components, and unless explicitly limited in this disclosure, methods and embodiments proposed in this disclosure may be applied to luma components and chroma components.

[0101] Example of coding hierarchy and structure

[0102] Video / images coded according to this disclosure may be processed, for example, according to the coding hierarchy and structure described below.

[0103] 6 is a diagram showing the hierarchical structure of a coded image. A coded image can be divided into a video coding layer (VCL) that handles the image decoding process and itself, a lower system that transmits and stores coded information, and a network abstraction layer (NAL) that exists between the VCL and the lower system and is responsible for network adaptation functions.

[0104] The VCL can generate VCL data containing compressed image data (slice data), or it can generate parameter sets containing information such as a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), and a Video Parameter Set (VPS), or an SEI (Supplemental Enhancement Information) message that is additionally required for image decoding processing.

[0105] In NAL, NAL units can be generated by adding header information (NAL unit header) to RBSP (Raw Byte Sequence Payload) generated by VCL. RBSP refers to slice data, parameter sets, SEI messages, etc. generated by VCL. The NAL unit header can include NAL unit type information identified by the RBSP data included in the corresponding NAL unit.

[0106] As shown in the drawing, NAL units can be classified into VCL NAL units and non-VCL NAL units according to the RBSP generated by the VCL. A VCL NAL unit can refer to a NAL unit containing information about an image (slice data), and a non-VCL NAL unit can refer to a NAL unit containing information necessary for decoding an image (parameter set or SEI message).

[0107] The VCL NAL unit and non-VCL NAL unit described above can be transmitted over a network with header information attached according to the data standard of the lower system. For example, the NAL unit can be transformed into a data format of a predetermined standard such as the H.266 / VVC file format, the Real-time Transport Protocol (RTP), or the Transport Stream (TS) and then transmitted over various networks.

[0108] As described above, the NAL unit type of an NAL unit can be identified according to the RBSP data structure included in the NAL unit, and information about such NAL unit type can be stored and signaled in the NAL unit header.

[0109] For example, NAL units can be broadly classified into VCL NAL unit types and non-VCL NAL unit types depending on whether they contain information about an image (slice data). VCL NAL unit types can be classified according to the nature and type of pictures contained in the VCL NAL unit, and non-VCL NAL unit types can be classified according to the type of parameter set.

[0110] Below is a list of examples of NAL unit types identified by the type of parameter set / information included in the non-VCL NAL unit type.

[0111] -DCI (Decoding capability information) NAL unit: Type for NAL units including DCI

[0112] -VPS (Video Parameter Set) NAL unit: Type for NAL units containing VPS

[0113] -SPS (Sequence Parameter Set) NAL unit: Type for NAL unit including SPS

[0114] -PPS (Picture Parameter Set) NAL unit: Type for NAL unit containing PPS

[0115] -APS (Adaptation Parameter Set) NAL unit: Type for NAL units containing APS

[0116] -PH(Picture header) NAL unit:Type for NAL unit including PH

[0117] The above-mentioned NAL unit type has syntax information for the NAL unit type, and the syntax information can be stored in a NAL unit header and signaled. For example, the syntax information can be nal_unit_type, and the NAL unit type can be specified as a value of nal_unit_type.

[0118] Meanwhile, as described above, one picture may include multiple slices, and one slice may include a slice header and slice data. In this case, one picture header may be added to multiple slices (slice header and slice data set) in one picture. The picture header (picture header syntax) may include information / parameters commonly applicable to the pictures.

[0119] The slice header (slice header syntax) may include information / parameters commonly applicable to the slices. The APS (APS syntax) or PPS (PPS syntax) may include information / parameters commonly applicable to one or more slices or pictures. The SPS (SPS syntax) may include information / parameters commonly applicable to one or more sequences. The VPS (VPS syntax) may include information / parameters commonly applicable to multiple layers. The DCI (DCI syntax) may include information / parameters commonly applicable to video in general. The DCI may include information / parameters related to decoding capability. In the present disclosure, a high level syntax (HLS) may include at least one of the APS syntax, PPS syntax, SPS syntax, VPS syntax, DCI syntax, picture header syntax, and slice header syntax. Meanwhile, in the present disclosure, the low level syntax (LLS) may include, for example, slice data syntax, CTU syntax, coding unit syntax, transform unit syntax, etc.

[0120] In the present disclosure, image / video information coded from a coding device to a decoding device and signaled in a bitstream format may include not only intra-picture partitioning-related information, intra / inter prediction information, residual information, in-loop filtering information, etc., but also the slice header information, the picture header information, the APS information, the PPS information, the SPS information, the VPS information, and / or the DCI information. In addition, the image / video information may further include general constraint information and / or NAL unit header information.

[0121] Picture Information Signaling Using NAL Units

[0122] Picture information can be signaled in units of NAL units. For example, picture information can be signaled as described below. A sub-layer is a temporal scalable layer of a temporal scalable bitstream that consists of VCL NAL units with a predetermined value of the variable TemporalId and associated non-VCL NAL units. Here, the variable TemporalId can be derived as follows:

[0123] [Formula 1]

[0124] TemporalId=nuh_temporal_id_plus1-1

[0125] The syntax element nuh_temproal_id_plus1 for signaling the value of the variable TemporalId can be signaled via the NAL unit header of the NAL unit. If the value of nal_unit_type in the NAL unit header is within the value range from IDR_W_RADL to RSV_IRAP_12, the value of TemporalId can be forced to 0. If the value of nal_unit_type is equal to STSA_NUT and the value of vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is 1, the value of TemporalId can be restricted not to be 0. The values ​​of TemporalId can be the same for all VCL NAL units of one AU. The value of TemporalId of a coded picture, PU, ​​or AU can be the value of TemporalId of the VCL NAL unit of the coded picture, PU, ​​or AU. The value of the TemporalId of a sublayer representation may be the largest value of the TemporalId of all VCL NAL units in one sublayer representation.

[0126] The value of TemporalId for non-VCL NAL units, but not for VCL NAL units, can be restricted as follows:

[0127] If -nal_unit_type is equal to DCI_NUT, VPS_NUT or SPS_NUT, TemporalId may be restricted to have a value of 0, and the value of TemporalId of the AU containing the NAL unit may be restricted to be 0.

[0128] Otherwise, if nal_unit_type is equal to PH_NUT, the value of TemporalId may be restricted to be equal to the TemporalId of the PU that contains the NAL unit.

[0129] Otherwise, if nal_unit_type is equal to EOS_NUT or EOB_NUT, TemporalID may be restricted to be equal to 0.

[0130] Otherwise, if nal_unit_type is equal to AUD_NUT, FD_NUT, PREFIX_SEI_NUT, or SUFFIX_SEI_NUT, TemporalId may be restricted to have the same value as the TemporalId of the AU that contains the NAL unit.

[0131] Otherwise, if nal_unit_type is equal to PPS_NUT, PREFIX_APS_NUT, or SUFFIX_APS_NUT, the value of TemporalId may be restricted to have a value equal to or greater than the TemporalId of the PU that contains the NAL unit.

[0132] For example, if the NAL unit is a non-VCL NAL, the value of TemporalId may be equal to the smallest TemporalId value among all AUs to which the non-VCL NAL unit is applied. If the value of nal_unit_type is equal to PPS_NUT, PREFIX_APS_NUT, or SUFFIX_APS_NUT, the value of TemporalId may be equal to or greater than the TemporalId of the AU that contains it. This is because all PPSs and APSs may be included at the beginning of the bitstream (e.g., such information is transmitted out-of-band, and the receiver places it at the beginning of the bitstream). Here, the first coded picture may have a TemporalId value of 0.

[0133] In one embodiment, coded pictures obtained from a bitstream based on NAL unit information can be signaled by a coding device and identified by a decoding device as follows: Note that this is just an example, and pictures can also be identified in other ways.

[0134] An IRAP (Intra random access point) picture is a coded picture in which all VCL NAL units have the same value for nal_unit_type, which ranges from IDR_W_RADL to CRA_NUT. In one embodiment, an IRAP picture may not reference any other pictures other than itself during its decoding process to perform inter prediction. An IRAP picture may be a CRA picture or an IDR picture, as described below. In decoding order, the first picture in the bitstream can be forced to be an IRAP or GDR picture. If a mandatory parameter set reference is required, the IRAP picture and all subsequent non-RASL pictures in the CVS in decoding order can be correctly decoded so that the parameter set is available. This can be done without performing the decoding process of any other pictures preceding the IRAP picture in decoding order.

[0135] A CRA (Clean Random Access) picture is an IRAP picture whose VCL NAL unit has the same nal_unit_type as CRA_NUT. For example, a CRA picture is a picture that does not reference any other picture other than itself for inter prediction during its decoding process. A CRA picture may be the first picture in a bitstream in decoding order, or it may appear as a later-ordered picture in the bitstream. A CRA picture may have an associated RADL or RASL picture. If a CRA picture has a NoIncorrectPicOutputFlag value of 1, the associated RASL picture may not be output by the decoding device. This is because the CRA picture may not be decodable for reasons such as not containing a reference to a picture not provided in the bitstream. In one embodiment, if an incomplete picture is not output during image decoding, the CRA picture may have a NoIncorrectPicOutputFlag value of 1 if the CRA picture is an incomplete picture.

[0136] An IDR (Instantaneous Decoding Refresh) picture is an IRAP picture whose individual VCL NAL unit has the same value as IDR_W_RADL or IDR_N_LP for nal_unit_type. For example, an IDR picture may not need to refer to any other picture other than itself for inter prediction during its decoding process. An IDR picture can appear as the first picture in a bitstream in decoding order. Alternatively, an IDR picture can appear as a later picture in the bitstream. Each IDR picture can be the first picture of a CVS in decoding order. If an IDR picture for each VCL NAL unit has the same value as IDR_W_RADL for nal_unit_type, the IDR picture can also have an associated RADL picture. If an IDR picture for a respective VCL NAL unit has the same value of nal_unit_type as IDR_N_LP, the IDR picture may not have an associated leading picture. An IDR picture may not have an associated RASL picture.

[0137] A RADL (Random Access Decodable Leading) picture may be a coded picture in which an individual VCL NAL unit has a value of RADL_NUT as the value of nal_unit_type. In one embodiment, all RADL pictures may be leading pictures. RADL pictures may not be used as reference pictures for the decoding process of trailing pictures of the same associated IRAP picture. If the value of the syntax element field_seq_flag obtained from the bitstream is 0, all existing RADL pictures may precede all non-leading pictures of the same associated IRAP picture in decoding order.

[0138] A Random Access Skipped Leading (RASL) picture may be a coded picture whose individual VCL NAL unit has RASL_NUT as the value of nal_unit_type. In one embodiment, all RASL pictures may be leading pictures of the associated CRA picture. If the associated CRA picture has NoIncorrectPicOutputFlag as the value 1, the RASL picture may not be output and may not be decoded successfully. This may be because the RASL picture contains a reference to a picture that is not provided in the bitstream. The RASL picture may not be used as a reference picture for the decoding process of a non-RASL picture. If the value of field_seq_flag is 0, all existing RASL pictures may precede all non-leading pictures of the same associated CRA picture in decoding order.

[0139] A trailing picture is a non-IRAP picture that follows the associated IRAP picture in output order but is not an STSA picture. A trailing picture associated with an IRAP picture can be located after the IRAP picture in decoding order. A picture that is located after the associated IRAP picture in output order but before the associated IRAP picture in decoding order is not allowed.

[0140] A GDR (Gradual decoding refresh) picture is a picture whose individual VCL NAL unit has a value of GDR_NUT as nal_unit_type.

[0141] An STSA (Step-wise temporal sublayer access) picture is a picture whose individual VCL NAL unit has STSA_NUT as the value of nal_unit_type. An STSA picture may not use a picture with the same TemporalId as the STSA picture for inter-prediction reference. A picture with the same TemporalId as the STSA picture and located after the STSA picture in decoding order may not use a picture with the same TemporalId as the STSA picture for inter-prediction reference and located before the STSA picture in decoding order.

[0142] An STSA picture can enable up-switching from a sublayer immediately below the sublayer containing the STSA picture to the sublayer containing the STSA picture. STSA pictures can be forced to have a TemporalId greater than 0.

[0143] In one embodiment, for single-layer or multi-layer bitstreams, at least one of the following restrictions may apply:

[0144] An individual picture that is not the first picture in the bitstream in decoding order can be considered to be related to the previous IRAP picture in decoding order.

[0145] - If it is the leading picture of an IRAP picture, it can be a RADL or RASL picture.

[0146] - If it is a trailing picture of an IRAP picture, it can be restricted to be a picture other than a RADL or RASL picture.

[0147] -It can be restricted so that RASL pictures are not provided in the bitstream associated with the IDR picture.

[0148] A bitstream associated with an IDR picture whose -nal_unit_type is IDR_N_LP may be restricted so that no RADL pictures are provided. (For example, if each parameter set is available in the bitstream or by external means when it is referenced, random access at the position of an IRAP PU can be performed by discarding all PUs before the IRAP PU, and the IRAP picture and subsequent non-RADL pictures in decoding order can be decoded correctly.)

[0149] -All pictures preceding the IRAP picture according to decoding order can be forced to precede the IRAP picture according to output order, and can be forced to precede all RADL pictures associated with the IRAP picture according to output order.

[0150] All RASL pictures associated with a CRA picture can be constrained to precede all RADL pictures associated with the CRA picture in the output order.

[0151] All RASL pictures associated with a CRA picture can be located after all IRAP pictures that precede the CRA picture in decoding order in output order.

[0152] If the value of -field_seq_flag is 0 and the current picture is a leading picture associated with an IRAP picture, the current picture can precede all non-leading pictures associated with the same IRAP picture in decoding order. Otherwise, for the first leading picture picA and the last leading picture picB in decoding order among the leading pictures associated with an IRAP picture, there can be at most one non-leading picture preceding picA in decoding order, and there can be no non-leading picture between picA and picB in decoding order.

[0153] Multi-layer based coding

[0154] Image / video coding according to the present disclosure may include multi-layer-based image / video coding. The multi-layer-based image / video coding may include scalable coding. In multi-layer-based coding or scalable coding, an input signal may be processed layer by layer. Depending on the layer, the input signal (input image / picture) may have different values ​​for at least one of resolution, frame rate, bit depth, color format, aspect ratio, and view. In this case, by utilizing the differences between layers (e.g., based on scalability), inter-layer prediction may be performed to reduce redundant transmission / processing of information and improve compression efficiency.

[0155] FIG. 7 shows a schematic block diagram of a multi-layer encoding device 700 to which embodiments of the present disclosure can be applied, where multi-layer based video / image signal encoding is performed.

[0156] The multi-layer encoding device 700 of Figure 7 may include the encoding device of Figure 2. Compared to Figure 2, the image dividing unit 110 and the adder 155 are omitted from the multi-layer encoding device 700 of Figure 7, but the multi-layer encoding device 700 may include the image dividing unit 110 and the adder 155. In one embodiment, the image dividing unit 110 and the adder 155 may be included in units of layers. The following description of Figure 7 will focus on multi-layer-based prediction. For example, in addition to the content described below, the multi-layer encoding device 700 of Figure 7 may include the technical concept of the encoding device previously described with reference to Figure 2.

[0157] For convenience of explanation, a multi-layer structure consisting of two layers is shown in Figure 7. However, the embodiments of the present disclosure are not limited to two layers, and the multi-layer structure to which the embodiments of the present disclosure are applied may include more than two layers.

[0158] 7, a coding apparatus 700 includes an encoder 700-1 for layer 1 and an encoder 700-0 for layer 0. Layer 0 may be a base layer, a reference layer, or a lower layer, and layer 1 may be an enhancement layer, a current layer, or a higher layer.

[0159] The layer 1 encoder 700-1 may include a predictor 720-1, a residual processor 730-1, a filter 760-1, a memory 770-1, an entropy encoder 740-1, and a multiplexer (MUX) 770. In one embodiment, the MUX may be included as an external component.

[0160] The coding unit 700-0 of layer 0 may include a prediction unit 720-0, a residual processing unit 730-0, a filtering unit 760-0, a memory 770-0, and an entropy coding unit 740-0.

[0161] The predictors 720-0 and 720-1 may perform prediction on the input image based on various prediction techniques as described above. For example, the predictors 720-0 and 720-1 may perform inter prediction and intra prediction. The predictors 720-0 and 720-1 may perform prediction in a predetermined processing unit. The prediction execution unit may be a coding unit (CU) or a transform unit (TU). A predicted block (including prediction samples) may be generated based on the prediction result, and the residual processor may derive a residual block (including residual samples) based on the predicted block.

[0162] In inter prediction, a prediction block can be generated by performing prediction based on information of at least one of a picture before and / or a picture after the current picture, and in intra prediction, a prediction block can be generated by performing prediction based on neighboring samples within the current picture.

[0163] The inter prediction mode or method may include the various prediction modes described above. In inter prediction, a reference picture may be selected for a current block to be predicted, and a reference block corresponding to the current block may be selected within the reference picture. The predictors 720-0 and 720-1 may generate predicted blocks based on the reference blocks.

[0164] Additionally, the predictor 720-1 may perform prediction for layer 1 using information of layer 0. In this disclosure, for convenience of explanation, a method of predicting information of a current layer using information of another layer is referred to as inter-layer prediction.

[0165] The information of the current layer that is predicted using information of other layers (e.g., predicted by inter-layer prediction) may be at least one of texture, motion information, unit information, and predetermined parameters (e.g., filtering parameters, etc.).

[0166] Furthermore, information of other layers used for prediction of the current layer (eg, used for inter-layer prediction) may be at least one of texture, motion information, unit information, and predetermined parameters (eg, filtering parameters, etc.).

[0167] In inter-layer prediction, the current block may be a block in a current picture in a current layer (e.g., layer 1) and may be a block to be coded. The reference block may be a block in a picture (reference picture) that belongs to the same access unit (AU) as the picture to which the current block belongs (current picture) in a layer (reference layer, e.g., layer 0) referenced for prediction of the current block, and may be a block corresponding to the current block.

[0168] An example of inter-layer prediction is inter-layer motion prediction, which predicts motion information of a current layer using motion information of a reference layer. According to inter-layer motion prediction, motion information of a current block can be predicted using motion information of a reference block. That is, when deriving motion information according to an inter-prediction mode described below, motion information candidates can be derived based on motion information of an inter-layer reference block instead of a temporally neighboring block.

[0169] When applying inter-layer motion prediction, the predictor 720-1 may scale and use motion information of a reference block of a reference layer (ie, an inter-layer reference block).

[0170] As another example of inter-layer prediction, inter-layer texture prediction can use the texture of a reconstructed reference block as a predicted value for the current block. In this case, the predictor 720-1 can scale the texture of the reference block by upsampling. Inter-layer texture prediction may be referred to as inter-layer (reconstructed) sample prediction or simply inter-layer prediction.

[0171] Another example of inter-layer prediction is inter-layer parameter prediction, in which derived parameters of a reference layer can be reused in the current layer, or parameters for the current layer can be derived based on parameters used in the reference layer.

[0172] In inter-layer residual prediction, which is another example of inter-layer prediction, residual information of another layer is used to predict the residual of the current layer, and prediction for the current block can be performed based on this.

[0173] Inter-layer differential prediction, which is another example of inter-layer prediction, can predict the current block using the difference between images obtained by upsampling or downsampling a reconstructed picture of the current layer and a reconstructed picture of the reference layer.

[0174] In inter-layer syntax prediction, which is another example of inter-layer prediction, the texture of a current block may be predicted or generated using syntax information of a reference layer, where the syntax information of the reference layer may include information on an intra-prediction mode and motion information.

[0175] A plurality of the above-described inter-layer prediction methods may be used when predicting a specific block.

[0176] Here, examples of inter-layer prediction have been described, such as inter-layer texture prediction, inter-layer motion prediction, inter-layer unit information prediction, inter-layer parameter prediction, inter-layer residual prediction, inter-layer differential prediction, and inter-layer syntax prediction, but the inter-layer prediction applicable to the present invention is not limited to these.

[0177] For example, inter-layer prediction may be applied as an extension of inter-prediction for the current layer. That is, a reference picture derived from a reference layer may be included in the reference pictures that can be referred to in inter-prediction of the current block, and inter-prediction may be performed for the current block.

[0178] In this case, the inter-layer reference picture may be included in a reference picture list for the current block. The predictor 720-1 may perform inter prediction on the current block using the inter-layer reference picture.

[0179] Here, the interlayer reference picture may be a reference picture constructed by sampling a reconstructed picture of a reference layer to correspond to a current layer. Therefore, when a reconstructed picture of a reference layer corresponds to a picture of a current layer, the reconstructed picture of the reference layer can be used as an interlayer reference picture without sampling. For example, when the width and height of the samples in the reconstructed picture of the reference layer and the reconstructed picture of the current layer are the same, and the offset between the upper left corner, upper right corner, lower left corner, and lower right corner of the picture of the reference layer and the upper left corner, upper right corner, lower left corner, and lower right corner of the picture of the current layer is 0, the reconstructed picture of the reference layer can be used as an interlayer reference picture of the current layer without being resampled.

[0180] Furthermore, the reconstructed picture of the reference layer from which the inter-layer reference picture is derived may be a picture that belongs to the same AU as the current picture to be coded.

[0181] When an inter-layer reference picture is included in a reference picture list to perform inter-prediction on a current block, the position of the inter-layer reference picture in the reference picture list may differ between reference picture lists L0 and L1. For example, in reference picture list L0, the inter-layer reference picture may be located after a short-term reference picture before the current picture, and in reference picture list L1, the inter-layer reference picture may be located at the end of the reference picture list.

[0182] Here, reference picture list L0 is a reference picture list used in inter prediction of a P slice or a reference picture list used as the first reference picture list in inter prediction of a B slice, and reference picture list L1 is a second reference picture list used in inter prediction of a B slice.

[0183] Therefore, the reference picture list L0 may be configured in the order of short-term reference pictures before the current picture, inter-layer reference pictures, short-term reference pictures after the current picture, and long-term reference pictures, while the reference picture list L1 may be configured in the order of short-term reference pictures after the current picture, short-term reference pictures before the current picture, long-term reference pictures, and inter-layer reference pictures.

[0184] Here, a P slice (predictive slice) is a slice in which intra prediction is performed or inter prediction is performed using up to one motion vector and reference picture index per prediction block. A B slice (bi-predictive slice) is a slice in which intra prediction is performed or prediction is performed using up to two motion vectors and reference picture indexes per prediction block. In this regard, an I slice (intra slice) is a slice to which only intra prediction is applied.

[0185] Furthermore, when inter-prediction is performed on the current block based on a reference picture list including inter-layer reference pictures, the reference picture list may include multiple inter-layer reference pictures derived from multiple layers.

[0186] When multiple inter-layer reference pictures are included, the inter-layer reference pictures may be inter-arranged in the reference picture lists L0 and L1. For example, assume that two inter-layer reference pictures, inter-layer reference picture ILRPi and inter-layer reference picture ILRPj, are included in the reference picture list used for inter-prediction of the current block. In this case, in the reference picture list L0, ILRPi may be located after the short-term reference picture before the current picture, and ILRPj may be located at the end of the list. Also, in the reference picture list L1, ILRPi may be located at the end of the list, and ILRPj may be located after the short-term reference picture after the current picture.

[0187] In this case, the reference picture list L0 may be configured in the order of short-term reference pictures before the current picture, interlayer reference picture ILRPi, short-term reference pictures after the current picture, long-term reference pictures, and interlayer reference picture ILRPj. The reference picture list L1 may be configured in the order of short-term reference pictures after the current picture, interlayer reference picture ILRPj, short-term reference pictures before the current picture, long-term reference pictures, and interlayer reference picture ILRPi.

[0188] Furthermore, one of the two interlayer reference pictures may be an interlayer reference picture derived from a scalable layer related to resolution, and the other may be an interlayer reference picture derived from a layer providing a different view. In this case, for example, if ILRPi is an interlayer reference picture derived from a layer providing a different resolution and ILRPj is an interlayer reference picture derived from a layer providing a different view, in the case of scalable video coding that supports only scalability excluding views, the reference picture list L0 may be configured in the order of short-term reference pictures before the current picture, interlayer reference picture ILRPi, short-term reference pictures after the current picture, and long-term reference pictures. The reference picture list L1 may be configured in the order of short-term reference pictures after the current picture, short-term reference pictures before the current picture, long-term reference pictures, and interlayer reference picture ILRPi.

[0189] On the other hand, in inter-layer prediction, the information of the inter-layer reference picture may use only sample values, only motion information (motion vectors), or both sample values ​​and motion information. When the reference picture index indicates an inter-layer reference picture, the predictor 720-1 may use only sample values ​​of the inter-layer reference picture, or only motion information (motion vectors) of the inter-layer reference picture, or both sample values ​​and motion information of the inter-layer reference picture, according to the information received from the encoding device.

[0190] When only sample values ​​of the inter-layer reference picture are used, the predictor 720-1 may derive samples of a block identified by a motion vector in the inter-layer reference picture as predicted samples of the current block. In the case of view-independent scalable video coding, the motion vector in inter prediction using the inter-layer reference picture (inter-layer prediction) may be set to a fixed value (e.g., 0).

[0191] When only motion information of an inter-layer reference picture is used, the predictor 720-1 can use a motion vector specified in the inter-layer reference picture as a motion vector predictor for deriving a motion vector of the current block. The predictor 720-1 can also use a motion vector specified in the inter-layer reference picture as a motion vector of the current block.

[0192] When using both samples and motion information from the interlayer reference picture, the prediction unit 720-1 can use samples of the area corresponding to the current block in the interlayer reference picture and motion information (motion vectors) identified in the interlayer reference picture to predict the current block.

[0193] When inter-layer prediction is applied, the encoding device can transmit to the decoding device a reference index that points to an inter-layer reference picture in a reference picture list, and can also transmit to the decoding device information that specifies which information (sample information, motion information, or both sample information and motion information) to use from the inter-layer reference picture, i.e., information that specifies the type of dependency for inter-layer prediction between two layers.

[0194] 8 illustrates a schematic block diagram of a decoding device to which an embodiment of the present disclosure can be applied, in which multi-layer-based video / image signal decoding is performed. The decoding device of FIG. 8 may include the decoding device of FIG. 3. The reordering unit shown in FIG. 8 may be omitted or may be included in the inverse quantization unit. The description of this drawing will focus on multi-layer-based prediction. The rest of the description may include the description of the decoding device described in FIG. 3.

[0195] 8, for convenience of explanation, a multi-layer structure consisting of two layers is described as an example. However, it should be noted that the embodiments of the present disclosure are not limited to this, and the multi-layer structure to which the embodiments of the present disclosure are applied may include two or more layers.

[0196] 8, a decoding device 800 may include a decoder 800-1 for layer 1 and a decoder 800-0 for layer 0. The decoder 800-1 for layer 1 may include an entropy decoder 810-1, a residual processor 820-1, a predictor 830-1, an adder 840-1, a filtering unit 850-1, and a memory 860-1. The decoder 800-0 for layer 0 may include an entropy decoder 810-0, a residual processor 820-0, a predictor 830-0, an adder 840-0, a filtering unit 850-0, and a memory 860-0.

[0197] When a bitstream including image information is transmitted from an encoding device, the DEMUX 805 can demultiplex the information for each layer and transmit it to a decoder for each layer.

[0198] The entropy decoding units 810-1 and 810-0 can perform decoding in accordance with the coding method used in the encoding device. For example, if CABAC is used in the encoding device, the entropy decoding units 810-1 and 810-0 can also perform entropy decoding using CABAC.

[0199] When the prediction mode for the current block is an intra prediction mode, the predictors 830-1 and 830-0 may perform intra prediction on the current block based on neighboring reconstructed samples in the current picture.

[0200] When the prediction mode for the current block is the inter prediction mode, the prediction units 830-1 and 830-0 may perform inter prediction for the current block based on information included in at least one of a picture preceding or following the current picture. Some or all of the motion information required for inter prediction may be determined by checking information received from the encoding device and derived accordingly.

[0201] When a skip mode is applied as an inter prediction mode, the residual is not transmitted from the encoding device, and the predicted block can be used as a reconstructed block.

[0202] On the other hand, the predictor 830-1 of layer 1 may perform inter prediction or intra prediction using only information in layer 1, or may perform inter layer prediction using information of another layer (layer 0).

[0203] The information of the current layer that is predicted using information of other layers (e.g., predicted by inter-layer prediction) may be at least one of texture, motion information, unit information, and predetermined parameters (e.g., filtering parameters, etc.).

[0204] Furthermore, information of another layer used for prediction of the current layer (for example, used for inter-layer prediction) may be at least one of texture, motion information, unit information, and predetermined parameters (for example, filtering parameters, etc.).

[0205] In inter-layer prediction, the current block may be a block in a current picture in a current layer (e.g., layer 1) and may be a block to be decoded. The reference block may be a block in a picture (reference picture) that belongs to the same access unit (AU) as the picture to which the current block belongs (current picture) in a layer (reference layer, e.g., layer 0) referenced for prediction of the current block and that corresponds to the current block.

[0206] The multi-layer decoding device 800 may perform inter-layer prediction as previously described for the multi-layer encoding device 700. For example, the multi-layer decoding device 800 may perform inter-layer texture prediction, inter-layer motion prediction, inter-layer unit information prediction, inter-layer parameter prediction, inter-layer residual prediction, inter-layer differential prediction, inter-layer syntax prediction, etc. as previously described for the multi-layer encoding device 700, and inter-layer prediction that can be applied in the present disclosure is not limited to these.

[0207] The predictor 830-1 may perform inter-layer prediction using an inter-layer reference picture when a reference picture index received from the encoding device or a reference picture index derived from a neighboring block points to an inter-layer reference picture in a reference picture list. For example, when the reference picture index points to an inter-layer reference picture, the predictor 830-1 may derive sample values ​​of an area identified by a motion vector in the inter-layer reference picture as a predicted block for the current block.

[0208] In this case, the inter-layer reference picture may be included in a reference picture list for the current block. The predictor 830-1 may perform inter prediction on the current block using the inter-layer reference picture.

[0209] As previously described in the multi-layer encoding apparatus 700, the inter-layer reference picture may be a reference picture constructed by sampling a reconstructed picture of a reference layer to correspond to a current layer in the operation of the multi-layer decoding apparatus 800. Processing for the case where a reconstructed picture of a reference layer corresponds to a picture of the current layer may be performed in the same manner as processing in the encoding process.

[0210] Also, as previously described in the multi-layer encoding device 700, in the operation of the multi-layer decoding device 800, the reconstructed picture of the reference layer from which the inter-layer reference picture is derived may be a picture belonging to the same AU as the current picture to be encoded.

[0211] Also, as previously described for the multi-layer encoding device 700, in the operation of the multi-layer decoding device 800, when an inter-layer reference picture is included in a reference picture list to perform inter-prediction on a current block, the position of the inter-layer reference picture within the reference picture list may differ between reference picture lists L0 and L1.

[0212] Also, as previously described in the multi-layer encoding device 700, in the operation of the multi-layer decoding device 800, when inter-prediction is performed on a current block based on a reference picture list including inter-layer reference pictures, the reference picture list may include multiple inter-layer reference pictures derived from multiple layers, and the arrangement of the inter-layer reference pictures may also be performed in a manner corresponding to that previously described in the encoding process.

[0213] Also, as previously described for the multi-layer encoding device 700, in the operation of the multi-layer decoding device 800, the information of the inter-layer reference picture may be such that only sample values ​​are used, only motion information (motion vectors) are used, or both sample values ​​and motion information are used.

[0214] The multi-layer decoding device 800 can receive a reference index indicating an inter-layer reference picture in a reference picture list from the multi-layer encoding device 700 and perform inter-layer prediction based thereon. The multi-layer decoding device 800 can also receive information indicating which information (sample information, motion information, or both sample information and motion information) to use from the inter-layer reference picture, i.e., information specifying the type of dependency for inter-layer prediction between two layers, from the multi-layer encoding device 700.

[0215] HLS (High level syntax) signaling and semantics

[0216] As described above, the HLS can be coded and / or signaled for video and / or image coding. As described above, the video / image information in the present disclosure can be included in the HLS. Then, the image / video coding method can be performed based on such image / video information.

[0217] Video Parameter Set signaling

[0218] A video parameter set (VPS) is a parameter set used for transmitting layer information. The layer information may include, for example, information about an output layer set (OLS), information about a profile tier level, information about the relationship between an OLS and a hypothetical reference decoder, information about the relationship between an OLS and a DPB, etc. The VPS may not be essential for decoding a bitstream. Before being referenced, the VPS raw byte sequence payload (RBSP) must be available to the decoding process by being included in at least one access unit (AU) with a TemporalID of 0 or by being provided via external means. All VPS NAL units with a particular value of vps_video_parameter_set_id in a CVS must have the same content.

[0219] 9 is a diagram illustrating an example of a syntax structure of a VPS according to an embodiment of the present disclosure. The syntax elements in FIG. 9 will be described below.

[0220] vps_video_parameter_set_id provides an identifier for the VPS. Other syntax elements can reference the VPS using vps_video_parameter_set_id. The value of vps_video_parameter_set_id must be greater than 0.

[0221] vps_max_layers_minus1 can indicate the maximum allowable number of layers that can exist in an individual CVS that references a VPS. For example, the value of vps_max_layers_minus1 plus 1 can indicate the maximum allowable number of layers that can exist in an individual CVS that references a VPS.

[0222] The value of vps_max_sublayers_minus1 plus 1 can indicate the maximum number of temporal sublayers that can exist in a layer in an individual CVS that references the VPS.

[0223] A value of 1 for vps_all_layers_same_num_sublayers_flag may indicate that the number of temporal sublayers is the same for all layers in an individual CVS that references the VPS. A value of 0 for vps_all_layers_same_num_sublayers_flag may indicate that the number of temporal sublayers may or may not be the same for layers in an individual CVS that references the VPS. If a value for vps_all_layers_same_num_sublayers_flag is not provided from the bitstream, the value of vps_all_layers_same_num_sublayers_flag may be set to 1.

[0224] A value of 1 for vps_all_independent_layers_flag can indicate that all layers belonging to the CVS are coded independently without using inter-layer prediction. A value of 0 for vps_all_independent_layers_flag can indicate that at least one layer belonging to the CVS is coded using inter-layer prediction.

[0225] vps_layer_id[i] may indicate the value of nuh_layer_id of the i-th layer. For any two non-negative integer values ​​m and n, if m is less than n, vps_layer_id[m] may be restricted to have a value less than vps_layer_id[n]. Here, nuh_layer_id is a syntax element signaled in the NAL unit header and may indicate an identifier of a NAL unit.

[0226] A value of 1 for vps_independent_layer_flag[i] may indicate that inter-layer prediction is not applied to the layer corresponding to index i. A value of 0 for vps_independent_layer_flag[i] may indicate that inter-layer prediction is applied to the layer corresponding to index i, and that the syntax element vps_direct_ref_layer_flag[i][j] is obtained from the VPS, where j may have a value from 0 to i-1. On the other hand, if the value of vps_independent_layer_flag[i] does not exist in the bitstream, its value may be set to 1.

[0227] A value of 0 for vps_direct_ref_layer_flag[i][j] may indicate that the layer with index j is not a direct reference layer of the layer with index i. A value of 1 for vps_direct_ref_layer_flag[i][j] may indicate that the layer with index j is a direct reference layer of the layer with index i. For i and j ranging from 0 to vps_max_layers_minus1, if the value of vps_direct_ref_layer_flag[i][j] is not obtained from the bitstream, the value may be set to 0. If the value of vps_independent_layer_flag[i] is 0, there may be at least one j that causes the value of vps_direct_ref_layer_flag[i][j] to be 1, and the value range of j may range from 0 to i-1.

[0228] In one embodiment, the variables NumDirectRefLayers[i], DirectRefLayerIdx[i][d], NumRefLayers[i], RefLayerIdx[i][r], and LayerUsedAsRefLayerFlagd[j] can be derived using the pseudocode of FIG. 10.

[0229] The variable GeneralLayerIdx[i] indicates the layer index of the layer whose value of nuh_layer_id is equal to vps_layer_id[i], and can be derived as follows:

[0230] [Formula 2]

[0231] for(i=0;i<=vps_max_layers_minus1;i++)

[0232] GeneralLayerIdx[vps_layer_id[i]]=i

[0233] A value of 1 for max_tid_ref_present_flag[i] may indicate that the syntax element max_tid_il_ref_pics_plus1[i] is provided in the bitstream. A value of 0 for max_tid_ref_present_flag[i] may indicate that the syntax element max_tid_il_ref_pics_plus1[i] is not provided in the bitstream.

[0234] A value of 0 for max_tid_il_ref_pics_plus1[i] may indicate that inter-layer prediction with non-IRAP pictures of the i-th layer is not used. A value greater than 0 for max_tid_il_ref_pics_plus1[i] may indicate that pictures with a temporal ID greater than max_tid_il_ref_pics_plus1[i]-1 for decoding pictures of the i-th layer are not used as inter-layer reference pictures (ILRPs). On the other hand, if the value of max_tid_il_ref_pics_plus1[i] is not obtained from the bitstream, its value may be induced to be 7.

[0235] A value of 1 for the syntax element each_layer_is_an_ols_flag may indicate that an individual OLS has only one layer, and that an individual layer belonging to a CVS referencing a VPS is an OLS with a single containing layer, which is the only output layer. A value of 0 for each_layer_is_an_ols_flag may indicate that an OLS may contain more than one layer. In one embodiment, if the value of vps_max_layers_minus1 is 0, the value of each_layer_is_an_ols_flag may be set to 1. Otherwise, if the value of vps_all_independent_layers_flag is 0, the value of each_layer_is_an_ols_flag may be set to 0.

[0236] A value of 0 for ols_mode_idc may indicate that the total number of OLSs specified by the VPS is equal to vps_max_layers_minus1 + 1. The i-th OLS may contain layers with layer indices from 0 to i. Then, for each OLS, the highest layer of the OLS may be output.

[0237] A value of 1 for ols_mode_idc may indicate that the total number of OLSs specified by the VPS is equal to vps_max_layers_minus1+1. The i-th OLS may contain layers with layer indices from 0 to i. And for each OLS, all layers of the OLS may be output.

[0238] A value of 2 for ols_mode_idc means that the total number of OLSs specified by the VPS is explicitly signaled, the output layers for each individual OLS are explicitly signaled, and other layers can be direct or reference layers of the output layers of the OLS.

[0239] The value of ols_mode_idc can have a value from 0 to 2. The value 3 of ols_mode_idc can be saved for future use. If the value of vps_all_independent_layers_flag is 1 and the value of each_layer_is_an_ols_flag is 0, the value of ols_mode_idc can be induced to 2.

[0240] When the value of ols_mode_idc is a predetermined value (for example, the value is 2), the value of the syntax element num_output_layer_sets_minus1 plus 1 can indicate the total number of OLSs specified by the VPS.

[0241] The variable TotalNumOlss, which indicates the total number of OLSs specified by the VPS, can be derived as shown in FIG.

[0242] A value of 1 for ols_output_layer_flag[i][j] can indicate that the layer with nuh_layer_id value equal to vps_layer_id[j] is the output layer of the i-th OLS when ols_mode_idc value is 2. A value of 0 for ols_output_layer_flag[i][j] can indicate that the layer with nuh_layer_id value equal to vps_layer_id[j] is not the output layer of the i-th OLS when ols_mode_idc value is 2.

[0243] The variable NumOutputLayersInOls[i] indicating the number of output layers in the i-th OLS, the variable NumSubLayersInLayerInOLS[i][j] indicating the number of sublayers in the j-th layer in the i-th OLS, the variable OutputLayerIdInOls[i][j] indicating the nuh_layer_id value of the j-th output layer in the i-th OLS, and the variable LayerUsedAsOutputLayerFlag[k] indicating whether the k-th layer is used as an output layer in at least one OLS can be derived as shown in the pseudocode of Figure 12.

[0244] For each value of i from 0 to vps_max_layers_minus1, the values ​​of LayerUsedAsRefLayerFlag[i] and LayerUsedAsOutputLayerFlag[i] can be forced to be non-zero. For example, it can be forced that there are no layers that are not the output layer of at least one OLS and are not a direct reference layer of another layer.

[0245] For each individual OLS, there can be forced to be at least one layer that is an output layer. For example, for each i value from 0 to TotalNumOlss-1, the value of NumOutputLayersInOls[i] can be forced to have a value greater than or equal to 1.

[0246] The variable NumLayersInOls[i] indicating the number of layers in the i-th OLS and the variable LayerIdInOls[i][j] indicating the value of nuh_layer_id of the j-th layer in the i-th OLS can be derived as shown in FIG. 13.

[0247] In one embodiment, the 0th OLS can include only the lowest layer. Here, the lowest layer may refer to the layer whose nuh_layer_id value is vps_layer_id[0]. And, the only layer included in the 0th OLS can be used as output.

[0248] The variable OlsLayerIdx[i][j], which indicates the OLS layer index of the layer whose value of nuh_layer_id is the same as LayerIdInOls[i][j], can be derived as shown in FIG.

[0249] The lowest layer present in each OLS can be restricted to be an independent layer. For example, for each i with a range of values ​​from 0 to TotalNumOlss-1, the value of vps_independent_layer_flag[GeneralLayerIdx[LayerIdInOls[i][0]]] can be forced to be 1.

[0250] Each layer can be forced to be included in at least one OLS specified by the VPS.

[0251] max_tid_il_ref_pics_plus1[i] signaling limit

[0252] The signaling regarding the syntax element max_tid_il_ref_pics_plus1[i] is somewhat problematic in its semantics and functionality.

[0253] For example, the semantics of max_tid_il_ref_pics_plus1[i] may indicate that if the value of max_tid_il_ref_pics_plus1[i] is greater than 0, a picture in the i-th layer uses only up to max_tid_il_ref_ref_pics_plus1[i]-1 sub-layers from the reference picture for its inter-layer prediction.

[0254] This indicates that a picture present in the i-th sub-layer can use up to max_tid_il_ref_ref_pics_plus1[i]-1 sub-layers from the reference picture to perform its inter-layer prediction.

[0255] The syntax element is designed to allow the derivation of the variable NumSubLayersInLayerInOLS[i][j], where the variable NumSubLayersInLayerInOLS[i][j] may indicate the number of sublayers present in the jth layer in the ith OLS, where OLS is an abbreviation for Output layer set and may refer to a set of at least one layer designated as an output layer.

[0256] The variable NumSubLayersInLayerInOLS[i][j] can be used in the bitstream extraction process to remove pictures that reside in sub-layers that reside in layers that are not the output layer of the extracted OLS.

[0257] Such a mechanism exhibits inefficiency when one layer uses more than one reference layer for inter-layer prediction and the number of sub-layers used in the individual reference layers is not the same.

[0258] For example, it may be assumed that layer 2 refers to layer 0 and layer 1 to perform inter-layer prediction. In layer 0, only two sub-layers are used to perform inter-layer prediction, whereas three sub-layers may be used in layer 1. In such an example, according to the above-mentioned signaling method, layer 2 uses three sub-layers to perform inter-layer prediction, which causes a problem that layer 0 cannot remove sub-layer 3 or more.

[0259] Improvement plan

[0260] The following embodiments provide solutions to the above-mentioned problems. Each embodiment may be implemented individually or in part or in whole in combination.

[0261] Improvement method 1. To signal the maximum number of sublayers used in inter-layer prediction for a layer, instead of signaling one value for all reference layers, the maximum number of sublayers used in the reference layer can be signaled for each individual reference layer of a layer.

[0262] Improvement Scheme 2: The maximum number of sublayers used in reference layer j for layer i can be provided only if layer j is the direct reference layer of layer i.

[0263] Improvement method 3: A flag indicating whether or not there is signaling of the maximum number of sublayers used for inter-layer prediction by one layer can be signaled for all layers. This can be signaled in the VPS.

[0264] a) The syntax element max_tid_ref_present_flag[i] can be changed to max_tid_ref_present_flag and used.

[0265] b) If all layers are independent layers, max_tid_ref_present_flag may not be provided. If max_tid_ref_present_flag is not provided, the value of max_tid_ref_present_flag may be set to 0.

[0266] Improvement method 4. NumSubLayersInLayerInOLS for direct reference layers and indirect reference layers can be derived as follows:

[0267] a) If the value of each_layer_is_an_ols_flag[i] is 1 (e.g., true), NumSubLayersInLayerInOLS[i][0] can be set to the value of vps_max_sub_layers_minus1+1.

[0268] b) If the value of each_layer_is_an_ols_flag[i] is 0 (e.g., false) and the value of vps_direct_ref_layer_flag[l][k] is true for an individual OLS, then NumSubLayersInLayerInOLS[i][GeneralLayerIdx[vps_layer_id[k]]] can be derived to max_tid_il_ref_pics_plus[l][k].

[0269] c) The improvement method 4 can be applied when the value of ols_mode_idc is 2.

[0270] Example 1

[0271] In one embodiment, improvements 1 and 2 can be implemented according to the modified VPS syntax shown in Figure 15. The following describes the syntax elements in Figure 15 that are modified relative to the VPS syntax described below.

[0272] A value of 1 for vps_independent_layer_flag[i] may indicate that the layer with index i does not use inter-layer prediction. A value of 0 for vps_independent_layer_flag[i] may indicate that the layer with index i can use inter-layer prediction, and the syntax element vps_direct_ref_layer_flag[i][j] is obtained from the VPS, where j may have a value from 0 to i-1. If the value of vps_independent_layer_flag[i] is not obtained from the bitstream, the value of vps_independent_layer_flag[i] may be set to 1.

[0273] A value of 1 for max_tid_ref_present_flag[i] may indicate that the syntax element max_tid_il_ref_pics_plus1[i][j] is provided in the bitstream. A value of 0 for max_tid_ref_present_flag[i] may indicate that the syntax element max_tid_il_ref_pics_plus1[i][j] is not provided in the bitstream.

[0274] A value of 0 for max_tid_il_ref_pics_plus1[i][j] may indicate that the j-th layer is not used as a reference layer for inter-layer prediction by a non-IRAP picture of the i-th layer. A value greater than 0 for max_tid_il_ref_pics_plus1[i][j] may indicate that a picture having a temporal ID (TemporalId) greater than max_tid_il_ref_pics_plus1[i][j]-1 in the j-th layer is not used as an inter-layer reference picture (ILRP) for decoding a picture of the i-th layer. On the other hand, if the value of max_tid_il_ref_pics_plus1[i][j] is not obtained from the bitstream, the value may be induced to be 7.

[0275] Then, using the syntax elements and variables modified according to the above description, the variables NumOutputLayersInOls[i], NumSubLayersInLayerInOLS[i][j], OutputLayerIdInOls[i][j], and LayerUsedAsOutputLayerFlag[k] can be determined as per the pseudocode of FIG. 16.

[0276] As in the above-described method, for each individual reference layer of a layer, the maximum number of sublayers used in the reference layer can be signaled, and the maximum number of sublayers used in reference layer j for layer i can be provided only if layer j is the direct reference layer of layer i.

[0277] Example 2

[0278] In one embodiment, improvement method 3 may be implemented according to the syntax of the modified VPS shown in Fig. 17. For example, a flag indicating whether or not there is signaling of the maximum number of sublayers used for inter-layer prediction by one layer, such as max_tid_ref_present_flag in Fig. 17, may be signaled for all layers.

[0279] For example, in the modified syntax of the VPS shown in Figure 17, a value of 1 for the syntax element max_tid_ref_present_flag can indicate that the syntax element max_tid_il_ref_pics_plus1[i] is provided in the bitstream, and a value of 0 for max_tid_ref_present_flag can indicate that the syntax element max_tid_il_ref_pics_plus1[i] is not provided in the bitstream.

[0280] Example 3

[0281] The derivation of NumOutputLayersInOls according to the above-described improvement schemes 1, 2, and 4 can be performed according to the pseudocode in Fig. 18. Fig. 18 illustrates pseudocode showing a method for deriving NumOutputLayersInOls for the direct reference layer and the indirect reference layer.

[0282] As shown in FIG. 18, when the value of ols_mode_idc is 2, if the value of each_layer_is_an_ols_flag[i] is true, NumSubLayersInLayerInOLS[i][0] can be set to the value of vps_max_sub_layers_minus1+1, and if the value of each_layer_is_an_ols_flag[i] is false and for an individual OLS, if the value of vps_direct_ref_layer_flag[l][k] is true, NumSubLayersInLayerInOLS[i][GeneralLayerIdx[vps_layer_id[k]]]] can be derived to max_tid_il_ref_pics_plus[l][k].

[0283] Encoding and Decoding Methods

[0284] Hereinafter, an image encoding method and a decoding method performed by an image encoding apparatus and an image decoding apparatus according to an embodiment will be described. Fig. 19 is a flowchart illustrating a method for determining the number of sub-layers of a current layer for an image encoding apparatus to encode an image and / or an image decoding apparatus to decode an image according to an embodiment.

[0285] An image decoding apparatus according to an embodiment includes a memory and a processor, and the decoding apparatus can perform decoding according to the embodiments described below through the operation of the processor. An image encoding apparatus according to an embodiment includes a memory and a processor, and the encoding apparatus can perform encoding according to the operation of the processor in a manner corresponding to the decoding of the decoding apparatus according to the embodiments described below. For convenience of the following description, the operation of the decoding apparatus will be described, but the following description can also be applied to the encoding apparatus.

[0286] In one embodiment, the number of sublayers of the current layer can be used to mean the number of sublayers belonging to the current layer, but in other embodiments, the number of sublayers of the current layer can also be used to mean the number of sublayers required for the current layer.

[0287] A decoding apparatus according to an embodiment may determine whether inter-layer direct reference is used (S1910). Here, whether inter-layer direct reference is used may be determined based on direct reference layer information acquired from a bitstream. The direct reference layer information may indicate whether inter-layer direct reference is used. For example, the direct reference layer information may be the above-mentioned syntax element vps_direct_ref_layer_flag[i][j].

[0288] The direct reference layer information may be obtained from the bitstream for a layer coded based on inter-layer prediction. Here, whether or not the layer is coded based on inter-layer prediction may be determined based on independent layer information obtained from the bitstream. For example, the independent layer information may be the above-mentioned syntax element vps_independent_layer_flag[i].

[0289] Next, the decoding apparatus can determine the number of sub-layers of the current layer based on whether the inter-layer direct reference is performed (S1920).

[0290] Here, the number of sub-layers of the current layer may be determined based on whether the current layer is an output layer. In one embodiment, whether the current layer is an output layer may be determined by "if(each_layer_is_an_ols_flag)" or "if(ols_output_layer_flag[i][k])" as shown in the pseudocode of FIG. 18.

[0291] In one embodiment, depending on whether the current layer is an output layer, the number of sub-layers of the current layer may be determined as the maximum available number of sub-layers, which can be achieved by "NumSubLayersInLayerInOLS[i][j]=vps_max_sub_layers_minus1+1", as shown in the pseudocode of FIG.

[0292] In one embodiment, if the current layer is not an output layer, the number of sub-layers of the current layer may be determined to a predetermined value determined based on whether or not there is direct inter-layer reference.

[0293] Here, the predetermined value may be determined based on maximum identifier information indicating a picture that can be referenced for performing inter-layer prediction. The maximum identifier information may be information indicating that a picture having a temporal identifier greater than a value identified by the maximum identifier information among a plurality of pictures of a first layer is not used as an inter-layer reference picture for decoding a current picture of a second layer. For example, the maximum identifier information may be the above-mentioned syntax element max_tid_il_ref_pics_plus1[i][j].

[0294] In one embodiment, the first layer may be the current layer, and the second layer may be a layer that can be used as a direct reference layer for the current layer.

[0295] For example, the maximum identifier information may be max_tid_il_ref_pics_plus1[l][k] shown in Figure 18. The relationship between the sublayer of the current layer and the layer that can directly reference it can be identified by "if(vps_direct_ref_layer_flag[l][k])" in Figure 18.

[0296] Whether the current layer is an output layer can be determined based on output layer set mode information obtained from the bitstream, where the output layer set mode information may be the above-mentioned syntax element ols_mode_idc.

[0297] In addition, whether the current layer is an output layer can be determined based on the output layer set mode information and output layer flag obtained from the bitstream, where the output layer flag may be the aforementioned syntax element ols_output_layer_flag[i][j].

[0298] Based on the number of sub-layers determined in this manner, the decoding device can decode the current layer by performing inter-layer prediction of the current layer, and the encoding device can encode the current layer by performing inter-layer prediction of the current layer.

[0299] For example, in one embodiment, the image decoding method includes the steps of obtaining a maximum allowable number of layers (e.g., vps_max_layers_minus1) from a bitstream, identifying a current layer having a first index (e.g., i) based on the maximum allowable number, obtaining an independent layer flag (e.g., vps_independent_layer_flag[i]) from the bitstream indicating whether the current layer is coded based on inter-layer prediction, obtaining a maximum temporal identifier signaling flag (e.g., vps_max_tid_ref_present_flag[i]) from the bitstream based on the independent layer flag, and The method may include the steps of: obtaining a direct reference layer flag (e.g., vps_direct_ref_layer_flag[i][j]) from the bitstream based on the first index and the second index, the direct reference layer flag indicating whether the reference layer having the second index is a direct reference layer of the current layer; obtaining maximum temporal identifier information (e.g., vps_max_tid_il_ref_pics_plus1[i][j]) from the bitstream based on the maximum temporal identifier signaling flag and the direct reference layer flag; and determining the number of sub-layers of the reference layer (e.g., NumSubLayersInLayerInOLS[i][k]) based on the maximum temporal identifier information.

[0300] Also, in one embodiment, the image encoding method may include a step of encoding a reference picture of a reference layer, a step of encoding a current picture of a current layer based on the reference picture, and a step of generating a bitstream including encoding information of the current picture.

[0301] Here, an independent layer flag (e.g., vps_independent_layer_flag[i]) indicating whether the current layer is coded based on inter-layer prediction may be included in the bitstream. Furthermore, a maximum temporal identifier signaling flag (e.g., vps_max_tid_ref_present_flag[i]) may be included in the bitstream based on whether the current layer is coded based on inter-layer prediction. In addition, a direct reference layer flag (e.g., vps_direct_ref_layer_flag[i][j]) indicating whether a reference layer is a direct reference layer of the current layer may be included in the bitstream.

[0302] In addition, maximum temporal identifier information (e.g., vps_max_tid_il_ref_pics_plus1[i][j]) can be included in the bitstream based on the maximum temporal identifier signaling flag and the direct reference flag. Furthermore, the number of sub-layers of the reference layer that can be referenced from the current layer (e.g., NumSubLayersInLayerInOLS[i][k]) can be coded based on the maximum temporal identifier information.

[0303] Here, the sub-layer may be referenced from the current layer, and the maximum temporal identifier flag may indicate whether the maximum temporal identifier information is obtained from the bitstream. Furthermore, among pictures of the reference layer, pictures having a temporal identifier value greater than the value identified by the maximum temporal identifier information may not be used as inter-layer reference pictures for decoding the current picture of the current layer.

[0304] Application example

[0305] Although the exemplary method of the present disclosure is expressed as a series of operations for clarity of explanation, this is not intended to limit the order in which the steps are performed, and the steps may be performed simultaneously or in a different order if necessary. To achieve the method according to the present disclosure, the steps illustrated may include other steps, or some steps may be omitted and the remaining steps may be included, or some steps may be omitted and additional other steps may be included.

[0306] In the present disclosure, an image encoding device or an image decoding device that performs a predetermined operation (step) can perform the operation (step) to check the execution conditions or situation of the operation (step). For example, if it is described that a predetermined operation is performed when a predetermined condition is satisfied, the image encoding device or the image decoding device can perform the predetermined operation after performing an operation to check whether the predetermined condition is satisfied.

[0307] The various embodiments of the present disclosure are not intended to enumerate all possible combinations, but are intended to describe representative aspects of the present disclosure, and the matters described in the various embodiments may be applied independently or in combination of two or more.

[0308] Additionally, various embodiments of the present disclosure may be implemented using hardware, firmware, software, or a combination thereof, etc. In the case of a hardware implementation, the implementation may be using one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), general processors, controllers, microcontrollers, microprocessors, etc.

[0309] In addition, an image decoding apparatus and an image encoding apparatus to which an embodiment of the present disclosure is applied may be included in a multimedia broadcast transmitting / receiving apparatus, a mobile communication terminal, a home cinema video apparatus, a digital cinema video apparatus, a surveillance camera, a video conversation apparatus, a real-time communication apparatus such as video communication, a mobile streaming apparatus, a storage medium, a camcorder, a video on demand (VoD) service providing apparatus, an over-the-top (OTT) video apparatus, an internet streaming service providing apparatus, a three-dimensional (3D) video apparatus, an image telephone video apparatus, a medical video apparatus, etc., and may be used to process a video signal or a data signal. For example, an over-the-top (OTT) video apparatus may include a game console, a Blu-ray player, an internet-connected TV, a home theater system, a smartphone, a tablet PC, a digital video recorder (DVR), etc.

[0310] FIG. 20 is a diagram illustrating a content streaming system to which an embodiment of the present disclosure can be applied.

[0311] As shown in FIG. 20, a content streaming system to which an embodiment of the present disclosure is applied can broadly include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.

[0312] The encoding server compresses content input from a multimedia input device such as a smartphone, camera, or camcorder into digital data to generate a bitstream and transmits the bitstream to the streaming server. As another example, if a multimedia input device such as a smartphone, camera, or video camera directly generates a bitstream, the encoding server can be omitted.

[0313] The bitstream can be generated by an image encoding method and / or image encoding device to which an embodiment of the present disclosure is applied, and the streaming server can temporarily store the bitstream during the process of transmitting or receiving the bitstream.

[0314] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server serves as an intermediary for informing the user of available services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, which then transmits the multimedia data to the user. In this case, the content streaming system may include a separate control server, which may control commands and responses between devices in the content streaming system.

[0315] The streaming server may receive content from a media storage and / or an encoding server. For example, when receiving content from the encoding server, the content may be received in real time. In this case, the streaming server may store the bitstream for a certain period of time to provide a smooth streaming service.

[0316] Examples of the user device include a mobile phone, a smartphone, a laptop computer, a digital broadcasting terminal, a personal digital assistant (PDA), a portable multimedia player (PMP), a navigation system, a slate PC, a tablet PC, an ultrabook, a wearable device such as a smartwatch, smart glass, a head mounted display (HMD), a digital TV, a desktop computer, and digital signage.

[0317] Each server in the content streaming system can be operated as a distributed server, in which case data received from each server can be processed in a distributed manner.

[0318] The scope of the present disclosure includes software or machine-executable commands (e.g., operating systems, applications, firmware, programs, etc.) that cause operations according to the methods of various embodiments to be performed on a device or computer, and non-transitory computer-readable medium on which such software or commands can be stored and executed on a device or computer. [Industrial Applicability]

[0319] Embodiments according to the present disclosure can be used to encode / decode images.

Claims

1. An image decoding method performed by an image decoding device, comprising: determining whether a first layer is a direct reference layer to a current layer, the first layer being a layer other than the current layer; determining the number of sub-layers of the current layer based on whether the first layer is the direct reference layer for the current layer; An image decoding method, wherein, based on the current layer not being an output layer, the number of sub-layers of the current layer is set to a predetermined value determined based on whether the first layer is the direct reference layer for the current layer.

2. The image decoding method according to claim 1 , wherein whether the first layer is the direct reference layer for the current layer is determined according to direct reference layer information obtained from a bitstream.

3. The image decoding method according to claim 2 , wherein the direct reference layer information is obtained from the bitstream for a layer that is coded based on inter-layer prediction.

4. The image decoding method according to claim 3 , wherein whether the layer is coded based on the inter-layer prediction is determined based on independent layer information obtained from the bitstream.

5. The image decoding method according to claim 1 , wherein the number of sub-layers of the current layer is determined to be the maximum available number of sub-layers based on whether the current layer is the output layer.

6. The image decoding method according to claim 1 , wherein the predetermined value is determined based on maximum identifier information indicating a picture that can be referenced for performing inter-layer prediction.

7. 7. The image decoding method of claim 6, wherein the maximum identifier information indicates that a picture among the plurality of pictures in the current layer having a temporal identifier greater than the value identified by the maximum identifier information is not used as an inter-layer reference picture for decoding a target picture in a second layer.

8. The image decoding method according to claim 7 , wherein the second layer is a layer that can be used as a direct reference layer for the current layer.

9. The image decoding method according to claim 1 , wherein whether the current layer is the output layer is determined based on output layer set mode information obtained from a bitstream.

10. The image decoding method according to claim 1 , wherein whether the current layer is the output layer is determined based on output layer set mode information and an output layer flag obtained from a bitstream.

11. An image coding method performed by an image coding device, comprising: determining whether a first layer is a direct reference layer to a current layer, the first layer being a layer other than the current layer; determining the number of sub-layers of the current layer based on whether the first layer is the direct reference layer for the current layer; An image encoding method, wherein, based on the current layer not being an output layer, the number of sub-layers of the current layer is determined to be a predetermined value determined based on whether the first layer is the direct reference layer for the current layer.

12. 1. A method for transmitting a bitstream, comprising: determining whether a first layer is a direct reference layer to a current layer, the first layer being a layer other than the current layer; determining the number of sub-layers of the current layer based on whether the first layer is the direct reference layer for the current layer; encoding information in the bitstream indicating whether the first layer is the direct reference layer for the current layer; transmitting the bitstream; A method in which, based on the current layer not being an output layer, the number of sub-layers of the current layer is determined to be a predetermined value determined based on whether the first layer is the direct reference layer for the current layer.