Video decoding method, video encoding method, and method for transmitting image data

By utilizing the sps_video_parameter_set_id syntax element in video decoding and encoding devices to infer the number of video parameter sets and output layer sets in a single-layer bitstream, the problem of efficient encoding of high-resolution images/videos is solved, and the encoding efficiency and compression effect are improved.

CN115552903BActive Publication Date: 2025-09-16LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202180034840.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-05-12
Filing Date
2021-05-11
Publication Date
2025-09-16
Estimated Expiration
2041-05-11

AI Technical Summary

Technical Problem

Existing image/video coding technologies have high transmission and storage costs when processing high-resolution, high-quality images/videos, and it is difficult to effectively compress image/video data with different characteristics.

Method used

Through a video decoding device and an encoding device, the sps_video_parameter_set_id syntax element is used to infer the number of video parameter sets and output layer sets in a single-layer bitstream, realize inter-frame or intra-frame prediction, and efficiently process the layer identifier of the parameter set when the bitstream is a single layer.

Benefits of technology

The coding efficiency of images/videos is improved, the processing capability of single-layer bit streams is enhanced, and the information of the output layer set can be obtained without including the video parameter set, thereby improving the compression efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115552903B_ABST
    Figure CN115552903B_ABST
Patent Text Reader

Abstract

According to the present disclosure, a method for decoding a video by a video decoding device includes the following steps: obtaining image information from a bitstream, the image information including a video coding layer (VCL) network abstraction layer (NAL) unit; performing inter-frame prediction or intra-frame prediction on a current block within a current picture based on the image information to generate prediction samples about the current block; and restoring the current block based on the prediction samples, wherein the image information includes an sps_video_parameter_set_id syntax element indicating the value of an identifier of a video parameter set; and based on the value of the sps_video_parameter_set_id syntax element being equal to 1, the value of the total number of output layer sets (OLS) specified by the video parameter set, and the value of the number of layers within the OLS can be derived as 1.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a method and apparatus for handling references of parameter sets within a single-layer bitstream when encoding / decoding image / video information in an image / video coding system. Background Art

[0002] Recently, the demand for high-resolution, high-quality images / videos (such as 4K, 8K, or even higher ultra-high-definition (UHD) images / videos) is increasing in various fields. As the resolution or quality of images / videos increases, more information or bits are transmitted compared to traditional image / video data. Therefore, if the image / video data is transmitted via a medium such as an existing wired / wireless broadband line or stored in a traditional storage medium, the transmission and storage costs will increase.

[0003] In addition, interest in and demand for virtual reality (VR) and artificial reality (AR) content and immersive media such as holograms is growing; and the broadcasting of images / videos that present image / video characteristics that are different from those of actual images / videos, such as game images / videos, is also growing.

[0004] Therefore, efficient image / video compression technology is needed to effectively compress and transmit, store, or play high-resolution, high-quality images / videos showing various characteristics as described above. Summary of the Invention

[0005] Technical issues

[0006] The technical purpose of the present disclosure is to provide a method and apparatus for improving the encoding efficiency of images / videos.

[0007] Another technical objective of the present disclosure is to provide a method and apparatus for efficiently processing a single-layer bitstream.

[0008] Yet another technical objective of the present disclosure is to provide a method and apparatus for deriving a layer identifier of a parameter set referenced by a VCL NAL unit when the bitstream is a single-layer bitstream.

[0009] Yet another technical objective of the present disclosure is to provide a method and apparatus for deriving information about an output layer set when a bitstream is a single-layer bitstream.

[0010] Technical Solution

[0011] According to an embodiment of the present specification, a video decoding method performed by a video decoding device is provided. The method may include the following steps: obtaining image information including a video coding layer (VCL) network abstraction layer (NAL) unit from a bitstream; generating a prediction sample of the current block by performing inter-frame prediction or intra-frame prediction on a current block in a current picture based on the image information; and reconstructing the current block based on the prediction sample, wherein the image information may include an sps_video_parameter_set_id syntax element indicating an identifier value of a video parameter set; and wherein, based on the value of the sps_video_parameter_set_id syntax element being equal to 0, the value of the total number of output layer sets (OLSs) specified by the video parameter set and the value of the number of layers within the OLS can be inferred to be equal to 1.

[0012] According to another embodiment of the present specification, a video encoding method performed by a video encoding device is provided herein. The method may include the following steps: performing inter-frame prediction or intra-frame prediction on a current block in a current picture; generating prediction information for the current block based on the inter-frame prediction or the intra-frame prediction; and encoding image information including the prediction information, wherein the image information may include a video coding layer (VCL) network abstraction layer (NAL) unit and an sps_video_parameter_set_id syntax element indicating an identifier value of a video parameter set; and wherein, based on the value of the sps_video_parameter_set_id syntax element being equal to 0, the total number of output layer sets (OLS) specified by the video parameter set and the number of layers within the OLS may be inferred to be equal to 1.

[0013] According to another embodiment of the present specification, a computer-readable digital recording medium is provided herein, which stores information that enables a video decoding device to perform a video decoding method, wherein the video decoding method may include the following steps: obtaining image information including a video coding layer VCL network abstraction layer NAL unit; generating prediction samples of the current block by performing inter-frame prediction or intra-frame prediction on a current block within a current picture based on the image information; and reconstructing the current block based on the prediction samples, wherein the image information may include an sps_video_parameter_set_id syntax element indicating an identifier value of a video parameter set; and wherein, based on the value of the sps_video_parameter_set_id syntax element being equal to 0, the value of the total number of output layer sets OLS specified by the video parameter set, and the value of the number of layers within the OLS can be inferred to be equal to 1.

[0014] Technical Effects

[0015] According to the embodiments of the present disclosure, the overall compression efficiency of images / videos can be enhanced.

[0016] According to the embodiments of the present disclosure, a single-layer bitstream can be efficiently handled.

[0017] According to an embodiment of the present disclosure, when the bitstream is a single-layer bitstream, a layer identifier of a parameter set referenced by a VCL NAL unit may be derived.

[0018] According to an embodiment of the present disclosure, constraints corresponding to a single-layer bitstream may be provided.

[0019] According to an embodiment of the present disclosure, even when a bitstream is a single-layer bitstream that does not include a video parameter set, information about an output layer set can be derived. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 An example of a video / image coding system to which embodiments of this document may be applied is schematically shown.

[0021] Figure 2 FIG2 is a diagram schematically showing a configuration of a video / image encoding device to which the embodiments of this document can be applied.

[0022] Figure 3 FIG. 1 is a diagram for schematically describing a configuration of a multi-layer based video / image encoding apparatus to which an embodiment of the present disclosure can be applied.

[0023] Figure 4 is a diagram for schematically explaining the configuration of a video / image decoding device to which an embodiment of the present disclosure can be applied.

[0024] Figure 5 FIG. 1 is a diagram for schematically describing a configuration of a multi-layer based video / image decoding apparatus to which an embodiment of the present disclosure can be applied.

[0025] Figure 6 A schematic example of a picture decoding process to which embodiments of the present disclosure may be applied is shown.

[0026] Figure 7 A schematic example of a picture encoding process to which embodiments of the present disclosure may be applied is shown.

[0027] Figure 8 The layered structure for the encoded image / video is shown as an example.

[0028] Figure 9 and Figure 10 General examples of video / image encoding methods and related components according to embodiments of the present disclosure are respectively shown.

[0029] Figure 11 and Figure 12 General examples of video / image decoding methods and related components according to embodiments of the present disclosure are respectively shown.

[0030] Figure 13 An example of a content streaming system to which embodiments of the present disclosure can be applied is shown. DETAILED DESCRIPTION

[0031] The disclosure of the present disclosure can be modified in various forms, and its specific embodiments will be described and illustrated in the accompanying drawings. The terms used in the present disclosure are only used to describe specific embodiments and are not intended to limit the methods disclosed in the present disclosure. The expression of a single number includes the expression of "at least one", as long as it is clearly read differently. Terms such as "including" and "having" are intended to indicate the presence of features, numbers, steps, operations, elements, components or combinations thereof used in the disclosure, and therefore should be understood that the possibility of the presence or addition of one or more different features, numbers, steps, operations, elements, components or combinations thereof is not excluded.

[0032] The present disclosure relates to video / image coding. For example, the methods / implementations disclosed in the present disclosure may be applied to methods disclosed in the Versatile Video Coding (VVC) standard. In addition, the methods / implementations disclosed in the present disclosure may be applied to methods disclosed in the Essential Video Coding (EVC) standard, the AO Media Video 1 (AV1) standard, the second-generation Audio Video Coding standard (AVS2), or next-generation video / image coding standards (e.g., H.267, H.268, etc.).

[0033] Various embodiments related to video / image coding are presented in the present disclosure, and unless otherwise stated, the embodiments can be combined with each other.

[0034] In addition, the various configurations of the drawings described in this disclosure are independent illustrations for explaining the functions of the features that are different from each other, and do not mean that the various configurations are implemented by different hardware or different software. For example, two or more configurations in a configuration can be combined to form one configuration, and one configuration can also be divided into multiple configurations. Without departing from the gist of the disclosed method of the present disclosure, embodiments of combining and / or separating configurations are included within the scope of the disclosure of the present disclosure.

[0035] In the present disclosure, the terms " / " and "," should be interpreted as indicating "and / or". For example, the expression "A / B" may mean "A and / or B". In addition, "A, B" may mean "A and / or B". In addition, "A / B / C" may mean "at least one of A, B, and / or C". In addition, "A / B / C" may mean "at least one of A, B, and / or C".

[0036] Furthermore, in the present disclosure, the word "or" should be interpreted as indicating "and / or". For example, the expression "A or B" may include: 1) only A, 2) only B, and / or 3) both A and B. In other words, the term "or" in the present disclosure should be interpreted as indicating "additionally or alternatively".

[0037] Furthermore, brackets used in this disclosure may mean "for example." Specifically, the phrase "prediction (intra-frame prediction)" may be used to indicate that "intra-frame prediction" is used as an example of "prediction." In other words, the term "prediction" in this disclosure is not limited to "intra-frame prediction," and "intra-frame prediction" is used as an example of "prediction." Furthermore, even when the phrase "prediction (i.e., intra-frame prediction)" is used, it may be used to indicate that "intra-frame prediction" is used as an example of "prediction."

[0038] In the present disclosure, technical features explained individually in one drawing may be implemented individually or simultaneously.

[0039] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. In addition, throughout the drawings, the same reference numerals are used to indicate the same elements, and the same description of the same elements may be omitted.

[0040] Figure 1 An example of a video / image encoding system to which the embodiments of the present disclosure can be applied is illustrated.

[0041] Reference Figure 1 The video / image coding system may include a first device (source device) and a second device (receiver). The source device may send coded video / image information or data to the receive device in the form of a file or stream via a digital storage medium or a network.

[0042] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.

[0043] The video source can obtain the video / image through a process of capturing, synthesizing, or generating the video / image. The video source may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, a video / image archive including previously captured videos / images, etc. The video / image generation device may include, for example, a computer, a tablet computer, and a smartphone, and may (electronically) generate the video / image. For example, a virtual video / image may be generated by a computer, etc. In this case, the video / image capture process may be replaced by a process of generating relevant data.

[0044] The encoding device can encode the input video / image. The encoding device can perform a series of processes such as prediction, transformation and quantization to achieve compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.

[0045] The transmitter can transmit the encoded image / image information or data in the form of a bitstream to a receiver of a receiving device via a digital storage medium or network in the form of a file or stream. Digital storage media can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter can include components for generating a media file in a predetermined file format and can also include components for transmitting via a broadcast / communication network. The receiver can receive / extract the bitstream and transmit the received bitstream to a decoding device.

[0046] The decoding device may decode the video / image by performing a series of processes such as dequantization, inverse transformation, and prediction corresponding to the operations of the encoding device.

[0047] The renderer can render the decoded video / image, and the rendered video / image can be displayed on a display.

[0048] In the present disclosure, video may refer to a series of images over time. A picture generally refers to a unit representing an image of a specific time frame, and a slice / tile refers to a unit that constitutes a part of a picture in terms of coding. A slice / tile may include one or more coding tree units (CTUs). A picture may be composed of one or more slices / tiles. A picture may be composed of one or more tile groups. A tile group may include one or more tiles. A tile may represent a rectangular area of ​​a CTU row within a tile in a picture. A tile may be divided into multiple tiles, each of which is composed of one or more CTU rows within the tile. A tile that is not divided into multiple tiles may also be referred to as a tile. Tile scanning is a specific sequential ordering of CTUs of a partitioned picture, where CTUs are ordered consecutively in a CTU raster scan within a tile, tiles within a tile are ordered consecutively in a raster scan of tiles within the tile, and tiles in a picture are ordered consecutively in a raster scan of tiles within the picture. A tile is a rectangular area of ​​CTUs within a specific tile column and a specific tile row in a picture. A tile column is a rectangular area of ​​multiple CTUs with a height equal to the height of the picture and a width specified by a syntax element in the picture parameter set. A tile scan is a specific sequential ordering of the CTUs that partition a picture, where the CTUs are ordered consecutively in a raster scan of the CTUs in the tile, and the tiles in the picture are ordered consecutively in a raster scan of the tiles of the picture. A slice includes an integer number of tiles of a picture that can be included exclusively in a single NAL unit. A slice can consist of multiple complete tiles or a consecutive sequence of only complete tiles of a tile. In this disclosure, tile group and slice can be used interchangeably. For example, in this disclosure, a tile group / tile group header can be referred to as a slice / slice header.

[0049] A pixel or a picture element (pel) may refer to the smallest unit constituting a picture (or image). In addition, "sample" may be used as a term corresponding to a pixel. A sample may generally represent a pixel or a pixel value, and may represent only a pixel / pixel value of a luminance component, or only a pixel / pixel value of a chrominance component.

[0050] A unit may represent a basic unit of image processing. A unit may include at least one of a specific area of ​​a picture and information related to the area. A unit may include a luminance block and two chrominance (e.g., cb, cr) blocks. In some cases, units may be used interchangeably with terms such as blocks or areas. In general, an M×N block may include M columns and N rows of samples (or sample arrays) or a set (or array) of transform coefficients. Alternatively, a sample may represent a pixel value in the spatial domain, and when such a pixel value is transformed into the frequency domain, it may represent a transform coefficient in the frequency domain.

[0051] In some cases, a unit may be used interchangeably with terms such as a block or region. Generally, an M×N block may represent a sample consisting of M columns and N rows or a set of transform coefficients. A sample may generally represent a pixel or a pixel value, and may also represent only a pixel / pixel value of a luma component, or only a pixel / pixel value of a chroma component. A sample may be used as an item corresponding to a pixel or picture element configuring a picture (or image).

[0052] Figure 2 FIG2 is a schematic diagram illustrating a configuration of a video / image encoding device to which an embodiment of the present disclosure can be applied. Hereinafter, a video encoding device may include an image encoding device.

[0053] Reference Figure 2 , the encoding device 200 may include and be configured with an image segmenter 210, a predictor 220, a residual processor 230 and an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may also include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstructed block generator. According to an embodiment, the image segmenter 210, the predictor 220, the residual processor 230, the entropy encoder 240, the adder 250, and the filter 260 described above may be composed of at least one hardware component (e.g., an encoder chipset or processor). In addition, the memory 270 may include a decoded picture buffer (DPB) or may be composed of a digital storage medium. The hardware components may also include the memory 270 as an internal / external component.

[0054] The image splitter 210 can split the input image (or picture or frame) input to the encoding device 200 into one or more processing units. For example, a processing unit can be referred to as a coding unit (CU). In this case, the coding unit can be recursively split from the coding tree unit (CTU) or the largest coding unit (LCU) according to a quadtree binary tree ternary tree (QTBTTT) structure. For example, a coding unit can be split into multiple coding units of greater depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. In this case, for example, the quadtree structure can be applied first, followed by the binary tree structure and / or the ternary tree structure. Alternatively, the binary tree structure can be applied first. The encoding process according to the present disclosure can be performed based on the final coding unit that is no longer split. In this case, the maximum coding unit can be used as the final coding unit based on coding efficiency according to image characteristics, or if necessary, the coding unit can be recursively split into coding units of greater depth and the coding unit with the optimal size can be used as the final coding unit. Here, the encoding process may include prediction, transformation, and reconstruction processes, which will be described later. As another example, the processing unit may also include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may be separated or partitioned from the final coding unit. The prediction unit may be a unit for sample prediction, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.

[0055] In the encoding device 200, the prediction signal (prediction block, prediction sample array) output from the inter-frame predictor 221 or the intra-frame predictor 222 can be subtracted from the input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is sent to the transformer 232. In this case, as shown in the figure, the unit in the encoding device 200 for subtracting the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) can be referred to as a subtractor 231. The predictor 220 can perform prediction on a block to be processed (hereinafter referred to as the current block) and generate a prediction block including prediction samples for the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction on a per-block or CU basis. As described later in the description of each prediction mode, the predictor 220 can generate various information related to the prediction, such as prediction mode information, and send the generated information to the entropy encoder 240. The information about the prediction can be encoded in the entropy encoder 240 and output in the form of a bitstream.

[0056] The intra-frame predictor 222 can predict the current block by referring to samples in the current picture. Depending on the prediction mode, the referenced samples can be located near the current block or far away from the current block. In intra-frame prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. The non-directional mode can include, for example, a DC mode and a planar mode. Depending on the level of detail of the prediction direction, the directional mode can include, for example, 33 directional prediction modes or 65 directional prediction modes. However, this is only an example, and more or fewer directional prediction modes can be used depending on the settings. The intra-frame predictor 222 can also determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring block.

[0057] The inter-frame predictor 221 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector in a reference picture. To reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can also include information about the inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, neighboring blocks can include spatially neighboring blocks in the current picture and temporally neighboring blocks in a reference picture. The reference picture containing the reference block and the reference picture containing the temporally neighboring block can be the same or different. Temporally neighboring blocks can be referred to as collocated reference blocks, collocated CUs (colCUs), etc., and the reference picture containing temporally neighboring blocks can be referred to as collocated pictures (colPics). For example, the inter-frame predictor 221 can configure a motion information candidate list based on the neighboring blocks and generate information indicating which candidate to use to derive the motion vector and / or reference picture index for the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the inter-frame predictor 221 can use the motion information of the neighboring block as the motion information of the current block. In skip mode, unlike merge mode, it may not be possible to send a residual signal. In the case of motion vector prediction (MVP) mode, the motion vector of the current block can be indicated by using the motion vector of the neighboring block as a motion vector predictor and signaling the motion vector difference.

[0058] The predictor 220 can generate a prediction signal based on various prediction methods described below. For example, the predictor 220 can not only apply intra prediction or inter prediction to predict a block, but can also apply both intra prediction and inter prediction at the same time. This can be called inter-frame intra-frame combined prediction (CIIP). In addition, the predictor can predict the block based on the intra block copy (IBC) prediction mode or based on the palette mode. The IBC prediction mode or palette mode can be used for image / video encoding of content such as games, for example, screen content coding (SCC). IBC basically performs prediction in the current picture, but can be performed similarly to inter prediction because the reference block is derived in the current picture. That is, IBC can use at least one of the inter prediction techniques described in this disclosure. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, the sample value within the picture can be signaled based on information about the palette table and the palette index.

[0059] The prediction signal generated by the predictor (including the inter-frame predictor 221 and / or the intra-frame predictor 222) may be used to generate a reconstructed signal or may be used to generate a residual signal.

[0060] The transformer 232 can generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a graph-based transform (GBT), or a conditional nonlinear transform (CNT). Here, GBT refers to a transform obtained from a graph when the relationship information between pixels is represented by a graph. CNT refers to a transform generated based on a prediction signal generated using all previously reconstructed pixels. In addition, the transform process can be applied to square pixel blocks of the same size, or can be applied to blocks of variable size rather than square.

[0061] The quantizer 233 may quantize the transform coefficients and transmit them to the entropy encoder 240, and the entropy encoder 240 may encode the quantized signal (information about the quantized transform coefficients) and output a bitstream. The information about the quantized transform coefficients may be referred to as residual information. The quantizer 233 may rearrange the quantized transform coefficients in block form into a one-dimensional vector form based on the coefficient scanning order, and generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form.

[0062] The entropy encoder 240 can implement various encoding methods, such as, for example, Exponential Golomb, Context-Adaptive Variable Length Coding (CAVLC), Context-Adaptive Binary Arithmetic Coding (CABAC), and the like. The entropy encoder 240 can encode information required for video / image reconstruction (e.g., syntax element values, etc.), in addition to quantized transform coefficients, together or separately. The encoded information (e.g., encoded video / image information) can be transmitted or stored in units of NALs (Network Abstraction Layers) in the form of a bitstream. The video / image information may also include information regarding various parameter sets, such as an Adaptation Parameter Set (APS), a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), or a Video Parameter Set (VPS). Furthermore, the video / image information may also include general constraint information. In the present disclosure, information and / or syntax elements transmitted / signaled from the encoding device to the decoding device may be included in the video / image information. The video / image information may be encoded through the above-described encoding process and included in the bitstream. The bitstream may be transmitted over a network or stored on a digital storage medium. The network may include a broadcast network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD. A transmission unit (not shown) that transmits a signal output from the entropy encoder 240 and / or a storage unit (not shown) that stores the signal may be included as an internal / external element of the encoding device 200, and alternatively, the transmitter may be included in the entropy encoder 240.

[0063] The quantized transform coefficients output from the quantizer 233 can be used to generate a prediction signal. For example, the residual signal (residual block or residual sample) can be reconstructed by applying dequantization and inverse transform to the quantized transform coefficients using the dequantizer 234 and the inverse transformer 235. The adder 250 adds the reconstructed residual signal to the prediction signal output from the inter-frame predictor 221 or the intra-frame predictor 222 to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). If the block to be processed has no residual (such as when skip mode is applied), the prediction block can be used as a reconstructed block. The adder 250 can be called a reconstruction unit or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current picture, and can be used for inter-frame prediction of the next picture through filtering as described below.

[0064] Furthermore, during the picture encoding and / or reconstruction process, luma mapping and chroma scaling (LMCS) may be applied.

[0065] The filter 260 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 260 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and store the modified reconstructed picture in the memory 270 (specifically, the DPB of the memory 270). Various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filter 260 can generate various information related to filtering and send the generated information to the entropy encoder 240, as described later in the description of various filtering methods. The information related to filtering can be encoded by the entropy encoder 240 and output in the form of a bit stream.

[0066] The modified reconstructed picture sent to the memory 270 may be used as a reference picture in the inter-frame predictor 221. When inter-frame prediction is applied by the encoding apparatus, prediction mismatch between the encoding apparatus 200 and the decoding apparatus may be avoided, and encoding efficiency may be improved.

[0067] The DPB of the memory 270 may store a modified reconstructed picture used as a reference picture in the inter-frame predictor 221. The memory 270 may store motion information of a block from which motion information in the current picture is derived (or encoded) and / or motion information of a reconstructed block in the picture. The stored motion information may be sent to the inter-frame predictor 221 and used as motion information of spatially neighboring blocks or motion information of temporally neighboring blocks. The memory 270 may store reconstructed samples of a reconstructed block in the current picture and may transmit the reconstructed samples to the intra-frame predictor 222.

[0068] In addition, the image / video coding according to the present disclosure may include image / video coding based on multiple layers. Image / video coding based on multiple layers may include scalable coding. Multi-layer coding or scalable coding can process input signals of various layers. The input signal (input image / picture) may differ in at least one of resolution, frame rate, bit depth, color format, aspect ratio, and view depending on the layer. In this case, repeated transmission / processing of information can be reduced and compression efficiency can be improved by performing prediction between layers using the difference between the layers (i.e., based on scalability).

[0069] Figure 3 is a block diagram of an encoding apparatus that performs multi-layer based encoding of a video / image signal according to an embodiment of the present disclosure.

[0070] Figure 3 The encoding device may include Figure 2 The encoding device. Figure 3In the embodiment, the image splitter and the adder are omitted. However, the encoding device may include the image splitter and the adder. In this case, the image splitter and the adder may be included in units of layers. The present disclosure mainly describes prediction based on multiple layers. For other descriptions, please refer to Figure 2 Description given.

[0071] For ease of description, Figure 3 The example of assuming a multilayer structure consisting of two layers. However, the embodiments of the present disclosure are not limited to the specific examples, and it should be noted that the multilayer structure to which the embodiments of the present disclosure are applied may include two or more layers.

[0072] Reference Figure 3 , the encoding apparatus 200 includes an encoder 200 - 1 for layer 1 and an encoder 200 - 0 for layer 0 .

[0073] Layer 0 can be a base layer, a reference layer, or a lower layer; layer 1 can be an enhancement layer, a current layer, or a higher layer.

[0074] The encoder 200-1 of layer 1 includes a predictor 220-1, a residual processor 230-1, a filter 260-1, a memory 270-1, an entropy encoder 240-1, and a multiplexer (MUX) 270. The MUX may be included as an external component.

[0075] The encoder 200 - 0 of layer 0 includes a predictor 220 - 0 , a residual processor 230 - 0 , a filter 260 - 0 , a memory 270 - 0 , and an entropy encoder 240 - 0 .

[0076] The predictors 220-0 and 220-1 can perform prediction on the input image based on the various prediction techniques described above. For example, the predictors 220-0 and 220-1 can perform inter-frame prediction and intra-frame prediction. The predictors 220-0 and 220-1 can perform prediction in predetermined processing units. The prediction unit can be a coding unit (CU) or a transform unit (TU). A prediction block (including prediction samples) can be generated based on the prediction result, and the residual processor can derive a residual block (including residual samples) based on the prediction block.

[0077] With inter-frame prediction, a prediction block may be generated by performing prediction based on information about at least one of a previous picture and / or a subsequent picture of a current picture. With intra-frame prediction, a prediction block may be generated by performing prediction based on neighboring samples within the current picture.

[0078] The various prediction modes and methods described above can be used in inter-frame prediction modes or methods. Inter-frame prediction can select a reference picture relative to the current block to be predicted and a reference block within the reference picture related to the current block. Predictors 220-0 and 220-1 can generate a prediction block based on the reference block.

[0079] Also, the predictor 220-1 may perform prediction on layer 1 using information of layer 0. In the present disclosure, for convenience of description, a method of predicting information of a current layer using information of another layer is referred to as inter-layer prediction.

[0080] The information of the current layer predicted based on information of another layer (ie, predicted by inter-layer prediction) includes at least one of texture, motion information, unit information, and predetermined parameters (eg, filtering parameters).

[0081] In addition, the information of another layer used for prediction of the current layer (ie, for inter-layer prediction) may include at least one of texture, motion information, unit information, and predetermined parameters (eg, filtering parameters).

[0082] In inter-layer prediction, the current block may be a block within the current picture of the current layer (e.g., layer 1) and may be a target block to be encoded. The reference block may be a block within a picture (reference picture) in a layer (reference layer, e.g., layer 0) referenced for prediction of the current block, belonging to the same access unit (AU) as the picture to which the current block belongs (current picture), and may be a block corresponding to the current block. Here, the access unit may be a group of picture units (PUs) including coded pictures associated with the same temporal output from different layers and DPBs. A picture unit may be a group of NAL units that are related to each other according to a specific classification rule, are continuous in decoding order, and contain only one coded picture. A coded video sequence (CVS) may be a group of AUs.

[0083] One example of inter-layer prediction is inter-layer motion prediction, which uses the motion information of a reference layer to predict the motion information of a current layer. According to inter-layer motion prediction, the motion information of a current block can be predicted based on the motion information of a reference block. In other words, when deriving motion information based on an inter-frame prediction mode (to be described later), motion information candidates can be derived using the motion information of an inter-layer reference block rather than the motion information of a temporally neighboring block.

[0084] When inter-layer motion prediction is applied, the predictor 220 - 1 may scale and use reference block (ie, inter-layer reference block) motion information of a reference layer.

[0085] In another example of inter-layer prediction, inter-layer texture prediction can use the texture of the reconstructed reference block as the prediction value of the current block. In this case, the predictor 220-1 can scale the texture of the reference block by upsampling. Inter-layer texture prediction can be referred to as inter-layer (reconstructed) sample prediction or simply inter-layer prediction.

[0086] In inter-layer parameter prediction (yet another example of inter-layer prediction), parameters derived from a reference layer may be reused in the current layer, or parameters of the current layer may be derived based on parameters used in the reference layer.

[0087] In inter-layer residual prediction (still another example of inter-layer prediction), the residual of the current layer may be predicted using residual information of another layer, and prediction of the current block may be performed based on the predicted residual.

[0088] In inter-layer differential prediction, which is yet another example of inter-layer prediction, prediction of the current block may be performed using a difference between images obtained by upsampling or downsampling a reconstructed picture of a current layer and a reconstructed picture of a reference layer.

[0089] In inter-layer syntax prediction (another example of inter-layer prediction), the texture of the current block can be predicted or generated using syntax information of the reference layer. In this case, the syntax information of the referenced reference layer may include information about intra prediction mode and motion information.

[0090] When predicting a specific block, multiple prediction methods using inter-layer prediction can use multiple layers

[0091] Here, as examples of inter-layer prediction, inter-layer texture prediction, inter-layer motion prediction, inter-layer unit information prediction, inter-layer parameter prediction, inter-layer residual prediction, inter-layer differential prediction, and inter-layer syntax prediction have been described; however, the inter-layer prediction applicable to the present disclosure is not limited to the above examples.

[0092] For example, inter-layer prediction may be applied as an extension of inter-frame prediction of the current layer. In other words, inter-frame prediction of the current block may be performed by including a reference picture derived from a reference layer in a reference picture that may be referenced for inter-frame prediction of the current block.

[0093] In this case, the inter-layer reference picture may be included in the reference picture list of the current block.Using the inter-layer reference picture, the predictor 220-1 may perform inter prediction on the current block.

[0094] Here, the inter-layer reference picture may be a reference picture constructed by sampling a reconstructed picture of a reference layer to correspond to the current layer. Therefore, when the reconstructed picture of the reference layer corresponds to a picture of the current layer, the reconstructed picture of the reference layer can be used as an inter-layer reference picture without sampling. For example, when the width and height of the sample in the reconstructed picture of the reference layer are the same as the width and height of the sample in the reconstructed picture of the current layer; and the offset between the upper left, upper right, lower left, and lower right of the picture of the reference layer and the upper left, upper right, lower left, and lower right of the picture of the current layer is 0, the reconstructed picture of the reference layer can be used as an inter-layer reference picture of the current layer without resampling.

[0095] In addition, the reconstructed picture of the reference layer used to derive the inter-layer reference picture may be a picture belonging to the same AU as the current picture to be encoded.

[0096] When inter-frame prediction of the current block is performed by including an inter-layer reference picture in a reference picture list, the positions of the inter-layer reference pictures in the reference picture lists L0 and L1 may be different from each other. For example, in the case of the reference picture list L0, the inter-layer reference picture may be located after the short-term reference picture before the current picture, and in the case of the reference picture list L1, the inter-layer reference picture may be located at the end of the reference picture list.

[0097] Here, reference picture list L0 is a reference picture list used for inter prediction of P slices or a reference picture list used as a first reference picture list in inter prediction of B slices. Reference picture list L1 is a second reference picture list used for inter prediction of B slices.

[0098] Therefore, the reference picture list L0 can be composed in the order of short-term reference pictures before the current picture, inter-layer reference pictures, short-term reference pictures after the current picture, and long-term reference pictures. The reference picture list L1 can be composed in the order of short-term reference pictures after the current picture, short-term reference pictures before the current picture, long-term reference pictures, and inter-layer reference pictures.

[0099] At this time, a prediction slice (P slice) is a slice on which intra prediction is performed or inter prediction is performed using up to one motion vector and reference picture index per prediction block. A bi-prediction slice (B slice) is a slice on which intra prediction is performed or prediction is performed using up to two motion vectors and reference picture indexes per prediction block. In this regard, an intra slice (I slice) is a slice to which only intra prediction is applied.

[0100] In addition, when inter prediction of the current block is performed based on a reference picture list including an inter-layer reference picture, the reference picture list may include a plurality of inter-layer reference pictures derived from a plurality of layers.

[0101] When the reference picture list includes multiple inter-layer reference pictures, the inter-layer reference pictures can be arranged alternately in the reference picture lists L0 and L1. For example, assuming that the reference picture list for inter-frame prediction of the current block includes two inter-layer reference pictures (inter-layer reference pictures ILRP i and inter-layer reference pictures ILRP j ). In this case, in the reference picture list L0, ILRP i It can be located after the short-term reference picture before the current picture, and ILRP j Can be located at the end of the list. Also, in the reference picture list L1, ILRP i can be at the end of the list, and ILRP j It can be located after the short-term reference picture after the current picture.

[0102] In this case, the reference picture list L0 can be a short-term reference picture before the current picture, an inter-layer reference picture ILRP, or a reference picture list L1. i , short-term reference pictures after the current picture, long-term reference pictures and inter-layer reference pictures ILRP j The reference picture list L1 can be composed of short-term reference pictures after the current picture, inter-layer reference pictures ILRP j , short-term reference pictures before the current picture, long-term reference pictures and inter-layer reference pictures ILRP i The order of composition.

[0103] In addition, one of the two inter-layer reference pictures may be an inter-layer reference picture derived from a resolution-dependent scalable layer, and the other may be an inter-layer reference picture derived from a layer providing a different view. In this case, for example, assuming ILRP i is an inter-layer reference picture derived from layers providing different resolutions, and ILRP j It is an inter-layer reference picture derived from a layer providing a different view. Then, in the case of scalable video coding that only supports scalability other than views, the reference picture list L0 can be a short-term reference picture before the current picture, an inter-layer reference picture ILRP, and a reference picture list L1. i , short-term reference pictures after the current picture, and long-term reference pictures. On the other hand, the reference picture list L1 can be composed of the short-term reference pictures after the current picture, the short-term reference pictures before the current picture, the long-term reference pictures, and the inter-layer reference pictures ILRP. j The order of composition.

[0104] In addition, for inter-layer prediction, the information of the inter-layer reference picture can be composed of only sample values, only motion information (motion vector), or both sample values ​​and motion information. When the reference picture index indicates an inter-layer reference picture, the predictor 220-1 uses only the sample values ​​of the inter-layer reference picture, the motion information (motion vector) of the inter-layer reference picture, or both the sample values ​​and motion information of the inter-layer reference picture according to the information received from the encoding device.

[0105] When only the sample values ​​of the inter-layer reference picture are used, the predictor 220-1 may derive the sample of the block specified by the motion vector in the inter-layer reference picture as the prediction sample of the current block. In the case of scalable video coding regardless of the view, the motion vector in inter-frame prediction (inter-layer prediction) using the inter-layer reference picture may be set to a fixed value (e.g., 0).

[0106] When only the motion information of the inter-layer reference picture is used, the predictor 220-1 can use the motion vector specified in the inter-layer reference picture as a motion vector predictor for deriving the motion vector of the current block. In addition, the predictor 220-1 can use the motion vector specified in the inter-layer reference picture as the motion vector of the current block.

[0107] When using both samples and motion information of an inter-layer reference picture, the predictor 220 - 1 may predict the current block using samples related to the current block in the inter-layer reference picture and motion information (motion vector) specified in the inter-layer reference picture.

[0108] When inter-layer prediction is applied, the encoding device may send a reference index indicating an inter-layer reference picture within a reference picture list to the decoding device, and also send information specifying which information (sample information, motion information, or sample information and motion information) from the inter-layer reference picture to use to the decoding device (i.e., information specifying a dependency type related to inter-layer prediction between two layers).

[0109] Figure 4 FIG. 1 is a diagram for schematically illustrating a configuration of a video / image decoding device to which an embodiment of the present disclosure can be applied.

[0110] Reference Figure 4 , the decoding device 300 may include and be configured with an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-frame predictor 332 and an intra-frame predictor 331. The residual processor 320 may include a dequantizer 321 and an inverse transformer 322. According to an embodiment, the entropy decoder 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 described above may be composed of hardware components (e.g., a decoder chipset or processor). In addition, the memory 360 may include a decoded picture buffer (DPB), or may be composed of a digital storage medium. The hardware component may also include the memory 360 as an internal / external component.

[0111] When a bit stream including video / image information is input, the decoding apparatus 300 can be used with Figure 2The image is reconstructed accordingly to the processing of the video / image information in the encoding device. For example, the decoding device 300 can derive the unit / block based on the block segmentation related information obtained from the bit stream. The decoding device 300 can perform decoding using the processing unit applied in the encoding device. Therefore, the processing unit of decoding can be, for example, a coding unit, and the coding unit can be divided from the coding tree unit or the maximum coding unit according to the quadtree structure, the binary tree structure and / or the ternary tree structure. One or more transform units can be derived from the coding unit. The reconstructed image signal decoded and output by the decoding device 300 can be reproduced by a reproduction device.

[0112] The decoding device 300 may receive the data in the form of a bit stream from Figure 2 The signal output by the encoding device can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can parse the bitstream to derive information required for image reconstruction (or picture reconstruction) (e.g., video / image information). The video / image information may also include information about various parameter sets such as an adaptive parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may also include general constraint information. The decoding device may also decode the picture based on the information about the parameter set and / or the general constraint information. The signaled / received information and / or syntax elements described later in this disclosure can be decoded through a decoding process and obtained from the bitstream. For example, the entropy decoder 310 decodes the information in the bitstream based on a coding method such as exponential Golomb coding, context-adaptive variable length coding (CAVLC), or context-adaptive arithmetic coding (CABAC), and outputs syntax elements required for image reconstruction and quantized values ​​of the transform coefficients of the residual. More specifically, the CABAC entropy decoding method can receive a bin corresponding to each syntax element in the bitstream, use the decoding target syntax element information, the decoding information of the decoding target block, or the information of the symbol / bin decoded in the previous stage to determine the context model, and arithmetically decode the bin by predicting the probability of occurrence of the bin according to the determined context model, and generate a symbol corresponding to the value of each syntax element. In this case, after determining the context model, the CABAC entropy decoding method can update the context model by using the information of the decoded symbol / bin for the context model of the next symbol / bin. The information related to the prediction among the information decoded by the entropy decoder 310 can be provided to the predictor (inter-frame predictor 332 and intra-frame predictor 331), and the residual value (that is, the quantized transform coefficient and related parameter information) on which entropy decoding is performed in the entropy decoder 310 can be input to the residual processor 320.

[0113] The residual processor 320 can derive a residual signal (residual block, residual sample, residual sample array). In addition, information about filtering among the information decoded by the entropy decoder 310 can be provided to the filter 350. In addition, a receiving unit (not shown) for receiving a signal output from the encoding device can be further configured as an internal / external element of the decoding device 300, or the receiving unit can be a component of the entropy decoder 310. In addition, the decoding device according to the present disclosure can be called a video / image / picture decoding device, and the decoding device can be classified into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include an entropy decoder 310, and the sample decoder may include at least one of a dequantizer 321, an inverse transformer 322, an adder 340, a filter 350, a memory 360, an inter-frame predictor 332, and an intra-frame predictor 331.

[0114] The dequantizer 321 can dequantize the quantized transform coefficients and output the transform coefficients. The dequantizer 321 can rearrange the quantized transform coefficients in the form of two-dimensional blocks. In this case, the rearrangement can be performed based on the coefficient scanning order performed in the encoding device. The dequantizer 321 can dequantize the quantized transform coefficients by using quantization parameters (e.g., quantization step size information) and obtain transform coefficients.

[0115] The inverse transformer 322 inversely transforms the transform coefficients to obtain a residual signal (residual block, residual sample array). In the present disclosure, at least one of quantization / dequantization and / or transform / inverse transform may be omitted. When quantization / dequantization is omitted, the quantized transform coefficients may be referred to as transform coefficients. When transform / inverse transform is omitted, the transform coefficients may be referred to as coefficients or residual coefficients, or may still be referred to as transform coefficients for consistency of expression.

[0116] In the present disclosure, quantized transform coefficients and transform coefficients may be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, residual information may include information about the transform coefficients, and the information about the transform coefficients may be signaled via residual coding syntax. The transform coefficients may be derived based on the residual information (or information about the transform coefficients), and the scaled transform coefficients may be derived by inverse transforming (scaling) the transform coefficients. Residual samples may be derived based on the inverse transform (transform) of the scaled transform coefficients. This may also be applied / expressed in other parts of the present disclosure.

[0117] The predictor 330 may perform prediction on the current block and generate a prediction block including prediction samples of the current block. The predictor may determine whether to apply intra prediction or inter prediction to the current block based on the prediction information output from the entropy decoder 310, and may determine a specific intra / inter prediction mode.

[0118] The predictor 330 can generate a prediction signal based on various prediction methods described below. For example, the predictor can not only apply intra prediction or inter prediction to predict a block, but also apply intra prediction and inter prediction at the same time. This can be called inter and intra combined prediction (CIIP). In addition, the predictor can predict the block based on the intra block copy (IBC) prediction mode or palette mode. The IBC prediction mode or palette mode can be used for content image / video coding of games, etc., for example, screen content coding (SCC). IBC basically performs prediction in the current picture, but can be performed similarly to inter prediction because the reference block is derived in the current picture. That is, IBC can use at least one of the inter prediction techniques described in this disclosure. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, the sample value within the picture can be signaled based on information about the palette table and the palette index.

[0119] The intra-frame predictor 331 can predict the current block by referencing samples in the current picture. Depending on the prediction mode, the referenced samples may be located near the current block or may be located far away from the current block. In intra-frame prediction, the prediction mode may include multiple non-directional modes and multiple directional modes. The intra-frame predictor 331 may determine the prediction mode to be applied to the current block by using the prediction modes applied to the neighboring blocks.

[0120] The inter-frame predictor 332 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector in a reference picture. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may also include information on the inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, the neighboring blocks may include spatial neighboring blocks in the current picture and temporal neighboring blocks in the reference picture. For example, the inter-frame predictor 332 may configure a motion information candidate list based on the neighboring blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and the prediction information may include information indicating the inter-frame prediction mode for the current block.

[0121] The adder 340 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor (including the inter-frame predictor 332 and / or the intra-frame predictor 331). If the block to be processed has no residual (for example, when skip mode is applied), the prediction block can be used as the reconstructed block.

[0122] The adder 340 may be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal may be used for intra-frame prediction of the next block to be processed in the current picture, may be output through filtering as described below, or may be used for inter-frame prediction of the next picture.

[0123] In addition, luminance mapping and chroma scaling (LMCS) can be applied during the picture decoding process.

[0124] The filter 350 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 350 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image, and store the modified reconstructed image in the memory 360 (specifically, the DPB of the memory 360). Various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.

[0125] The (modified) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter-frame predictor 332. The memory 360 can store the motion information of the block from which the motion information in the current picture is derived (or decoded) and / or the motion information of the reconstructed block in the picture. The stored motion information can be sent to the inter-frame predictor 332 to be used as the motion information of the spatially neighboring block or the motion information of the temporally neighboring block. The memory 360 can store the reconstructed samples of the reconstructed block in the current picture and can transmit the reconstructed samples to the intra-frame predictor 331.

[0126] In the present disclosure, the embodiments described in the filter 260 , the inter predictor 221 , and the intra predictor 222 of the encoding apparatus 200 may be the same as or correspond to the filter 350 , the inter predictor 332 , and the intra predictor 331 .

[0127] Figure 5 FIG. 1 is a diagram for schematically describing a configuration of a multi-layer based video / image decoding apparatus to which an embodiment of the present disclosure can be applied.

[0128] Figure 5 The decoding device may include Figure 4 The decoding device. Figure 5The rearranger can be omitted or included in the dequantizer. The figure will be described mainly in terms of multi-layer based prediction. The remainder may include Figure 4 The description content.

[0129] For ease of description, Figure 5 The example of assuming a multilayer structure consisting of two layers. However, the embodiments of the present disclosure are not limited to the specific examples, and it should be noted that the multilayer structure to which the embodiments of the present disclosure are applied may include two or more layers.

[0130] Reference Figure 5 , the decoding apparatus 500 includes a decoder 500 - 1 for layer 1 and a decoder 500 - 0 for layer 0 .

[0131] The decoder 500 - 1 of layer 1 may include an entropy decoder 510 - 1 , a residual processor 520 - 1 , a predictor 530 - 1 , an adder 540 - 1 , a filter 550 - 1 , and a memory 560 - 1 .

[0132] The decoder 500 - 0 of layer 0 may include an entropy decoder 510 - 0 , a residual processor 520 - 0 , a predictor 530 - 0 , an adder 540 - 0 , a filter 550 - 0 , and a memory 560 - 0 .

[0133] When a bitstream including image information is transmitted from the encoding device, the demultiplexer 505 may demultiplex information of each layer and deliver the information to the decoding device for each layer.

[0134] The entropy decoders 510-1 and 510-0 may perform decoding according to the encoding method used in the encoding apparatus. For example, when CABAC is used in the encoding apparatus, the entropy decoders 510-1 and 510-0 may also perform entropy decoding based on CABAC.

[0135] When the prediction mode for the current block is the intra prediction mode, the predictors 530 - 1 , 530 - 0 may perform intra prediction on the current block based on neighboring reconstructed samples within the current picture.

[0136] When the prediction mode for the current block is the inter prediction mode, the predictors 530-1 and 530-0 may perform inter prediction on the current block based on information included in at least one of a picture before the current picture or a picture after the current picture. Information received from the encoding device may be checked, and part or all of the motion information required for inter prediction may be derived based on the checked information.

[0137] When the skip mode is applied as the inter prediction mode, the residual may not be transmitted from the encoding apparatus, and the prediction block may be used as the reconstructed block.

[0138] Also, the predictor 530 - 1 of layer 1 may perform inter prediction or intra prediction using only information within layer 1 , or may perform inter-layer prediction using information of another layer (layer 0 ).

[0139] The information of the current layer predicted using information of a different layer (ie, predicted by inter-layer prediction) includes at least one of texture, motion information, unit information, and predetermined parameters (eg, filtering parameters).

[0140] In addition, the information of the different layers used for prediction of the current layer (ie, for inter-layer prediction) may include at least one of texture, motion information, unit information, and predetermined parameters (eg, filtering parameters).

[0141] In inter-layer prediction, the current block may be a block in the current picture of the current layer (e.g., layer 1) and may be a target block to be decoded. The reference block may be a block in a picture (reference picture) in the layer referenced for prediction of the current block (reference layer, e.g., layer 0) that belongs to the same access unit (AU) as the picture to which the current block belongs (current picture) and may be a block corresponding to the current block.

[0142] One example of inter-layer prediction is inter-layer motion prediction, which uses the motion information of a reference layer to predict the motion information of a current layer. According to inter-layer motion prediction, the motion information of a current block can be predicted based on the motion information of a reference block. In other words, when deriving motion information based on an inter-frame prediction mode (to be described later), motion information candidates can be derived using the motion information of an inter-layer reference block rather than the motion information of a temporally neighboring block.

[0143] When inter-layer motion prediction is applied, the predictor 530 - 1 may scale and use reference block (ie, inter-layer reference block) motion information of a reference layer.

[0144] In another example of inter-layer prediction, inter-layer texture prediction can use the texture of the reconstructed reference block as the prediction value of the current block. In this case, the predictor 530-1 can scale the texture of the reference block by upsampling. Inter-layer texture prediction can be called inter-layer (reconstructed) sample prediction or simply inter-layer prediction.

[0145] In inter-layer parameter prediction (yet another example of inter-layer prediction), parameters derived from a reference layer may be reused in the current layer, or parameters of the current layer may be derived based on parameters used in the reference layer.

[0146] In inter-layer residual prediction (still another example of inter-layer prediction), the residual of the current layer may be predicted using residual information of another layer, and prediction of the current block may be performed based on the predicted residual.

[0147] In inter-layer differential prediction, which is yet another example of inter-layer prediction, prediction of the current block may be performed using a difference between images obtained by upsampling or downsampling a reconstructed picture of a current layer and a reconstructed picture of a reference layer.

[0148] In inter-layer syntax prediction (another example of inter-layer prediction), the texture of the current block can be predicted or generated using syntax information of the reference layer. In this case, the syntax information of the referenced reference layer may include information about intra prediction mode and motion information.

[0149] When predicting a specific block, multiple prediction methods using inter-layer prediction can use multiple layers

[0150] Here, as examples of inter-layer prediction, inter-layer texture prediction, inter-layer motion prediction, inter-layer unit information prediction, inter-layer parameter prediction, inter-layer residual prediction, inter-layer differential prediction, and inter-layer syntax prediction have been described; however, the inter-layer prediction applicable to the present disclosure is not limited to the above examples.

[0151] For example, inter-layer prediction can be applied as an extension of inter-frame prediction of the current layer. In other words, inter-frame prediction of the current block can be performed by including a reference picture derived from a reference layer in a reference picture that can be referenced for inter-frame prediction of the current block.

[0152] When the reference picture index received from the encoding device or the reference picture index derived from the neighboring block indicates an inter-layer reference picture in the reference picture list, the predictor 530-1 can use the inter-layer reference picture to perform inter-layer prediction. For example, when the reference picture index indicates an inter-layer reference picture, the predictor 530-1 can derive the sample value of the area specified by the motion vector in the inter-layer reference picture as the prediction block of the current block.

[0153] In this case, the inter-layer reference picture may be included in the reference picture list of the current block.Using the inter-layer reference picture, the predictor 530-1 may perform inter prediction on the current block.

[0154] Here, the inter-layer reference picture may be a reference picture constructed by sampling a reconstructed picture of a reference layer to correspond to the current layer. Therefore, when the reconstructed picture of the reference layer corresponds to a picture of the current layer, the reconstructed picture of the reference layer can be used as an inter-layer reference picture without sampling. For example, when the width and height of the sample in the reconstructed picture of the reference layer are the same as the width and height of the sample in the reconstructed picture of the current layer; and the offset between the upper left, upper right, lower left, and lower right of the picture of the reference layer and the upper left, upper right, lower left, and lower right of the picture of the current layer is 0, the reconstructed picture of the reference layer can be used as an inter-layer reference picture of the current layer without resampling.

[0155] In addition, the reconstructed picture of the reference layer used to derive the inter-layer reference picture may be a picture belonging to the same AU as the current picture to be encoded. When inter-frame prediction of the current block is performed by including the inter-layer reference picture in the reference picture list, the positions of the inter-layer reference picture in the reference picture lists L0 and L1 may be different from each other. For example, in the case of reference picture list L0, the inter-layer reference picture may be located after the short-term reference picture preceding the current picture, and in the case of reference picture list L1, the inter-layer reference picture may be located at the end of the reference picture list.

[0156] Here, reference picture list L0 is a reference picture list used for inter prediction of P slices or a reference picture list used as a first reference picture list in inter prediction of B slices. Reference picture list L1 is a second reference picture list used for inter prediction of B slices.

[0157] Therefore, the reference picture list L0 can be composed in the order of short-term reference pictures before the current picture, inter-layer reference pictures, short-term reference pictures after the current picture, and long-term reference pictures. The reference picture list L1 can be composed in the order of short-term reference pictures after the current picture, short-term reference pictures before the current picture, long-term reference pictures, and inter-layer reference pictures.

[0158] At this time, a prediction slice (P slice) is a slice on which intra prediction is performed or inter prediction is performed using up to one motion vector and reference picture index per prediction block. A bi-prediction slice (B slice) is a slice on which intra prediction is performed or prediction is performed using up to two motion vectors and reference picture indexes per prediction block. In this regard, an intra slice (I slice) is a slice to which only intra prediction is applied.

[0159] In addition, when inter prediction of the current block is performed based on a reference picture list including an inter-layer reference picture, the reference picture list may include a plurality of inter-layer reference pictures derived from a plurality of layers.

[0160] When the reference picture list includes multiple inter-layer reference pictures, the inter-layer reference pictures can be arranged alternately in the reference picture lists L0 and L1. For example, assuming that two inter-layer reference pictures (inter-layer reference pictures ILRP i and inter-layer reference pictures ILRP j ) is included in the reference picture list for inter-frame prediction of the current block. In this case, in the reference picture list L0, ILRP i It can be located after the short-term reference picture before the current picture, and ILRP j Can be located at the end of the list. Also, in the reference picture list L1, ILRP i can be at the end of the list, and ILRP jIt can be located after the short-term reference picture after the current picture.

[0161] In this case, the reference picture list L0 can be a short-term reference picture before the current picture, an inter-layer reference picture ILRP, or a reference picture list L1. i , short-term reference pictures after the current picture, long-term reference pictures and inter-layer reference pictures ILRP j The reference picture list L1 can be composed of short-term reference pictures after the current picture, inter-layer reference pictures ILRP j , short-term reference pictures before the current picture, long-term reference pictures and inter-layer reference pictures ILRP i The order of composition.

[0162] In addition, one of the two inter-layer reference pictures may be an inter-layer reference picture derived from a resolution-dependent scalable layer, and the other may be an inter-layer reference picture derived from a layer providing a different view. In this case, for example, assuming ILRP i is an inter-layer reference picture derived from layers providing different resolutions, and ILRP j It is an inter-layer reference picture derived from a layer providing a different view. Then, in the case of scalable video coding that only supports scalability other than views, the reference picture list L0 can be a short-term reference picture before the current picture, an inter-layer reference picture ILRP, and a reference picture list L1. i , short-term reference pictures after the current picture, and long-term reference pictures. On the other hand, the reference picture list L1 can be composed of the short-term reference pictures after the current picture, the short-term reference pictures before the current picture, the long-term reference pictures, and the inter-layer reference pictures ILRP. j The order of composition.

[0163] In addition, for inter-layer prediction, the information of the inter-layer reference picture can be composed of only sample values, only motion information (motion vector), or both sample values ​​and motion information. When the reference picture index indicates an inter-layer reference picture, the predictor 530-1 uses only the sample values ​​of the inter-layer reference picture, the motion information (motion vector) of the inter-layer reference picture, or both the sample values ​​and motion information of the inter-layer reference picture according to the information received from the encoding device.

[0164] When only the sample values ​​of the inter-layer reference picture are used, the predictor 530-1 may derive the sample of the block specified by the motion vector in the inter-layer reference picture as the prediction sample of the current block. In the case of scalable video coding regardless of the view, the motion vector in inter-frame prediction (inter-layer prediction) using the inter-layer reference picture may be set to a fixed value (e.g., 0).

[0165] When only the motion information of the inter-layer reference picture is used, the predictor 530-1 can use the motion vector specified in the inter-layer reference picture as a motion vector predictor for deriving the motion vector of the current block. In addition, the predictor 530-1 can use the motion vector specified in the inter-layer reference picture as the motion vector of the current block.

[0166] When using both samples and motion information of an inter-layer reference picture, the predictor 530 - 1 may predict the current block using samples related to the current block in the inter-layer reference picture and motion information (motion vector) specified in the inter-layer reference picture.

[0167] The decoding device may receive a reference index indicating an inter-layer reference picture within a reference picture list from the encoding device, and perform inter-layer prediction based on the received reference index. In addition, the decoding device may receive information specifying which information (sample information, motion information, or both sample information and motion information) to use from the inter-layer reference picture, that is, information specifying a dependency type related to inter-layer prediction between two layers, from the encoding device.

[0168] In addition, in the video / image encoding according to the present disclosure, the image processing unit may have a hierarchical structure. A picture may be divided into one or more tiles, tiles, slices, and / or tile groups. A slice may include one or more tiles. A tile may include one or more CTU rows within a tile. A slice may include an integer number of tiles of a picture. A tile group may include one or more tiles. A tile may include one or more CTUs. A CTU may be divided into one or more CUs. A tile represents a rectangular area of ​​a CTU within a specific tile column and a specific tile row in a picture. A tile group may include an integer number of tiles raster scanned according to the tiles in the picture. A slice header may carry information / parameters that may be applicable to the corresponding slice (tile in the slice). When the encoding / decoding device has a multi-core processor, the encoding / decoding process for tiles, slices, tiles, and / or tile groups may be processed in parallel. In the present disclosure, slices or tile groups may be used interchangeably. That is, the tile group header may be referred to as a slice header. Here, the slice may have one of the slice types including intra (I) slice, prediction (P) slice, and bi-prediction (B) slice. When predicting a tile in an I slice, inter prediction may not be used, and only intra prediction may be used. Of course, even in this case, notification can be performed by encoding the original sample value without prediction. With respect to tiles in a P slice, intra prediction or inter prediction may be used, and in the case of using inter prediction, only uni-prediction may be used. In addition, with respect to tiles in a B slice, intra prediction or inter prediction may be used, and in the case of using inter prediction, up to bi-prediction may be used to the maximum extent.

[0169] The encoding device can determine the block / block group, tile, slice, and maximum and minimum coding unit sizes taking into account coding efficiency or parallel processing or according to the characteristics of the video image (e.g., resolution), and can include its information used for the content or information capable of obtaining the content in the bitstream.

[0170] The decoding device can obtain information indicating whether the tiles / tile groups, tiles, slices, and CTUs in the current picture are divided into multiple coding units. Efficiency can be improved by only obtaining (sending) such information under specific conditions.

[0171] In addition, as described above, a picture may include multiple slices, and a slice may include a slice header and slice data. In this case, a picture header may be further added to the multiple slices (slice header and slice data set) in a picture. The picture header (picture header syntax) may include information / parameters that are generally applicable to the picture. The slice header (slice header syntax) may include information / parameters that may be commonly applied to the slice. The APS (APS syntax) or PPS (PPS syntax) may include information / parameters that may be commonly applied to one or more slices or pictures. The SPS (SPS syntax) may include information / parameters that may be commonly applied to one or more sequences. The VPS (VPS syntax) may include information / parameters that may be commonly applied to multiple layers. The DCI may include information / parameters related to decoding capabilities.

[0172] The high-level syntax (HLS) in this specification may include at least one of APS syntax, PPS syntax, SPS syntax, VPS syntax, DCI syntax, picture header syntax, and slice header syntax.

[0173] In addition, for example, information on partitioning and configuration of tiles / tile groups / tiles / slices and the like may be configured in the encoding device based on a high-level syntax and then delivered (or transmitted) to the decoding device in a bitstream format.

[0174] Figure 6 A schematic example of a picture decoding process to which embodiments of the present disclosure may be applied is shown.

[0175] In image / video encoding, pictures configuring an image / video can be encoded / decoded according to a decoding order. A picture order corresponding to the output order of decoded pictures can be configured differently from the decoding order. Furthermore, when performing inter-frame prediction based on the configured picture order, forward prediction as well as backward prediction can be performed.

[0176] exist Figure 6 In the S600, the above Figure 4S600 may include the information decoding process described in this specification, S610 may include the inter / intra prediction process described in this specification, S620 may include the residual processing process described in this specification, S630 may include the block / picture reconstruction process described in this specification, and S640 may include the loop filtering process described in this specification.

[0177] Reference Figure 6 , as mentioned above Figure 4 As described in [ 6 ], the picture decoding process generally includes a process of obtaining image / video information from a bitstream (via decoding) ( S600 ), a picture reconstruction process ( S610 to S630 ), and a loop filtering process ( S640 ) for reconstructing the picture. The picture reconstruction process can be performed based on prediction samples and residual samples obtained by performing an inter / intra prediction process ( S610 ) and a residual processing (or processing) process ( S620 , a process of dequantizing and inverse transforming quantized transform coefficients). By performing a loop filtering process on the reconstructed picture generated by the picture reconstruction process, a modified reconstructed picture can be generated. The modified reconstructed picture can be output as a decoded picture and then stored in a decoded picture buffer or memory 360 of the decoding device for use as a reference picture during the inter-frame prediction process when decoding the picture in a later process. In some cases, the loop filtering process can be skipped. In this case, the reconstructed picture can be output as a decoded picture and then stored in a decoded picture buffer or memory 360 of the decoding device for use as a reference picture during the inter-frame prediction process when decoding the picture in a later process. As described above, the loop filtering process (S640) may include a deblocking filtering process, a sample adaptive offset (SAO) process, an adaptive loop filter (ALF) process, and / or a bilateral filter process, and may skip part or all of the loop filtering process. In addition, one or part of the deblocking filtering process, the sample adaptive offset (SAO) process, the adaptive loop filter (ALF) process, and the bilateral filter process may be applied sequentially, or all of the deblocking filtering process, the sample adaptive offset (SAO) process, the adaptive loop filter (ALF) process, and the bilateral filter process may be applied sequentially. For example, the SAO process may be performed after the deblocking filtering process is applied to the reconstructed picture. Alternatively, for example, the ALF process may be performed after the deblocking filtering process is applied to the reconstructed picture. This may also be performed in the encoding device.

[0178] Figure 7A schematic example of a picture encoding process to which embodiments of the present disclosure may be applied is shown.

[0179] exist Figure 7 In the S700, the above Figure 2 S700 may include the inter / intra prediction process described in this specification, S710 may include the residual processing process described in this specification, and S720 may include the information encoding process described in this specification.

[0180] Reference Figure 7 , as mentioned above Figure 2 As described in [ 701 ], the picture encoding process generally includes a process of encoding information used for picture reconstruction (e.g., prediction information, residual information, segmentation information, etc.) and outputting the encoded information in the form of a bitstream, a process of generating a reconstructed picture for the current picture, and a process of applying loop filtering to the reconstructed picture (optional). The encoding device can derive residual samples (which are modified) from the quantized transform coefficients through the dequantizer 234 and the inverse transformer 235, and then the encoding device can generate a reconstructed picture based on the prediction samples and the (modified) residual samples as output of S700. The reconstructed picture generated as described above can be the same as the above-mentioned reconstructed picture generated in the decoding device. The modified reconstructed picture can be generated by performing a loop filtering process on the reconstructed picture, and the modified reconstructed picture is then stored in the decoded picture buffer or memory 270 of the decoding device. Moreover, as in the decoding device, the modified reconstructed picture can be used as a reference picture during the inter-frame prediction process when encoding the picture. As described above, in some cases, part or all of the loop filtering process can be skipped. When performing a loop filtering process, (loop) filtering related information (parameters) can be encoded in the entropy encoder 240 and then sent in the form of a bit stream, and the decoding device can perform a loop filtering process based on the filtering related information by using the same method as the encoding device.

[0181] By performing the above-described loop filtering process, noise (such as blocking artifacts and ringing artifacts) that occurs when encoding images / moving pictures can be reduced, and subjective / objective visual quality can be enhanced. In addition, by having both the encoding device and the decoding device perform the loop filtering process, the encoding device and the decoding device can obtain the same prediction result, increasing the reliability of picture encoding and reducing the size (or amount) of data to be transmitted for picture encoding.

[0182] As described above, the picture reconstruction process can be performed in a decoding device as well as an encoding device. A reconstructed block can be generated for each block unit based on intra prediction / inter prediction, and a reconstructed picture including the reconstructed block can be generated. When the current picture / slice / patchwork group is an I picture / slice / patchwork group, the blocks included in the current picture / slice / patchwork group can be reconstructed based only on intra prediction. In addition, when the current picture / slice / patchwork group is a P or B picture / slice / patchwork group, the blocks included in the current picture / slice / patchwork group can be reconstructed based on intra prediction or inter prediction. In this case, inter prediction can be applied to part of the blocks within the current picture / slice / patchwork group, and intra prediction can be applied to the remaining blocks. The color components of the picture may include a luminance component and a chrominance component. Furthermore, unless explicitly limited (or constrained) in this specification, the methods and embodiments proposed in this specification may be applied to the luminance component and the chrominance component.

[0183] Figure 8 The layered structure for the encoded image / video is shown as an example.

[0184] Reference Figure 8 The coded image / video is divided into the VCL (Video Coding Layer) responsible for the image / video decoding process and itself, the subsystem for sending and storing the coded information, and the Network Abstraction Layer (NAL) that exists between the VCL and the subsystem and is responsible for the network adaptation function.

[0185] VCL can generate VCL data including compressed image data (slice data), or generate parameter sets including picture parameter sets (picture parameter set: PPS), sequence parameter sets (sequence parameter set: SPS), video parameter sets (video parameter set: VPS), etc., or supplementary enhancement information (SEI) messages additionally required for the image decoding process.

[0186] In NAL, a NAL unit can be generated by adding header information (NAL unit header) to a raw byte sequence payload (RBSP) generated in the VCL. In this case, the RBSP refers to slice data, parameter sets, SEI messages, etc. generated in the VCL. The NAL unit header may include NAL unit type information, which is specified based on the RBSP data included in the corresponding NAL unit.

[0187] As shown in the figure, NAL units can be divided into VCL NAL units and non-VCL NAL units according to the RBSP generated in the VCL. A VCL NAL unit may refer to a NAL unit including information about an image (slice data), and a non-VCL NAL unit may refer to a NAL unit containing information required for decoding an image (parameter set or SEI message).

[0188] The VCL NAL unit and the non-VCL NAL unit described above can be sent over a network by attaching header information according to the data standard of the subsystem. For example, the NAL unit can be converted into a predetermined standard data form (such as H.266 / VVC file format, Real-time Transport Protocol (RTP), Transport Stream (TS), etc.) and sent over various networks.

[0189] As described above, in a NAL unit, a NAL unit type may be specified according to an RBSP data structure included in a corresponding NAL unit, and information about the NAL unit type may be stored and signaled in a NAL unit header.

[0190] For example, NAL units can be roughly classified into VCL NAL unit types and non-VCL NAL unit types depending on whether the NAL unit includes information about the image (slice data). VCL NAL unit types can be classified according to the properties and types of pictures included in the VCL NAL unit, and non-VCL NAL unit types can be classified according to the type of parameter set.

[0191] The following are examples of NAL unit types specified according to the types of parameter sets included in non-VCL NAL unit types.

[0192] -DCI (Decoding Capability Information) NAL unit: Type of NAL unit containing DCI

[0193] -VPS (Video Parameter Set) NAL unit: Type of NAL unit containing VPS

[0194] -SPS (Sequence Parameter Set) NAL unit: the type of NAL unit that includes the SPS

[0195] -PPS (Picture Parameter Set) NAL unit: Type of NAL unit including PPS

[0196] -APS (Adaptation Parameter Set) NAL unit: type of NAL unit including APS

[0197] - PH (Picture Header) NAL unit: the type of NAL unit including PH

[0198] The NAL unit type described above has syntax information for the NAL unit type, and the syntax information can be stored and signaled in the NAL unit header. For example, the syntax information can be nal_unit_type, and the NAL unit type can be specified by the nal_unit_type value.

[0199] In addition, as described above, a picture may include multiple slices, and a slice may include a slice header and slice data. In this case, a picture header may be further added to the multiple slices (slice header and slice data set) in a picture. The picture header (picture header syntax) may include information / parameters that are commonly applicable to the picture. The slice header (slice header syntax) may include information / parameters that are commonly applicable to the slice. The APS (APS syntax) or PPS (PPS syntax) may include information / parameters that are commonly applicable to one or more slices or pictures. The SPS (SPS syntax) may include information / parameters that are commonly applicable to one or more sequences. The VPS (VPS syntax) may include information / parameters that are commonly applicable to multiple layers. The DCI (DCI syntax) may include information / parameters related to decoding capabilities.

[0200] In this specification, high-level syntax (HLS) may include at least one of APS syntax, PPS syntax, SPS syntax, VPS syntax, DCI syntax, picture header syntax, and slice header syntax. In addition, in this specification, low-level syntax (LLS) may include, for example, slice data syntax, CTU syntax, transform unit syntax, etc.

[0201] In this specification, image / video information encoded from an encoding device to a decoding device and then signaled in a bitstream format may include not only information related to intra-picture segmentation, intra-frame / inter-frame prediction information, residual information, loop filtering information, etc., but also information from slice headers, picture headers, APS information, PPS information, SPS information, VPS information, and / or DCI information. In addition, the image / video information may further include general constraint information and / or NAL unit header information.

[0202] In addition, as described above, the video / image information of the present specification may include high-layer signaling, and the video / image encoding method may be performed based on the video / image information.

[0203] A coded picture may include one or more slices. Parameters describing the coded picture may be signaled in a picture header, and parameters describing the slice may be signaled in a slice header. The picture header (PH) is carried in its own NAL unit type. The slice header is present at the beginning of a NAL unit that includes a slice payload (slice data).

[0204] In addition, the coded picture may include slices of another NAL unit type.A picture shall refer to a picture parameter set including the mixed_nalu_type_in_pic_flag syntax element.

[0205] If the mixed_nalu_type_in_pic_flag value is equal to 1, this indicates that each picture of the referenced PPS has one or more VCL NAL units, the VCL NAL units do not have the same value as nal_unit_type, and the picture is not an IRAP picture. And, if the mixed_nalu_type_in_pic_flag value is equal to 0, this indicates that each picture of the referenced PPS has one or more VCL NAL units, and the VCL NAL units of each picture of the referenced PPS have the same value as nal_unit_type.

[0206] When the value of no_mixed_nalu_type_in_pic_constraint_flag is equal to 1, the value of mixed_nalu_type_in_pic_flag is equal to 0.

[0207] For each slice with a nal_unit_type value of nalUnitTypeA, in the range from IDR_W_RADL to CRA_NUT, in a picture picA including one or more slices with another value of nal_unit_type (ie, the mixed_nalu_type_in_pic_flag value of picture picA is equal to 1), the following holds.

[0208] - The slice shall belong to sub-picture subpicA with corresponding subpic_treat_as_pic_flag[i] value equal to 1.

[0209] - A slice shall not belong to a sub-picture of picA that includes VCL NAL units with nal_unit_type different from nalUnitTypeA.

[0210] - For all PUs in the following PU within a coding layer video sequence (CLVS), the RefPicList[0] or RefPicList[1] of the slice within subpicA shall not include pictures that precede picA in decoding order within valid entries.

[0211] Additionally, the following applies to VCL NAL units for specific pictures.

[0212] - If the value of mixed_nalu_type_in_pic_flag is equal to 0, the nal_unit_type value shall be the same for all coded slice NAL units within a picture or PU with the same NAL unit type as the coded slice NAL units of that picture or PU.

[0213] Otherwise (if mixed_nalu_type_in_pic_flag value is equal to 1), one or more VCL NAL units shall all have a specific value of nal_unit_type in the range of IDR_W_RADL to CRA_NUT, and other VCL NAL units shall all have a specific value of nal_unit_type in the range of TRAIL_NUT to RSV_VCL_6.

[0214] In the following, the signaling of multiple layers of information within a video parameter set will be described in detail.

[0215] The following shows available layer sets (output layer set (OLS)), profile, tier and level (PTL), information about OLS, DPB information, HRD information, etc.) that can be decoded for a multi-layer bitstream.

[0216] [Table 1]

[0217]

[0218]

[0219] The VPS RBSP of Table 1 should be available to the decoding process before being referenced, and the VPS RBSP should include at least one AU having a temporary identifier (temporalId) equal to 0 or should be provided through external means.

[0220] VPSNAL units within a coded video sequence (CVS) that each have vps_video_parameter_set_id with a specific value shall all have the same content.

[0221] vps_video_parameter_set_id provides an identifier for the VPS so that it can be referenced by another syntax element. The vps_video_parameter_set_id value should be greater than 0.

[0222] vps_max_layers_minus1+1 indicates the maximum number of allowed layers within each CVS of the reference VPS.

[0223] vps_max_sublayer_minus1 plus 1 indicates the maximum number of temporal sublayers that can be present in a layer within each CVS of the reference VPS. The value of vps_max_sublayer_minus1 should be in the range of 0 to 6.

[0224] When the vps_all_layers_same_num_sublayer_flag value is equal to 1, this indicates that the number of temporal sublayers is the same for all layers within each CVS of the reference VPS. And, when the vps_all_layers_same_num_sublayer_flag value is equal to 0, this indicates that the layers within each CVS of the reference VPS may or may not have the same number of temporal sublayers. When the vps_all_layers_same_num_sublayer_flag syntax element is not present in the VPS syntax, the vps_all_layers_same_num_sublayer_flag value is inferred to be equal to 1 (or is deduced to 1).

[0225] If the value of vps_all_independent_layers_flag is equal to 1, this indicates that all layers within the CVS are independently coded without using inter-layer prediction. If the vps_all_independent_layers_flag value is equal to 1, this indicates that one or more layers within the CVS may use inter-layer prediction. When the vps_all_independent_layers_flag syntax element is not present in the VPS syntax, the vps_all_independent_layers_flag value is inferred to be equal to 1.

[0226] vps_layer_id[i] indicates the nuh_layer_id value of layer i. For two non-negative integer values ​​between m and n, where m is less than n, the vps_layer_id[m] value shall be less than the vps_layer_id[n] value.

[0227] When the vps_independent_layer_flag[i] value is equal to 1, this indicates that the layer with index i does not use inter-layer prediction. And, when the vps_independent_layer_flag[i] value is equal to 0, this indicates that the layer with index i can use inter-layer prediction, and the syntax element vps_direct_ref_layer_flag[i][j] is present in the VPS, j ranges from 0 to i-1 (inclusive). When the vps_independent_layer_flag syntax element is not present in the VPS syntax, the vps_independent_layer_flag value is inferred to be equal to 1.

[0228] When the vps_max_tid_ref_present_flag[i] value is equal to 1, this indicates that the syntax element vps_max_tid_il_ref_pics_plus1[i][j] is present. And, when the vps_max_tid_ref_present_flag[i] value is equal to 0, this indicates that the syntax element vps_max_tid_il_ref_pics_plus1[i][j] is not present.

[0229] If the vps_direct_ref_layer_flag[i][j] value is equal to 0, this indicates that the layer with index j is not the direct reference layer of the layer with index i. And, if the vps_direct_ref_layer_flag[i][j] value is equal to 1, this indicates that the layer with index j is the direct reference layer of the layer with index i. If vps_direct_ref_layer_flag[i][j] is not present for i and j in the range from 0 to vps_max_layers_minus1, the value of vps_direct_ref_layer_flag[i][j] is inferred to be equal to 0. When the vps_independent_layer_flag[i] value is equal to 0, one or more values ​​of j in the range from 0 to i-1 (inclusive) should be present, such that vps_direct_ref_layer_flag[i][j] is allowed to be equal to 1.

[0230] As described below, NumDirectRefLayers[i], DirectRefLayerIdx[i][d], NumRefLayers[i], RefLayerIdx[i][r], and LayerUsedAsRefLayerFlag[j] are derived (or inferred).

[0231] If the vps_max_tid_il_ref_pics_plus1[i][j] value is equal to 0, this indicates that a picture of the jth layer that is neither an IRAP picture nor a GDR picture with a ph_recovery_poc_cnt value of 0 is not used as an ILRP for picture decoding of the i-th layer picture. When the value of vps_max_tid_il_ref_pics_plus1[i][j] is greater than 0, this indicates that a picture of the jth layer with a TemporalId greater than vps_max_tid_il_ref_pics_plus1[i][j]-1 is not used as an ILRP when decoding the i-th layer picture. When the vps_max_tid_il_ref_pics_plus1 syntax element is not present in the VPS, the vps_max_tid_il_ref_pics_plus1[i][j] value is inferred to be equal to vps_max_sublayer_minus1+1.

[0232] If the vps_each_layer_is_an_ols_flag value is equal to 1, this indicates that each OLS includes only one layer, and the layer included in each layer within the CVS of the reference VPS is an OLS, which is the only output layer. If vps_each_layer_is_an_ols_flag is equal to 0, this indicates that one or more OLSs include two or more layers. If the vps_max_layers_minus1 value is equal to 0, the value of vps_each_layer_is_an_ols_flag can be inferred to be equal to 1. Otherwise, if vps_all_independent_layers_flag is equal to 0, the value of vps_each_layer_is_an_ols_flag can be inferred to be equal to 0.

[0233] If the value of vps_ols_mode_idc is equal to 0, this indicates that the total number of OLSs indicated by the VPS is equal to vps_max_layers_minus1+1, and this also indicates that the i-th OLS includes layers each having a layer index ranging from 0 to i, and for each OLS only the highest layer within the OLS is the output layer.

[0234] If the value of vps_ols_mode_idc is equal to 1, this indicates that the total number of OLSs indicated by the VPS is equal to vps_max_layers_minus1+1, and this also indicates that the i-th OLS includes layers each having a layer index ranging from 0 to i, and all layers of the OLS are output layers for each OLS.

[0235] If the value of vps_ols_mode_idc is equal to 2, this indicates that the total number of OLSs indicated by the VPS is explicitly signaled, the output layer is explicitly signaled for each OLS, and other layers are direct or indirect reference layers of the output layer of the OLS.

[0236] The value of vps_ols_mode_idc should be in the range of 0 to 2.

[0237] If the vps_all_independent_layers_flag value is equal to 1, and if the vps_each_layer_is_an_ols_flag value is equal to 0, then the value of vps_ols_mode_idc is inferred to be equal to 2.

[0238] vps_num_output_layer_sets_minus1+1 indicates the total number of OLSs indicated by the VPS when the vps_ols_mode_idc value is equal to 2.

[0239] The variable TotalNumOlss, which indicates the total number of OLSs indicated by VPS, is derived (or inferred) as follows.

[0240] [Table 2]

[0241]

[0242] When the value of vps_ols_output_layer_flag[i][j] is equal to 1, this indicates that when the vps_ols_mode_idc value is equal to 2, the layer with nuh_layer_id equal to vps_layer_id[j] is the output layer of the i-th OLS. When the value of vps_ols_output_layer_flag[i][j] is equal to 0, this indicates that when the vps_ols_mode_idc value is equal to 2, the layer with nuh_layer_id equal to vps_layer_id[j] is not the output layer of the i-th OLS.

[0243] As shown below, the variable NumOutputLayersInOls[i] indicating the number of output layers in the i-th OLS, the variable NumSubLayersInLayerInLayerInOLS[i][j] indicating the number of sublayers of the j-th layer in the i-th OLS, the variable OutputLayerIdInOls[i][j] indicating the nuh_layer_id value of the j-th output layer in the i-th OLS, and the variable LayerUsedAsOutputLayerFlag[k] indicating whether the k-th layer is used as an output layer in at least one OLS are derived.

[0244] [Table 3]

[0245]

[0246]

[0247] For each value of i in the range from 0 to vps_max_layers_minus1, the values ​​of LayerUsedAsRefLayerFlag[i] and LayerUsedAsOutputLayerFlag[i] are both different from 0. That is, there should be no layer that is not an output layer of at least one OLS or a direct reference layer of another layer. There should be one or more layers that are output layers for each OLS. That is, for values ​​of i in the range from 0 to TotalNumOls1-1 (inclusive), the value of NumOutputLayersInOls[i] should be equal to or greater than 1.

[0248] As shown below, the variable NumLayersInOls[i] indicating the number of layers within the i-th OLS, the variable LayerIdInOls[i][j] indicating the nuh_layer_id value of the j-th layer within the i-th OLS, the variable NumMultiLayerOlss indicating the number of multi-layer OLSs (i.e., OLSs including two or more layers), and the variable MultiLayerOlsIdx[i] indicating the index used for the multi-layer OLS list or the i-th OLS when NumLayersInOls[i] is greater than 0 are derived.

[0249] [Table 4]

[0250]

[0251] The 0th OLS includes only the lowest layer (ie, the layer with nuh_layer_id equal to vps_layer_id[0]), and in the case of the 0th OLS, the only included layer is output.

[0252] As shown below, a variable OlsLayerIdx[i][j] is derived (or inferred) which indicates the OLS layer index of the layer with nuh_layer_id equal to LayerIdInOls[i][j].

[0253] [Table 5]

[0254]

[0255] The lowest layer of each OLS shall be an independent layer. That is, for each value of i in the range from 0 to TotalNumOlss-1 (inclusive), the value of vps_independent_layer_flag[GeneralLayerIdx[LayerIdInOls[i][0]]] shall be equal to 1.

[0256] Each layer should be included in one or more OLSs indicated by the VPS. That is, there should be one or more pairs of i and j values ​​such that for each layer with a particular nuh_layer_idnuhLayerId equal to one of the vps_layer_id[k] values, the LayerIdInOls[i][j] value can be equal to nuhLayerId, with k in the range from 0 to vps_max_layers_minus1. Here, i ranges from 0 to TotalNumOlss-1, and j ranges from 0 to NumLayersInOls[i]-1 (inclusive).

[0257] vps_num_ptls_minus1+1 indicates the number of profile_tier_level() syntax structures within the VPS. The vps_num_ptls_minus1 value should be less than TotalNumOlss.

[0258] If the value of vps_pt_present_flag[i] is equal to 1, this indicates that the profile, tier, and general constraint information is present in the i-th profile_tier_level() syntax structure within the VPS. If the value of vps_pt_present_flag[i] is equal to 0, this indicates that the profile, tier, and general constraint information is not present in the i-th profile_tier_level() syntax structure within the VPS. The vps_pt_present_flag[0] value is inferred to be equal to 1. If vps_pt_present_flag[i] is equal to 0, the profile, tier, and general constraint information of the i-th profile_tier_level() syntax structure within the VPS is derived (or inferred) to be the same as the i-1-th profile_tier_level() syntax structure within the VPS.

[0259] vps_ptl_max_temporal_id[i] indicates the TemporalId of the highest sublayer representation, where level information is present in the i-th profile_tier_level() syntax structure. The vps_ptl_max_temporal_id[i] value shall be in the range of 0 to vps_max_sublayer_minus1. When the vps_ptl_max_temporal_id syntax element is not present in the VPS, the vps_ptl_max_temporal_id[i] value is inferred to be equal to vps_max_sublayer_minus1.

[0260] vps_ptl_alignment_zero_bit shall be equal to 0.

[0261] vps_ols_ptl_idx[i] specifies the index of the profile_tier_level() syntax structure that applies to the i-th OLS for the list of profile_tier_level() syntax structures within the VPS. When the vps_ols_ptl_idx syntax element is present in the VPS, the vps_ols_ptl_idx[i] value shall be in the range from 0 to vps_num_ptls_minus1, inclusive.

[0262] When the vps_ols_ptl_idx syntax element is not present in the VPS, the vps_ols_ptl_idx[i] value is derived (or inferred) as described below.

[0263] - If the vps_num_ptls_minus1 value is equal to 0, then the vps_ols_ptl_idx[i] value is inferred to be equal to 0.

[0264] Otherwise (the vps_num_ptls_minus1 value is greater than 0 and vps_num_ptls_minus1+1 is equal to TotalNumOlss), the vps_ols_ptl_idx[i] value is inferred to be equal to i.

[0265] If the NumLayersInOls[i] value is equal to 1, the profile_tier_level() syntax structure applied to the i-th OLS is also present in the SPS referenced by layers within the i-th OLS. When the NumLayersInOls[i] value is equal to 1, the profile_tier_level() syntax structure signaled in the VPS and the profile_tier_level() syntax structure signaled in the SPS for the i-th OLS shall be the same according to the bitstream conformance requirements.

[0266] Each profile_tier_level() syntax structure within a VPS shall be referenced by at least one vps_ols_ptl_idx[i] value, where i ranges from 0 to TotalNumOlss-1 (inclusive).

[0267] (When present) vps_num_dpb_params_minus1+1 indicates the number of dpb_parameters() syntax structures within the VPS. The vps_num_dpb_params_minus1 value shall be in the range from 0 to NumMultiLayerOlss-1, inclusive.

[0268] As shown below, a variable VpsNumDpbParams is derived (or inferred) which indicates the number of dpb_parameters() syntax structures within the VPS.

[0269] [Table 6]

[0270]

[0271] The vps_sublayer_dpb_params_present_flag is used to control the presence of the max_dec_pic_buffering_minus1[], max_num_reorder_pics[], and max_latency_increase_plus1[] syntax elements in the dpb_parameters() syntax element within the VPS. If not present, the vps_sub_dpb_params_info_present_flag value is inferred to be equal to 0.

[0272] vps_dpb_max_temporal_id[i] indicates the TemporalId of the highest sublayer representation, where DPB parameters may be present in the i-th dpb_parameters() syntax structure within the VPS. vps_dpb_max_temporal_id[i] values ​​shall be in the range of 0 to vps_max_sublayer_minus1. When not present, the vps_dpb_max_temporal_id[i] value is inferred to be equal to vps_max_sublayer_minus1.

[0273] vps_ols_dpb_pic_width[i] indicates the width of each picture storage buffer for the i-th multi-layer OLS in units of luma samples.

[0274] vps_ols_dpb_pic_height[i] indicates the height of each picture storage buffer for the i-th multi-layer OLS in units of luma samples.

[0275] vps_ols_dpb_chroma_format[i] indicates the maximum allowed value of sps_chroma_format_idc for all SPSs referenced by the CLVS within the CVS for the i-th multi-layer OLS.

[0276] vps_ols_dpb_bitdepth_minus8[i] indicates the maximum allowed value of sps_bit_depth_minus8 for all SPSs referenced by the CLVS within the CVS for the i-th multi-layer OLS.

[0277] To decode the i-th multi-layer OLS, the decoding apparatus may safely allocate memory to the DPB according to syntax element values ​​of vps_ols_dpb_pic_width[i], vps_ols_dpb_pic_height[i], vps_ols_dpb_chroma_format[i], and vps_ols_dpb_bitdepth_ols_dpb_bitdepth_ols_dpb_bitdepth.

[0278] vps_ols_dpb_params_idx[i] indicates the index of the dpb_parameters() syntax structure applied to the i-th layer of OLS for the list of dpb_parameters() syntax structures within the VPS. When present, vps_ols_dpb_params_idx[i] values ​​shall be in the range from 0 to VpsNumDpbParams-1, inclusive.

[0279] If not present, vps_ols_dpb_params_idx[i] is inferred as follows.

[0280] -When VpsNumDpbParams is equal to 1, the value of vps_ols_dpb_params_idx[i] is equal to 0.

[0281] Otherwise (VpsNumDpbParams is greater than 1 and equal to NumMultiLayerOlss), the vps_ols_dpb_params_idx[i] value is inferred to be equal to i.

[0282] In the case of a single-layer OLS, the applicable dpb_parameters() syntax structure is present in the SPS referenced by the layers within the OLS.

[0283] Each dpb_parameters() syntax structure within a VPS shall be referenced by at least one vps_ols_dpb_params_idx[i] value, where i ranges from 0 to NumMultiLayerOlss-1 (inclusive).

[0284] If the vps_general_hrd_params_present_flag value is equal to 1, this indicates that the VPS includes the general_hrd_parameters() syntax structure and other HRD parameters. If the vps_general_hrd_params_present_flag value is equal to 0, this indicates that the VPS does not include the general_hrd_parameters() syntax structure nor other HRD parameters. When not present, the vps_general_hrd_params_present_flag value is inferred to be equal to 0.

[0285] When the value of NumLayersInOls[i] is equal to 1, the general_hrd_parameters() syntax structure and the ols_hrd_parameters() syntax structure applied to the i-th OLS exist in the SPS referenced by the layers within the i-th OLS.

[0286] If the vps_sublayer_cpb_params_present_flag value is equal to 1, this indicates that the i-th ols_hrd_parameters() syntax structure within the VPS includes HRD parameters for sublayer representations with TemporalIds ranging from 0 to vps_hrd_max_tid[i] (inclusive). And, if the vps_sublayer_cpb_params_present_flag value is equal to 0, this indicates that the i-th ols_hrd_parameters() syntax structure within the VPS includes only HRD parameters for sublayer representations with TemporalIds equal to vps_hrd_max_tid[i]. If the vps_max_sublayer_minus1 value is equal to 0, the vps_sublayer_cpb_params_present_flag value is inferred to be equal to 0.

[0287] When the vps_sublayer_cpb_params_present_flag value is equal to 0, the HRD parameters for the sublayer representation with a TemporalId ranging from 0 to vps_hrd_max_tid[i]-1 (inclusive) are inferred to be equal to the sublayer representation with a TemporalId equal to vps_hrd_max_tid[i]. This includes the HRD parameters starting from the fixed_pic_rate_general_flag[i] syntax element to the sublayer_hrd_parameters(i) syntax structure immediately following the "if (general_vcl_hrd_params_present_flag)" condition within the ols_hrd_parameters syntax structure.

[0288] vps_num_ols_hrd_params_minus1 plus 1 indicates the number of ols_hrd_parameters() syntax structures present within the VPS when the vps_general_hrd_params_present_flag value is equal to 1. The vps_num_ols_hrd_params_minus1 value shall be in the range from 0 to NumMultiLayerOlss-1, inclusive.

[0289] vps_hrd_max_tid[i] indicates the TemporalId of the highest sublayer representation with HRD parameters included in the i-th ols_hrd_parameters() syntax structure. The value of vps_hrd_max_tid[i] shall be in the range of 0 to vps_max_sublayer_minus1. When not present, the vps_hrd_max_tid[i] value is inferred to be equal to vps_max_sublayer_minus1.

[0290] vps_ols_hrd_idx[i] indicates the index of the ols_hrd_parameters() syntax structure applied to the i-th multi-layer OLS with respect to the list of ols_hrd_parameters() syntax structures within the VPS. The vps_ols_hrd_idx[i] value should be in the range of 0 to vps_num_ols_hrd_params_minus1.

[0291] If not present, vps_ols_hrd_idx[i] is inferred (or deduced) as follows.

[0292] If the vps_num_ols_hrd_params_minus1 value is equal to 0, then the vps_ols_hrd_idx[i] value is inferred to be equal to 0.

[0293] Otherwise (vps_num_ols_hrd_params_minus1+1 is greater than 1 and equal to NumMultiLayerOlss), the vps_ols_hrd_idx[i] value is inferred to be equal to i.

[0294] In the case of a single-layer OLS, the applicable ols_hrd_parameters() syntax structure exists in the SPS referenced by the layers within the OLS.

[0295] Each ols_hrd_parameters() syntax structure within a VPS shall be referenced by at least one vps_ols_hrd_idx[i] value, with i ranging from 0 to NumMultiLayerOlss-1 (inclusive).

[0296] If the vps_extension_flag value is equal to 0, this indicates that the vps_extension_data_flag syntax element is not present in the VPS RBSP syntax structure. And, if the vps_extension_flag value is equal to 1, this indicates that the vps_extension_data_flag syntax element is present in the VPS RBSP syntax structure.

[0297] vps_extension_data_flag may have a random value.

[0298] In case of a bitstream with a single layer, the presence of a VPS is optional.If no VPS is present, the value of the sps_video_parameter_set_id syntax element is equal to 0, and part of the parameter value is inferred as described below.

[0299] If the value of sps_video_parameter_set_id is equal to 0, the following applies.

[0300] - The SPS does not reference the VPS, and when decoding each CLVS that references the SPS, the VPS is not referenced.

[0301] The -vps_max_layers_minus1 value is inferred to be equal to 0.

[0302] -vps_max_sublayer_minus1 value is inferred to be equal to 6.

[0303] - A CVS shall include only one layer (ie, all VCL NAL units within a CVS shall have the same nuh_layer_id value).

[0304] - The GeneralLayerIdx[nuh_layer_id] value is inferred to be equal to 0.

[0305] -vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] value is inferred to be equal to 1.

[0306] When performing video / image coding, parameter sets (PPS, SPS, VPS, etc.) can be shared between layers. That is, a VCL NAL unit in a specific layer can reference a parameter set in another layer. When using the parameter set sharing function, the following constraints exist.

[0307] spsLayerId is set to the nuh_layer_id of a specific (or particular) SPS NAL unit, and vclLayerId is set to the nuh_layer_id value of a specific VCL NAL unit. A specific VCL NAL unit does not reference a specific SPS NAL unit except when spsLayerId is less than or not equal to vclLayerId and when not all OLSs indicated by a VPS including layers each having the same nuh_layer_id as vclLayerId include a layer with the same nuh_layer_id as spslayerId.

[0308] ppsLayerId is set to the nuh_layer_id of the specific PPS NAL unit, and vclLayerId is set to the nuh_layer_id value of the specific VCL NAL unit. A specific VCL NAL unit shall not reference a specific PPS NAL unit except when ppsLayerId is less than or not equal to vclLayerId and when not all OLSs indicated by a VPS including layers each having the same nuh_layer_id as vclLayerId include a layer having the same nuh_layer_id as ppslayerId.

[0309] apsLayerId is set to the nuh_layer_id of the specific APS NAL unit, and vclLayerId is set to the nuh_layer_id value of the specific VCL NAL unit. A specific VCL NAL unit shall not reference a specific APS NAL unit except when apsLayerId is less than or not equal to vclLayerId and when not all OLSs indicated by the VPS including layers each having the same nuh_layer_id as vclLayerId include a layer with the same nuh_layer_id as apslayerId.

[0310] Furthermore, when the VPS is not present in the CVS (i.e., when the bitstream is a single-layer bitstream), all VCL NAL units in the bitstream should have the same nuh_layer_id, and all VCL NAL units should reference parameter sets in the same layer. However, since such constraints are not included in the aforementioned constraints, additional complexity may arise in the case of a decoding device designed to handle only single-layer bitstreams.

[0311] In addition, if the VPS does not exist in the CVS, the OLS information is not created. Therefore, some parameters required for decoding, such as TotalNumOlss, NumLayersInOls[], NumOutputLayersInOls[], etc., cannot be derived (or inferred) or initialized. This causes problems in the operation of the decoding device.

[0312] To describe the detailed examples of this specification, the following drawings are illustrated. The detailed terms of the devices (or equipment) or the detailed terms of the signals / information specified in the drawings are only exemplary. Therefore, the technical features of this specification will not be limited to the detailed terms used in the following drawings.

[0313] This specification provides the following methods to solve the above problems. The content of each method can be applied independently or in combination.

[0314] For example, when no VPS exists for a CVS (ie, when the sps_video_parameter_set_id value is equal to 0), the following constraints may apply.

[0315] a) The layer identifiers (nuh_layer_id) of all VCL NAL units within a CVS that references an SPS each have the same value as the layer identifier (nuh_layer_id) of the SPS.

[0316] b) The layer identifier (nuh_layer_id) value of all VCL NAL units is the same as the value of the layer identifier (nuh_layer_id) of the parameter set referenced by all VCL NAL units.

[0317] In this document, the parameter set referenced by the VCL NAL unit includes a parameter set for decoding the video / image information disclosed in this specification. For example, the parameter set may include APS, PPS, SPS, VPS, etc.

[0318] Alternatively, the aforementioned constraints can be expressed as follows.

[0319] a) All VCL NAL units within a CVS have the same layer identifier (nuh_layer_id) value.

[0320] b) The layer identifier (nuh_layer_id) of the parameter set referenced by each VCL NAL unit within the CVS is the same as the layer identifier (nuh_layer_id) of the VCL NAL unit.

[0321] Alternatively, the aforementioned constraints can be expressed as follows.

[0322] a) All VCL NAL units within a CVS and parameter sets referenced by VCL NAL units have the same layer identifier (nuh_layer_id) value.

[0323] Alternatively, when a VCL NAL unit refers to a parameter having a layer identifier (nuh_layer_id) different from that of the VCL NAL unit, the aforementioned constraint may be expressed (or indicated) such that the sps_video_parameter_set_id value is equal to 0.

[0324] In addition, for example, when the VPS does not exist in the CVS (i.e., when the sps_video_parameter_set_id value is equal to 0), only one output layer set exists within the CVS, and the output layer set includes only one layer within the CVS, and the layer can be derived (or inferred) as the output layer of the output layer set.

[0325] In the absence of a VPS for the CVS, values ​​for the following parameters can be derived (or inferred) as described below.

[0326] a) TotalNumOlss is inferred to be equal to 1.

[0327] b) NumLayersInOls[0] is inferred to be equal to 1.

[0328] c) NumOutputLayersInOls[0] is inferred to be equal to 1.

[0329] d)OutputLayerIdInOls[0][0] is inferred to be equal to the nuh_layer_id of the SPS.

[0330] According to an embodiment, Table 7 shown below may be applicable to the sps_video_parameter_set_id syntax element.

[0331] [Table 7]

[0332]

[0333] 7, when the value of the sps_video_parameter_set_id syntax element is equal to 0 or greater (ie, when the bitstream including the sps_video_parameter_set_id syntax element is a multi-layer bitstream), the sps_video_parameter_set_id syntax element indicates the value of the vps_video_parameter_set_id syntax element of the VPS referenced by the SPS within the corresponding bitstream.

[0334] When the sps_video_parameter_set_id syntax element is equal to 0 (i.e., when the bitstream including the sps_video_parameter_set_id syntax element is a single-layer bitstream), the corresponding bitstream may not include a VPS. Therefore, the SPS in the corresponding bitstream does not reference the VPS, and when decoding each CLVS that references the SPS, the VPS is not referenced.

[0335] Also, the syntax element (vps_max_layers_minus1) indicating the maximum number of layers within a CVS is derived (or inferred) to be equal to 0, and the value of the syntax element (vps_max_sublayer_minus1) indicating the number of temporal sublayers that may exist in a CVS may be inferred to be equal to 6.

[0336] In addition, when the value of the sps_video_parameter_set_id syntax element is equal to 0, the CVS includes only one layer, and the following applies to the CVS.

[0337] - The value of the nuh_layer_id syntax element of all VCL NAL units that reference the SPS that references this CVS is the same as the value of the nuh_layer_id syntax element of the SPS.

[0338] - The value of the nuh_layer_id syntax element of all VCL NAL units within this CVS is the same as the value of the nuh_layer_id syntax element of the parameter set referenced by the VCL NAL unit.

[0339] Herein, the parameter set may include APS, PPS, SPS, VPS, etc. Therefore, when CVS includes only one layer, the nuh_layer_id (apsLayerId) of VPS, the nuh_layer_id (ppsLayerId) of PPS, the nuh_layer_id (spsLayerId) of SPS, and the nuh_layer_id (vps_layer_id) of VPS are the same as the nuh_layer_id of the VCL NAL unit.

[0340] In addition, the value of the syntax element related to inter-layer prediction (vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]]) is inferred to be equal to 1. That is, inter-layer prediction is not used.

[0341] When the value of vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 1, the SPS referenced by the CLVS of nuhLayerId having a specific nuh_layer_id value has the same nuh_layer_id as the nuhLayerId.

[0342] The value of sps_video_parameter_set_id is the same for all SPSs referenced by a CLVS within a CVS.

[0343] According to another embodiment, Table 8 shown below may be applied to the sps_video_parameter_set_id syntax element.

[0344] [Table 8]

[0345]

[0346]

[0347] Referring to Table 8, when the value of the sps_video_parameter_set_id syntax element is equal to 0 or greater, the sps_video_parameter_set_id syntax element indicates the value of the vps_video_parameter_set_id syntax element for the VPS referenced by the SPS within the corresponding bitstream.

[0348] When the sps_video_parameter_set_id syntax element is equal to 0, the SPS in the corresponding bitstream does not reference the VPS, and the VPS is not referenced when decoding each CLVS that references the SPS.

[0349] Also, the syntax element (vps_max_layers_minus1) indicating the maximum number of layers within a CVS is derived (or inferred) to be equal to 0, and the value of the syntax element (vps_max_sublayer_minus1) indicating the number of temporal sublayers that may exist in a CVS may be inferred to be equal to 6.

[0350] In addition, when the value of the sps_video_parameter_set_id syntax element is equal to 0, the CVS includes only one layer, and the following applies to the CVS.

[0351] - All VCL NAL units within this CVS have the same nuh_layer_id value.

[0352] - The nuh_layer_id of the parameter set referenced by each VCL NAL unit within this CVS is the same as the nuh_layer_id of the VCL NAL unit.

[0353] In addition, the value of the syntax element related to inter-layer prediction (vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]]) is inferred to be equal to 1.

[0354] When the value of vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 1, the SPS referenced by the CLVS of nuhLayerId having a specific nuh_layer_id value has the same nuh_layer_id as the nuhLayerId.

[0355] The value of sps_video_parameter_set_id is the same for all SPSs referenced by a CLVS within a CVS.

[0356] According to yet another embodiment, Table 9 shown below may be applicable to the sps_video_parameter_set_id syntax element.

[0357] [Table 9]

[0358]

[0359] Referring to Table 9, when the value of the sps_video_parameter_set_id syntax element is equal to 0 or greater, the sps_video_parameter_set_id syntax element indicates the value of the vps_video_parameter_set_id syntax element of the VPS referenced by the SPS within the corresponding bitstream.

[0360] When the sps_video_parameter_set_id syntax element is equal to 0, the SPS in the corresponding bitstream does not reference the VPS, and the VPS is not referenced when decoding each CLVS that references the SPS.

[0361] Also, the syntax element (vps_max_layers_minus1) indicating the maximum number of layers within a CVS is derived (or inferred) to be equal to 0, and the value of the syntax element (vps_max_sublayer_minus1) indicating the number of temporal sublayers that may exist in a CVS may be inferred to be equal to 6.

[0362] In addition, when the value of the sps_video_parameter_set_id syntax element is equal to 0, the CVS includes only one layer, and all VCL NAL units within the CVS and parameter sets referenced by the VCL NAL units have the same nuh_layer_id value.

[0363] In addition, the value of the syntax element related to inter-layer prediction (vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]]) is inferred to be equal to 1.

[0364] When the value of vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 1, the SPS referenced by the CLVS having a specific nuh_layer_id value nuhLayerId has the same nuh_layer_id as nuhLayerId.

[0365] The value of sps_video_parameter_set_id is the same for all SPSs referenced by a CLVS within a CVS.

[0366] As yet another embodiment, when the CVS includes only one layer, the following constraints may apply, as shown in Table 10 below.

[0367] [Table 10]

[0368]

[0369] That is, spsLayerId is set to the nuh_layer_id of a specific (or concrete) SPS NAL unit, and vclLayerId is set to the nuh_layer_id value of a specific VCL NAL unit. A specific VCL NAL unit does not reference a specific SPS NAL unit except when spsLayerId is less than or not equal to vclLayerId, when the sps_video_parameter_set_id value is not equal to 0, and when not all OLSs indicated by a VPS including layers each having the same nuh_layer_id as vclLayerId include a layer having the same nuh_layer_id as spslayerId.

[0370] ppsLayerId is set to the nuh_layer_id of the specific PPS NAL unit, and vclLayerId is set to the nuh_layer_id value of the specific VCL NAL unit. A specific VCL NAL unit shall not reference a specific PPS NAL unit except when ppsLayerId is less than or not equal to vclLayerId, when the sps_video_parameter_set_id value is not equal to 0, and when not all OLSs indicated by the VPS including layers each having the same nuh_layer_id as vclLayerId include a layer having the same nuh_layer_id as spslayerId.

[0371] apsLayerId is set to the nuh_layer_id of the specific APS NAL unit, and vclLayerId is set to the nuh_layer_id value of the specific VCL NAL unit. A specific VCL NAL unit shall not reference a specific APS NAL unit except when apsLayerId is less than or not equal to vclLayerId, and when not all OLSs indicated by the VPS including layers each having the same nuh_layer_id as vclLayerId include a layer having the same nuh_layer_id as apslayerId.

[0372] Additionally, when the CVS includes only one layer, the following constraints may apply, as shown below in Table 11.

[0373] [Table 11]

[0374]

[0375]

[0376] Referring to Table 11, when the sps_video_parameter_set_id syntax element is equal to 0, the SPS in the corresponding bitstream does not reference the VPS, and when each CLVS that references the SPS is decoded, the VPS is not referenced.

[0377] Additionally, a syntax element indicating the maximum number of layers within a CVS (vps_max_layers_minus1) is derived (or inferred) to be equal to 0, and a value of a syntax element indicating the number of temporal sublayers that may be present in a CVS (vps_max_sublayer_minus1) is inferred to be equal to 6.

[0378] In addition, a CVS includes only one layer, and the following applies to this CVS. That is, all VCL NAL units within this CVS have the same nuh_layer_id value.

[0379] Additionally, TotalNumOlss indicating the total number of OLSs specified by the VPS is inferred to be equal to 1, and NumLayersInOls[0] indicating the number of layers within the 0th OLS is inferred to be equal to 1.

[0380] Additionally, NumOutputLayersInOls[0], indicating the number of output layers within the 0th OLS, is inferred to be equal to 1, and OutputLayerIdInOls[0][0], indicating the nuh_layer_id value of the 0th output layer within the 0th OLS, is inferred to be equal to 1.

[0381] Additionally, the GeneralLayerIdx[nuh_layer_id] value is inferred to be equal to nuh_layer_id, and the vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] value is inferred to be equal to 1. When the vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] value is equal to 1, the SPS referenced by the CLSV with a specific nuh_layer_id value nuhLayerId has the same nuh_layer_id as nuhLayerId.

[0382] The sps_video_parameter_set_id value is the same in all SPSs referenced by the CLVS within the CVS.

[0383] Additionally, as described below, the variable PictureOutputFlag for the current picture may be derived (or inferred).

[0384] -PictureOutputFlag is set to 0 when the current layer is not an output layer (i.e., when nuh_layer_id is not the same as OutputLayerIdInOls[TargetOlsIdx][i] for values ​​of i in the range of 0 to NumOutputLayersInOls[TargetOlsIdx]-1 (inclusive), or when one of the following conditions is true.

[0385] - The current picture is a RASL picture and the NoOutputBeforeRecoveryFlag of the related IRAP picture is equal to 1.

[0386] - The current picture is a GDR picture with NoOutputBeforeRecoveryFlag equal to 1, or the current picture is a reconstructed picture of a GDR picture with NoOutputBeforeRecoveryFlag equal to 1.

[0387] - Otherwise, PictureOutputFlag is set to the same as ph_pic_output_flag.

[0388] In addition, the decoding device may output pictures that do not belong to the output layer. For example, when the AU cannot use pictures of the output layer, when there is only one output layer (for example, due to loss or layer down switching), among all pictures that can be used by the AU, for the picture with the highest nuh_layer_id value and ph_pic_output_flag equal to 1, the decoding device may set PictureOutputFlag to 1. For all other pictures that can be used by the AU, the decoding device may set PictureOutputFlag to 0.

[0389] Figure 9 and Figure 10 General examples of video / image encoding methods and related components according to embodiments of the present disclosure are respectively shown.

[0390] Figure 9 The video / image encoding method disclosed in Figure 2 、 Figure 3 and Figure 10 More specifically, for example, Figure 9S900 and S910 may be performed by the predictor 220 of the encoding device 200 , and S920 may be performed by the entropy encoder 240 of the encoding device 200 . Figure 9 The video / image encoding method disclosed in the disclosure may include the embodiments described above in this specification.

[0391] More specifically, refer to Figure 9 and Figure 10 , the predictor 220 of the encoding device may perform at least one of inter prediction or intra prediction on the current block within the current picture ( S900 ), and then generate a prediction sample (prediction block) and prediction information about the current block based on the prediction ( S910 ).

[0392] When performing intra prediction, the predictor 220 may predict the current block by referring to samples within the current picture (neighboring samples of the current block). The predictor 220 may determine a prediction mode to be applied to the current block by using prediction modes applied to neighboring samples.

[0393] When performing inter-frame prediction, the predictor 220 can generate prediction information and a block predicted for the current block by performing inter-frame prediction based on the motion information of the current block. The prediction information described above may include information related to the prediction mode, information related to the motion information, and the like. The information related to the motion information may include candidate selection information (e.g., merge index, mvp_flag, or mvp_index), which is information used to derive a motion vector. Furthermore, the information related to the motion information may include the aforementioned information regarding the motion vector difference (MVD) and / or reference picture index information. Furthermore, the information related to the motion information may include information indicating whether L0 prediction, L1 prediction, or bi-prediction is applied. For example, the predictor 220 may derive motion information for the current block within the current picture based on motion estimation. To this end, using the original block within the original picture corresponding to the current block, the predictor 220 may search for similar reference blocks with high correlation in units of fractional pixels within a determined search range within the reference picture. The predictor 220 may then derive motion information from the searched reference blocks. Block similarity may be derived based on phase-based differences between sample values. For example, block similarity can be calculated based on the sum of absolute differences (SAD) between the current block (or current block template) and the reference block (or reference block template). In this case, motion information can be derived based on the reference block with the minimum SAD within the search area. The derived motion information can be signaled to the decoding device using various methods based on the inter-frame prediction mode.

[0394] The residual processor 230 of the encoding device can generate residual samples and residual information based on the predicted samples generated by the predictor 220 and the original picture (original block, original sample). In this article, the residual information is information related to the residual sample, and the residual information may include information related to the (quantized) transform coefficient of the residual sample.

[0395] The adder (or reconstructor) of the encoding apparatus may generate reconstructed samples (reconstructed picture, reconstructed block, reconstructed sample array) by adding the residual samples generated in the residual processor 230 and the prediction samples generated in the predictor 220 .

[0396] The entropy encoder 240 of the encoding device may encode image information including prediction information generated in the predictor 220, residual information generated in the residual processor 230, and the like (S920). In this article, the image information may further include information about the VCL NAL unit and information about the HLS, and may be delivered (or transmitted) to the decoding device in the form of a bitstream. A bitstream is a bit sequence configured in the form of a NAL unit stream or a byte stream forming an expression of an access unit (AU) constituting one or more CVSs. In the case of a single-layer bitstream, the bitstream may be formed of one CVS, and in this case, CVS may be used as the same meaning as a bitstream.

[0397] The HLS-related information may include information / syntax related to parameter sets used to decode image / video information. For example, parameter sets may include APS, PPS, SPS, VPS, etc. SPS may include an sps_video_parameter_set_id syntax element.

[0398] In this embodiment, the layer identifier of the VCL NAL unit included in the bitstream and the layer identifier of the parameter set referenced by the VCL NAL unit may be derived based on the value of the sps_video_parameter_set_id syntax element.

[0399] For example, when the value of the sps_video_parameter_set_id syntax element is greater than 0, that is, when the sps_video_parameter_set_id syntax element value is not equal to 0, the sps_video_parameter_set_id syntax element may indicate the value of the identifier (vps_video_parameter_set_id syntax element) of the VPS referenced by the SPS.

[0400] When the value of the sps_video_parameter_set_id syntax element is equal to 0, the value of the nuh_layer_id syntax element of all VCL NAL units of the SPS within the CVS may be equal to the value of the nuh_layer_id syntax element of the SPS. And the value of the nuh_layer_id syntax element of the VCL NAL unit may be equal to the value of the nuh_layer_id syntax element of the parameter set referenced by the VCL NAL unit.

[0401] Alternatively, when the sps_video_parameter_set_id syntax element value is equal to 0, all VLS NAL units within the CVS can have the same NAL unit header layer identifier (nuh_layer_id) value, and the NAL unit header layer identifier (nuh_layer_id) of the parameter set referenced by each VCL NAL unit within the CVS can be the same as the NAL unit header layer identifier (nuh_layer_id) of the VCL NAL unit.

[0402] Alternatively, when the sps_video_parameter_set_id syntax element value is equal to 0, all VCL NAL units within the CVS and parameter sets referenced by the VCL NAL units have the same NAL unit header layer identifier (nuh_layer_id).

[0403] In addition, according to this embodiment, information related to OLS can be obtained based on the value of the sps_video_parameter_set_id syntax element. The information related to OLS may include the above-mentioned TotalNumOlss, NumLayersInOls[i], NumOutputLayersInOls[i], OutputLayerIdInOls[i][j], etc. In this document, TotalNumOlss indicates the total number of OLSs specified by the video parameter set. NumOutputLayersInOls[i] indicates the number of output layers within the i-th OLS. OutputLayerIdInOls[i][j] indicates the value of the layer identifier (nuh_layer_id) of the NAL unit header of the j-th output layer within the i-th OLS.

[0404] For example, when the sps_video_parameter_set_id syntax element value is equal to 0, the value of at least one of TotalNumOlss, NumLayersInOls[0], NumOutputLayersInOls[0], and OutputLayerIdInOls[0][0] may be inferred to be equal to 1. Additionally, the GeneralLayerIdx[nuh_layer_id] value may be inferred to be the same as nuh_layer_id.

[0405] Therefore, according to the present specification, even in the case of a single-layer bitstream (VSP does not exist in the CVS), since the layer identifier of the VPS referenced by the VCL NAL unit can be derived (or inferred), the coding efficiency of a decoding device designed to handle only a single-layer bitstream can be improved. In addition, even if the VSP does not exist in the CVS, since information related to the OLS can be derived (or inferred) or initialized, problems that may occur during the decoding process when the VSP does not exist in the CVS can be prevented.

[0406] Figure 11 and Figure 12 General examples of video / image decoding methods and related components according to embodiments of the present disclosure are respectively shown.

[0407] Figure 11 The video / image decoding method disclosed in Figure 4 、 Figure 5 and Figure 12 More specifically, for example, Figure 11 S1100 may be performed by the entropy decoder 310 of the decoding apparatus. S1110 may be performed by the predictor 330 of the decoding apparatus, and S1120 may be performed by the adder 340 of the decoding apparatus. Figure 11 The video / image decoding method disclosed in the disclosure may include the embodiments described above in this specification.

[0408] Reference Figure 11 and Figure 12, the entropy decoder 310 of the decoding device can obtain image information including VCL NAL units from the bitstream (S1100). In addition to information related to the VCL NAL unit, the image information may also include prediction information, residual information, information related to HLS, information related to loop filtering, etc. The prediction information may include inter / intra prediction distinction information, intra prediction mode related information, inter prediction mode related information, etc. The information related to HLS may include information / syntax related to the parameter set for decoding the image / video information. In this article, the parameter set may include APS, PPS, SPS, VPS, etc. The SPS may include an sps_video_parameter_set_id syntax element.

[0409] The entropy decoder 310 of the decoding apparatus may derive (or infer) a layer identifier of a VCL NAL unit, a layer identifier of a parameter set referenced by the VCL NAL unit, and / or information related to the OLS based on the sps_video_parameter_set_id syntax element value.

[0410] For example, when the value of the sps_video_parameter_set_id syntax element parsed from the bitstream is greater than 0, the entropy decoder 310 of the decoding device may derive (or infer) that the value of the sps_video_parameter_set_id syntax element is equal to the value of the identifier of the VPS (vps_video_parameter_set_id syntax element) referenced by the SPS. However, if the value of the sps_video_parameter_set_id syntax element is equal to 0, the entropy decoder 310 of the decoding device may derive (or infer) that the value of the nuh_layer_id syntax element of all VCL NAL units of the SPS within the CVS is the same as the value of the nuh_layer_id syntax element of the SPS, and may derive (or infer) that the value of the nuh_layer_id syntax element of the VCL NAL unit is the same as the value of the nuh_layer_id syntax element of the parameter set referenced by the VCL NAL unit.

[0411] As another example, if the value of the sps_video_parameter_set_id syntax element parsed from the bitstream is greater than 0, the entropy decoder 310 of the decoding device can derive (or infer) that all VLS NAL units within the CVS are the same as the NAL unit header layer identifier (nuh_layer_id) value, and can derive (or infer) that the NAL unit header layer identifier (nuh_layer_id) of the parameter set referenced by each VCL NAL unit within the CVS is the same as the NAL unit header layer identifier (nuh_layer_id) of the VCL NAL unit.

[0412] As another example, if the value of the sps_video_parameter_set_id syntax element parsed from the bitstream is greater than 0, the entropy decoder 310 of the decoding device can derive (or infer) that all VCL NAL units within the CVS have the same NAL unit header layer identifier (nuh_layer_id) as the parameter set referenced by the VCL NAL unit.

[0413] In addition, the entropy decoder 310 of the decoding device can derive information related to the OLS based on the value of the sps_video_parameter_set_id syntax element parsed from the bitstream. The information related to the OLS may include TotalNumOlss indicating the total number of OLSs specified by the video parameter set, NumLayersInOls[i] indicating the number of layers within the i-th OLS, NumOutputLayersInOls[i] indicating the number of output layers within the i-th OLS, and OutputLayerIdInOls[i][j] indicating the value of the layer identifier (nuh_layer_id) of the NAL unit header of the j-th output layer within the i-th OLS.

[0414] For example, if the value of the sps_video_parameter_set_id syntax element parsed from the bitstream is equal to 0, then the value of at least one of TotalNumOlss, NumLayersInOls[0], NumOutputLayersInOls[0], and OutputLayerIdInOls[0][0] can be inferred to be equal to 1. In addition, the GeneralLayerIdx[nuh_layer_id] value can be inferred to be the same as nuh_layer_id.

[0415] More specifically, the predictor 330 of the decoding device may perform inter-frame prediction and / or intra-frame prediction on the current block within the current picture based on the prediction information obtained from the bitstream to generate a prediction sample of the current block (S1110). Thereafter, the residual processor 320 of the decoding device may generate a residual sample based on the residual information obtained from the bitstream. The adder 340 of the decoding device may generate a reconstructed sample based on the prediction sample generated in the predictor 330 and the residual sample generated in the residual processor 320, and then may generate a reconstructed picture (reconstructed block) based on the reconstructed sample (S1120).

[0416] Thereafter, a loop filtering process (such as deblocking filtering, SAO and / or ALF process) may be applied to the reconstructed picture as needed to enhance the subjective / objective picture quality.

[0417] Although the method has been described based on a flowchart that lists steps or blocks in the sequence in the above-described embodiment, the steps of the present disclosure are not limited to a specific order, and specific steps may be performed in a different order or in a different order or simultaneously with respect to the steps described above. In addition, it will be understood by those skilled in the art that the steps of the flowchart are not exclusive and another step may be included therein, or one or more steps in the flowchart may be deleted without affecting the scope of the present disclosure.

[0418] The aforementioned method according to the present disclosure may be in the form of software, and the encoding device and / or decoding device according to the present disclosure may be included in a device for performing image processing, such as a TV, a computer, a smart phone, a set-top box, a display device, etc.

[0419] When the embodiments of the present disclosure are implemented by software, the above methods can be implemented by modules (processes or functions) that perform the above functions. The modules can be stored in a memory and executed by a processor. The memory can be installed inside or outside the processor and can be connected to the processor via various well-known methods. The processor may include an application-specific integrated circuit (ASIC), other chipsets, logic circuits and / or data processing devices. The memory may include a read-only memory (ROM), a random access memory (RAM), a flash memory, a memory card, a storage medium and / or other storage devices. In other words, according to the embodiments of the present disclosure, it can be implemented and executed on a processor, a microprocessor, a controller or a chip. For example, the functional units shown in the corresponding figures can be implemented and executed on a computer, a processor, a microprocessor, a controller or a chip. In this case, information about the implementation (for example, information about instructions) or the algorithm can be stored in a digital storage medium.

[0420] In addition, the decoding device and encoding device to which the embodiments of the present disclosure are applied may be included in multimedia broadcast transceivers, mobile communication terminals, home movie video devices, digital movie video devices, surveillance cameras, video chat devices, and real-time communication devices, such as video communication, mobile streaming devices, storage media, cameras, video on demand (VoD) service providers, over-the-top (OTT) video devices, Internet streaming service providers, 3D video devices, virtual reality (VR) devices, augmented reality (AR) devices, image phone video devices, vehicle terminals (e.g., vehicle (including autonomous vehicle) terminals, aircraft terminals, or ship terminals), and medical video devices; and may be used to process image signals or data. For example, OTT video devices may include game consoles, Blu-ray players, Internet-connected TVs, home theater systems, smartphones, tablet PCs, and digital video recorders (DVRs).

[0421] In addition, the processing method of the embodiment of the present disclosure can be generated in the form of a program executed by a computer and can be stored in a computer-readable recording medium. Multimedia data having a data structure according to the embodiment of the present disclosure can also be stored in a computer-readable recording medium. Computer-readable recording media include all types of storage devices and distributed storage devices that store computer-readable data. Computer-readable recording media may include, for example, Blu-ray discs (BDs), universal serial buses (USBs), ROMs, PROMs, EPROMs, EEPROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, and optical data storage devices. Computer-readable recording media also include media embodied in the form of carrier waves (e.g., transmission over the Internet). In addition, the bit stream generated by the encoding method can be stored in a computer-readable recording medium or sent over a wired or wireless communication network.

[0422] In addition, the embodiments of the present disclosure may be embodied as a computer program product based on program code, and the program code may be executed on a computer according to the embodiments of the present disclosure. The program code may be stored on a computer-readable carrier.

[0423] Figure 13 An example of a content streaming system to which embodiments of the present disclosure can be applied is shown.

[0424] Reference Figure 13 A content streaming system to which the embodiments of the present disclosure are applied may generally include an encoding server, a streaming server, a network server, a media storage device, a user device, and a multimedia input device.

[0425] The encoding server is used to compress content input from a multimedia input device such as a smartphone, a camera, or a camcorder into digital data to generate a bitstream and transmit the bitstream to the streaming server. As another example, when the multimedia input device such as a smartphone, a camera, or a camcorder directly generates the bitstream, the encoding server can be omitted.

[0426] A bitstream may be generated by an encoding method or a bitstream generating method to which an embodiment of the present disclosure is applied, and a streaming server may temporarily store the bitstream in the process of transmitting or receiving the bitstream.

[0427] The streaming server transmits multimedia data to user devices via a network server based on user requests, and the network server serves as an intermediary for notifying users of services. When a user requests a desired service from the network server, the network server transmits the request to the streaming server, which then transmits the multimedia data to the user. In this case, the content streaming system may include a separate control server. In this case, the control server is used to control commands and responses between devices within the content streaming system.

[0428] The streaming server can receive content from a media storage and / or encoding server. For example, when receiving content from an encoding server, the content can be received in real time. In this case, in order to provide a stable streaming service, the streaming server can store the bitstream for a predetermined period of time.

[0429] For example, user devices may include mobile phones, smart phones, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation, tablet PCs, tablet PCs, ultrabooks, wearable devices (e.g., watch-type terminals (smart watches), glasses-type terminals (smart glasses), head-mounted displays (HMDs)), digital TVs, desktop computers, digital signage, etc.

[0430] Each of the servers in the content streaming system may operate as a distributed server, and in this case, data received by the each server may be processed in a distributed manner.

Claims

1. A video decoding method performed by a video decoding device, the video decoding method comprising the following steps: Obtain image information including video coding layer VCL network abstraction layer NAL unit from the bitstream; Generate a prediction sample for the current block by performing inter-frame prediction or intra-frame prediction on the current block in the current picture based on the image information; and reconstructing the current block based on the prediction samples, The image information includes the sps_video_parameter_set_id syntax element, When the value of the sps_video_parameter_set_id syntax element is greater than 0, the sps_video_parameter_set_id syntax element indicates the value of the identifier of the video parameter set referenced by the sequence parameter set. wherein, based on the value of the sps_video_parameter_set_id syntax element being equal to 0, i) the maximum number of allowed layers within each coded video sequence CVS is inferred to be equal to 1, and ii) the total number of output layer sets OLS specified by the video parameter set is inferred to be equal to 1, and Wherein, based on the value of the sps_video_parameter_set_id syntax element being equal to 0, the number of layers within the 0th OLS is inferred to be equal to 1 without referring to a parameter specifying whether at least one OLS includes one or more layers.

2. The video decoding method according to claim 1, wherein: The bit stream is a single-layer bit stream.

3. The video decoding method according to claim 1, wherein: Based on the value of the sps_video_parameter_set_id syntax element being equal to 0, the number of output layers within the OLS is inferred to be equal to 1.

4. The video decoding method according to claim 1, wherein: Based on the value of the sps_video_parameter_set_id syntax element, a layer identifier of the VCL NAL unit and a layer identifier of a parameter set referenced by the VCL NAL unit are inferred.

5. The video decoding method according to claim 1, wherein: Based on the value of the sps_video_parameter_set_id syntax element being equal to 0, the value of the layer identifier of all VCL NAL units referenced by the sequence parameter set is the same as the value of the layer identifier of the sequence parameter set.

6. The video decoding method according to claim 4, wherein: Based on the value of the sps_video_parameter_set_id syntax element being equal to 0, the value of the layer identifier of the VCL NAL unit and the value of the layer identifier of the parameter set are the same.

7. The video decoding method according to claim 1, wherein: Based on the value of the sps_video_parameter_set_id syntax element being equal to 0, a layer identifier of a NAL unit header of a parameter set referenced by each VCL NAL unit within the image information is the same as a layer identifier of a NAL unit header of the VCL NAL unit.

8. The video decoding method according to claim 1, wherein: The value of the sps_video_parameter_set_id syntax element is inferred to be different from 0 based on a VCL NAL unit within the VCL NAL unit referencing a parameter set having a layer identifier different from the layer identifier of the VCL NAL unit.

9. A video encoding method performed by a video encoding device, the video encoding method comprising the following steps: Performing inter-frame prediction or intra-frame prediction on a current block in a current picture; generating prediction information for the current block based on the inter-frame prediction or the intra-frame prediction; as well as encoding the image information including the prediction information, The image information includes the video coding layer VCL network abstraction layer NAL unit and the sps_video_parameter_set_id syntax element. When the value of the sps_video_parameter_set_id syntax element is greater than 0, the sps_video_parameter_set_id syntax element indicates the value of the identifier of the video parameter set referenced by the sequence parameter set. wherein, based on the value of the sps_video_parameter_set_id syntax element being equal to 0, i) the maximum number of allowed layers within each coded video sequence CVS is inferred to be equal to 1, and ii) the total number of output layer sets OLS specified by the video parameter set is inferred to be equal to 1, and Wherein, based on the value of the sps_video_parameter_set_id syntax element being equal to 0, the number of layers within the 0th OLS is inferred to be equal to 1 without referring to a parameter specifying whether at least one OLS includes one or more layers.

10. The video encoding method according to claim 9, wherein: The image information includes only one layer.

11. The video encoding method according to claim 9, wherein: Based on the value of the sps_video_parameter_set_id syntax element, a layer identifier of the VCL NAL unit and a layer identifier of a parameter set referenced by the VCL NAL unit are inferred.

12. The video encoding method according to claim 11, wherein: Based on the value of the sps_video_parameter_set_id syntax element being equal to 0, the value of the layer identifier of the VCL NAL unit and the value of the layer identifier of the parameter set are the same.

13. A method for transmitting data for image information, the method comprising the steps of: generating a bitstream for the image information, wherein the bitstream is generated based on the following steps: performing inter-frame prediction or intra-frame prediction on a current block in a current picture, generating prediction information for the current block based on the inter-frame prediction or intra-frame prediction, and encoding the image information including the prediction information; and sending said data comprising said bitstream, The image information includes a video coding layer (VCL) network abstraction layer (NAL) unit and an sps_video_parameter_set_id syntax element; When the value of the sps_video_parameter_set_id syntax element is greater than 0, the sps_video_parameter_set_id syntax element indicates the value of the identifier of the video parameter set referenced by the sequence parameter set. wherein, based on the value of the sps_video_parameter_set_id syntax element being equal to 0, i) the maximum number of allowed layers within each coded video sequence CVS is inferred to be equal to 1, and ii) the total number of output layer sets OLS specified by the video parameter set is inferred to be equal to 1, and Wherein, based on the value of the sps_video_parameter_set_id syntax element being equal to 0, the number of layers within the 0th OLS is inferred to be equal to 1 without referring to a parameter specifying whether at least one OLS includes one or more layers.