Video information-based video decoding method and apparatus including an OLS DPB parameter index

The video decoding method enhances coding efficiency by using DPB parameters for OLS to manage pictures adaptively, addressing the increased costs associated with high-resolution images.

JP7715857B2Active Publication Date: 2025-07-30LG ELECTRONICS INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024022279
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-12-30
Filing Date
2024-02-16
Publication Date
2025-07-30
Estimated Expiration
2040-12-29

AI Technical Summary

Technical Problem

The increasing demand for high-resolution and high-quality images has led to a surge in transmission and storage costs due to the increased amount of information, necessitating a more efficient video coding technology.

Method used

A video decoding method and apparatus that utilizes DPB parameters mapped to OLS (Output Layer Set) for adaptive picture management, enabling efficient coding and decoding processes through the use of OLS DPB parameter indices.

Benefits of technology

This approach allows for adaptive updating of DPB parameters, improving overall coding efficiency and reducing transmission and storage costs for high-resolution and high-quality images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007715857000016
    Figure 0007715857000016
  • Figure 0007715857000017
    Figure 0007715857000017
  • Figure 0007715857000018
    Figure 0007715857000018
Patent Text Reader

Abstract

To provide an image decoding method.SOLUTION: An image decoding method performed by a decoding apparatus, according to the present document includes the steps of: acquiring image information; performing a picture management process for pictures of a DPB (Decoded Picture Buffer) on the basis of the image information; and decoding a current picture on the basis of the pictures. The image information includes an OLS (Output Layer Set) DPB parameter index for a target OLS, and the picture management process is performed on the basis of DPB parameter information for the target OLS derived on the basis of the OLS DPB parameter index.SELECTED DRAWING: Figure 9
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This document relates to video coding technology, and more particularly, to a video decoding method and apparatus for coding video information including DPB parameters mapped to OLS in a video coding system.

Background Art

[0002] Recently, the demand for high-resolution and high-quality images such as HD (High Definition) images and UHD (Ultra High Definition) images has been increasing in various fields. As the image data becomes higher in resolution and quality, the amount of information or bits to be transmitted relatively increases compared to existing image data. Therefore, when transmitting image data using a medium such as an existing wired or wireless broadband line, or storing image data using an existing recording medium, the transmission cost and storage cost increase.

[0003] Accordingly, in order to effectively transmit, store, and reproduce information of high-resolution and high-quality images, a highly efficient image compression technology is required.

Summary of the Invention

Problems to be Solved by the Invention

[0004] The technical problem of this document is to provide a method and apparatus for increasing video coding efficiency.

[0005] Another technical problem of this document is to provide a method and apparatus for deriving DPB parameters for OLS.

Means for Solving the Problems

[0006] According to an embodiment of this document, a video decoding method executed by a decoding device is provided. The method includes steps of acquiring video information, executing a picture management process for pictures in a DPB (Decoded Picture Buffer) based on the video information, and decoding a current picture based on the pictures, where the video information includes an OLS DPB parameter index for a target OLS (Output Layer Set), and the picture management process is executed based on DPB parameter information for the target OLS derived based on the OLS DPB parameter index.

[0007] According to another embodiment of this document, a decoding device that executes video decoding is provided. The decoding device includes an entropy decoding unit that acquires video information, a DPB that executes a picture management process for pictures in a DPB (Decoded Picture Buffer) based on the video information, and a prediction unit that decodes a current picture based on the pictures, where the picture management process is executed based on DPB parameter information for the target OLS derived based on the OLS DPB parameter index.

[0008] According to another embodiment of this document, a video encoding method executed by an encoding device is provided. The method includes steps of executing a picture management process for pictures in a DPB (Decoded Picture Buffer) based on DPB parameter information for a target OLS (Output Layer Set), decoding a current picture based on the pictures, and encoding video information, where the video information includes the DPB parameter information for the target OLS and an OLS DPB parameter index for the target OLS.

[0009] According to another embodiment of the present document, a video encoding apparatus is provided. The encoding apparatus includes a DPB that executes a picture management process for pictures in the DPB based on DPB parameter information for a target OLS (Output Layer Set), a prediction unit that decodes a current picture based on the picture, and an entropy encoding unit that encodes video information. The video information includes the DPB parameter information for the target OLS and an OLS DPB parameter index for the target OLS.

[0010] According to another embodiment of the present document, a computer-readable digital storage medium storing a bitstream including video information for executing a video decoding method is provided. In the computer-readable digital storage medium, the video decoding method includes steps of acquiring video information, executing a picture management process for pictures in a DPB (Decoded Picture Buffer) based on the video information, and decoding a current picture based on the picture. The video information includes an OLS DPB parameter index for a target OLS (Output Layer Set), and the picture management process is executed based on DPB parameter information for the target OLS derived based on the OLS DPB parameter index.

Advantages of the Invention

[0011] According to the present document, DPB parameters for an OLS can be signaled, whereby the DPB can be adaptively updated for the OLS, and the overall coding efficiency can be improved.

[0012] According to this document, index information indicating DPB parameters for OLS can be signaled, whereby the DPB parameters can be adaptively derived for OLS, and based on the derived DPB parameters, the DPB for OLS can be updated to improve the overall coding efficiency.

Brief Description of the Drawings

[0013]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Embodiments for Carrying Out the Invention

[0014] This document can be modified in various ways and can have various embodiments. Specific embodiments will be illustrated in the drawings and described in detail. However, this is not intended to limit this document to specific embodiments. The terms commonly used in this specification are merely used to describe specific embodiments and are not used with the intention of limiting the technical idea of this document. Singular expressions include plural expressions unless the context clearly indicates otherwise. Terms such as "including" or "having" in this specification are intended to specify that features, numbers, steps, operations, components, parts, or combinations thereof described in the specification exist, and it should be understood that the existence or addition possibility of one or more other features, numbers, steps, operations, components, parts, or combinations thereof, etc., is not precluded in advance.

[0015] On the other hand, each configuration in the drawings described in this document is independently illustrated for the convenience of explaining different characteristic functions, and it does not mean that each configuration is realized by separate hardware or separate software. For example, among the configurations, two or more configurations can be combined to form one configuration, and one configuration can also be divided into multiple configurations. Embodiments in which each configuration is integrated and / or separated are included in the scope of rights of this document as long as they do not deviate from the essence of this document.

[0016] Hereinafter, with reference to the accompanying drawings, preferred embodiments of this document will be described in more detail. Hereinafter, the same reference numerals will be used for the same components in the drawings, and overlapping descriptions of the same components can be omitted.

[0017] FIG. 1 schematically shows an example of a video / image coding system to which an embodiment of this document can be applied.

[0018] As shown in FIG. 1, a video / image coding system can include a first device (source device) and a second device (receiving device). The source device can transmit encoded video / image information or data in a file or streaming form to the receiving device via a digital recording medium or a network.

[0019] The source device can include a video source, an encoding device, and a transmitting unit. The receiving device can include a receiving unit, a decoding device, and a renderer. The encoding device can be called a video / image encoding device, and the decoding device can be called a video / image decoding device. A transmitter can be included in the encoding device. A receiver can be included in the decoding device. The renderer can include a display unit, and the display unit can also be composed of a separate device or an external component.

[0020] The video source can obtain video / images through processes such as video / image capture, synthesis, or generation. The video source can include a video / image capture device and / or a video / image generation device. The video / image capture device can include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. The video / image generation device can include, for example, a computer, a tablet, and a smartphone, etc., and can (electronically) generate video / images. For example, virtual video / images can be generated via a computer, etc., and in this case, the video / image capture process can be replaced during the process of generating related data.

[0021] The encoding device can encode input video / images. The encoding device can execute a series of procedures such as prediction, transformation, quantization, etc. for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.

[0022] The transmitting unit can transmit the encoded video / image information or data output in the form of a bitstream to the receiving unit of the receiving device via a digital recording medium or a network in file or streaming form. The digital recording medium can include various recording media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit can include elements for generating a media file via a predetermined file format and can include elements for transmission via a broadcast / communication network. The receiving unit can receive / extract the bitstream and transmit it to the decoding device.

[0023] The decoding device can decode video / images by executing a series of procedures such as inverse quantization, inverse transformation, prediction, etc. corresponding to the operation of the encoding device.

[0024] The renderer can render the decoded video / images. The rendered video / images can be displayed via the display unit.

[0025] This document relates to video / image coding. For example, the methods / embodiments disclosed in this document can be applied to the methods disclosed in the VVC (versatile video coding) standard, EVC (essential video coding) standard, AV1 (AOMedia Video 1) standard, AVS2 (2nd generation of audio video coding standard), or next-generation video / image coding standards (e.g., H.267 or H.268, etc.).

[0026] This document presents various embodiments related to video / image coding, and unless otherwise stated, the embodiments can also be executed in combination with each other.

[0027] In this document, "video" may mean a set of a series of "images" over time. "Picture" generally means a unit indicating one image in a specific time period, and "subpicture / slice / tile" is a unit constituting a part of a picture in coding. A subpicture / slice / tile may include one or more CTUs (Coding Tree Units). One picture may be composed of one or more subpictures / slices / tiles. One picture may be composed of one or more groups of tiles. One tile group may include one or more tiles. A brick represents a rectangular region of CTU rows within a tile in a picture. A tile may be partitioned into multiple bricks, each of which consisting of one or more CTU rows within the tile. A tile that is not partitioned into multiple bricks may be also referred to as a brick.A brick scan shows a specific sequential ordering of CTUs partitioning a picture in which the CTUs are ordered consecutively in CTU raster scan in a brick, bricks within a tile are ordered consecutively in a raster scan of the bricks of the tile, and tiles in a picture are ordered consecutively in a raster scan of the tiles of the picture. Also, a subpicture represents a rectangular region of one or more slices within a picture. That is, a subpicture contains one or more slices that collectively cover a rectangular region of a picture. A tile is a rectangular region of CTUs within a particular tile column and a particular tile row in a picture.The tile column is a rectangular region of CTUs having a height equal to the height of the picture and a width specified by syntax elements in the picture parameter set. The tile row is a rectangular region of CTUs having a height specified by syntax elements in the picture parameter set and a width equal to the width of the picture. A tile scan is a specific sequential ordering of CTUs partitioning a picture in which the CTUs are ordered consecutively in CTU raster scan in a tile whereas tiles in a picture are ordered consecutively in a raster scan of the tiles of the picture.A slice includes an integer number of bricks of a picture that maybe exclusively contained in a single NAL unit. A slice may consists of either a number of complete tiles or only a consecutive sequence of complete bricks of one tile. In this document, tile groups and slices may be used interchangeably. For example, in this document, a tile group / tile group header may be referred to as a slice / slice header.

[0028] A pixel or pel can mean the smallest unit that makes up a picture (or image). Also, the term "sample" can be used as the term corresponding to a pixel. A sample can generally indicate a pixel or the value of a pixel, and can also indicate only the pixel / pixel value of the luma component, or only the pixel / pixel value of the chroma component.

[0029] A unit can indicate the basic unit of image processing. A unit can include at least one of a specific region of a picture and the information related to that region. One unit can include one luma block and two chroma (e.g., cb, cr) blocks. A unit can, in some cases, be used interchangeably with terms such as block or area. In general, an M×N block can include a set (or array) of samples (or sample array) consisting of M columns and N rows, or a set (or array) of transform coefficients.

[0030] In this specification, "A or B" may mean "only A", "only B", or "both A and B". In other words, in this specification, "A or B" may be construed as "A and / or B". For example, in this specification, "A, B, or C" may mean "only A", "only B", "only C", or "any combination of A, B, and C".

[0031] The slashes ( / ) and commas used in this specification may mean "and / or". For example, "A / B" may mean "A and / or B". Thus, "A / B" may mean "only A", "only B", or "both A and B". For example, "A, B, C" may mean "A, B, or C".

[0032] In this specification, "at least one of A and B" may mean "only A", "only B", or "both A and B". Also, in this specification, expressions such as "at least one of A or B" and "at least one of A and / or B" may be construed in the same way as "at least one of A and B".

[0033] Also, in this specification, "at least one of A, B and C" can mean "only A", "only B", "only C" or "any combination of A, B and C". Also, "at least one of A, B or C" and "at least one of A, B and / or C" can mean "at least one of A, B and C".

[0034] Also, the parentheses used in this specification can mean "for example". Specifically, when it is displayed as "prediction (intra prediction)", "intra prediction" can be proposed as an example of "prediction". In other words, "prediction" in this specification is not limited to "intra prediction", and "intra prediction" can be proposed as an example of "prediction". Also, when it is displayed as "prediction (that is, intra prediction)", "intra prediction" can be proposed as an example of "prediction".

[0035] The technical features separately described within one drawing in this specification may be realized separately or simultaneously.

[0036] The following drawings are created to explain a specific example of this specification. Since the names of the specific devices and the names of the specific signals / messages / fields described in the drawings are presented exemplarily, the technical features of this specification are not limited to the specific names used in the following drawings.

[0037] Figure 2 is a diagram schematically explaining the configuration of a video / image encoding device to which the embodiment of this document can be applied. Hereinafter, the video encoding device can include an image encoding device.

[0038] As shown in FIG. 2, the encoding apparatus 200 can be configured to include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 can include an inter-predictor 221 and an intra-predictor 222. The residual processor 230 can include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 can further include a subtractor 231. The adder 250 can be called a reconstructor or a reconstructed block generator. The above-described image partitioner 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 can be configured by one or more hardware components (e.g., an encoder chipset or a processor) according to an embodiment. Also, the memory 270 can include a DPB (decoded picture buffer) and can be configured by a digital recording medium. The hardware component can further include the memory 270 as an internal / external component.

[0039] The image segmentation unit 210 can divide an input image (or picture, frame) input to the encoding device 200 into one or more processing units. As an example, the processing unit can be called a coding unit (CU). In this case, the coding unit can be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) by a QTBTTT (Quad-tree binary-tree ternary-tree) structure. For example, one coding unit can be divided into a plurality of coding units with a deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure can be applied first, and the binary-tree structure and / or the ternary structure can be applied later. Or the binary-tree structure can also be applied first. The coding procedure according to this document can be executed based on the final coding unit that is no longer divided. In this case, based on the coding efficiency according to the image characteristics, etc., the largest coding unit can be immediately used as the final coding unit, or if necessary, the coding unit can be recursively divided into coding units with a deeper depth, and the coding unit with the optimal size can be used as the final coding unit. Here, the coding procedure can include procedures such as prediction, transformation, and restoration described later. As another example, the processing unit can further include a prediction unit (PU: Prediction Unit) or a transform unit (TU: Transform Unit). In this case, the prediction unit and the transform unit can each be divided or partitioned from the final coding unit described above.The prediction unit is a unit of sample prediction, and the conversion unit is a unit for deriving a conversion coefficient and / or a unit for deriving a residual signal from the conversion coefficient.

[0040] The term "unit" can, in some cases, be used interchangeably with terms such as "block" or "area". In general, an M×N block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luma component, or only the pixel / pixel value of the chroma component. A sample can be used as a term corresponding to a pixel or a pel in one picture (or image).

[0041] The encoding device 200 can subtract a prediction signal (predicted block, predicted sample array) output from the inter prediction unit 221 or the intra prediction unit 222 from an input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is transmitted to the conversion unit 232. In this case, as shown in the figure, the unit that subtracts the prediction signal (predicted block, predicted sample array) from the input image signal (original block, original sample array) in the encoder 200 can be called the subtraction unit 231. The prediction unit can perform prediction on a block to be processed (hereinafter referred to as the current block) and generate a predicted block including predicted samples for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied in units of the current block or CU. The prediction unit can generate various pieces of information related to prediction, such as prediction mode information, and transmit them to the entropy encoding unit 240 as described later in the description of each prediction mode. The information related to prediction can be encoded by the entropy encoding unit 240 and output in the form of a bit stream.

[0042] The intra prediction unit 222 can predict the current block by referring to samples within the current picture. The samples to be referred to can be located in the neighborhood of the current block according to the prediction mode, or can also be located remotely. In intra prediction, the prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes can include, for example, the DC mode and the Planar Mode. The directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes according to the fineness of the prediction direction. However, this is only an illustration, and more or fewer directional prediction modes can be used according to the setting. The intra prediction unit 222 can also determine the prediction mode to be applied to the current block by using the prediction mode applied to the adjacent block.

[0043] The inter prediction unit 221 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between adjacent blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the adjacent blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block can be called by names such as a collocated reference block and a collocated CU (colCU), and the reference picture including the temporal neighboring block can also be called a collocated picture (colPic). For example, the inter prediction unit 221 can construct a motion information candidate list based on adjacent blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction can be performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the inter prediction unit 221 can use the motion information of adjacent blocks as the motion information of the current block. In the case of the skip mode, unlike the merge mode, a residual signal may not be transmitted.In the case of the motion information prediction (motion vector prediction, MVP) mode, the motion vector of an adjacent block can be used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.

[0044] The prediction unit 220 can generate a prediction signal based on various prediction methods described below. For example, for the prediction of one block, the prediction unit can not only apply intra prediction or inter prediction, but also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). In addition, the prediction unit can also be based on the intra block copy (IBC) prediction mode or the palette mode for the prediction of the block. The IBC prediction mode or the palette mode can be used for content image / video coding such as games, for example, like SCC (screen content coding). IBC basically performs prediction within the current picture, but can be executed similarly to inter prediction in terms of deriving a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this document. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, the sample values in the picture can be signaled based on the information regarding the palette table and the palette index.

[0045] The prediction signal generated via the prediction unit (including the inter-prediction unit 221 and / or the intra-prediction unit 222) can be used to generate a restored signal or can be used to generate a residual signal. The conversion unit 232 can generate transform coefficients by applying a conversion technique to the residual signal. For example, the conversion technique can include at least one of DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), KLT (Karhunen-Loeve Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, when the GBT represents the relationship information between pixels as a graph, it means the conversion obtained from this graph. The CNT means the conversion obtained based on generating a prediction signal using all previously reconstructed pixels. Also, the conversion process can be applied to a pixel block having the same size of a square, and can also be applied to a non-square, variable-size block.

[0046] The quantization unit 233 quantizes the transform coefficients and transmits them to the entropy encoding unit 240. The entropy encoding unit 240 can encode the quantized signal (information regarding the quantized transform coefficients) and output it as a bitstream. The information regarding the quantized transform coefficients can be called residual information. The quantization unit 233 can reorder the quantized transform coefficients in block form into a one-dimensional vector form based on the coefficient scan order, and can also generate the information regarding the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. The entropy encoding unit 240 can execute various encoding methods such as, for example, exponential Golomb, CAVLC (context-adaptive variable length coding), CABAC (context-adaptive binary arithmetic coding), etc. The entropy encoding unit 240 can encode, together or separately, in addition to the quantized transform coefficients, information necessary for video / image restoration (for example, values of syntax elements, etc.). The encoded information (for example, encoded video / image information) can be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer) units. The video / image information can further include information regarding various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Also, the video / image information can further include general constraint information. In this document, the information and / or syntax elements transmitted / signaled from the encoding device to the decoding device can be included in the video / image information. The video / image information can be encoded through the above-described encoding procedure and included in the bitstream.The bitstream can be transmitted via a network or stored in a digital recording medium. Here, the network can include a broadcast network and / or a communication network, etc., and the digital recording medium can include various recording media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The signal output from the entropy encoding unit 240 can be configured as an internal / external element of the encoding device 200 by a transmission unit (not shown) for transmission and / or a storage unit (not shown) for storage, or the transmission unit can also be included in the entropy encoding unit 240.

[0047] The quantized transform coefficients output from the quantization unit 233 can be used to generate a prediction signal. For example, a residual signal (residual block or residual sample) can be restored by applying inverse quantization and inverse transformation to the quantized transform coefficients via the inverse quantization unit 234 and the inverse transform unit 235. The addition unit 250 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the restored residual signal to the prediction signal output from the inter prediction unit 221 or the intra prediction unit 222. When there is no residual for the block to be processed, as in the case where the skip mode is applied, the predicted block can be used as the reconstructed block. The addition unit 250 can be called a restoration unit or a reconstructed block generation unit. The generated reconstructed signal can be used for intra prediction of the next block to be processed within the current picture and, as will be described later, can also be used for inter prediction of the next picture after passing through filtering.

[0048] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied in the picture encoding and / or restoration process.

[0049] The filtering unit 260 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 260 can apply various filtering methods to the restored picture to generate a modified restored picture, and store the modified restored picture in the memory 270, specifically, in the DPB of the memory 270. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit 260 can generate various information related to filtering and transmit it to the entropy encoding unit 240, as will be described later in the description of each filtering method. The information related to filtering can be encoded by the entropy encoding unit 240 and output in the form of a bit stream.

[0050] The modified restored picture transmitted to the memory 270 can be used as a reference picture in the inter prediction unit 221. When inter prediction is applied through this, the encoding device can avoid prediction mismatches between the encoding device 200 and the decoding device 300, and can also improve the encoding efficiency.

[0051] The memory 270 DPB can store the modified restored picture for use as a reference picture in the inter prediction unit 221. The memory 270 can store the motion information of the block where the motion information in the current picture was derived (or encoded) and / or the motion information of the block in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 221 for utilization as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 270 can store the restored samples of the restored blocks in the current picture and transmit them to the intra prediction unit 222.

[0052] FIG. 3 is a diagram schematically illustrating the configuration of a video / image decoding apparatus to which an embodiment of this document can be applied.

[0053] As shown in FIG. 3, the decoding apparatus 300 can be configured to include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 can include an inter-predictor 331 and an intra-predictor 332. The residual processor 320 can include a dequantizer 321 and an inverse transformer 322. The entropy decoder 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 described above can be configured by one hardware component (e.g., a decoder chipset or a processor) according to an embodiment. Also, the memory 360 can include a DPB (decoded picture buffer) and can also be configured by a digital recording medium. The hardware component can further include the memory 360 as an internal / external component.

[0054] When a bitstream including video / image information is input, the decoding device 300 can restore an image corresponding to the process in which the video / image information was processed by the encoding device in FIG. 2. For example, the decoding device 300 can derive units / blocks based on the block splitting related information obtained from the bitstream. The decoding device 300 can perform decoding using the processing units applied in the encoding device. Therefore, the processing unit for decoding is, for example, a coding unit, and the coding unit can be split according to a quad tree structure, a binary tree structure, and / or a ternary tree structure from a coding tree unit or a maximum coding unit. One or more transform units can be derived from the coding unit. Then, the restored image signal decoded and output via the decoding device 300 can be reproduced via a reproducing device.

[0055] The decoding device 300 can receive the signal output from the encoding device in FIG. 2 in the form of a bitstream, and the received signal can be decoded via the entropy decoding unit 310. For example, the entropy decoding unit 310 can parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information can further include information regarding various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Also, the video / image information can further include general constraint information. The decoding device can further decode a picture based on the information regarding the parameter set and / or the general constraint information. The signaling / received information and / or syntax elements described later in this document can be decoded via the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 310 can decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output values of syntax elements necessary for image restoration, quantized values of transform coefficients regarding residuals, etc. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element in the bitstream, determines a context model using the syntax element information to be decoded, the surrounding and decoded information of the block to be decoded, or the information of symbols / bins decoded in the previous step, predicts the occurrence probability of a bin based on the determined context model, and performs arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element. At this time, after determining the context model, the CABAC entropy decoding method can update the context model using the information of the symbol / bin decoded for the context model of the next symbol / bin.Of the information decoded by the entropy decoding unit 310, the information related to prediction is provided to the prediction units (inter prediction unit 332 and intra prediction unit 331), and the residual value for which entropy decoding is executed by the entropy decoding unit 310, that is, the quantized transform coefficient and related parameter information, can be input to the residual processing unit 320. The residual processing unit 320 can derive a residual signal (residual block, residual sample, residual sample array). Also, of the information decoded by the entropy decoding unit 310, the information related to filtering can be provided to the filtering unit 350. On the other hand, a receiving unit (not shown) that receives the signal output from the encoding device can be further configured as an internal / external element of the decoding device 300, or the receiving unit is a component of the entropy decoding unit 310. On the other hand, the decoding device according to this document can be called a video / image / picture decoding device, and the decoding device can be classified into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder can include the entropy decoding unit 310, and the sample decoder can include at least one of the inverse quantization unit 321, inverse transform unit 322, addition unit 340, filtering unit 350, memory 360, inter prediction unit 332, and intra prediction unit 331.

[0056] In the inverse quantization unit 321, the quantized transform coefficient can be inverse quantized to output a transform coefficient. The inverse quantization unit 321 can reorder the quantized transform coefficients in a two-dimensional block form. In this case, the reordering can be executed based on the coefficient scan order executed by the encoding device. The inverse quantization unit 321 can use a quantization parameter (for example, quantization step size information) to perform inverse quantization on the quantized transform coefficient and obtain a transform coefficient.

[0057] In the inverse conversion unit 322, the conversion coefficient is inversely converted to obtain a residual signal (residual block, residual sample array).

[0058] The prediction unit can perform prediction on the current block and generate a predicted block including predicted samples for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied to the current block based on the information regarding the prediction output from the entropy decoding unit 310, and can determine a specific intra / inter prediction mode.

[0059] The prediction unit 320 can generate a prediction signal based on various prediction methods described later. For example, the prediction unit can apply not only intra prediction or inter prediction for prediction of one block, but also can apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). Also, the prediction unit can be based on the intra block copy (IBC) prediction mode or the palette mode for prediction of a block. The IBC prediction mode or the palette mode can be used for content image / video coding such as games, for example, like SCC (screen content coding). IBC basically performs prediction within the current picture, but can be executed similarly to inter prediction in terms of deriving a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this document. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, information regarding the palette table and the palette index can be included in and signaled in the video / image information.

[0060] The intra prediction unit 331 can predict the current block by referring to samples within the current picture. The samples to be referred to can be located in the neighborhood of the current block or at a distance therefrom depending on the prediction mode. In intra prediction, the prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The intra prediction unit 331 can also determine the prediction mode to be applied to the current block by using the prediction mode applied to an adjacent block.

[0061] The inter prediction unit 332 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted from the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the adjacent block and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the adjacent blocks can include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit 332 can construct a motion information candidate list based on the adjacent blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction can be executed based on various prediction modes, and the information regarding the prediction can include information indicating the mode of inter prediction for the current block.

[0062] The adder 340 can generate a restored signal (restored picture, restored block, restored sample array) by adding the acquired residual signal to the predicted signal (predicted block, predicted sample array) output from the prediction unit (including the inter-prediction unit 332 and / or the intra-prediction unit 331). When there is no residual for the block to be processed, such as when the skip mode is applied, the predicted block can be used as the restored block.

[0063] The adder 340 can be referred to as a restoration unit or a restored block generation unit. The generated restored signal can be used for intra-prediction of the next block to be processed within the current picture, and as will be described later, it can be output after filtering, or can also be used for inter-prediction of the next picture.

[0064] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied during the picture decoding process.

[0065] The filtering unit 350 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 350 can apply various filtering methods to the restored picture to generate a modified restored picture, and can send the modified restored picture to the memory 360, specifically, to the DPB of the memory 360. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.

[0066] The (corrected) restored picture stored in the DPB of the memory 360 can be used as a reference picture in the inter prediction unit 332. The memory 360 can store the motion information of the block for which the motion information in the current picture has been derived (or decoded) and / or the motion information of the block in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 260 for utilization as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 360 can store the restored samples of the restored blocks in the current picture and can transmit them to the intra prediction unit 331.

[0067] In this specification, the embodiments described in the filtering unit 260, the inter prediction unit 221, and the intra prediction unit 222 of the encoding device 200 can be applied to the filtering unit 350, the inter prediction unit 332, and the intra prediction unit 331 of the decoding device 300 in the same or corresponding manner, respectively.

[0068] In this document, at least one of quantization / inverse quantization and / or transform / inverse transform can be omitted. When the quantization / inverse quantization is omitted, the quantized transform coefficients can be called transform coefficients. When the transform / inverse transform is omitted, the transform coefficients can be called coefficients or residual coefficients, or can still be called transform coefficients for the sake of uniformity of expression.

[0069] In this document, the quantized transform coefficients and the transform coefficients can each be referred to as the transform coefficients and the scaled transform coefficients, respectively. In this case, the residual information can include information regarding the transform coefficients (etc.), and the information regarding the transform coefficients (etc.) can be signaled via a residual coding syntax. The transform coefficients can be derived based on the residual information (or the information regarding the transform coefficients (etc.)), and the scaled transform coefficients can be derived via an inverse transform (scaling) with respect to the transform coefficients. The residual samples can be derived based on an inverse transform (transformation) with respect to the scaled transform coefficients. This can be applied / expressed similarly in other parts of this document.

[0070] As described above, the encoding device can execute various encoding methods such as, for example, exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. Further, the decoding device can decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output values of syntax elements necessary for image restoration, quantized values of transform coefficients regarding the residual, etc.

[0071] For example, the coding methods described above can be performed as described later.

[0072] FIG. 4 exemplarily shows context-adaptive binary arithmetic coding (CABAC) for encoding a syntax element. For example, in the encoding process of CABAC, when the input signal is a syntax element that is not a binary value, the encoder can binarize the value of the input signal to convert the input signal into a binary value. Also, when the input signal is already a binary value (i.e., when the value of the input signal is a binary value), binarization is not performed and it can be bypassed. Here, each binary number 0 or 1 that constitutes a binary value can be called a bin. For example, when the binary string after binarization is 110, each of 1, 1, and 0 is called one bin. The bin (etc.) for one syntax element can represent the value of the syntax element.

[0073] Thereafter, the binarized bins etc. of the syntax element can be input as a regular coding engine or a bypass coding engine. The regular coding engine of the encoder can assign a context model that reflects a probability value to the bin, and can encode the bin based on the assigned context model. The regular coding engine of the encoder can update the context model for the bin after encoding each bin. The bins encoded as described above can be represented as context-coded bins.

[0074] On the one hand, when the binary bins of the syntax elements are input to the bypass encoding engine, they can be coded as follows. For example, the bypass encoding engine of the encoding device omits the procedure of estimating the probability for the input bin and the procedure of updating the probability model applied to the bin after encoding. When bypass encoding is applied, the encoding device can apply a uniform probability distribution instead of assigning a context model to encode the input bin, thereby improving the encoding speed. The bin encoded as described above can be referred to as a bypass bin.

[0075] Entropy decoding can be represented as a process of performing the same process as the above-described entropy encoding in reverse order.

[0076] For example, when the syntax element is decoded based on the context model, the decoding device can receive the bin corresponding to the syntax element via the bit stream, and use the decoding information of the decoding target block or the peripheral block or the information of the symbol / bin decoded in the previous step and the syntax element to determine the context model. The probability of occurrence of the received bin can be predicted by the determined context model, and arithmetic decoding of the bin can be performed to derive the value of the syntax element. Thereafter, the context model of the bin to be decoded next can be updated in the determined context model.

[0077] Also, for example, when a syntax element is bypass decoded, the decoding device can receive a bin corresponding to the syntax element via a bit stream and apply a uniform probability distribution to decode the input bin. In this case, the procedure for deriving the context model of the syntax element and the procedure for updating the context model applied to the bin after decoding can be omitted.

[0078] As described above, residual samples and the like can be derived into quantized transform coefficients and the like through the conversion and quantization processes. The quantized transform coefficients and the like can also be referred to as transform coefficients and the like. In this case, the transform coefficients and the like within the block can be signaled in the form of residual information. The residual information can include a residual coding syntax. That is, the encoding device can construct a residual coding syntax as residual information, encode this, and output it in the form of a bit stream. The decoding device can decode the residual coding syntax from the bit stream to derive residual (quantized) transform coefficients and the like. The residual coding syntax can include syntax elements such as whether a transform has been applied to the block, where the position of the last valid transform coefficient within the block is, whether there are valid transform coefficients within the sub-block, and what the magnitude / symbol of the valid transform coefficient is, as will be described later.

[0079] On one hand, the aforementioned DPB (Decoded Picture Buffer) can conceptually be composed of sub DPBs, and the sub DPB can include a picture storage buffer for storing one layer of decoded pictures. The picture storage buffer can include decoded pictures marked as "used for reference" or held for future output.

[0080] Also, for multilayer bitstreams, the DPB parameters cannot be assigned separately for each OLS (Output Layer Set, OLS), but instead can be assigned for each layer. For example, up to two DPB parameters can be assigned for each layer. One can be assigned when the layer is an output layer (i.e., for example, when the layer can be used for reference and future output), and the other can be assigned when the layer is not an output layer but is used as a reference layer (for example, when the layer can only be used as a reference for pictures / slices / blocks of the output layer in the case of no layer switching). This is considered simpler when compared with the DPB parameters for the multilayer bitstreams of the HEVC layered extension where each layer of the OLS has its own DPB parameters.

[0081] For example, the signaling of the DPB parameters is as follows in terms of the following syntax and semantic.

[0082] [Table 1]

[0083] For example, Table 1 described above can show a VPS (Video Parameter Set) including syntax elements for DPB parameters to be signaled.

[0084] The semantics for the syntax elements shown in Table 1 described above are as follows.

[0085] [Table 2-1]

[0086] [Table 2-2]

[0087] For example, the syntax element vps_num_dpb_params can indicate the number of dpb_parameters() syntax structures in the VPS. For example, the value of vps_num_dpb_params is in the range of 0 to 16. Also, when the syntax element vps_num_dpb_params does not exist, the value of the syntax element vps_num_dpb_params can be regarded as the same as 0.

[0088] Also, for example, the syntax element same_dpb_size_output_or_nonoutput_flag can indicate whether the syntax element layer_nonoutput_dpb_params_idx[i] can exist in the VPS. For example, when the value of the syntax element same_dpb_size_output_or_nonoutput_flag is 1, the syntax element same_dpb_size_output_or_nonoutput_flag can indicate that the syntax element layer_nonoutput_dpb_params_idx[i] does not exist in the VPS, and when the value of the syntax element same_dpb_size_output_or_nonoutput_flag is 0, the syntax element same_dpb_size_output_or_nonoutput_flag can indicate that the syntax element layer_nonoutput_dpb_params_idx[i] can exist in the VPS.

[0089] Also, for example, the syntax element vps_sublayer_dpb_params_present_flag can be used when controlling the existence of the syntax elements max_dec_pic_buffering_minus1[], max_num_reorder_pics[], and max_latency_increase_plus1[] in the dpb_parameters() syntax structure of the VPS. Also, when the syntax element vps_sublayer_dpb_params_present_flag does not exist, the value of the syntax element vps_sublayer_dpb_params_present_flag can be considered to be the same as 0.

[0090] Also, for example, the syntax element dpb_size_only_flag[i] can indicate whether the syntax elements max_num_reorder_pics[] and max_latency_increase_plus1[] can exist in the i-th dpb_parameters() syntax structure of the VPS. For example, when the value of the syntax element dpb_size_only_flag[i] is 1, the syntax element dpb_size_only_flag[i] can indicate that the syntax elements max_num_reorder_pics[] and max_latency_increase_plus1[] do not exist in the i-th dpb_parameters() syntax structure of the VPS, and when the value of the syntax element dpb_size_only_flag[i] is 0, the syntax element dpb_size_only_flag[i] can indicate that the syntax elements max_num_reorder_pics[] and max_latency_increase_plus1[] can exist in the i-th dpb_parameters() syntax structure of the VPS.

[0091] Also, for example, the syntax element dpb_max_temporal_id[i] can indicate the TemporalId of the highest sublayer representation of the DPB parameters that can exist in the i-th dpb_parameters() syntax structure in the VPS. Also, the value of dpb_max_temporal_id[i] is in the range of 0 to vps_max_sublayers_minus1. Also, for example, if the value of vps_max_sublayers_minus1 is 0, the value of dpb_max_temporal_id[i] can be regarded as 0. Also, for example, if the value of vps_max_sublayers_minus1 is greater than 0 and vps_all_layers_same_num_sublayers_flag is 1, the value of dpb_max_temporal_id[i] can be regarded as the same as vps_max_sublayers_minus1.

[0092] Also, for example, the syntax element layer_output_dpb_params_idx[i] can specify the index of the dpb_parameters() syntax structure applied to the i-th layer which is the output layer of the OLS in the list of the dpb_parameters() syntax structures of the VPS. If the syntax element layer_output_dpb_params_idx[i] exists, the value of the syntax element layer_output_dpb_params_idx[i] is in the range of 0 to vps_num_dpb_params - 1.

[0093] For example, if vps_independent_layer_flag[i] is 1, it is the dpb_parameters() syntax structure that exists in the SPS referred to by the dpb_parameters() syntax structure layer applied to the i-th layer which is the output layer.

[0094] Or, for example, when vps_independent_layer_flag[i] is 0, the following content can be applied.

[0095] - When vps_num_dpb_params is 1, the value of layer_output_dpb_params_idx[i] can be regarded as 0.

[0096] - It is a requirement of bitstream conformance that the value of layer_output_dpb_params_idx[i] be such that the value of dpb_size_only_flag[layer_output_dpb_params_idx[i]] is 0.

[0097] Also, for example, the syntax element layer_nonoutput_dpb_params_idx[i] can specify the index of the dpb_parameters() syntax structure applied to the i-th layer, which is a non-output layer of OLS, in the list of dpb_parameters() syntax structures of the VPS. If the syntax element layer_nonoutput_dpb_params_idx[i] exists, the value of the syntax element layer_nonoutput_dpb_params_idx[i] is in the range from 0 to vps_num_dpb_params - 1.

[0098] For example, when same_dpb_size_output_or_nonoutput_flag is 1, the following content can be applied.

[0099] - When vps_independent_layer_flag[i] is 1, the layer of the dpb_parameters() syntax structure applied to the i-th layer, which is a non-output layer, refers to the dpb_parameters() syntax structure in the SPS.

[0100] When -vps_independent_layer_flag[i] is 0, the value of layer_nonoutput_dpb_params_idx[i] can be regarded as the same as that of layer_output_dpb_params_idx[i].

[0101] Or, for example, when same_dpb_size_output_or_nonoutput_flag is 0 and vps_num_dpb_params is 1, the value of layer_output_dpb_params_idx[i] can be regarded as 0.

[0102] On the other hand, for example, the dpb_parameters() syntax structure is the same as the following syntax and semantic.

[0103] [Table 3]

[0104] Referring to Table 3, the dpb_parameters() syntax structure can provide information on the DPB size, the maximum picture reorder number, and the maximum latency for each CLVS of CVS. The dpb_parameters() syntax structure can also be represented by information on DPB parameters or DPB parameter information.

[0105] When the dpb_parameters() syntax structure is included in the VPS, the OLS to which the dpb_parameters() syntax structure is applied can be specified by the VPS. Also, when the dpb_parameters() syntax structure is included in the SPS, the dpb_parameters() syntax structure can be applied to the OLS that includes only the lowest layer among the layers referring to the SPS, where the lowest layer is an independent layer.

[0106] The semantics for the syntax elements shown in Table 3 above are as follows.

[0107]

Table 4

[0108] For example, the value obtained by adding 1 to the syntax element max_dec_pic_buffering_minus1[i] can represent the maximum required size of the DPB in picture buffer units for each CLVS of the CVS when the Htid is the same as i. For example, max_dec_pic_buffering_minus1[i] is information regarding the DPB size. For example, the value of the syntax element max_dec_pic_buffering_minus1[i] is in the range from 0 to MaxDpbSize - 1. Also, for example, when i is greater than 0, max_dec_pic_buffering_minus1[i] is greater than or equal to max_dec_pic_buffering_minus1[i - 1]. Also, for example, if max_dec_pic_buffering_minus1[i] for i in the range from 0 to maxSubLayersMinus1 - 1 does not exist, since subLayerInfoFlag is 0, the value of the syntax element max_dec_pic_buffering_minus1[i] can be regarded as the same as max_dec_pic_buffering_minus1[maxSubLayersMinus1].

[0109] Also, for example, the syntax element max_num_reorder_pics[i] can indicate the maximum allowed number of pictures of a CLVS that can precede all pictures of the CLVS in decoding order for each CLVS of the CVS, and when Htid is the same as i, the corresponding picture can follow in the output order. For example, max_num_reorder_pics[i] is information regarding the maximum picture reorder number of the DPB. The value of max_num_reorder_pics[i] is in the range from 0 to max_dec_pic_buffering_minus1[i]. Also, for example, when i is greater than 0, max_num_reorder_pics[i] is greater than or equal to max_num_reorder_pics[i - 1]. Also, for example, if there is no max_num_reorder_pics[i] for i within the range from 0 to maxSubLayersMinus1 - 1, since subLayerInfoFlag is 0, the syntax element max_num_reorder_pics[i] can be regarded as the same as max_num_reorder_pics[maxSubLayersMinus1].

[0110] Also, for example, the syntax element max_latency_increase_plus1[i] whose value is not 0 can be used when calculating the value of MaxLatencyPictures[i]. The MaxLatencyPictures[i] can indicate the maximum number of pictures of a CLVS that can precede all pictures of the CLVS in output order for each CLVS of the CVS, and when Htid is the same as i, the corresponding picture can follow in the decoding order. For example, max_latency_increase_plus1[i] is information regarding the maximum latency of the DPB.

[0111] For example, when max_latency_increase_plus1[i] is not 0, the value of MaxLatencyPictures[i] can be derived as follows.

[0112]

Equation

[0113] On the other hand, for example, when max_latency_increase_plus1[i] is 0, the corresponding restriction is not displayed. The value of max_latency_increase_plus1[i] is in the range of 0 to 2 32 -2. Also, for example, when max_latency_increase_plus1[i] for i within the range of 0 to maxSubLayersMinus1 - 1 does not exist, since subLayerInfoFlag is 0, the syntax element max_latency_increase_plus1[i] can be regarded as the same as max_latency_increase_plus1[maxSubLayersMinus1].

[0114] On the other hand, the DPB parameters can be used for the output and removal of the picture process as shown in the following table.

[0115]

Table 5-1

[0116]

Table 5-2

[0117] On the other hand, the DPB parameter signaling design in the existing VVC standard has at least the following problems.

[0118] First, although the VVC draft text considered the concept of the sub-DPB, a physical decoding device can have only one DPB for decoding multilayer bitstreams. Therefore, the decoding device must know the DPB size requirements before decoding the OLS with a given multilayer bitstream, but it is not clearly disclosed in the existing VVC draft text how such information can be known.

[0119] For example, the DPB size required for the OLS in the bitstream is not simply derived from the sub-DPB sizes of each layer of the OLS. That is, the DPB size required for the OLS is not simply derived as the sum of the max_dec_pic_buffering_minus1[] + 1 values of the layers within the OLS. For example, the sum of max_dec_pic_buffering_minus1[] + 1 for each layer in the OLS is larger than the actual DPB size. For example, in a specific access unit, each layer within the DPB can have a different reference picture list structure, and the number of reconstructed pictures for each layer within the DPB is not the maximum, and therefore, the DPB size required for the OLS is not simply derived as the sum of the max_dec_pic_buffering_minus1[] + 1 values of the layers within the OLS.

[0120] For example, the following table exemplarily shows the pictures required when there are two spatial scalability layers, the GOP size is 16, and there are no temporal sub-layers for a bitstream where each sub-DPB exists.

[0121]

Table 6

[0122] Referring to Table 6, the base layer (i.e., layer 0) has a more complex RPL structure than layer 1, and the size of sub-DPB0 can include more reference pictures than sub-DPB1 considering the picture size between the two layers. Also, for example, as shown in Table 6, the maximum number of reference pictures of the two layers (i.e., 12) is larger than the number of actual overall pictures of the DPB (i.e., 11).

[0123] Second, the bumping process is not called when it is actually needed. Using the example described above, after the first slice header of the picture with POC (picture order count) 37 of the second layer is decoded, the number of pictures in sub-DPB1 does not reach the maximum sub-DPB size, so the bumping process is not called. That is, when including the picture with POC 37, the number of pictures in sub-DPB1 is 4, and the maximum number of pictures in the said sub-DPB1 can increase to 5. However, since the maximum number of pictures in the DPB has already been reached, the bumping process must be called at that point. Such a problem can occur because only the DPB parameters of the current layer are checked for the condition to call the bumping process. Here, for example, the said bumping process can mean the process of deriving the pictures necessary for output among the pictures in the DPB and removing the pictures not used for reference from the DPB.

[0124] Therefore, this document proposes a solution to the above-mentioned problem. The proposed embodiments can be applied individually or in combination.

[0125] As an example, in addition to signaling the DPB parameters mapped to each layer, a solution is proposed to signal the DPB parameters mapped to OLS.

[0126] Also, as an example, the signaling of DPB parameters mapped to OLS is an option. When there are no DPB parameters mapped to OLS, a method is proposed to derive the value of max_dec_pic_buffering_minus1[i] in the same way as the value obtained by subtracting 1 from the sum of the values obtained by adding 1 to max_dec_pic_buffering_minus1[i] for all layers in OLS. The method proposed in this embodiment can be executed based on a flag indicating whether there are DPB parameters mapped to OLS. For example, when the value of the flag is 1, the flag can indicate that there are DPB parameter indexes for all OLSs including at least one or more layers. Otherwise, that is, when the value of the flag is 0, the flag can indicate that there are no DPB parameters (i.e., DPB parameter indexes for OLS) mapped to OLS. On the other hand, for example, the flag can also exist for each OLS.

[0127] Also, as an example, a method can be proposed to ensure that the value of max_dec_pic_buffering_minus1[hightest temporal sublayer] for each OLS is not greater than the value obtained by subtracting 1 from MaxDpbSize and the value obtained by subtracting 1 from the sum of the values obtained by adding 1 to Imax_dec_pic_buffering_minus1[hightes temporal sublayer] for the layers in OLS.

[0128] Also, as an example, a method can be proposed where the DPB parameters assigned to OLS include only the DPB size.

[0129] Also, as an example, a method can be proposed to update the conditions for calling the bumping process by considering the number of pictures in the DPB and the value of max_dec_pic_buffering_minus1[i] of the OLS being processed in the decoding device.

[0130] On the one hand, for example, an embodiment (etc.) can be applied according to the following procedure.

[0131] FIG. 5 exemplarily shows an encoding procedure according to an embodiment of this document.

[0132] Referring to FIG. 5, an encoding device can decode a (restored) picture (S500). The encoding device can update the DPB based on DPB parameters (S510). For example, the decoded picture can basically be inserted into the DPB, and the decoded picture can be used as a reference picture for inter prediction. Also, a picture decoded by the DPB can be deleted based on the DPB parameters. Further, the encoding device can encode video information including the DPB parameters (S520). Also, although not shown, the encoding device can further decode the current picture based on the DPB updated after step S510. Also, the decoded current picture can be inserted into the DPB, and the DPB including the decoded current picture can be further updated based on the DPB parameters before decoding the picture next in the decoding order.

[0133] FIG. 6 exemplarily shows a decoding procedure according to an embodiment of this document.

[0134] Referring to FIG. 6, a decoding device can obtain video information including information on DPB parameters from a bitstream (S600). The decoding device can output a picture decoded by the DPB based on the information on the DPB parameters (S605). On the other hand, when a layer related to the DPB (or DPB parameters) is a reference layer that is not an output layer, the step S605 can also be omitted.

[0135] Also, the decoding device can update the DPB based on the information regarding the DPB parameters (S610). The decoded picture can basically be inserted into the DPB. Thereafter, the DPB can be updated before decoding the current picture. For example, the picture decoded in the DPB based on the information regarding the DPB parameters can also be deleted. Here, DPB updating can also be referred to as DPB management.

[0136] The information regarding the DPB parameters can include the information / syntax elements disclosed in Table 1 and Table 3 described above. Also, for example, depending on whether the current layer is the output layer or the reference layer, other DPB parameters (etc.) can be signaled, or depending on whether the DPB (or the DPB parameters) is for the OLS (mapped to the OLS) as in the embodiments proposed in this document, other DPB parameters (etc.) can be signaled.

[0137] On the other hand, the decoding device can decode the current picture based on the DPB (S620). For example, the decoding device can decode the current picture based on the inter prediction for the blocks / slices of the current picture using the pictures decoded in the DPB (before the current picture) as reference pictures.

[0138] On the other hand, although not shown in the figure, the encoding device can decode the current picture based on the DPB updated after the step S510 described above. Also, the decoded current picture can be inserted into the DPB, and the DPB including the decoded current picture can be further updated based on the DPB parameters before decoding the next picture.

[0139] The syntax and DPB management process to which the embodiments proposed in this document are applied are as follows.

[0140] As an example, the signaled VPS (video parameter set) syntax is as follows.

[0141]

Table 7

[0142] Referring to Table 7, the VPS can include the syntax elements vps_num_dpb_params, same_dpb_size_output_or_nonoutput_flag, vps_sublayer_dpb_params_present_flag, dpb_size_only_flag[i], dpb_max_temporal_id[i], layer_output_dpb_params_idx[i] and / or layer_nonoutput_dpb_params_idx[i].

[0143] Also, referring to Table 7, the VPS can further include the syntax elements vps_ols_dpb_params_present_flag and / or ols_dpb_params_idx[i].

[0144] For example, the syntax element vps_ols_dpb_params_present_flag can indicate whether ols_dpb_params_idx[] can exist. For example, when the value of vps_ols_dpb_params_present_flag is 1, vps_ols_dpb_params_present_flag can indicate that ols_dpb_params_idx[] can exist, and when the value of vps_ols_dpb_params_present_flag is 0, vps_ols_dpb_params_present_flag can indicate that ols_dpb_params_idx[] does not exist. On the other hand, when vps_ols_dpb_params_present_flag does not exist, the value of vps_ols_dpb_params_present_flag can be regarded as 0.

[0145] Also, for example, when i is smaller than TotalNumOlss, vps_ols_dpb_params_present_flag is 1, and vps_num_dpb_params is greater than 1, if NumLayersInOls[i] is greater than 1, the syntax element ols_dpb_params_idx[i] can be signaled. The ols_dpb_params_idx[i] can also be represented as vps_ols_dpb_params_idx[i].

[0146] For example, when NumLayersInOls[i] is greater than 1, the syntax element ols_dpb_params_idx[i] can specify the index of the dpb_parameters() syntax structure applied to the i-th OLS in the list of the dpb_parameters() syntax structure of the VPS. That is, for example, the syntax element ols_dpb_params_idx[i] can indicate the dpb_parameters() syntax structure of the VPS for the target OLS (i.e., the i-th OLS). When ols_dpb_params_idx[i] exists, the value of ols_dpb_params_idx[i] is in the range of 0 to vps_num_dpb_params - 1.

[0147] Also, for example, when NumLayersInOls[i] is the same as 1, the dpb_parameters() syntax structure applied to the i-th OLS can exist in the SPS referred to by the layer in the i-th OLS.

[0148] On the other hand, according to this embodiment, OlsMaxDecPicBufferingMinus1[Htid] can be defined as follows.

[0149] [Table 8]

[0150] For example, referring to Table 8, the value of OlsMaxDecPicBufferingMinus1[Htid] for the target OLS can be derived as follows.

[0151] For example, when the value of vps_ols_dpb_params_present_flag is 1, OlsMaxDecPicBufferingMinus1[Htid] can be derived in the same way as the value of max_dec_pic_buffering_minus1[Htid] in ols_dpb_params_idx[opOlsIdx].

[0152] Also, for example, in cases other than , i.e., when the value of vps_ols_dpb_params_present_flag is 0, OlsMaxDecPicBufferingMinus1[Htid] can be derived as the sum of max_dec_pic_buffering_minus1[Htid]+1 of each layer in the target OLS minus 1.

[0153] Also, according to this embodiment, the picture output and removal process (i.e., the DPB management process) can be defined as follows.

[0154] [Table 9-1]

[0155] [Table 9-2]

[0156] For example, referring to Table 9, the number of pictures in the sub-DPB is greater than or equal to max_dec_pic_buffering_minus1[Htid]+1. Also, for example, the number of pictures in the DPB is greater than or equal to OlsMaxDecPicBufferingMinus1[Htid]+1.

[0157] Also, according to this embodiment, the constraint on the maximum number of pictures in a DPB (i.e., the maximum number of pictures in a DPB) can be updated as follows: Here, the maximum number of pictures in a DPB can also be expressed as the maximum DPB size.

[0158] [Table 10]

[0159] For example, referring to Table 10, when the level is not level 8.5, the value of OlsMaxDecPicBufferingMinus1[Htid]+1 is less than or equal to MaxDpbSize.

[0160] Alternatively, as an example, the signaling VPS (video parameter set) syntax is as follows.

[0161]

Table 11

[0162] Referring to Table 11, the VPS can include the syntax elements vps_num_dpb_params, same_dpb_size_output_or_nonoutput_flag, vps_sublayer_dpb_params_present_flag, dpb_size_only_flag[i], dpb_max_temporal_id[i], layer_output_dpb_params_idx[i], and / or layer_nonoutput_dpb_params_idx[i].

[0163] Also, referring to Table 11, the VPS can further include the syntax elements vps_ols_dpb_params_present_flag and / or ols_dpb_params_idx[i].

[0164] For example, when i is smaller than TotalNumOlss and vps_num_dpb_params is greater than 1, if NumLayersInOls[i] is greater than 1, the syntax element vps_ols_dpb_params_present_flag can be signaled. Different from the case where vps_ols_dpb_params_present_flag is signaled without separate conditions in the embodiment shown in Table 7 above, vps_ols_dpb_params_present_flag can be signaled only when i is smaller than TotalNumOlss and vps_num_dpb_params is greater than 1.

[0165] For example, the syntax element vps_ols_dpb_params_present_flag can indicate whether ols_dpb_params_idx[] can exist. For example, when the value of vps_ols_dpb_params_present_flag is 1, vps_ols_dpb_params_present_flag can indicate that ols_dpb_params_idx[] can exist, and when the value of vps_ols_dpb_params_present_flag is 0, vps_ols_dpb_params_present_flag can indicate that ols_dpb_params_idx[] does not exist. On the other hand, when vps_ols_dpb_params_present_flag does not exist, the value of vps_ols_dpb_params_present_flag can be regarded as 0.

[0166] Also, for example, when vps_ols_dpb_params_present_flag is 1, the syntax element ols_dpb_params_idx[i] can be signaled.

[0167] For example, the syntax element ols_dpb_params_idx[i] can specify the index of the dpb_parameters() syntax structure applied to the i-th OLS in the list of the dpb_parameters() syntax structure of the VPS when NumLayersInOls[i] is greater than 1. That is, for example, the syntax element ols_dpb_params_idx[i] can indicate the dpb_parameters() syntax structure of the VPS for the target OLS (i.e., the i-th OLS). When ols_dpb_params_idx[i] exists, the value of ols_dpb_params_idx[i] is in the range of 0 to vps_num_dpb_params - 1.

[0168] FIG. 7 schematically shows a video encoding method by an encoding apparatus according to this document. The method disclosed in FIG. 7 can be executed by the encoding apparatus disclosed in FIG. 2. Specifically, for example, S700 in FIG. 7 can be executed by the DPB of the encoding apparatus, S710 in FIG. 7 can be executed by the prediction unit and the residual processing unit of the encoding apparatus, and S720 in FIG. 7 can be executed by the entropy encoding unit of the encoding apparatus.

[0169] The encoding device executes a picture management process for the pictures in the DPB based on the DPB (Decoded Picture Buffer) parameter information for the target OLS (Output Layer Set) (S700). For example, the encoding device can execute a picture management process for the pictures in the DPB based on the DPB parameter information for the target OLS. For example, the encoding device can generate and encode the DPB parameter information for the OLS including the DPB parameter information for the target OLS. The video information can include the DPB parameter information for the OLS. For example, the VPS (Video Parameter Set, VPS) syntax can include the DPB parameter information for the OLS.

[0170] For example, the DPB parameter information for the target OLS can include information on the DPB size for the target OLS, information on the maximum picture reordering number of the DPB for the target OLS, and / or information on the maximum latency of the DPB for the target OLS. Here, the DPB size can indicate the maximum number of pictures that the DPB can contain.

[0171] The syntax element of the information on the DPB size for the target OLS is the aforementioned max_dec_pic_buffering_minus1[i], the syntax element of the information on the maximum picture reordering number of the DPB for the target OLS is the aforementioned max_num_reorder_pics[i], and the syntax element of the information on the maximum latency of the DPB for the target OLS is the aforementioned max_latency_increase_plus1[i].

[0172] On one hand, for example, it can be determined whether a bumping process for pictures in the DPB is executed based on the number of pictures in the DPB and information regarding the DPB size for the target OLS. For example, if the number of pictures in the DPB is greater than or equal to a value derived based on the information regarding the DPB size, the bumping process can be executed; if the number of pictures in the DPB is less than the value derived based on the information regarding the DPB size, the bumping process is not executed. Here, for example, the value derived based on the information regarding the DPB size is a value obtained by adding 1 to the value of the information regarding the DPB size.

[0173] The encoding device decodes a current picture based on the picture (S710). For example, the encoding device can decode the current picture based on the pictures in the DPB where the picture management process has been executed. That is, the encoding device can decode the current picture based on the pictures in the updated DPB. For example, the encoding device can perform an inter prediction for blocks in the current picture based on the reference pictures in the updated DPB to derive prediction samples, and generate restored samples and / or a restored picture for the current picture based on the prediction samples. On the other hand, for example, the encoding device can derive residual samples for blocks in the current picture, and generate restored samples and / or a restored picture through the addition of the prediction samples and the residual samples.

[0174] On one hand, for example, an encoding device can generate and encode prediction information for a block of the current picture. In this case, various prediction methods disclosed in this document, such as inter prediction or intra prediction, can be applied. For example, the encoding device can determine whether to perform inter prediction or intra prediction on the block, and can determine a specific inter prediction mode or a specific intra prediction mode based on the RD cost. According to the determined mode, the encoding device can derive prediction samples for the current chroma block. The prediction information can include prediction mode information for the current chroma block. The video information can include the prediction information.

[0175] Also, for example, an encoding device can encode residual information for a block of the picture.

[0176] For example, the encoding device can derive the residual samples through subtraction of the original samples and the prediction samples for the block.

[0177] Thereafter, for example, the encoding device can quantize the residual samples to derive quantized residual samples, can derive transform coefficients based on the quantized residual samples, and can generate and encode the residual information based on the transform coefficients. Or, for example, the encoding device can quantize the residual samples to derive quantized residual samples, can transform the quantized residual samples to derive transform coefficients, and can generate and encode the residual information based on the transform coefficients. The video information can include the residual information. Also, for example, the encoding device can encode the video information and output it in the form of a bitstream.

[0178] The encoding device can generate a restored sample and / or a restored picture through addition of the prediction sample and the residual sample. Thereafter, as described above, in-loop filtering procedures such as deblocking filtering, SAO, and / or ALF procedures can be applied to the restored sample, if necessary, to improve subjective / objective picture quality.

[0179] The encoding device encodes video information (S720). The encoding device can encode video information. For example, the video information can include the DPB parameter information for the target OLS and the OLS DPB parameter index for the target OLS.

[0180] For example, the encoding device can generate and encode the DPB parameter information for the OLS including the DPB parameter information for the target OLS. The video information can include the DPB parameter information for the OLS. For example, the VPS (Video Parameter Set) syntax can include the DPB parameter information for the OLS.

[0181] For example, the DPB parameter information for the target OLS may include information on the DPB size for the target OLS, information on the maximum picture order number of the DPB for the target OLS, and / or information on the maximum latency of the DPB for the target OLS. Here, the DPB size can indicate the maximum number of pictures that the DPB can contain. The syntax element of the information on the DPB size for the target OLS is the aforementioned max_dec_pic_buffering_minus1[i], the syntax element of the information on the maximum picture order number of the DPB for the target OLS is the aforementioned max_num_reorder_pics[i], and the syntax element of the information on the maximum latency of the DPB for the target OLS is the aforementioned max_latency_increase_plus1[i].

[0182] Also, for example, the encoding device can generate and encode an OLS DPB parameter index for the DPB parameter information of the target OLS. The video information can include the OLS DPB parameter index for the DPB parameter information of the target OLS. For example, the VPS syntax can include the OLS DPB parameter index.

[0183] For example, the OLS DPB parameter index for the target OLS can point to the DPB parameter information for the target OLS. For example, the OLS DPB parameter index for the target OLS can point to the DPB parameter information for the target OLS among the DPB parameter information for the OLS. The syntax element of the OLS DPB parameter index is the aforementioned vps_ols_dpb_params_idx[i] or ols_dpb_params_idx[i].

[0184] On the one hand, for example, an encoding device can generate and encode an OLS DPB parameter flag regarding whether the DPB parameter information for the OLS exists. For example, video information can include the OLS DPB parameter flag. Also, for example, the VPS syntax can include the OLS DPB parameter flag. For example, the OLS DPB parameter flag can indicate whether the DPB parameter information for the OLS exists. For example, when the value of the OLS DPB parameter flag is 1, the OLS DPB parameter flag can indicate that the DPB parameter information for the OLS can exist, and when the value of the OLS DPB parameter flag is 0, the OLS DPB parameter flag can indicate that the DPB parameter information for the OLS does not exist. Also, for example, an OLS DPB parameter index can be generated and encoded based on the OLS DPB parameter flag. For example, when the value of the OLS DPB parameter flag is 1, the OLS DPB parameter index may be generated / encoded / signaled, and when the value of the OLS DPB parameter flag is 0, the OLS DPB parameter index may not be generated / encoded / signaled. The syntax element of the OLS DPB parameter flag is the aforementioned vps_ols_dpb_params_present_flag.

[0185] On the one hand, for example, an encoding device can encode prediction information and residual information for blocks within the current picture.

[0186] For example, the encoding device can determine whether to perform inter prediction or intra prediction on the block, and can determine a specific inter prediction mode or a specific intra prediction mode based on the RD cost. According to the determined mode, the encoding device can derive prediction samples for the block. The prediction information can include prediction mode information for the block.

[0187] In addition, the encoding device can generate and encode a reference picture index indicating a reference picture for the block. For example, the prediction information can include the reference picture index. Also, the encoding device can derive motion information for the block, and can generate and encode information related to the motion information. For example, the prediction information can include the reference picture index and information related to the motion information.

[0188] Also, for example, the encoding device can encode residual information for a block of the picture. For example, the encoding device can derive the residual samples through subtraction of the original samples and the prediction samples for the block.

[0189] Thereafter, for example, the encoding device can quantize the residual sample to derive a quantized residual sample, can derive a transform coefficient based on the quantized residual sample, and can generate and encode the residual information based on the transform coefficient. Or, for example, the encoding device can quantize the residual sample to derive a quantized residual sample, can transform the quantized residual sample to derive a transform coefficient, and can generate and encode the residual information based on the transform coefficient. The video information can include the residual information. Also, for example, the encoding device can encode the video information and output it in the form of a bitstream.

[0190] On the other hand, the bitstream including the video information can be transmitted to the decoding device via a network or a (digital) storage medium. Here, the network can include a broadcast network and / or a communication network, etc., and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.

[0191] FIG. 8 schematically shows an encoding device that executes the video encoding method according to this document. The method disclosed in FIG. 7 can be executed by the encoding device disclosed in FIG. 8. Specifically, for example, the DPB of the encoding device in FIG. 8 can execute S700, the prediction unit and the residual processing unit of the encoding device in FIG. 8 can execute S710, and the entropy encoding unit of the encoding device in FIG. 8 can execute S720.

[0192] FIG. 9 schematically shows a video decoding method by a decoding apparatus according to this document. The method disclosed in FIG. 9 can be executed by the decoding apparatus disclosed in FIG. 3. Specifically, for example, S900 in FIG. 9 can be executed by the entropy decoding unit of the decoding apparatus, S910 in FIG. 9 can be executed by the DPB of the decoding apparatus, and S920 in FIG. 9 can be executed by the prediction unit of the decoding apparatus.

[0193] The decoding apparatus acquires video information (S900). The decoding apparatus can acquire the video information from a bitstream. For example, the video information can include an OLS DPB parameter index for a target OLS (Output Layer Set).

[0194] For example, the decoding apparatus can acquire a VPS (Video Parameter Set, VPS) syntax from the bitstream. The video information can include the VPS syntax. The video information can be received in the bitstream. The VPS syntax can include the OLS DPB parameter index for the target OLS. That is, for example, the decoding apparatus can acquire the OLS DPB parameter index for the target OLS with the VPS syntax.

[0195] For example, the OLS DPB parameter index for the target OLS can indicate the DPB parameter information for the target OLS. On the other hand, for example, the video information can include the DPB parameter information for the OLS. The VPS syntax can include the DPB parameter information for the OLS. That is, for example, the decoding device can obtain the DPB parameter information for the OLS from the VPS syntax. For example, the DPB parameter information for the OLS can include the DPB parameter information for the target OLS, and the OLS DPB parameter index for the target OLS can indicate the DPB parameter information for the target OLS with the DPB parameter information for the OLS. The syntax element of the OLS DPB parameter index is the aforementioned vps_ols_dpb_params_idx[i] or ols_dpb_params_idx[i].

[0196] For example, the DPB parameter information for the target OLS can include information on the DPB size for the target OLS, information on the maximum picture order number of the DPB for the target OLS, and / or information on the maximum latency of the DPB for the target OLS. Here, the DPB size can indicate the maximum number of pictures that the DPB can include. The syntax element of the information on the DPB size for the target OLS is the aforementioned max_dec_pic_buffering_minus1[i], the syntax element of the information on the maximum picture order number of the DPB for the target OLS is the aforementioned max_num_reorder_pics[i], and the syntax element of the information on the maximum latency of the DPB for the target OLS is the aforementioned max_latency_increase_plus1[i].

[0197] On the one hand, for example, the decoding device can obtain an OLS DPB parameter flag regarding whether there is DPB parameter information for OLS. For example, the video information can include the OLS DPB parameter flag. Also, for example, the VPS syntax can include the OLS DPB parameter flag. For example, the OLS DPB parameter flag can indicate whether there is the DPB parameter information for the OLS. For example, when the value of the OLS DPB parameter flag is 1, the OLS DPB parameter flag can indicate that the DPB parameter information for the OLS can exist, and when the value of the OLS DPB parameter flag is 0, the OLS DPB parameter flag can indicate that the DPB parameter information for the OLS does not exist. Also, for example, the OLS DPB parameter index can be obtained based on the OLS DPB parameter flag. For example, when the value of the OLS DPB parameter flag is 1, the OLS DPB parameter index may be signaled / obtained, and when the value of the OLS DPB parameter flag is 0, the OLS DPB parameter index may not be signaled / obtained. The syntax element of the OLS DPB parameter flag is the aforementioned vps_ols_dpb_params_present_flag.

[0198] The decoding device executes a picture management process for the pictures in the DPB (Decoded Picture Buffer) based on the video information (S910). For example, the decoding device can execute a picture management process for the pictures in the DPB (Decoded Picture Buffer) based on the video information. For example, the picture management process can be executed based on the DPB parameter information for the target OLS derived based on the OLS DPB parameter index.

[0199] For example, the decoding device can derive DPB parameter information for the target OLS among the DPB parameter information for the OLS based on the OLS DPB parameter index. For example, the decoding device can derive DPB parameter information for the target OLS pointed to by the OLS DPB parameter index among the DPB parameter information for the OLS.

[0200] Also, for example, the decoding device can execute a picture management process for the pictures in the DPB based on the DPB parameter information for the target OLS. The decoding device can update the DPB based on the DPB parameter information. For example, the decoding device can execute a picture management process for the (decoded) pictures in the DPB based on the DPB parameter information. For example, the decoding device can add the decoded picture to the DPB or remove the decoded picture in the DPB. For example, the decoded picture in the DPB can be used as a reference picture for inter prediction for the current picture, or the decoded picture in the DPB can be used as an output picture. The decoded picture can mean a picture decoded before the current picture in the decoding order at the target OLS.

[0201] For example, the decoding device can determine whether a bumping process for the pictures in the DPB is to be executed based on the number of pictures in the DPB and the information regarding the DPB size with respect to the target OLS, and can execute the bumping process for the pictures in the DPB based on the determination result. For example, when the number of pictures in the DPB is greater than or equal to the value derived based on the information regarding the DPB size, the bumping process may be executed, and when the number of pictures in the DPB is less than the value derived based on the information regarding the DPB size, the bumping process may not be executed. Here, for example, the value derived based on the information regarding the DPB size is a value obtained by adding 1 to the value of the information regarding the DPB size.

[0202] The decoding device decodes the current picture based on the picture (S920). The decoding device can decode the current picture based on the pictures in the DPB for which the picture management process has been executed. That is, the decoding device can decode the current picture based on the pictures in the updated DPB. For example, the decoding device can perform inter prediction for blocks in the current picture based on the reference pictures in the DPB to derive prediction samples, and can generate restored samples and / or a restored picture for the current picture based on the prediction samples. On the other hand, for example, the decoding device can derive residual samples for blocks in the current picture based on the residual information received via the bitstream, and can generate restored samples and / or a restored picture through the addition of the prediction samples and the residual samples.

[0203] On one hand, for example, the decoding device can obtain the prediction information and the residual information for the block of the current picture. The video information can include the prediction information and the residual information. For example, the prediction information can include a reference picture index indicating a reference picture for the block. Also, for example, the prediction information can include information related to motion information for the block.

[0204] Also, for example, the residual information can include information such as value information of (quantized) transform coefficients, position information, transform technique, transform kernel, quantization parameter, etc. for the block. Also, for example, the residual information can include a transform skip flag. The transform skip flag can indicate whether transform is applicable to the block.

[0205] As described above, hereinafter, if necessary, in-loop filtering procedures such as deblocking filtering, SAO and / or ALF procedures can be applied to the restored samples to improve subjective / objective image quality.

[0206] FIG. 10 schematically shows a decoding device that executes the video decoding method according to this document. The method disclosed in FIG. 9 can be executed by the decoding device disclosed in FIG. 10. Specifically, for example, the entropy decoding unit of the decoding device in FIG. 10 can execute S900 in FIG. 9, the DPB of the decoding device in FIG. 10 can execute S910 in FIG. 9, and the prediction unit of the decoding device in FIG. 10 can execute S920 in FIG. 9.

[0207] According to the detailed description in this document, DPB parameters for OLS can be signaled, whereby the DPB can be adaptively updated for OLS, and the overall coding efficiency can be improved.

[0208] Also, according to this document, index information indicating DPB parameters for OLS can be signaled, whereby the DPB parameters can be adaptively derived for OLS, and the DPB for OLS can be updated based on the derived DPB parameters to improve the overall coding efficiency.

[0209] In the foregoing embodiments, the method has been described based on a flowchart in a series of steps or blocks, but this document is not limited to the order of the steps, and a certain step can occur in a different order or simultaneously with steps different from the foregoing. Also, those skilled in the art can understand that the steps shown in the flowchart are not exclusive, other steps are included, or one or more of the steps in the flowchart can be deleted without affecting the scope of this document.

[0210] The embodiments described in this document can be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units illustrated in each drawing can be implemented and executed on a computer, processor, microprocessor, controller, or chip. In this case, information for implementation (e.g., information on instructions) or algorithms can be stored on a digital recording medium.

[0211] In addition, the decoding device and the encoding device to which the embodiments of this document are applied can be included in multimedia broadcast transmission / reception devices, mobile communication terminals, home cinema video devices, digital cinema video devices, surveillance cameras, video conferencing devices, real-time communication devices such as video communication, mobile streaming devices, recording media, camcorders, pay-per-view (VoD) service providing devices, over-the-top (OTT) video devices, Internet streaming service providing devices, three-dimensional (3D) video devices, picture phone video devices, transportation means terminals (e.g., vehicle terminals, airplane terminals, ship terminals, etc.), and medical video devices, etc., and can be used to process video signals or data signals. For example, as an over-the-top (OTT) video device, it can be equipped with a game console, Blu-ray player, Internet-connected TV, home theater system, smartphone, tablet PC, digital video recorder (DVR), etc.

[0212] In addition, the processing method to which the embodiments of this document are applied can be produced in the form of a program executed by a computer and can be stored in a computer-readable recording medium. Multimedia data having the data structure according to this document can also be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices in which data that can be read by a computer is stored. The computer-readable recording medium can include, for example, Blu-ray Disc (BD), universal serial bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disk, and optical data storage devices. In addition, the computer-readable recording medium includes media realized in the form of a carrier wave (e.g., transmission via the Internet). Also, the bitstream generated by the encoding method can be stored in a computer-readable recording medium or transmitted via a wired or wireless communication network.

[0213] In addition, the embodiments of this document can be implemented by a computer program product with program code, and the program code can be executed by a computer according to the embodiments of this document. The program code can be stored on a computer-readable carrier.

[0214] FIG. 11 exemplarily shows a structural diagram of a content streaming system to which the embodiments of this document are applied.

[0215] The content streaming system to which the embodiments of this document are applied can generally include an encoding server, a streaming server, a web server, a media repository, a user device, and a multimedia input device.

[0216] The encoding server compresses the content input from a multimedia input device such as a smartphone, a camera, or a camcorder into digital data to generate a bitstream and transmits it to the streaming server. As another example, when a multimedia input device such as a smartphone, a camera, or a camcorder directly generates a bitstream, the encoding server can be omitted.

[0217] The bitstream can be generated by an encoding method or a bitstream generation method to which the embodiments of this document are applied, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.

[0218] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server serves as a medium for informing the user of what services are available. When the user requests a desired service from the web server, the web server transmits this to the streaming server, and the streaming server transmits multimedia data to the user. At this time, the content streaming system may include another control server, and in this case, the control server plays a role in controlling commands / responses between each device within the content streaming system.

[0219] The streaming server can receive content from a media repository and / or an encoding server. For example, when it comes to receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.

[0220] Examples of the user device include mobile phones, smartphones, laptop computers, digital broadcast terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation devices, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head-mounted displays (HMDs)), digital TVs, desktop computers, digital signage, etc. Each server within the content streaming system can be operated as a distributed server, and in this case, the data received by each server can be processed distributively.

[0221] The claims described in this specification can be combined in various ways. For example, the technical features of the method claims in this specification can be combined and implemented as an apparatus, and the technical features of the apparatus claims in this specification can be combined and implemented as a method. Also, the technical features of the method claims in this specification and the technical features of the apparatus claims can be combined and implemented as an apparatus, and the technical features of the method claims in this specification and the technical features of the apparatus claims can be combined and implemented as a method.

Claims

1. In a method for video decoding executed by a decoding device, a step of acquiring video information; a step of executing a picture management process for pictures in a DPB (Decoded Picture Buffer) based on the video information; a step of decoding a current picture based on the picture, and includes: the video information includes an OLS DPB parameter index for a target OLS (Output Layer Set); the picture management process is executed based on DPB parameter information for the target OLS derived based on the OLS DPB parameter index and the number of pictures in the DPB, the method.

2. The method according to claim 1, wherein the video information includes an OLS DPB parameter flag for whether DPB parameter information for the OLS exists.

3. When the value of the OLS DPB parameter flag is 1, the OLS DPB parameter flag indicates that the DPB parameter information for the OLS exists; The method according to claim 2, wherein when the value of the OLS DPB parameter flag is 0, the OLS DPB parameter flag indicates that the DPB parameter information for the OLS does not exist.

4. The method according to claim 3, wherein the OLS DPB parameter index is acquired based on the OLS DPB parameter flag.

5. The method according to claim 4, wherein when the value of the OLS DPB parameter flag is 1, the OLS DPB parameter index is acquired.

6. The method according to claim 2, wherein the OLS DPB parameter index and the OLS DPB parameter flag are included in a VPS (Video Parameter Set) syntax.

7. The method according to claim 1, wherein the DPB parameter information for the target OLS includes information regarding the DPB size, information regarding the maximum picture order number of the DPB for the target OLS, and information regarding the maximum latency of the DPB.

8. The method according to claim 7, wherein the DPB parameter information for the target OLS is included in a VPS syntax.

9. In a method of video encoding executed by an encoding device, executing a picture management process for pictures in a DPB (Decoded Picture Buffer) based on DPB parameter information for a target OLS (Output Layer Set) and the number of pictures in the DPB; decoding a current picture based on the picture; encoding video information, the method comprising: wherein the video information includes the DPB parameter information for the target OLS and an OLS DPB parameter index for the target OLS. **Claim 10** The method according to claim 9, wherein the video information includes an OLS DPB parameter flag indicating whether there is DPB parameter information for the OLS. **Claim 11** When the value of the OLS DPB parameter flag is 1, the OLS DPB parameter flag indicates that there is DPB parameter information for the OLS, and when the value of the OLS DPB parameter flag is 0, the OLS DPB parameter flag indicates that there is no DPB parameter information for the OLS, according to the method of claim 10. **Claim 12** The method according to claim 11, wherein the OLS DPB parameter index is encoded based on the OLS DPB parameter flag. **Claim 13** The method according to claim 12, wherein when the value of the OLS DPB parameter flag is 1, the OLS DPB parameter index is encoded. **Claim 14** The method according to claim 9, wherein the DPB parameter information for the target OLS includes information regarding the DPB size, information regarding the maximum picture reorder number of the DPB for the target OLS, and information regarding the maximum latency of the DPB. **Claim 15** A method for transmitting data for video, comprising: generating a bitstream of video information including DPB (Decoded Picture Buffer) parameter information and an OLS DPB parameter index for a target OLS (Output Layer Set); A step of transmitting the data including the bitstream of the video information including the DPB parameter information and the OLS DPB parameter index for the target OLS, is included. A method, wherein a picture management process for a picture of the DPB is executed based on the DPB parameter information for the target OLS and the number of pictures in the DPB.

Citation Information

Patent Citations

  • Signalization change in output layer set

    JP2016518763A

  • Scaling list signaling and parameter set activation

    JP2016530734A

  • Signaling for sub-decoded picture buffer (sub-DPB) based DPB operation in video coding.

    JP2016539537A