Image decoding method and apparatus for coding DPB parameters

By deriving and updating DPB parameters based on OLS indices, the method improves video coding efficiency, addressing the increased costs associated with high-resolution images.

JP2026012339AActive Publication Date: 2026-01-23LG ELECTRONICS INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025183310
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-12-30
Filing Date
2025-10-30
Publication Date
2026-01-23
Estimated Expiration
2040-12-29

AI Technical Summary

Technical Problem

The increasing demand for high-resolution, high-quality images leads to higher transmission and storage costs due to the increased amount of information, necessitating more efficient video coding techniques.

Method used

A method and apparatus for deriving and updating Decoded Picture Buffer (DPB) parameters based on Output Layer Set (OLS) DPB parameter indices to improve video coding efficiency.

Benefits of technology

Adaptive updating of DPB parameters enhances overall coding efficiency by optimizing video decoding processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026012339000001_ABST
    Figure 2026012339000001_ABST
Patent Text Reader

Abstract

To provide a video decoding method.SOLUTION: A video decoding method performed by a decoding device according to the present disclosure includes acquiring video information including DecodedPictureBuffer (DPB) parameter information and an OutputLayerset (OLS) DPB parameter index for a target OLS, deriving DPB parameter information for the target OLS based on the OLSDPB parameter index, updating a DPB based on the DPB parameter information for the target OLS, and decoding a current picture based on the updated DPB.SELECTED DRAWING: Figure 8
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This document relates to video coding technology, and more particularly to a video decoding method and apparatus for coding video information including DPB parameters mapped to an OLS in a video coding system. [Background technology]

[0002] Recently, demand for high-resolution, high-quality images such as HD (High Definition) images and UHD (Ultra High Definition) images has been increasing in various fields. As the resolution and quality of image data increases, the amount of information or bits to be transmitted increases relatively compared to existing image data. Therefore, when image data is transmitted using a medium such as an existing wired or wireless broadband line or when image data is stored using an existing recording medium, transmission costs and storage costs increase.

[0003] Therefore, highly efficient image compression techniques are required to effectively transmit, store and reproduce high-resolution, high-quality image information. Summary of the Invention [Problem to be solved by the invention]

[0004] The technical problem of this document is to provide a method and apparatus for increasing video coding efficiency.

[0005] Another technical problem of this document is to provide a method and apparatus for deriving DPB parameters for OLS. [Means for solving the problem]

[0006] According to one embodiment of the present document, there is provided a video decoding method performed by a decoding device, the method including the steps of: acquiring video information including Decoded Picture Buffer (DPB) parameter information and an OLS DPB parameter index for a target Output Layer Set (OLS), deriving DPB parameter information for the target OLS based on the OLS DPB parameter index, updating a DPB based on the DPB parameter information for the target OLS, and decoding a current picture based on the updated DPB.

[0007] According to another embodiment of the present document, there is provided a decoding device for performing video decoding, including an entropy decoding unit that acquires video information including Decoded Picture Buffer (DPB) parameter information and an OLS (Output Layer Set) DPB parameter index for a target OLS, a DPB that derives DPB parameter information for the target OLS based on the OLS DPB parameter index and updates the DPB based on the DPB parameter information for the target OLS, and a prediction unit that decodes a current picture based on the updated DPB.

[0008] According to another embodiment of the present document, there is provided a video encoding method performed by an encoding device, the method including the steps of generating Decoded Picture Buffer (DPB) parameter information, generating an OLS (Output Layer Set) DPB parameter index for the DPB parameter information of a target OLS, and encoding video information including the DPB parameter information and the OLS DPB parameter index.

[0009] According to another embodiment of the present document, there is provided a video encoding apparatus, which includes an entropy encoding unit that generates DPB (Decoded Picture Buffer) parameter information, generates an OLS (Output Layer Set) DPB parameter index for the DPB parameter information of a target OLS, and encodes video information including the DPB parameter information and the OLS DPB parameter index.

[0010] According to another embodiment of the present document, there is provided a computer-readable digital storage medium storing a bitstream including video information for performing a video decoding method, the video decoding method in the computer-readable digital storage medium including the steps of: acquiring video information including Decoded Picture Buffer (DPB) parameter information and an OLS DPB parameter index for a target Output Layer Set (OLS), deriving DPB parameter information for the target OLS based on the OLS DPB parameter index, updating a DPB based on the DPB parameter information for the target OLS, and decoding a current picture based on the updated DPB. [Effects of the Invention]

[0011] According to this document, DPB parameters for OLS can be signaled, which allows DPB to be updated adaptively to OLS, thereby improving overall coding efficiency.

[0012] According to this document, index information pointing to DPB parameters for OLS can be signaled, thereby enabling DPB parameters to be adaptively derived for OLS, and the DPB for OLS can be updated based on the derived DPB parameters to improve overall coding efficiency. [Brief explanation of the drawings]

[0013] [Figure 1] 1 illustrates schematically an example of a video / image coding system to which embodiments of the present document may be applied. [Figure 2] 1 is a diagram illustrating a schematic configuration of a video / image encoding device to which embodiments of the present document can be applied; [Figure 3] 1 is a diagram illustrating the configuration of a video / image decoding device to which the embodiments of the present document can be applied; [Figure 4] 1 illustrates an exemplary encoding procedure according to an embodiment of the present document. [Figure 5] 1 illustrates an exemplary decoding procedure according to an embodiment of the present document. [Figure 6] 1 illustrates a schematic diagram of an image encoding method using an encoding device according to the present document. [Figure 7] 1 shows a schematic diagram of an encoding device for performing the image encoding method according to the present document; [Figure 8] 1 illustrates an image decoding method using a decoding device according to the present document. [Figure 9] 1 shows a schematic diagram of a decoding device for performing the image decoding method according to the present document; [Figure 10] 1 exemplarily illustrates a structural diagram of a content streaming system to which an embodiment of the present document is applied. DETAILED DESCRIPTION OF THE INVENTION

[0014] This document may be modified in various ways and may have various embodiments. Specific embodiments will be illustrated in the drawings and described in detail. However, this is not intended to limit this document to the specific embodiment. Common terms used in this document are used merely to describe specific embodiments and are not intended to limit the technical ideas of this document. A singular expression includes a plural expression unless the context clearly dictates otherwise. In this specification, terms such as "comprise" or "have" are intended to specify the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, and should be understood not to preclude the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0015] Meanwhile, each component in the drawings described in this document is illustrated independently for the convenience of explaining the different characteristic functions, and does not mean that each component is realized by separate hardware or software. For example, two or more components may be combined to form a single component, or a single component may be divided into multiple components. Embodiments in which each component is integrated and / or separated are also within the scope of this document as long as they do not deviate from the essence of this document.

[0016] Hereinafter, preferred embodiments of the present invention will be described in more detail with reference to the accompanying drawings. Hereinafter, the same reference numerals will be used for the same components in the drawings, and duplicated descriptions of the same components may be omitted.

[0017] FIG. 1 illustrates schematically an example of a video / image coding system in which embodiments of the present document may be applied.

[0018] As shown in Figure 1, a video / image coding system may include a first device (source device) and a second device (receiving device). The source device may transmit encoded video / image information or data to the receiving device in a file or streaming format via a digital recording medium or a network.

[0019] The source device may include a video source, an encoding device, and a transmitting unit. The receiving device may include a receiving unit, a decoding device, and a renderer. The encoding device may be called a video / image encoding device, and the decoding device may be called a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, which may be a separate device or an external component.

[0020] A video source can acquire video / images through a video / image capture, synthesis, or generation process. A video source can include a video / image capture device and / or a video / image generation device. A video / image capture device can include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. A video / image generation device can include, for example, a computer, a tablet, a smartphone, etc., and can (electronically) generate video / images. For example, virtual video / images can be generated via a computer, etc., in which case the video / image capture process can be replaced by a process in which the associated data is generated.

[0021] An encoding device can encode input video / images. The encoding device can perform a series of steps such as prediction, transformation, and quantization for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.

[0022] The transmitter can transmit the encoded video / image information or data output in the form of a bitstream to a receiver of a receiving device via a digital recording medium or a network in the form of a file or streaming. The digital recording medium can include various recording media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The transmitter can include elements for generating a media file in a predetermined file format and elements for transmission via a broadcasting / communication network. The receiver can receive / extract the bitstream and transmit it to a decoding device.

[0023] The decoding device can decode the video / image by performing a series of steps such as inverse quantization, inverse transform, prediction, etc., which correspond to the operations of the encoding device.

[0024] The renderer can render the decoded video / image, and the rendered video / image can be displayed via a display unit.

[0025] This document relates to video / image coding. For example, the methods / embodiments disclosed in this document can be applied to methods disclosed in the versatile video coding (VVC) standard, the essential video coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the second generation of audio video coding standard (AVS2), or next generation video / image coding standards (e.g., H.267 or H.268).

[0026] This document presents various embodiments relating to video / image coding, which, unless otherwise stated, may also be implemented in combination with one another.

[0027] In this document, video may refer to a collection of a series of images over time. A picture generally refers to a unit that shows an image at a specific time, and a subpicture, slice, or tile is a unit that constitutes part of a picture in coding. A subpicture, slice, or tile may contain one or more coding tree units (CTUs). A picture may be composed of one or more subpictures, slices, or tiles. A picture may be composed of one or more groups of tiles. A tile group may contain one or more tiles. A brick may represent a rectangular region of CTU rows within a tile in a picture. A tile may be partitioned into multiple bricks, each consisting of one or more CTU rows within the tile. A tile that is not partitioned into multiple bricks may also be referred to as a brick.A brick scan refers to a specific sequential ordering of CTUs partitioning a picture in which the CTUs are ordered consecutively in CTU raster scan in a brick, bricks within a tile are ordered consecutively in a raster scan of the bricks of the tile, and tiles in a picture are ordered consecutively in a raster scan of the tiles of the picture. Also, a subpicture may represent a rectangular region of one or more slices within a picture. That is, a subpicture contains one or more slices that collectively cover a rectangular region of a picture. A tile is a rectangular region of CTUs within a particular tile column and a particular tile row in a picture.The tile column is a rectangular region of CTUs having a height equal to the height of the picture and a width specified by syntax elements in the picture parameter set. The tile row is a rectangular region of CTUs having a height specified by syntax elements in the picture parameter set and a width equal to the height of the picture. A tile scan refers to a specific sequential ordering of CTUs partitioning a picture in which the CTUs are ordered consecutively in a CTU raster scan in a tile whereas tiles in a picture are ordered consecutively in a raster scan of the tiles of the picture.A slice includes an integer number of bricks of a picture that maybe exclusively contained in a single NAL unit. A slice may consist of either a number of complete tiles or only a consecutive sequence of complete bricks of one tile. In this document, the terms tile group and slice may be used interchangeably. For example, in this document, tile group / tile group header may be called slice / slice header.

[0028] A pixel or a pel may refer to the smallest unit that constitutes one picture (or image). A "sample" may also be used as a term corresponding to a pixel. A sample may generally refer to a pixel or a pixel value, or may refer to only a pixel / pixel value of a luma component, or may refer to only a pixel / pixel value of a chroma component.

[0029] A unit may refer to a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to that region. One unit may include one luma block and two chroma (e.g., cb, cr) blocks. The term unit may be used interchangeably with terms such as block or area. In a general case, an M×N block may include samples (or a sample array) consisting of M columns and N rows, or a set (or an array) of transform coefficients.

[0030] As used herein, "A or B" may mean "A only," "B only," or "both A and B." In other words, as used herein, "A or B" may be interpreted as "A and / or B." For example, as used herein, "A, B, or C" may mean "A only," "B only," "C only," or "any combination of A, B, and C."

[0031] As used herein, a slash ( / ) or a comma may mean "and / or." For example, "A / B" may mean "A and / or B." Thus, "A / B" may mean "A only," "B only," or "both A and B." For example, "A, B, C" may mean "A, B, or C."

[0032] As used herein, "at least one of A and B" can mean "A only," "B only," or "both A and B." Furthermore, as used herein, the expressions "at least one of A or B" and "at least one of A and / or B" can be interpreted in the same way as "at least one of A and B."

[0033] Furthermore, in this specification, "at least one of A, B and C" can mean "A only," "B only," "C only," or "any combination of A, B and C." Furthermore, "at least one of A, B or C" and "at least one of A, B and / or C" can mean "at least one of A, B and C."

[0034] Furthermore, parentheses used in this specification may mean "for example." Specifically, when "prediction (intra prediction)" is displayed, "intra prediction" may be proposed as an example of "prediction." In other words, "prediction" in this specification is not limited to "intra prediction," and "intra prediction" may be proposed as an example of "prediction." Furthermore, when "prediction (i.e., intra prediction)" is displayed, "intra prediction" may be proposed as an example of "prediction."

[0035] Technical features described separately in one drawing in this specification may be realized separately or simultaneously.

[0036] The following drawings are created to illustrate a specific example of the present specification. The names of specific devices and names of specific signals / messages / fields shown in the drawings are provided for illustrative purposes only, and the technical features of the present specification are not limited to the specific names used in the following drawings.

[0037] 2 is a diagram for explaining the configuration of a video / image encoding device to which the embodiments of this document can be applied. Hereinafter, the video encoding device may include an image encoding device.

[0038] As shown in FIG. 2, the encoding device 200 may include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter predictor 221 and an intra predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstructed block generator. The image dividing unit 210, the predicting unit 220, the residual processing unit 230, the entropy encoding unit 240, the adding unit 250, and the filtering unit 260 may be configured by one or more hardware components (e.g., an encoder chipset or a processor) depending on the embodiment. Also, the memory 270 may include a decoded picture buffer (DPB) and may be configured by a digital recording medium. The hardware components may further include the memory 270 as an internal / external component.

[0039] The image division unit 210 may divide an input image (or picture, frame) input to the encoding device 200 into one or more processing units. For example, the processing units may be called coding units (CUs). In this case, the coding units may be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) according to a quad-tree, binary-tree, ternary-tree (QTBTTT) structure. For example, one coding unit may be divided into multiple coding units of deeper depths based on a quad-tree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first, and then the binary tree structure and / or the ternary structure may be applied later. Alternatively, the binary tree structure may be applied first. The coding procedure according to this document may be performed based on the final coding unit that is not further divided. In this case, the largest coding unit may be immediately used as the final coding unit based on coding efficiency according to image characteristics, or the coding unit may be recursively divided into coding units of lower depths as needed, and the coding unit of the optimal size may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration, which will be described later. As another example, the processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may each be divided or partitioned from the final coding unit.The prediction unit is a unit of sample prediction, and the transform unit is a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.

[0040] The term "unit" can be used interchangeably with terms such as "block" or "area." In general, an MxN block can refer to a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally refer to a pixel or pixel value, and can refer to only a pixel / pixel value of the luma component, or only a pixel / pixel value of the chroma component. A sample can also be used as a term corresponding to one pixel or pel of a picture (or image).

[0041] The encoding apparatus 200 may generate a residual signal (residual block, residual sample array) by subtracting a prediction signal (predicted block, prediction sample array) output from the inter prediction unit 221 or the intra prediction unit 222 from an input image signal (original block, original sample array), and the generated residual signal is transmitted to the conversion unit 232. In this case, as shown in the figure, a unit in the encoder 200 that subtracts the prediction signal (predicted block, prediction sample array) from the input image signal (original block, original sample array) may be referred to as a subtraction unit 231. The prediction unit may perform prediction on a current block to be processed (hereinafter, referred to as a current block) and generate a predicted block including prediction samples for the current block. The prediction unit may determine whether intra prediction or inter prediction is applied on a current block or CU basis. The prediction unit may generate various information related to prediction, such as prediction mode information, and transmit the information to the entropy encoding unit 240, as will be described later in the description of each prediction mode. The prediction information can be encoded by the entropy encoding unit 240 and output in the form of a bitstream.

[0042] The intra prediction unit 222 may predict the current block by referring to samples in the current picture. The referenced samples may be located in the neighborhood of the current block or may be located far away, depending on the prediction mode. In intra prediction, prediction modes may include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes may include, for example, DC mode and planar mode. The directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the granularity of the prediction direction. However, this is merely an example, and more or less directional prediction modes may be used depending on the settings. The intra prediction unit 222 may also determine the prediction mode to be applied to the current block using the prediction modes applied to neighboring blocks.

[0043] The inter prediction unit 221 may derive a predicted block for a current block based on a reference block (reference sample array) identified by a motion vector on a reference picture. To reduce the amount of motion information transmitted in inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on an inter prediction direction (such as L0 prediction, L1 prediction, or Bi prediction). In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks in the current picture and temporal neighboring blocks in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block may be called a collocated reference block, a collocated CU (colCU), or the like, and the reference picture including the temporal neighboring block may be called a collocated picture (colPic). For example, the inter predictor 221 may configure a motion information candidate list based on neighboring blocks and generate information indicating which candidates are used to derive a motion vector and / or a reference picture index for the current block. Inter prediction may be performed based on various prediction modes, and for example, in the case of a skip mode or a merge mode, the inter predictor 221 may use motion information of neighboring blocks as motion information of the current block. In the case of the skip mode, unlike in the merge mode, a residual signal may not be transmitted.In the case of motion vector prediction (MVP) mode, the motion vector of the neighboring block is used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.

[0044] The predictor 220 may generate a prediction signal based on various prediction methods, which will be described later. For example, the predictor may apply intra prediction or inter prediction for prediction of a block, or may simultaneously apply intra prediction and inter prediction. This may be referred to as combined inter and intra prediction (CIIP). The predictor may also use an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode may be used for content image / video coding, such as games, such as screen content coding (SCC). IBC basically performs prediction within a current picture, but may be performed similarly to inter prediction in deriving a reference block within the current picture. That is, IBC may use at least one of the inter prediction techniques described herein. The palette mode may be seen as an example of intra coding or intra prediction. When the palette mode is applied, sample values ​​within a picture may be signaled based on information about a palette table and a palette index.

[0045] The prediction signal generated by the prediction unit (including the inter prediction unit 221 and / or the intra prediction unit 222) may be used to generate a reconstructed signal or a residual signal. The transform unit 232 may generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loeve transform (KLT), a graph-based transform (GBT), or a conditionally non-linear transform (CNT). Here, GBT refers to a transform obtained from a graph representing inter-pixel relationship information. CNT refers to a transform obtained based on a prediction signal generated using all previously reconstructed pixels. The transform process may be applied to pixel blocks having the same square size or non-square blocks of variable size.

[0046] The quantizer 233 quantizes the transform coefficients and transmits them to the entropy encoder 240. The entropy encoder 240 encodes the quantized signal (information about the quantized transform coefficients) and outputs it as a bitstream. The information about the quantized transform coefficients may be referred to as residual information. The quantizer 233 may rearrange the quantized transform coefficients in a block form into a one-dimensional vector form based on a coefficient scan order, and may generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. The entropy encoder 240 may perform various encoding methods, such as exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. In addition to the quantized transform coefficients, the entropy encoder 240 may also encode information required for video / image restoration (e.g., values ​​of syntax elements, etc.) together with or separately from the quantized transform coefficients. The encoded information (e.g., encoded video / image information) may be transmitted or stored in the form of a bitstream in units of network abstraction layer (NAL) units. The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / image information may also include general constraint information. Information and / or syntax elements transmitted / signaled from an encoding device to a decoding device in this document may be included in the video / image information. The video / image information may be encoded through the above-described encoding procedure and included in the bitstream.The bitstream may be transmitted via a network or stored on a digital recording medium. Here, the network may include a broadcasting network and / or a communication network, and the digital recording medium may include various recording media such as a USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) for transmitting the signal output from the entropy encoding unit 240 and / or a storage unit (not shown) for storing the signal may be configured as an internal / external element of the encoding device 200, or the transmitter may be included in the entropy encoding unit 240.

[0047] The quantized transform coefficients output from the quantization unit 233 may be used to generate a prediction signal. For example, a residual signal (residual block or residual sample) may be reconstructed by applying inverse quantization and inverse transform to the quantized transform coefficients via the inverse quantization unit 234 and the inverse transform unit 235. The adder 250 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter prediction unit 221 or the intra prediction unit 222. When there is no residual for the current block, such as when skip mode is applied, a predicted block may be used as the reconstructed block. The adder 250 may be referred to as a reconstruction unit or a reconstructed block generator. The generated reconstructed signal may be used for intra prediction of the next block to be processed in the current picture, or may be used for inter prediction of the next picture after filtering, as described below.

[0048] Meanwhile, luma mapping with chroma scaling (LMCS) can be applied during picture encoding and / or reconstruction.

[0049] The filtering unit 260 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 260 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture and store the modified reconstructed picture in the memory 270, specifically, in the DPB of the memory 270. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, etc. The filtering unit 260 may generate various information related to filtering and transmit it to the entropy encoding unit 240, as will be described later in connection with each filtering method. The filtering information may be encoded by the entropy encoding unit 240 and output in the form of a bitstream.

[0050] The modified reconstructed picture transmitted to the memory 270 can be used as a reference picture in the inter prediction unit 221. When inter prediction is applied through this, the encoding apparatus can avoid prediction mismatch between the encoding apparatus 200 and the decoding apparatus 300 and can also improve encoding efficiency.

[0051] The memory 270DPB may store modified reconstructed pictures for use as reference pictures in the inter predictor 221. The memory 270 may store motion information of blocks from which motion information in the current picture is derived (or encoded) and / or motion information of blocks in already reconstructed pictures. The stored motion information may be transmitted to the inter predictor 221 to be used as motion information of spatially neighboring blocks or temporally neighboring blocks. The memory 270 may store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra predictor 222.

[0052] FIG. 3 is a diagram illustrating the configuration of a video / image decoding device to which the embodiments of this document can be applied.

[0053] As shown in FIG. 3, the decoding device 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-predictor 331 and an intra-predictor 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 322. The entropy decoder 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 may be configured as a single hardware component (e.g., a decoder chipset or processor) according to an embodiment. The memory 360 may include a decoded picture buffer (DPB) or may be configured as a digital recording medium. The hardware components may further include a memory 360 as an internal / external component.

[0054] When a bitstream including video / image information is input, the decoding device 300 can reconstruct an image corresponding to the process by which the video / image information was processed by the encoding device of FIG. 2. For example, the decoding device 300 can derive units / blocks based on block division-related information obtained from the bitstream. The decoding device 300 can perform decoding using a processing unit applied by the encoding device. Accordingly, the processing unit for decoding is, for example, a coding unit, and the coding unit can be divided from a coding tree unit or a maximal coding unit according to a quad tree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units can be derived from the coding unit. The reconstructed image signal decoded and output by the decoding device 300 can be reproduced through a playback device.

[0055] The decoding device 300 may receive a signal output from the encoding device of FIG. 2 in the form of a bitstream, and the received signal may be decoded via the entropy decoding unit 310. For example, the entropy decoding unit 310 may parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / image information may also include general constraint information. The decoding device may further decode pictures based on the information on the parameter sets and / or the general constraint information. Signaling / received information and / or syntax elements, which will be described later in this document, may be decoded via the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 310 may decode information in a bitstream based on a coding method such as Exponential-Golomb coding, CAVLC, or CABAC, and output values ​​of syntax elements required for image restoration, quantized values ​​of transform coefficients related to residuals, etc. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element in the bitstream, determines a context model using information on the syntax element to be decoded and decoded information on neighboring and current blocks, or information on symbols / bins decoded in previous steps, predicts the occurrence probability of the bins according to the determined context model, and performs arithmetic decoding of the bins to generate symbols corresponding to the values ​​of each syntax element. After determining the context model, the CABAC entropy decoding method may update the context model using information on the decoded symbols / bins for the context model of the next symbol / bin.Among the information decoded by the entropy decoding unit 310, information related to prediction is provided to a prediction unit (inter prediction unit 332 and intra prediction unit 331), and residual values ​​entropy decoded by the entropy decoding unit 310, i.e., quantized transform coefficients and related parameter information, may be input to a residual processing unit 320. The residual processing unit 320 may derive a residual signal (residual block, residual sample, residual sample array). In addition, among the information decoded by the entropy decoding unit 310, information related to filtering may be provided to a filtering unit 350. Meanwhile, a receiving unit (not shown) that receives a signal output from the encoding device may be further configured as an internal / external element of the decoding device 300, or the receiving unit may be a component of the entropy decoding unit 310. Meanwhile, the decoding device according to this document may be called a video / image / picture decoding device, and the decoding device may be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoding unit 310, and the sample decoder may include at least one of the inverse quantization unit 321, the inverse transform unit 322, the addition unit 340, the filtering unit 350, the memory 360, the inter prediction unit 332, and the intra prediction unit 331.

[0056] The inverse quantization unit 321 may inverse quantize the quantized transform coefficients and output the transform coefficients. The inverse quantization unit 321 may rearrange the quantized transform coefficients in a two-dimensional block format. In this case, the rearrangement may be performed based on the coefficient scanning order performed in the encoding device. The inverse quantization unit 321 may perform inverse quantization on the quantized transform coefficients using a quantization parameter (e.g., quantization step size information) to obtain transform coefficients.

[0057] The inverse transform unit 322 performs inverse transform on the transform coefficients to obtain a residual signal (residual block, residual sample array).

[0058] The prediction unit may perform prediction on a current block and generate a predicted block including prediction samples for the current block. The prediction unit may determine whether intra prediction or inter prediction is applied to the current block based on information about the prediction output from the entropy decoding unit 310, and may determine a specific intra / inter prediction mode.

[0059] The predictor 320 may generate a prediction signal based on various prediction methods, which will be described later. For example, the predictor may apply intra prediction or inter prediction for predicting a block, or may simultaneously apply intra prediction and inter prediction. This may be referred to as combined inter and intra prediction (CIIP). The predictor may also use an intra block copy (IBC) prediction mode or a palette mode for predicting a block. The IBC prediction mode or palette mode may be used for content image / video coding, such as games, such as screen content coding (SCC). IBC basically performs prediction within a current picture, but may be performed similarly to inter prediction in deriving a reference block within the current picture. That is, IBC may use at least one of the inter prediction techniques described in this document. The palette mode may be seen as an example of intra coding or intra prediction. When the palette mode is applied, information regarding a palette table and a palette index may be included in the video / image information and signaled.

[0060] The intra prediction unit 331 may predict a current block by referring to samples in a current picture. The referenced samples may be located in the neighborhood of the current block or may be located far away from the current block depending on the prediction mode. In intra prediction, prediction modes may include a plurality of non-directional modes and a plurality of directional modes. The intra prediction unit 331 may also determine a prediction mode to be applied to the current block using prediction modes applied to neighboring blocks.

[0061] The inter prediction unit 332 may derive a predicted block for the current block based on a reference block (reference sample array) identified by a motion vector on a reference picture. To reduce the amount of motion information transmitted from the inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on the inter prediction direction (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks in the current picture and temporal neighboring blocks in the reference picture. For example, the inter prediction unit 332 may construct a motion information candidate list based on the neighboring blocks and derive a motion vector and / or a reference picture index for the current block based on received candidate selection information. Inter prediction may be performed based on various prediction modes, and the prediction information may include information indicating the inter prediction mode for the current block.

[0062] The adder 340 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the acquired residual signal to a predicted signal (predicted block, predicted sample array) output from a prediction unit (including the inter prediction unit 332 and / or the intra prediction unit 331). When there is no residual for the current block, such as when a skip mode is applied, the predicted block may be used as a reconstructed block.

[0063] The adder 340 may be referred to as a reconstruction unit or a reconstruction block generator. The generated reconstruction signal may be used for intra prediction of a next block to be processed in the current picture, may be output after filtering as described below, or may be used for inter prediction of a next picture.

[0064] Meanwhile, LMCS (luma mapping with chroma scaling) can be applied during the picture decoding process.

[0065] The filtering unit 350 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 350 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture, and may transmit the modified reconstructed picture to the memory 360, specifically, to the DPB of the memory 360. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, etc.

[0066] The (modified) reconstructed picture stored in the DPB of the memory 360 may be used as a reference picture in the inter predictor 332. The memory 360 may store motion information of a block from which motion information in the current picture is derived (or decoded) and / or motion information of a block in an already reconstructed picture. The stored motion information may be transmitted to the inter predictor 260 to be used as motion information of a spatially neighboring block or a temporally neighboring block. The memory 360 may store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra predictor 331.

[0067] In this specification, the embodiments described for the filtering unit 260, inter prediction unit 221, and intra prediction unit 222 of the encoding device 200 can also be applied identically or correspondingly to the filtering unit 350, inter prediction unit 332, and intra prediction unit 331 of the decoding device 300, respectively.

[0068] In this document, at least one of quantization / dequantization and / or transform / inverse transform may be omitted. When the quantization / dequantization is omitted, the quantized transform coefficients may be referred to as transform coefficients. When the transform / inverse transform is omitted, the transform coefficients may be referred to as coefficients or residual coefficients, or may still be referred to as transform coefficients for consistency of expression.

[0069] In this document, quantized transform coefficients and transform coefficients may be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, residual information may include information about the transform coefficient(s), and the information about the transform coefficient(s) may be signaled via a residual coding syntax. Transform coefficients may be derived based on the residual information (or information about the transform coefficient(s)), and scaled transform coefficients may be derived through an inverse transform (scaling) of the transform coefficient(s). Residual samples may be derived based on an inverse transform (transform) of the scaled transform coefficient(s). This may be similarly applied / expressed in other parts of this document.

[0070] Meanwhile, the above-mentioned Decoded Picture Buffer (DPB) can be conceptually composed of sub-DPBs, and each sub-DPB can include a picture storage buffer for storing decoded pictures of one layer. The picture storage buffer can contain decoded pictures that are indicated as "used for reference" or that are retained for future output.

[0071] Also, for multilayer bitstreams, DPB parameters are not assigned per Output Layer Set (OLS), but instead can be assigned for each layer. For example, up to two DPB parameters can be assigned for each layer. One can be assigned if the layer is an output layer (i.e., for example, if the layer can be used for reference and future output), and the other can be assigned if the layer is not an output layer but is used as a reference layer (e.g., if there is no layer switching and the layer can only be used as a reference for the picture / slice / block of the output layer). This is considered simpler when compared with DPB parameters for multilayer bitstreams of HEVC layered extension, where each layer of an OLS has its own DPB parameters.

[0072] For example, the signaling of DPB parameters has the following syntax and semantics:

[0073] [Table 1]

[0074] For example, Table 1 above may show a Video Parameter Set (VPS) that includes syntax elements for signaled DPB parameters.

[0075] The semantics for the syntax elements shown in Table 1 above are as follows:

[0076] [Table 2-1]

[0077] [Table 2-2]

[0078] For example, the syntax element vps_num_dpb_params may indicate the number of dpb_parameters() syntax structures in the VPS. For example, the value of vps_num_dpb_params may range from 0 to 16. Also, if the syntax element vps_num_dpb_params is not present, the value of the syntax element vps_num_dpb_params may be considered to be equal to 0.

[0079] Also, for example, the syntax element same_dpb_size_output_or_nonoutput_flag can indicate whether the syntax element layer_nonoutput_dpb_params_idx[i] can be present in the VPS. For example, if the value of the syntax element same_dpb_size_output_or_nonoutput_flag is 1, the syntax element same_dpb_size_output_or_nonoutput_flag can indicate that the syntax element layer_nonoutput_dpb_params_idx[i] does not exist in the VPS, and if the value of the syntax element same_dpb_size_output_or_nonoutput_flag is 0, the syntax element same_dpb_size_output_or_nonoutput_flag can indicate that the syntax element layer_nonoutput_dpb_params_idx[i] can be present in the VPS.

[0080] Also, for example, the syntax element vps_sublayer_dpb_params_present_flag can be used to control the presence of the syntax elements max_dec_pic_buffering_minus1[], max_num_reorder_pics[], and max_latency_increase_plus1[] in the dpb_parameters() syntax structure of the VPS. Also, if the syntax element vps_sublayer_dpb_params_present_flag does not exist, the value of the syntax element vps_sublayer_dpb_params_present_flag can be considered to be the same as 0.

[0081] Also, for example, the syntax element dpb_size_only_flag[i] may indicate whether the syntax elements max_num_reorder_pics[] and max_latency_increase_plus1[] can be present in the i-th dpb_parameters() syntax structure of the VPS. For example, if the value of the syntax element dpb_size_only_flag[i] is 1, the syntax element dpb_size_only_flag[i] can indicate that the syntax elements max_num_reorder_pics[] and max_latency_increase_plus1[] are not present in the i-th dpb_parameters() syntax structure of the VPS, and if the value of the syntax element dpb_size_only_flag[i] is 0, the syntax element dpb_size_only_flag[i] can indicate that the syntax elements max_num_reorder_pics[] and max_latency_increase_plus1[] can be present in the i-th dpb_parameters() syntax structure of the VPS.

[0082] Also, for example, the syntax element dpb_max_temporal_id[i] may indicate the TemporalId of the highest sublayer representation in which a DPB parameter can exist in the i-th dpb_parameters() syntax structure in the VPS. Also, the value of dpb_max_temporal_id[i] is in the range of 0 to vps_max_sublayers_minus1. Also, for example, if the value of vps_max_sublayers_minus1 is 0, the value of dpb_max_temporal_id[i] may be considered to be 0. Also, for example, if the value of vps_max_sublayers_minus1 is greater than 0 and vps_all_layers_same_num_sublayers_flag is 1, the value of dpb_max_temporal_id[i] may be considered to be the same as vps_max_sublayers_minus1.

[0083] Also, for example, the syntax element layer_output_dpb_params_idx[i] can specify the index of the dpb_parameters() syntax structure that applies to the i-th layer, which is the output layer of the OLS, in the list of dpb_parameters() syntax structures of the VPS. If the syntax element layer_output_dpb_params_idx[i] is present, the value of the syntax element layer_output_dpb_params_idx[i] is in the range of 0 to vps_num_dpb_params-1.

[0084] For example, if vps_independent_layer_flag[i] is 1, the dpb_parameters() syntax structure applied to the i-th layer, which is the output layer, is the dpb_parameters() syntax structure present in the SPS referenced by the layer.

[0085] Or, for example, if vps_independent_layer_flag[i] is 0, the following may apply:

[0086] -If vps_num_dpb_params is 1, the value of layer_output_dpb_params_idx[i] can be considered to be 0.

[0087] -It is a bitstream conformance requirement that the value of layer_output_dpb_params_idx[i] is such that the value of dpb_size_only_flag[layer_output_dpb_params_idx[i]] is 0.

[0088] Also, for example, the syntax element layer_nonoutput_dpb_params_idx[i] can specify the index of the dpb_parameters() syntax structure that applies to the i-th layer, which is a non-output layer of the OLS, in the list of dpb_parameters() syntax structures of the VPS. If the syntax element layer_nonoutput_dpb_params_idx[i] is present, the value of the syntax element layer_nonoutput_dpb_params_idx[i] is in the range of 0 to vps_num_dpb_params-1.

[0089] For example, if same_dpb_size_output_or_nonoutput_flag is 1, the following may apply:

[0090] - If vps_independent_layer_flag[i] is 1, the dpb_parameters() syntax structure in the SPS referenced by the layer in the dpb_parameters() syntax structure applied to the i-th layer, which is a non-output layer.

[0091] -If vps_independent_layer_flag[i] is 0, the value of layer_nonoutput_dpb_params_idx[i] can be considered to be the same as layer_output_dpb_params_idx[i].

[0092] Or, for example, if same_dpb_size_output_or_nonoutput_flag is 0, and vps_num_dpb_params is 1, the value of layer_output_dpb_params_idx[i] can be considered to be 0.

[0093] Meanwhile, for example, the dpb_parameters() syntax structure is the same as the following syntax and semantics:

[0094] [Table 3]

[0095] Referring to Table 3, the dpb_parameters() syntax structure can provide information on the DPB size, maximum picture reorder number, and maximum latency for each CLV of the CVS. The dpb_parameters() syntax structure can also be expressed as information on DPB parameters or DPB parameter information.

[0096] If the VPS includes a dpb_parameters() syntax structure, the OLS to which the dpb_parameters() syntax structure applies can be specified by the VPS. Also, if the dpb_parameters() syntax structure is included in the SPS, the dpb_parameters() syntax structure can be applied to an OLS that includes only the lowest layer among the layers that reference the SPS, where the lowest layer is an independent layer.

[0097] The semantics for the syntax elements shown in Table 3 above are as follows:

[0098] [Table 4]

[0099] For example, the value of the syntax element max_dec_pic_buffering_minus1[i] plus 1 can represent the maximum required size of the DPB in picture storage buffer units for each CLVS of the CVS when Htid is the same as i. For example, max_dec_pic_buffering_minus1[i] is information on the DPB size. For example, the value of the syntax element max_dec_pic_buffering_minus1[i] ranges from 0 to MaxDpbSize-1. Also, for example, when i is greater than 0, max_dec_pic_buffering_minus1[i] is greater than or equal to max_dec_pic_buffering_minus1[i-1]. Also, for example, if there is no max_dec_pic_buffering_minus1[i] for i in the range of 0 to maxSubLayersMinus1-1, then the value of the syntax element max_dec_pic_buffering_minus1[i] can be considered to be the same as max_dec_pic_buffering_minus1[maxSubLayersMinus1], since subLayerInfoFlag is 0.

[0100] Also, for example, the syntax element max_num_reorder_pics[i] may indicate the maximum allowed number of pictures of a CLVS that can precede all pictures of the CLVS in decoding order for each CLVS of a CVS and can follow the corresponding picture in output order when Htid is the same as i. For example, max_num_reorder_pics[i] is information on the maximum number of picture reorders of a DPB. The value of max_num_reorder_pics[i] ranges from 0 to max_dec_pic_buffering_minus1[i]. Also, for example, if i is greater than 0, max_num_reorder_pics[i] is greater than or equal to max_num_reorder_pics[i-1]. Also, for example, if there is no max_num_reorder_pics[i] for i in the range of 0 to maxSubLayersMinus1-1, then the syntax element max_num_reorder_pics[i] can be considered to be the same as max_num_reorder_pics[maxSubLayersMinus1], since subLayerInfoFlag is 0.

[0101] Also, for example, a syntax element max_latency_increase_plus1[i] whose value is not 0 can be used when calculating the value of MaxLatencyPictures[i]. MaxLatencyPictures[i] can indicate the maximum number of pictures of a CLVS that can precede all pictures of the CLVS in output order for each CLVS of a CVS and can follow the corresponding picture in decoding order when Htid is the same as i. For example, max_latency_increase_plus1[i] is information on the maximum latency of a DPB.

[0102] For example, if max_latency_increase_plus1[i] is not 0, the value of MaxLatencyPictures[i] can be derived as follows:

[0103]

number

[0104] On the other hand, if max_latency_increase_plus1[i] is 0, the corresponding limit is not displayed. The value of max_latency_increase_plus1[i] ranges from 0 to 2. 32 Also, for example, if there is no max_latency_increase_plus1[i] for i in the range 0 to maxSubLayersMinus1-1, then the syntax element max_latency_increase_plus1[i] can be considered to be the same as max_latency_increase_plus1[maxSubLayersMinus1], since subLayerInfoFlag is 0.

[0105] Meanwhile, the DPB parameters can be used for the output and removal of the picture process as shown in the table below.

[0106] [Table 5-1]

[0107] [Table 5-2]

[0108] On the other hand, the DPB parameter signaling design in the existing VVC standard has at least the following problems:

[0109] First, although the VVC draft text considered the concept of sub-DPBs, a physical decoding device can have only one DPB for decoding multi-layer bitstreams. Therefore, a decoding device must know the DPB size requirements before decoding an OLS in a given multi-layer bitstream, but the existing VVC draft text does not clearly disclose how such information is known.

[0110] For example, the DPB size required for the OLS in a bitstream is not simply derived from the sub-DPB size of each layer of the OLS. That is, the DPB size required for the OLS is not simply derived as the sum of the max_dec_pic_buffering_minus1[]+1 values ​​of the layers in the OLS. For example, the sum of max_dec_pic_buffering_minus1[]+1 of each layer in the OLS is larger than the actual DPB size. For example, in a particular access unit, each layer in the DPB may have a different reference picture list structure, and the number of reconstructed pictures of each layer in the DPB is not maximum. Therefore, the DPB size required for the OLS is not simply derived as the sum of the max_dec_pic_buffering_minus1[]+1 values ​​of the layers in the OLS.

[0111] For example, the following table shows an example of the pictures required when each sub-DPB exists for a bitstream with two spatial scalability layers, a GOP (group of pictures) size of 16, and no temporal sub-layers:

[0112] [Table 6]

[0113] Referring to Table 6, the base layer (i.e., layer 0) has a more complex RPL structure than layer 1, and the size of sub-DPB0 can include more reference pictures than sub-DPB1, taking into account the picture size between the two layers. Also, for example, as shown in Table 6, the maximum number of reference pictures in the two layers (i.e., 12) is greater than the actual total number of pictures in the DPB (i.e., 11).

[0114] Second, the bumping process is not invoked when it is actually needed. Using the above example, after the first slice header of the picture having POC (picture order count) 37 in the second layer is decoded, the number of pictures in sub-DPB1 does not reach the maximum sub-DPB size, so the bumping process is not invoked. That is, if a picture having POC 37 is included, the number of pictures in sub-DPB1 is four, and the maximum number of pictures in sub-DPB1 can be increased to five. However, since the maximum number of pictures in the DPB has already been reached, the bumping process must be invoked at that time. This problem occurs because only the DPB parameters of the current layer are checked as a condition for invoking the bumping process. Here, for example, the bumping process may refer to a process of deriving pictures required for output from the pictures in the DPB and removing pictures not used for reference from the DPB.

[0115] Therefore, this document proposes solutions to the above-mentioned problems. The proposed embodiments can be applied individually or in combination.

[0116] As an example, a method is proposed for signaling DPB parameters mapped to OLS in addition to signaling DPB parameters mapped to each layer.

[0117] Also, as an example, signaling of DPB parameters mapped to an OLS is optional. If there are no DPB parameters mapped to an OLS, a solution is proposed in which the value of max_dec_pic_buffering_minus1[i] is derived as a value obtained by subtracting 1 from the sum of the values ​​of max_dec_pic_buffering_minus1[i] for all layers in the OLS plus 1. The solution proposed in this embodiment may be performed based on a flag indicating whether there are DPB parameters mapped to an OLS. For example, if the flag has a value of 1, the flag may indicate that DPB parameter indexes for all OLSs including at least one layer exist; otherwise, if the flag has a value of 0, the flag may indicate that there are no DPB parameters mapped to an OLS (i.e., DPB parameter indexes for an OLS). Meanwhile, for example, the flag may exist for each OLS.

[0118] As an example, a solution may be proposed in which the value of max_dec_pic_buffering_minus1[highest temporal sublayer] for each OLS is not greater than the sum of MaxDpbSize minus 1 and Imax_dec_pic_buffering_minus1[highest temporal sublayer] plus 1 for the layer within the OLS minus 1.

[0119] Also, as an example, a method may be proposed in which the DPB parameters assigned to the OLS include only the DPB size.

[0120] As another example, a method can be proposed in which the conditions for invoking the bumping process are updated taking into account the number of pictures in the DPB and the value of max_dec_pic_buffering_minus1[i] of the OLS being processed in the decoding device.

[0121] Meanwhile, for example, the embodiment (etc.) can be applied by the following procedure.

[0122] FIG. 4 exemplarily illustrates an encoding procedure according to an embodiment of the present document.

[0123] Referring to FIG. 4, an encoding apparatus may decode a (reconstructed) picture (S400). The encoding apparatus may update the DPB based on DPB parameters (S410). For example, a decoded picture may basically be inserted into the DPB, and the decoded picture may be used as a reference picture for inter-prediction. A decoded picture may also be deleted from the DPB based on the DPB parameters. The encoding apparatus may also encode video information including the DPB parameters (S420). Although not shown, the encoding apparatus may further decode a current picture based on the DPB updated after step S410. The decoded current picture may also be inserted into the DPB, and the DPB including the decoded current picture may further be updated based on the DPB parameters before decoding the next picture in decoding order.

[0124] FIG. 5 exemplarily illustrates a decoding procedure according to an embodiment of the present document.

[0125] 5, a decoding device may obtain video information including information on DPB parameters from a bitstream (S500). The decoding device may output a DPB-decoded picture based on the information on the DPB parameters (S505). Meanwhile, if a layer associated with the DPB (or DPB parameters) is a reference layer that is not an output layer, step S505 may be omitted.

[0126] The decoding device may also update the DPB based on information about the DPB parameters (S510). A decoded picture may basically be inserted into the DPB. Thereafter, the DPB may be updated before decoding the current picture. For example, a decoded picture may be deleted from the DPB based on information about the DPB parameters. Here, DPB updating may also be referred to as DPB management.

[0127] The information for the DPB parameters may include the information / syntax elements disclosed in the above-described Tables 1 and 3. In addition, for example, other DPB parameters (etc.) may be signaled depending on whether the current layer is an output layer or a reference layer, or other DPB parameters (etc.) may be signaled depending on whether the DPB (or DPB parameters) are for OLS (mapped to OLS) as in the embodiments proposed in this document.

[0128] Meanwhile, the decoding device may decode the current picture based on the DPB (S520). For example, the decoding device may decode the current picture based on inter prediction for a block / slice of the current picture using a picture decoded (before the current picture) of the DPB as a reference picture.

[0129] Meanwhile, although not shown, the encoding apparatus may decode the current picture based on the DPB updated after step S410. The decoded current picture may be inserted into the DPB, and the DPB including the decoded current picture may be further updated based on the DPB parameters before decoding the next picture.

[0130] The syntax and DPB management process applied in the embodiments proposed in this document are described below.

[0131] As an example, the signaled VPS (video parameter set) syntax is as follows:

[0132] [Table 7]

[0133] Referring to Table 7, the VPS may include the syntax elements vps_num_dpb_params, same_dpb_size_output_or_nonoutput_flag, vps_sublayer_dpb_params_present_flag, dpb_size_only_flag[i], dpb_max_temporal_id[i], layer_output_dpb_params_idx[i] and / or layer_nonoutput_dpb_params_idx[i].

[0134] Also, referring to Table 7, the VPS may further include syntax elements vps_ols_dpb_params_present_flag and / or ols_dpb_params_idx[i].

[0135] For example, the syntax element vps_ols_dpb_params_present_flag may indicate whether ols_dpb_params_idx[] can be present. For example, if the value of vps_ols_dpb_params_present_flag is 1, vps_ols_dpb_params_present_flag may indicate that ols_dpb_params_idx[] can be present, and if the value of vps_ols_dpb_params_present_flag is 0, vps_ols_dpb_params_present_flag may indicate that ols_dpb_params_idx[] does not exist. On the other hand, if vps_ols_dpb_params_present_flag does not exist, the value of vps_ols_dpb_params_present_flag may be considered to be 0.

[0136] Also, for example, if i is less than TotalNumOlss, vps_ols_dpb_params_present_flag is 1, and vps_num_dpb_params is greater than 1, the syntax element ols_dpb_params_idx[i] may be signaled if NumLayersInOls[i] is greater than 1. The ols_dpb_params_idx[i] may also be expressed as vps_ols_dpb_params_idx[i].

[0137] For example, the syntax element ols_dpb_params_idx[i] may specify the index of the dpb_parameters() syntax structure that applies to the i-th OLS in the list of dpb_parameters() syntax structures of the VPS when NumLayersInOls[i] is greater than 1. That is, for example, the syntax element ols_dpb_params_idx[i] may indicate the dpb_parameters() syntax structure of the VPS for the target OLS (i.e., the i-th OLS). If ols_dpb_params_idx[i] is present, the value of ols_dpb_params_idx[i] is in the range of 0 to vps_num_dpb_params-1.

[0138] Also, for example, if NumLayersInOls[i] is equal to 1, then the dpb_parameters() syntax structure that applies to the ith OLS can exist in the SPS that the layer in the ith OLS references.

[0139] Meanwhile, according to this embodiment, OlsMaxDecPicBufferingMinus1[Htid] can be defined as follows:

[0140] [Table 8]

[0141] For example, referring to Table 8, the value of OlsMaxDecPicBufferingMinus1[Htid] for the target OLS can be derived as follows:

[0142] For example, if the value of vps_ols_dpb_params_present_flag is 1, OlsMaxDecPicBufferingMinus1[Htid] can be derived the same as the value of max_dec_pic_buffering_minus1[Htid] in ols_dpb_params_idx[opOlsIdx].

[0143] Also, for example, in cases other than , i.e., when the value of vps_ols_dpb_params_present_flag is 0, OlsMaxDecPicBufferingMinus1[Htid] can be derived as the sum of max_dec_pic_buffering_minus1[Htid]+1 of each layer in the target OLS minus 1.

[0144] Also, according to this embodiment, the picture output and removal process (i.e., the DPB management process) can be defined as follows.

[0145] [Table 9-1]

[0146] [Table 9-2]

[0147] For example, referring to Table 9, the number of pictures in the sub-DPB is greater than or equal to max_dec_pic_buffering_minus1[Htid]+1. Also, for example, the number of pictures in the DPB is greater than or equal to OlsMaxDecPicBufferingMinus1[Htid]+1.

[0148] Also, according to this embodiment, the constraint on the maximum number of pictures in a DPB (i.e., the maximum number of pictures in a DPB) can be updated as follows: Here, the maximum number of pictures in a DPB can also be expressed as the maximum DPB size.

[0149] [Table 10]

[0150] For example, referring to Table 10, if the level is not level 8.5, the value of OlsMaxDecPicBufferingMinus1[Htid]+1 is less than or equal to MaxDpbSize.

[0151] Alternatively, as an example, the signaled VPS (video parameter set) syntax is as follows:

[0152] [Table 11]

[0153] Referring to Table 11, the VPS may include the syntax elements vps_num_dpb_params, same_dpb_size_output_or_nonoutput_flag, vps_sublayer_dpb_params_present_flag, dpb_size_only_flag[i], dpb_max_temporal_id[i], layer_output_dpb_params_idx[i] and / or layer_nonoutput_dpb_params_idx[i].

[0154] Also, referring to Table 11, the VPS may further include syntax elements vps_ols_dpb_params_present_flag and / or ols_dpb_params_idx[i].

[0155] For example, if i is less than TotalNumOlss, vps_num_dpb_params is greater than 1, and NumLayersInOls[i] is greater than 1, the syntax element vps_ols_dpb_params_present_flag can be signaled. Unlike the embodiment shown in Table 7, in which vps_ols_dpb_params_present_flag is signaled without any additional condition, vps_ols_dpb_params_present_flag can be signaled only when i is less than TotalNumOlss and vps_num_dpb_params is greater than 1.

[0156] For example, the syntax element vps_ols_dpb_params_present_flag may indicate whether ols_dpb_params_idx[] can be present. For example, if the value of vps_ols_dpb_params_present_flag is 1, vps_ols_dpb_params_present_flag may indicate that ols_dpb_params_idx[] can be present, and if the value of vps_ols_dpb_params_present_flag is 0, vps_ols_dpb_params_present_flag may indicate that ols_dpb_params_idx[] does not exist. On the other hand, if vps_ols_dpb_params_present_flag does not exist, the value of vps_ols_dpb_params_present_flag may be considered to be 0.

[0157] Also, for example, if vps_ols_dpb_params_present_flag is 1, the syntax element ols_dpb_params_idx[i] can be signaled.

[0158] For example, the syntax element ols_dpb_params_idx[i] may specify the index of the dpb_parameters() syntax structure that applies to the i-th OLS in the list of dpb_parameters() syntax structures of the VPS when NumLayersInOls[i] is greater than 1. That is, for example, the syntax element ols_dpb_params_idx[i] may indicate the dpb_parameters() syntax structure of the VPS for the target OLS (i.e., the i-th OLS). If ols_dpb_params_idx[i] is present, the value of ols_dpb_params_idx[i] is in the range of 0 to vps_num_dpb_params-1.

[0159] Figure 6 schematically illustrates a video encoding method by an encoding device according to the present document. The method disclosed in Figure 6 may be performed by the encoding device disclosed in Figure 2. Specifically, for example, steps S600 to S620 in Figure 6 may be performed by an entropy encoding unit of the encoding device. Also, although not shown, the process of updating the DPB may be performed by the DPB of the encoding device, and the process of decoding the current picture may be performed by a prediction unit and a residual processing unit of the encoding device.

[0160] An encoding apparatus generates DPB (Decoded Picture Buffer) parameter information (S600). The encoding apparatus can generate and encode the DPB (Decoded Picture Buffer) parameter information. Video information can include the DPB (Decoded Picture Buffer) parameter information. For example, a Video Parameter Set (VPS) syntax can include the DPB parameter information.

[0161] For example, the DPB (Decoded Picture Buffer) parameter information may include DPB parameter information for a target Output Layer Set (OLS). For example, the DPB parameter information for the target OLS may include information on a DPB size for the target OLS, information on a maximum number of picture reorders in the DPB for the target OLS, and / or information on a maximum latency of the DPB for the target OLS. Here, the DPB size may indicate the maximum number of pictures that the DPB can include.

[0162] The syntax element for information regarding the DPB size for the target OLS is the aforementioned max_dec_pic_buffering_minus1[i], the syntax element for information regarding the maximum picture reorder number of the DPB for the target OLS is the aforementioned max_num_reorder_pics[i], and the syntax element for information regarding the maximum latency of the DPB for the target OLS is the aforementioned max_latency_increase_plus1[i].

[0163] Meanwhile, whether or not to perform a bumping process for pictures in the DPB can be determined based on, for example, the number of pictures in the DPB and information on the DPB size for the target OLS. For example, if the number of pictures in the DPB is greater than or equal to a value derived based on the information on the DPB size, the bumping process can be performed, and if the number of pictures in the DPB is less than the value derived based on the information on the DPB size, the bumping process is not performed. Here, for example, the value derived based on the information on the DPB size is a value obtained by adding 1 to the value of the information on the DPB size.

[0164] The encoding apparatus generates an OLS DPB parameter index for DPB parameter information of a target OLS (Output Layer Set) (S610). The encoding apparatus can generate and encode the OLS DPB parameter index for the DPB parameter information of the target OLS. The video information can include the OLS DPB parameter index for the DPB parameter information of the target OLS. For example, the VPS syntax can include the OLS DPB parameter index.

[0165] For example, the OLS DPB parameter index for the target OLS may point to DPB parameter information for the target OLS. For example, the OLS DPB parameter index for the target OLS may point to DPB parameter information for the target OLS in the DPB parameter information. The syntax element of the OLS DPB parameter index is the above-mentioned vps_ols_dpb_params_idx[i] or ols_dpb_params_idx[i].

[0166] Meanwhile, for example, the encoding apparatus may generate and encode an OLS DPB parameter flag indicating whether the DPB parameter information for the target OLS exists. For example, video information may include the OLS DPB parameter flag. Also, for example, the VPS syntax may include the OLS DPB parameter flag. For example, the OLS DPB parameter flag may indicate whether the DPB parameter information for the target OLS exists. For example, if the OLS DPB parameter flag has a value of 1, the OLS DPB parameter flag may indicate that the DPB parameter information for the target OLS exists. If the OLS DPB parameter flag has a value of 0, the OLS DPB parameter flag may indicate that the DPB parameter information for the target OLS does not exist. Also, for example, the OLS DPB parameter index may be generated and encoded based on the OLS DPB parameter flag. For example, if the value of the OLS DPB parameter flag is 1, the OLS DPB parameter index may be generated / encoded / signaled, and if the value of the OLS DPB parameter flag is 0, the OLS DPB parameter index may not be generated / encoded / signaled. The syntax element of the OLS DPB parameter flag is the above-mentioned vps_ols_dpb_params_present_flag.

[0167] An encoding apparatus encodes video information including the DPB parameter information and the OLS DPB parameter index (S620). The encoding apparatus may encode the DPB parameter information and the OLS DPB parameter index. The video information may include the DPB parameter information and the OLS DPB parameter index. The video information may also include the OLS DPB parameter flag.

[0168] Meanwhile, the encoding device may decode a picture of the target OLS. For example, the encoding device may update the DPB based on the DPB parameter information for the target OLS. For example, the encoding device may perform a picture management process for a decoded picture of the DPB based on the DPB parameter information. For example, the encoding device may add a decoded picture to the DPB or remove a decoded picture from the DPB. For example, the decoded picture from the DPB may be used as a reference picture for inter-prediction for the current picture, or the decoded picture from the DPB may be used as an output picture. The decoded picture may refer to a picture decoded before the current picture in the target OLS in decoding order.

[0169] Furthermore, for example, an encoding device may decode the current picture of the target OLS based on the updated DPB. For example, the encoding device may derive predicted samples by performing inter prediction on blocks in the current picture based on reference pictures of the updated DPB, and may generate reconstructed samples and / or reconstructed pictures for the current picture based on the predicted samples. Meanwhile, for example, the encoding device may derive residual samples for blocks in the current picture, and may generate reconstructed samples and / or reconstructed pictures by adding the predicted samples and the residual samples.

[0170] Meanwhile, for example, an encoding apparatus may generate and encode prediction information for a block of the current picture. In this case, various prediction methods disclosed herein, such as inter prediction or intra prediction, may be applied. For example, the encoding apparatus may determine whether to perform inter prediction or intra prediction on the block, and may determine a specific inter prediction mode or a specific intra prediction mode based on an RD cost. Depending on the determined mode, the encoding apparatus may derive prediction samples for the current chroma block. The prediction information may include prediction mode information for the current chroma block. The video information may include the prediction information.

[0171] Also, for example, the encoding device may encode residual information for a block of the picture.

[0172] For example, the encoding apparatus may derive the residual samples through subtraction of the original samples and the predicted samples for the block.

[0173] Thereafter, for example, the encoding apparatus may quantize the residual samples to derive quantized residual samples, derive transform coefficients based on the quantized residual samples, and generate and encode the residual information based on the transform coefficients. Alternatively, for example, the encoding apparatus may quantize the residual samples to derive quantized residual samples, transform the quantized residual samples to derive transform coefficients, and generate and encode the residual information based on the transform coefficients. The video information may include the residual information. Also, for example, the encoding apparatus may encode the video information and output it in the form of a bitstream.

[0174] The encoding apparatus may generate reconstructed samples and / or reconstructed pictures by adding the predicted samples and the residual samples. As described above, in-loop filtering procedures such as deblocking filtering, SAO, and / or ALF procedures may be applied to the reconstructed samples as needed to improve subjective / objective image quality.

[0175] Meanwhile, the bitstream containing the video information can be transmitted to the decoding device via a network or a (digital) storage medium, where the network can include a broadcasting network and / or a communication network, and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.

[0176] Figure 7 schematically illustrates an encoding device that performs the video encoding method according to the present document. The method disclosed in Figure 6 can be performed by the encoding device disclosed in Figure 7. Specifically, for example, the entropy encoding unit of the encoding device in Figure 7 can perform S600 to S620. Also, although not shown, the process of updating the DPB can be performed by the DPB of the encoding device, and the process of decoding the current picture can be performed by a prediction unit and a residual processing unit of the encoding device.

[0177] Figure 8 illustrates a video decoding method by a decoding device according to the present document. The method disclosed in Figure 8 may be performed by the decoding device disclosed in Figure 3. Specifically, for example, S800 of Figure 8 may be performed by an entropy decoding unit of the decoding device, S810 to S820 of Figure 8 may be performed by a DPB of the decoding device, and S830 of Figure 8 may be performed by a prediction unit and a residual processing unit of the decoding device.

[0178] A decoding apparatus acquires video information including DPB (Decoded Picture Buffer) parameter information and an OLS DPB parameter index for a target OLS (Output Layer Set) (S800). The decoding apparatus can acquire video information including DPB (Decoded Picture Buffer) parameter information and an OLS DPB parameter index for a target OLS (Output Layer Set).

[0179] For example, a decoding device may obtain a Video Parameter Set (VPS) syntax from a bitstream. Video information may include the VPS syntax. The video information may be received as a bitstream. The VPS syntax may include the DPB parameter information and the OLS DPB parameter index for the target OLS. That is, for example, a decoding device may obtain the DPB parameter information and the OLS DPB parameter index for the target OLS from the VPS syntax.

[0180] For example, the OLS DPB parameter index for the target OLS may point to DPB parameter information for the target OLS. For example, the DPB parameter information may include DPB parameter information for the target OLS, and the OLS DPB parameter index for the target OLS may point to the DPB parameter information for the target OLS in the DPB parameter information. The syntax element of the OLS DPB parameter index is vps_ols_dpb_params_idx[i] or ols_dpb_params_idx[i], as described above.

[0181] Meanwhile, for example, the decoding apparatus may acquire an OLS DPB parameter flag indicating whether the DPB parameter information for the target OLS exists. For example, the video information may include the OLS DPB parameter flag. Also, for example, the VPS syntax may include the OLS DPB parameter flag. For example, the OLS DPB parameter flag may indicate whether the DPB parameter information for the target OLS exists. For example, if the OLS DPB parameter flag has a value of 1, the OLS DPB parameter flag may indicate that the DPB parameter information for the target OLS exists. If the OLS DPB parameter flag has a value of 0, the OLS DPB parameter flag may indicate that the DPB parameter information for the target OLS does not exist. Also, for example, the OLS DPB parameter index may be acquired based on the OLS DPB parameter flag. For example, if the value of the OLS DPB parameter flag is 1, the OLS DPB parameter index can be signaled / obtained, and if the value of the OLS DPB parameter flag is 0, the OLS DPB parameter index is not signaled / obtained. The syntax element of the OLS DPB parameter flag is vps_ols_dpb_params_present_flag, which has been described in detail.

[0182] The decoding apparatus derives DPB parameter information for the target OLS based on the OLS DPB parameter index (S810). The decoding apparatus may derive the DPB parameter information for the target OLS based on the OLS DPB parameter index. For example, the decoding apparatus may derive DPB parameter information for the target OLS indicated by the OLS DPB parameter index. The DPB parameter information may include the DPB parameter information for the target OLS.

[0183] For example, the DPB parameter information for the target OLS may include information on a DPB size for the target OLS, information on a maximum picture reorder number of a DPB for the target OLS, and / or information on a maximum latency of a DPB for the target OLS, where the DPB size may indicate the maximum number of pictures that the DPB can include.

[0184] The syntax element for information regarding the DPB size for the target OLS is the detailed max_dec_pic_buffering_minus1[i], the syntax element for information regarding the maximum picture reorder number of the DPB for the target OLS is the detailed max_num_reorder_pics[i], and the syntax element for information regarding the maximum latency of the DPB for the target OLS is the detailed max_latency_increase_plus1[i].

[0185] The decoding device updates the DPB based on the DPB parameter information for the target OLS (S820). The decoding device may update the DPB based on the DPB parameter information. For example, the decoding device may perform a picture management process for decoded pictures in the DPB based on the DPB parameter information. For example, the decoding device may add a decoded picture to the DPB or remove a decoded picture in the DPB. For example, the decoded picture in the DPB may be used as a reference picture for inter prediction for the current picture, or the decoded picture in the DPB may be used as an output picture. The decoded picture may refer to a picture decoded before the current picture in the target OLS in decoding order.

[0186] For example, the decoding device may determine whether to perform a bumping process on a picture in the DPB based on the number of pictures in the DPB and information on the DPB size for the target OLS, and may perform the bumping process on the picture in the DPB based on the determination result. For example, if the number of pictures in the DPB is greater than or equal to a value derived based on the information on the DPB size, the bumping process may be performed, and if the number of pictures in the DPB is less than the value derived based on the information on the DPB size, the bumping process may not be performed. Here, for example, the value derived based on the information on the DPB size is a value obtained by adding 1 to the value of the information on the DPB size.

[0187] The decoding device decodes the current picture based on the updated DPB (S830). The decoding device may decode the current picture based on the updated DPB. For example, the decoding device may derive prediction samples by performing inter prediction on blocks in the current picture based on reference pictures of the updated DPB, and may generate reconstructed samples and / or reconstructed pictures for the current picture based on the prediction samples. Meanwhile, for example, the decoding device may derive residual samples for blocks in the current picture based on residual information received via a bitstream, and may generate reconstructed samples and / or reconstructed pictures by adding the prediction samples and the residual samples.

[0188] As mentioned above, in-loop filtering procedures such as deblocking filtering, SAO and / or ALF procedures can be applied to the reconstructed samples to improve the subjective / objective image quality, if necessary.

[0189] Figure 9 schematically illustrates a decoding device that performs the video decoding method according to the present document. The method disclosed in Figure 8 can be performed by the decoding device disclosed in Figure 9. Specifically, for example, the entropy decoding unit of the decoding device of Figure 9 can perform S800 of Figure 8, the DPB of the decoding device of Figure 9 can perform S810 to S820 of Figure 8, and the prediction unit and residual processing unit of the decoding device of Figure 9 can perform S830 of Figure 8.

[0190] According to the detailed description in this document, DPB parameters for OLS can be signaled, which allows the DPB to be updated adaptively to OLS, thereby improving overall coding efficiency.

[0191] In addition, according to this document, index information indicating DPB parameters for OLS can be signaled, thereby enabling DPB parameters to be adaptively derived for OLS, and the DPB for OLS can be updated based on the derived DPB parameters to improve overall coding efficiency.

[0192] In the above-described embodiments, the method is described based on a flow chart as a series of steps or blocks, but this document is not limited to the order of steps, and some steps may occur in a different order or simultaneously with other steps than those described above. Furthermore, those skilled in the art will understand that the steps shown in the flow chart are not exclusive, and other steps may be included, or one or more steps of the flow chart may be deleted without affecting the scope of this document.

[0193] The embodiments described herein may be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units illustrated in the drawings may be implemented and executed on a computer, processor, microprocessor, controller, or chip. In this case, information (e.g., information on instructions) or algorithms for implementation may be stored in a digital recording medium.

[0194] In addition, the decoding device and encoding device to which the embodiments of this document are applied may be included in a multimedia broadcast transmitting / receiving device, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video interaction device, a real-time communication device such as video communication, a mobile streaming device, a recording medium, a camcorder, a custom video (VoD) service providing device, an over-the-top (OTT) video (over-the-top) device, an internet streaming service providing device, a three-dimensional (3D) video device, an image telephone video device, a vehicle terminal (e.g., a vehicle terminal, an airplane terminal, a ship terminal, etc.), a medical video device, etc., and may be used to process video signals or data signals. For example, an over-the-top (OTT) video (over-the-top) device may include a game console, a Blu-ray player, an internet-connected TV, a home theater system, a smartphone, a tablet PC, a digital video recorder (DVR), etc.

[0195] In addition, a processing method to which an embodiment of this document is applied may be produced in the form of a computer-executable program and stored in a computer-readable recording medium. Multimedia data having a data structure according to this document may also be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices in which computer-readable data is stored. The computer-readable recording medium may include, for example, a Blu-ray Disc (BD), a Universal Serial Bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. The computer-readable recording medium may also include media implemented in the form of a carrier wave (e.g., transmission via the Internet). The bitstream generated by the encoding method may be stored in a computer-readable recording medium or transmitted via a wired or wireless communication network.

[0196] Furthermore, the embodiments of the present document may be implemented in a computer program product by program code, which may be executed by a computer in accordance with the embodiments of the present document. The program code may be stored on a computer-readable carrier.

[0197] FIG. 10 exemplarily illustrates a structural diagram of a content streaming system to which the embodiments of this document are applied.

[0198] A content streaming system to which the embodiments of this document are applied can broadly include an encoding server, a streaming server, a web server, a media repository, a user device, and a multimedia input device.

[0199] The encoding server compresses content input from a multimedia input device such as a smartphone, camera, camcorder, etc. into digital data to generate a bitstream and transmits the bitstream to the streaming server. As another example, if a multimedia input device such as a smartphone, camera, camcorder, etc. directly generates a bitstream, the encoding server may be omitted.

[0200] The bitstream can be generated by an encoding method or a bitstream generation method to which an embodiment of this document is applied, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.

[0201] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server acts as an intermediary for informing the user of available services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, which then transmits the multimedia data to the user. The content streaming system may include a separate control server, which controls commands and responses between devices in the content streaming system.

[0202] The streaming server can receive content from a media repository and / or an encoding server. For example, if content is received from the encoding server, the content can be received in real time. In this case, the streaming server can store the bitstream for a certain period of time to provide a smooth streaming service.

[0203] Examples of the user devices include mobile phones, smartphones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, and head mounted displays (HMDs)), digital TVs, desktop computers, digital signage, etc. Each server in the content streaming system can be operated as a distributed server, and in this case, data received by each server can be processed in a distributed manner.

[0204] The claims described in this specification may be combined in various ways. For example, the technical features of the method claims in this specification may be combined to be realized as an apparatus, and the technical features of the apparatus claims in this specification may be combined to be realized as a method. Furthermore, the technical features of the method claims and the technical features of the apparatus claims in this specification may be combined to be realized as an apparatus, and the technical features of the method claims and the technical features of the apparatus claims in this specification may be combined to be realized as a method.

Claims

1. A video decoding method performed by a decoding device, comprising: acquiring video information including Decoded Picture Buffer (DPB) parameter information for one or more Output Layer Sets (OLSs) including a target OLS, and an OLS DPB parameter index for the target OLS; deriving DPB parameter information for the target OLS based on the OLS DPB parameter index; updating a DPB based on the DPB parameter information for the target OLS; decoding the current picture based on the updated DPB; The method, wherein the OLS DPB parameter index specifies an index for the DPB parameter information that applies to the target OLS, which is a multi-layer OLS.

2. The method of claim 1 , wherein the OLS DPB parameter index specifies the DPB parameter information for the target OLS.

3. The method of claim 1 , wherein the DPB parameter information and the OLS DPB parameter index are included in a Video Parameter Set (VPS) syntax.

4. The method of claim 1 , wherein an OLS DPB parameter flag for whether the DPB parameter information for the target OLS exists is obtained.

5. The method of claim 4 , wherein the OLS DPB parameter index is obtained based on the OLS DPB parameter flag.

6. The method of claim 5 , wherein if the value of the OLS DPB parameter flag is 1, the OLS DPB parameter index is obtained.

7. The method of claim 5 , wherein the OLS DPB parameter flag is included in a Video Parameter Set (VPS) syntax.

8. The DPB parameter information for the target OLS includes information on a DPB size, information on a maximum picture reorder number of the DPB for the target OLS, and information on a maximum latency of the DPB; the information regarding the maximum picture reorder number of the DPB for the target OLS specifies a maximum allowable number of pictures in a Coded Layer Video Sequence (CLVS), and the pictures in the CLVS can precede any picture in the CLVS in decoding order and can follow the same picture in output order; 2. The method of claim 1, wherein the information regarding the maximum latency of the DPB specifies a maximum number of pictures in the CLVS that can precede any picture in the CLVS in output order and follow that same picture in decoding order.

9. A video encoding method performed by an encoding device, comprising: generating DPB (Decoded Picture Buffer) parameter information for one or more OLSs (Output Layer Sets) including a target OLS; generating an OLS DPB parameter index for the DPB parameter information of the target OLS; encoding video information including the OLS DPB parameter index and the DPB parameter information; The method, wherein the OLS DPB parameter index specifies an index for the DPB parameter information that applies to the target OLS, which is a multi-layer OLS.

10. The method of claim 9 , wherein the DPB parameter information and the OLS DPB parameter index are included in a Video Parameter Set (VPS) syntax.

11. The method of claim 9 , wherein an OLS DPB parameter flag is encoded to indicate whether the DPB parameter information for the target OLS exists.

12. The method of claim 11 , wherein the OLS DPB parameter index is generated based on the OLS DPB parameter flag.

13. The method of claim 12 , wherein the OLS DPB parameter index is generated if the OLS DPB parameter flag has a value of 1.

14. A method for transmitting data for video, comprising: generating a bitstream of video information including DPB (Decoded Picture Buffer) parameter information for one or more OLSs (Output Layer Sets) including a target OLS, and an OLS DPB parameter index for the DPB parameter information of the target OLS; transmitting the data including the bitstream of the video information including the DPB parameter information and the OLS DPB parameter index; The method, wherein the OLS DPB parameter index specifies an index for the DPB parameter information that applies to the target OLS, which is a multi-layer OLS.

Citation Information

Patent Citations

  • Image decoding method and apparatus for coding DPB parameters

    JP7439267B2

  • Video decoding method and apparatus for coding DPB parameters

    JP7769021B2