Image decoding method and device for encoding DPB parameters
By adaptively updating the decoded picture buffer (DPB) parameters of the output layer set (OLS), the problems of high resolution, high quality image transmission and storage costs are solved, and encoding efficiency is improved.
Patent Information
- Application Number
- CN202080097709.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-30
- Filing Date
- 2020-12-29
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2040-12-29
AI Technical Summary
The prior art when transmitting and storing high-resolution and high-quality images, the increase in the amount of information leads to high costs and requires improving image encoding efficiency.
By obtaining and updating the decoded picture buffer (DPB) parameter index and information of the output layer set (OLS), the DPB is adaptively updated to improve coding efficiency.
By signaling DPB parameters, adaptive update of OLS is realized and the overall coding efficiency is improved.
Smart Images

Figure CN115244932B_ABST
Abstract
Description
Technical Field
[0001] The present document relates to image coding technology, and more particularly, to an image decoding method and apparatus for encoding image information including DPB parameters mapped to an OLS in an image coding system. Background Art
[0002] Recently, demand for high-resolution, high-quality images, such as HD (High Definition) and UHD (Ultra High Definition), has been growing in various fields. Because image data has high resolution and high quality, the amount of information or bits to be transmitted has increased compared to conventional image data. Consequently, when image data is transmitted using media such as conventional wired / wireless broadband lines or stored using existing storage media, transmission and storage costs increase.
[0003] Therefore, there is a need for efficient image compression technology for effectively transmitting, storing, and reproducing information of high-resolution and high-quality images. Summary of the Invention
[0004] Technical issues
[0005] The present disclosure provides a method and apparatus for improving image coding efficiency.
[0006] The present disclosure also provides a method and apparatus for deriving decoded picture buffer (DPB) parameters of an output layer set (OLS).
[0007] Technical Solution
[0008] According to an embodiment of the present document, an image decoding method performed by a decoding device is provided. The method includes the following steps: obtaining image information including an OLS decoded picture buffer (DPB) parameter index and DPB parameter information for a target output layer set (OLS); deriving DPB parameter information for the target OLS based on the OLS DPB parameter index; updating the DPB based on the DPB parameter information for the target OLS; and decoding a current picture based on the updated DPB.
[0009] According to another embodiment of the present document, a decoding device for performing image decoding is provided. The decoding device includes: an entropy decoder configured to obtain image information including an OLS decoded picture buffer (DPB) parameter index and DPB parameter information for a target output layer set (OLS); a DPB configured to derive DPB parameter information for the target OLS based on the OLS DPB parameter index and update the DPB based on the DPB parameter information for the target OLS; and a predictor configured to decode a current picture based on the updated DPB.
[0010] According to another embodiment of the present document, a video encoding method performed by an encoding device is provided. The method includes the following steps: generating decoded picture buffer (DPB) parameter information; generating an OLSDPB parameter index for the DPB parameter information of a target output layer set (OLS); and encoding image information including the OLSDPB parameter index and the DPB parameter information.
[0011] According to another embodiment of the present document, a video encoding device is provided. The encoding device includes an entropy encoder configured to generate decoded picture buffer (DPB) parameter information, generate an OLSDPB parameter index for the DPB parameter information of a target output layer set (OLS), and encode image information including the OLSDPB parameter index and the DPB parameter information.
[0012] According to another embodiment of the present document, a computer-readable digital storage medium storing a bitstream including image information is provided, wherein the bitstream causes a decoding device to perform an image decoding method. In the computer-readable digital storage medium, the image decoding method includes the following steps: obtaining image information including an OLS decoded picture buffer (DPB) parameter index and DPB parameter information for a target output layer set (OLS); deriving DPB parameter information for the target OLS based on the OLS DPB parameter index; updating the DPB based on the DPB parameter information for the target OLS; and decoding a current picture based on the updated DPB.
[0013] Technical Effects
[0014] According to this document, DPB parameters for OLS can be signaled, and thus, DPB can be adaptively updated for OLS and overall coding efficiency can be improved.
[0015] According to this document, index information indicating DPB parameters for OLS can be signaled, and thus, DPB parameters can be adaptively derived for OLS, and overall coding efficiency can be improved by updating the DPB for OLS based on the derived DPB parameters. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 An example of a video / image encoding device to which the embodiments of the present disclosure can be applied is briefly illustrated.
[0017] Figure 2 is a schematic diagram illustrating a configuration of a video / image encoding device to which an embodiment of the present disclosure can be applied.
[0018] Figure 3 FIG. 1 is a schematic diagram illustrating a configuration of a video / image decoding device to which an embodiment of the present disclosure can be applied.
[0019] Figure 4 The encoding process according to the embodiment of this document is exemplified.
[0020] Figure 5 The decoding process according to the embodiment of this document is exemplified.
[0021] Figure 6 The image encoding method performed by the encoding device according to the present disclosure is briefly illustrated.
[0022] Figure 7 An encoding device for performing an image encoding method according to the present disclosure is briefly illustrated.
[0023] Figure 8 The image decoding method performed by the decoding device according to the present disclosure is briefly illustrated.
[0024] Figure 9 A decoding device for performing an image decoding method according to the present disclosure is briefly illustrated.
[0025] Figure 10 A structural diagram of a content streaming system to which the present disclosure is applied is illustrated. DETAILED DESCRIPTION
[0026] The present disclosure can be modified in various forms, and specific embodiments thereof will be described and illustrated in the accompanying drawings. However, the embodiments are not intended to limit the present disclosure. The terms used in the following description are only used to describe specific embodiments and are not intended to limit the present disclosure. As long as it is clearly understood in different ways, singular expressions include plural expressions. Terms such as "including" and "having" are intended to indicate the presence of features, quantities, steps, operations, elements, components, or combinations thereof used in the following description, so it should be understood that the possibility of the presence or addition of one or more different features, quantities, steps, operations, elements, components, or combinations thereof is not excluded.
[0027] Furthermore, the elements in the drawings described in this disclosure are drawn independently for the purpose of explaining different specific functions, and do not imply that these elements are specifically implemented by independent hardware or independent software. For example, two or more elements in the drawings may be combined to form a single element, or an element may be divided into multiple elements. Implementations of combining and / or dividing elements belong to this disclosure and do not depart from the concepts of this disclosure.
[0028] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. In addition, throughout the drawings, like reference numerals are used to indicate like elements, and the same description of like elements will be omitted.
[0029] Figure 1An example of a video / image encoding device to which the embodiments of the present disclosure can be applied is briefly illustrated.
[0030] Reference Figure 1 The video / image coding system may include a first device (source device) and a second device (receiver device). The source device may transmit coded video / image information or data to the receive device in the form of a file or stream via a digital storage medium or a network.
[0031] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.
[0032] The video source can obtain the video / image by capturing, synthesizing, or generating a video / image. The video source may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, a video / image archive including previously captured videos / images, etc. The video / image generation device may include, for example, a computer, a tablet computer, and a smartphone, and may (electronically) generate the video / image. For example, a virtual video / image may be generated by a computer, etc. In this case, the video / image capture process may be replaced by a process that generates relevant data.
[0033] The encoding device can encode the input video / image. The encoding device can perform a series of processes such as prediction, transformation, and quantization to achieve compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0034] The transmitter can transmit encoded video / image information or data in the form of a bitstream to a receiver of a receiving device via a digital storage medium or network in the form of a file or stream. Digital storage media can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter can include components for generating media files in a predetermined file format and can also include components for transmitting via a broadcast / communication network. The receiver can receive / extract the bitstream and transmit the received bitstream to a decoding device.
[0035] The decoding device may decode a video / image by performing a series of processes such as dequantization, inverse transformation, and prediction corresponding to the operations of the encoding device.
[0036] The renderer may render the decoded video / image, and the rendered video / image may be displayed on a display.
[0037] The present disclosure relates to video / image coding. For example, the methods / implementations disclosed in the present disclosure can be applied to methods disclosed in Versatile Video Coding (VVC), EVC (Essential Video Coding) standards, AOMedia Video 1 (AV1) standards, the second generation Audio Video Coding standard (AVS2), or next-generation video / image coding standards (e.g., H.267 or H.268).
[0038] The present disclosure presents various embodiments of video / image encoding, and unless otherwise mentioned, the embodiments may be performed in combination with each other.
[0039] In the present disclosure, video may refer to a series of images over time. Generally, a picture refers to a unit representing an image in a specific time region, and a sub-picture / slice / tile is a unit that constitutes a part of a picture in encoding. A sub-picture / slice / tile may include one or more coding tree units (CTUs). A picture may be composed of one or more sub-pictures / slices / tiles. A picture may be composed of one or more tile groups. A tile group may include one or more tiles. A tile may represent a rectangular area of a CTU row within a tile in a picture. A tile may be partitioned into multiple tiles, each tile consisting of one or more CTU rows within the tile. A tile that is not partitioned into multiple tiles may also be referred to as a tile. Tile scanning is a specific sequential ordering of CTUs that partition a picture, wherein CTUs are continuously ordered in a tile by a CTU raster scan, tiles within a tile are continuously ordered by a raster scan of the tiles of the tile, and tiles in the picture are continuously ordered by a raster scan of the tiles of the picture. In addition, a sub-picture can represent a rectangular area of one or more slices within a picture. That is, a sub-picture contains one or more slices that together cover a rectangular area of the picture. A tile is a rectangular area of a CTU within a specific tile column and a specific tile row in a picture. A tile column is a rectangular area of a CTU whose height is equal to the height of the picture and whose width is specified by a syntax element in the picture parameter set. A tile row is a rectangular area of a CTU whose height is specified by a syntax element in the picture parameter set and whose width is equal to the width of the picture. Tile scan is a specific sequential ordering of CTUs that partition a picture, where CTUs can be ordered continuously within a tile according to a CTU raster scan, while tiles in a picture can be ordered continuously according to a raster scan of the tiles of the picture. A slice includes an integer number of tiles of a picture that can be exclusively contained in a single NAL unit. A slice can consist of multiple complete tiles or only a continuous sequence of complete tiles of a tile. In this disclosure, tile group and slice can be used interchangeably. For example, in this disclosure, a tile / tile header may be referred to as a slice / slice header.
[0040] A pixel or picture element (pel) may represent the smallest unit constituting a picture (or image). In addition, "sample" may be used as a term corresponding to a pixel. A sample may generally represent a pixel or a pixel value, and may represent only a pixel / pixel value of a luminance component or only a pixel / pixel value of a chrominance component.
[0041] A unit may represent a basic unit of image processing. A unit may include at least one of a specific area of a picture and information related to the area. A unit may include a luminance block and two chrominance (e.g., CB, CR) blocks. In some cases, a unit may be used interchangeably with terms such as block or area. In general, an M×N block may include M columns and N rows of samples (or sample arrays) or a set (or array) of transform coefficients.
[0042] In this specification, "A or B" may mean "only A," "only B," or "A and B." In other words, in this specification, "A or B" may be interpreted as "A and / or B." For example, "A, B, or C" herein means "only A," "only B," "only C," or "any one and any combination of A, B, and C."
[0043] As used herein, a slash ( / ) or a comma may mean "and / or." For example, "A / B" may mean "A and / or B." Thus, "A / B" may mean "only A," "only B," or "A and B." For example, "A,B,C" may mean "A, B, or C."
[0044] In this specification, “at least one of A and B” may mean “only A”, “only B”, or “both A and B”. In addition, in this specification, the expression “at least one of A or B” or “at least one of A and / or B” may be interpreted as the same as “at least one of A and B”.
[0045] In addition, in this specification, "at least one of A, B, and C" means "only A", "only B", "only C", or "any combination of A, B, and C". In addition, "at least one of A, B, or C" or "at least one of A, B, and / or C" may mean "at least one of A, B, and C".
[0046] Furthermore, brackets used in this specification may refer to "for example." Specifically, when "prediction (intra-frame prediction)" is indicated, "intra-frame prediction" may be provided as an example of "prediction." In other words, "prediction" in this specification is not limited to "intra-frame prediction," and "intra-frame prediction" may be provided as an example of "prediction." Furthermore, even when "prediction (i.e., intra-frame prediction)" is indicated, "intra-frame prediction" may be provided as an example of "prediction."
[0047] In this specification, technical features described separately in one drawing may be implemented separately or simultaneously.
[0048] The following figures are created to explain specific examples of this specification. Since the names of specific devices or the names of specific signals / messages / fields described in the figures are presented by way of example, the technical features of this specification are not limited to the specific names used in the following figures.
[0049] Figure 2 1 is a schematic diagram illustrating a configuration of a video / image encoding device to which an embodiment of the present disclosure can be applied. Hereinafter, a video encoding device may include an image encoding device.
[0050] Reference Figure 2 , the encoding device 200 includes an image segmenter 210, a predictor 220, a residual processor 230 and an entropy encoder 240, an adder 250, a filter 260 and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234 and an inverse transformer 235. The residual processor 230 may also include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstructed block generator. According to an embodiment, the image segmenter 210, the predictor 220, the residual processor 230, the entropy encoder 240, the adder 250 and the filter 260 may be composed of at least one hardware component (e.g., an encoder chipset or processor). In addition, the memory 270 may include a decoded picture buffer (DPB) or may be composed of a digital storage medium. The hardware components may also include the memory 270 as an internal / external component.
[0051] The image splitter 210 can split the input image (or picture or frame) input to the encoding device 200 into one or more processors. For example, the processor can be referred to as a coding unit (CU). In this case, the coding unit can be recursively split from the coding tree unit (CTU) or the largest coding unit (LCU) according to the quadtree binary tree ternary tree (QTBTTT) structure. For example, a coding unit can be split into multiple coding units with a deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quadtree structure can be applied first, and then the binary tree structure and / or the ternary structure can be applied. Alternatively, the binary tree structure can be applied first. The encoding process according to the present disclosure can be performed based on the final coding unit that is no longer split. In this case, the maximum coding unit can be used as the final coding unit based on coding efficiency according to image characteristics, or if necessary, the coding unit can be recursively split into coding units with a deeper depth and the coding unit with the optimal size can be used as the final coding unit. Here, the encoding process may include prediction, transformation, and reconstruction processes, which will be described later. As another example, the processor may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may be separated or partitioned from the final coding unit. The prediction unit may be a unit for sample prediction, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.
[0052] In some cases, the term "unit" can be used interchangeably with terms such as "block" or "region." In general, an M×N block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or pixel value, and can represent only the pixel / pixel value of the luma component or only the pixel / pixel value of the chroma component. A sample can be used as a term corresponding to a picture (or image) of pixels or picture elements.
[0053] In the encoding device 200, the prediction signal (prediction block, prediction sample array) output from the inter-frame predictor 221 or the intra-frame predictor 222 is subtracted from the input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is sent to the transformer 232. In this case, as shown in the figure, the unit in the encoding device 200 for subtracting the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) can be referred to as a subtractor 231. The predictor can perform prediction on the block to be processed (hereinafter referred to as the current block) and generate a prediction block including prediction samples of the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction on a current block or CU basis. As described later in the description of each prediction mode, the predictor can generate various information related to the prediction (such as prediction mode information) and send the generated information to the entropy encoder 240. The information about the prediction can be encoded in the entropy encoder 240 and output in the form of a bitstream.
[0054] The intra-frame predictor 222 can predict the current block by referencing samples in the current picture. Depending on the prediction mode, the referenced samples can be located near the current block or can be far away from the current block. In intra-frame prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. The non-directional mode can include, for example, a DC mode and a planar mode. Depending on the level of detail of the prediction direction, the directional mode can include, for example, 33 directional prediction modes or 65 directional prediction modes. However, this is only an example, and more or fewer directional prediction modes can be used depending on the settings. The intra-frame predictor 222 can determine the prediction mode applied to the current block by using the prediction mode applied to the adjacent block.
[0055] The inter-frame predictor 221 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. To reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can also include information about the inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, neighboring blocks can include spatially neighboring blocks in the current picture and temporally neighboring blocks in the reference picture. The reference picture including the reference block and the reference picture including the temporally neighboring block can be the same or different. Temporally neighboring blocks can be referred to as collocated reference blocks, collocated CUs (colCUs), etc., and the reference picture including temporally neighboring blocks can be referred to as collocated pictures (colPics). For example, the inter-frame predictor 221 can configure a motion information candidate list based on neighboring blocks and generate information indicating which candidate to use to derive the motion vector and / or reference picture index for the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the inter-frame predictor 221 can use the motion information of the adjacent block as the motion information of the current block. In skip mode, unlike merge mode, it may not be possible to send a residual signal. In the case of motion vector prediction (MVP) mode, the motion vector of the adjacent block can be used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.
[0056] The predictor 220 can generate a prediction signal based on various prediction methods described below. For example, the predictor can not only apply intra prediction or inter prediction to predict a block, but also apply both intra prediction and inter prediction at the same time. This can be called inter-frame intra-frame combined prediction (CIIP). In addition, the predictor can predict the block based on the intra block copy (IBC) prediction mode or palette mode. The IBC prediction mode or palette mode can be used for content image / video encoding of games, etc., for example, screen content coding (SCC). IBC basically performs prediction in the current picture, but can be performed similarly to inter prediction because the reference block is derived in the current picture. That is, IBC can use at least one of the inter prediction techniques described in this disclosure. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, the sample value within the picture can be signaled based on information about the palette table and the palette index.
[0057] The prediction signal generated by the predictor (including the inter-frame predictor 221 and / or the intra-frame predictor 222) can be used to generate a reconstruction signal or to generate a residual signal. The transformer 232 can generate a transform coefficient by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loève transform (KLT), a graph-based transform (GBT), or a conditional nonlinear transform (CNT). Here, GBT refers to a transform obtained from a graph when the relationship information between pixels is represented by a graph. CNT refers to a transform generated based on a prediction signal generated using all previously reconstructed pixels. In addition, the transform process can be applied to square pixel blocks of the same size, or can be applied to blocks of variable size other than square.
[0058] The quantizer 233 can quantize the transform coefficients and send them to the entropy encoder 240. The entropy encoder 240 can encode the quantized signal (information about the quantized transform coefficients) and output a bitstream. The information about the quantized transform coefficients can be called residual information. The quantizer 233 can rearrange the block-type quantized transform coefficients into a one-dimensional vector form based on the coefficient scanning order, and generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. Information about the transform coefficients can be generated. The entropy encoder 240 can perform various encoding methods, such as, for example, exponential Golomb, context-adaptive variable-length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. The entropy encoder 240 can encode information required for video / image reconstruction (e.g., syntax element values, etc.) in addition to the quantized transform coefficients, together or separately. The encoded information (e.g., encoded video / image information) can be transmitted or stored in the form of a bitstream in units of NAL (Network Abstraction Layer). The video / image information may also include information about various parameter sets such as an adaptive parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may also include general constraint information. In the present disclosure, information and / or syntax elements sent / signaled from the encoding device to the decoding device may be included in the video / image information. The video / image information may be encoded through the above-mentioned encoding process and included in the bitstream. The bitstream may be transmitted over a network or may be stored in a digital storage medium. The network may include a broadcast network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) that transmits the signal output from the entropy encoder 240 and / or a storage unit (not shown) that stores the signal may be included as internal / external elements of the encoding device 200. Alternatively, the transmitter may be included in the entropy encoder 240.
[0059] The quantized transform coefficients output from the quantizer 233 can be used to generate a prediction signal. For example, the residual signal (residual block or residual sample) can be reconstructed by applying dequantization and inverse transformation to the quantized transform coefficients using the dequantizer 234 and the inverse transformer 235. The adder 250 adds the reconstructed residual signal to the prediction signal output from the inter-frame predictor 221 or the intra-frame predictor 222 to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). If the block to be processed has no residual (such as when skip mode is applied), the prediction block can be used as a reconstructed block. The adder 250 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current picture, and can be used for inter-frame prediction of the next picture through filtering as described below.
[0060] Furthermore, during picture coding and / or reconstruction, luma mapping and chroma scaling (LMCS) may be applied.
[0061] The filter 260 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 260 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and store the modified reconstructed picture in the memory 270 (specifically, the DPB of the memory 270). Various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filter 260 can generate various information related to filtering and send the generated information to the entropy encoder 240, as described later in the description of various filtering methods. The information related to filtering can be encoded by the entropy encoder 240 and output in the form of a bit stream.
[0062] The modified reconstructed picture sent to the memory 270 may be used as a reference picture in the inter predictor 221. When inter prediction is applied by the encoding apparatus, prediction mismatch between the encoding apparatus 200 and the decoding apparatus may be avoided, and encoding efficiency may be improved.
[0063] The DPB of the memory 270 can store a modified reconstructed picture used as a reference picture in the inter-frame predictor 221. The memory 270 can store motion information of a block from which motion information in the current picture is derived (or encoded) and / or motion information of a reconstructed block in the picture. The stored motion information can be sent to the inter-frame predictor 221 and used as motion information of spatially neighboring blocks or motion information of temporally neighboring blocks. The memory 270 can store reconstructed samples of the reconstructed blocks in the current picture and can transmit the reconstructed samples to the intra-frame predictor 222.
[0064] Figure 3 FIG. 1 is a schematic diagram illustrating a configuration of a video / image decoding device to which an embodiment of the present disclosure can be applied.
[0065] Reference Figure 3 , the decoding device 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-frame predictor 332 and an intra-frame predictor 331. The residual processor 320 may include a dequantizer 321 and an inverse transformer 322. According to an embodiment, the entropy decoder 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 may be composed of hardware components (e.g., a decoder chipset or processor). In addition, the memory 360 may include a decoded picture buffer (DPB) or may be composed of a digital storage medium. The hardware components may also include the memory 360 as an internal / external component.
[0066] When a bit stream including video / image information is input, the decoding device 300 can be used with Figure 2 The image is reconstructed accordingly to the processing of the video / image information in the encoding device. For example, the decoding device 300 can derive the unit / block based on the block segmentation related information obtained from the bit stream. The decoding device 300 can perform decoding using a processor applied in the encoding device. Therefore, the decoding processor can be, for example, a coding unit, and the coding unit can be divided from the coding tree unit or the maximum coding unit according to the quadtree structure, the binary tree structure and / or the ternary tree structure. One or more transform units can be derived from the coding unit. The reconstructed image signal decoded and output by the decoding device 300 can be reproduced by a reproduction device.
[0067] The decoding device 300 can receive the data in the form of a bit stream from Figure 2The signal output by the encoding device of the encoding device can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can parse the bitstream to derive information required for image reconstruction (or picture reconstruction) (for example, video / image information). The video / image information may also include information about various parameter sets such as an adaptive parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may also include general constraint information. The decoding device may also decode the picture based on the information about the parameter set and / or the general constraint information. The signaled / received information and / or syntax elements described later in this disclosure can be decoded through a decoding process and obtained from the bitstream. For example, the entropy decoder 310 decodes the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and outputs the syntax elements required for image reconstruction and the quantized values of the transform coefficients of the residual. More specifically, the CABAC entropy decoding method can receive a bin corresponding to each syntax element in the bitstream, use the decoding target syntax element information, the decoding information of the decoding target block, or the information of the symbol / bin decoded in the previous stage to determine the context model, and arithmetically decode the bin by predicting the probability of occurrence of the bin according to the determined context model, and generate a symbol corresponding to the value of each syntax element. In this case, after determining the context model, the CABAC entropy decoding method can update the context model by using the decoded symbol / bin information for the context model of the next symbol / bin. The information related to prediction among the information decoded by the entropy decoder 310 can be provided to the predictor (inter-frame predictor 332 and intra-frame predictor 331), and the residual value (that is, quantized transform coefficient and related parameter information) on which entropy decoding is performed in the entropy decoder 310 can be input to the residual processor 320. The residual processor 320 can derive a residual signal (residual block, residual sample, residual sample array). In addition, the information about filtering among the information decoded by the entropy decoder 310 can be provided to the filter 350. In addition, a receiver (not shown) for receiving a signal output from the encoding device may be further configured as an internal / external element of the decoding device 300, or the receiver may be a component of the entropy decoder 310. In addition, the decoding device according to the present disclosure may be referred to as a video / image / picture decoding device, and the decoding device may be classified into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoder 310, and the sample decoder may include at least one of a dequantizer 321, an inverse transformer 322, an adder 340, a filter 350, a memory 360, an inter-frame predictor 332, and an intra-frame predictor 331.
[0068] The dequantizer 321 may dequantize the quantized transform coefficients and output the transform coefficients. The dequantizer 321 may rearrange the quantized transform coefficients in the form of two-dimensional blocks. In this case, the rearrangement may be performed based on the coefficient scanning order performed in the encoding device. The dequantizer 321 may dequantize the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain the transform coefficients.
[0069] The inverse transformer 322 performs an inverse transform on the transform coefficients to obtain a residual signal (residual block, residual sample array).
[0070] The predictor may perform prediction on the current block and generate a prediction block including prediction samples of the current block. The predictor may determine whether to apply intra prediction or inter prediction to the current block based on the information on prediction output from the entropy decoder 310, and may determine a specific intra / inter prediction mode.
[0071] The predictor 330 can generate a prediction signal based on various prediction methods described below. For example, the predictor can not only apply intra prediction or inter prediction to predict a block, but also apply intra prediction and inter prediction at the same time. This can be called inter-frame intra-frame combined prediction (CIIP). In addition, the predictor can predict the block based on the intra block copy (IBC) prediction mode or palette mode. The IBC prediction mode or palette mode can be used for content image / video encoding of games, etc., for example, screen content coding (SCC). IBC basically performs prediction in the current picture, but can be performed similarly to inter prediction because the reference block is derived in the current picture. That is, IBC can use at least one of the inter prediction techniques described in this disclosure. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, the sample value within the picture can be signaled based on information about the palette table and the palette index.
[0072] The intra-frame predictor 331 can predict the current block by referencing samples in the current picture. Depending on the prediction mode, the referenced samples may be located near the current block or may be located far away from the current block. In intra-frame prediction, the prediction mode may include multiple non-directional modes and multiple directional modes. The intra-frame predictor 331 may determine the prediction mode to be applied to the current block by using the prediction modes applied to the neighboring blocks.
[0073] The inter-frame predictor 332 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can also include information about the inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, neighboring blocks can include spatial neighboring blocks in the current picture and temporal neighboring blocks in the reference picture. For example, the inter-frame predictor 332 can configure a motion information candidate list based on the neighboring blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and the prediction information can include information indicating the inter-frame prediction mode for the current block.
[0074] The adder 340 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor (including the inter-frame predictor 332 and / or the intra-frame predictor 331). If the block to be processed has no residual (for example, when the skip mode is applied), the prediction block can be used as the reconstructed block.
[0075] The adder 340 may be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal may be used for intra-frame prediction of the next block to be processed in the current picture, may be output through filtering as described below, or may be used for inter-frame prediction of the next picture.
[0076] In addition, luma mapping and chroma scaling (LMCS) can be applied during picture decoding.
[0077] The filter 350 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 350 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and store the modified reconstructed picture in the memory 360 (specifically, the DPB of the memory 360). Various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0078] The (modified) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter-frame predictor 332. The memory 360 can store the motion information of the block from which the motion information in the current picture is derived (or decoded) and / or the motion information of the reconstructed block in the picture. The stored motion information can be sent to the inter-frame predictor 332 to be used as the motion information of the spatially adjacent blocks or the motion information of the temporally adjacent blocks. The memory 360 can store the reconstructed samples of the reconstructed blocks in the current picture and can transmit the reconstructed samples to the intra-frame predictor 331.
[0079] In the present disclosure, the embodiments described in the filter 260, the inter-frame predictor 221, and the intra-frame predictor 222 of the encoding device 200 may be the same as the filter 350, the inter-frame predictor 332, and the intra-frame predictor 331 of the decoding device 300 or may be applied to correspond to the filter 350, the inter-frame predictor 332, and the intra-frame predictor 331 of the decoding device 300, respectively. The same contents can also be applied to the inter-frame predictor 332 and the intra-frame predictor 331.
[0080] In the present disclosure, at least one of quantization / inverse quantization and / or transform / inverse transform may be omitted. When quantization / inverse quantization is omitted, the quantized transform coefficient may be referred to as a transform coefficient. When transform / inverse transform is omitted, the transform coefficient may be referred to as a coefficient or a residual coefficient, or, for uniformity of expression, may still be referred to as a transform coefficient.
[0081] In the present disclosure, quantized transform coefficients and transform coefficients may be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, residual information may include information about the transform coefficients, and information about the transform coefficients may be signaled via residual coding syntax. Transform coefficients may be derived based on the residual information (or information about the transform coefficients), and scaled transform coefficients may be derived by inversely transforming (scaling) the transform coefficients. Residual samples may be derived based on inversely transforming (transforming) the scaled transform coefficients. This may also be applied / expressed in other parts of the present disclosure.
[0082] In addition, the above-mentioned decoded picture buffer (DPB) can be conceptually constructed from sub-DPBs, and the sub-DPB can include a picture storage buffer for storing decoded pictures of a layer. The picture storage buffer can include decoded pictures marked as "for reference" or reserved for future output.
[0083] In addition, for multi-layer bitstreams, instead of assigning DPB parameters to each output layer set (OLS), DPB parameters can be assigned to each layer. For example, up to two PB parameters can be assigned to each layer. One can be assigned when the layer is an output layer (i.e., for example, when the layer can be used for reference and future output), and another can be assigned when the layer is not an output layer but is used as a reference layer (e.g., when there is no layer switching, and when the layer can only be used as a reference for a picture / slice / block of an output layer). This is considered to be simpler than the DPB parameters of the multi-layer bitstream of the HEVC layered extension (where each layer of the OLS has its own DPB parameters).
[0084] For example, the signaling of DPB parameters may be as follows: syntax and semantics.
[0085] [Table 1]
[0086]
[0087] For example, the above Table 1 may represent a video parameter set (VPS) including syntax elements of signaled DPB parameters.
[0088] The semantics of the syntax elements shown in Table 1 above may be as follows.
[0089] [Table 2]
[0090]
[0091]
[0092]
[0093] For example, the syntax element vps_num_dpb_params may indicate the number of dpb_parameters() syntax structures in the VPS. For example, the value of vps_num_dpb_params may be in the range of 0 to 16. Additionally, when the syntax element vps_num_dpb_params is not present, the value of the syntax element vps_num_dpb_params may be inferred to be equal to 0.
[0094] For example, the syntax element same_dpb_size_output_or_nonoutput_flag may also indicate whether the syntax element layer_nonoutput_dpb_params_idx[i] may be present in the VPS. For example, when the value of the syntax element same_dpb_size_output_or_nonoutput_flag is 1, the syntax element same_dpb_size_output_or_nonoutput_flag may indicate that the syntax element layer_nonoutput_dpb_params_idx[i] is not present in the VPS, and when the value of the syntax element same_dpb_size_output_or_nonoutput_flag is 0, the syntax element same_dpb_size_output_or_nonoutput_flag may indicate that the syntax element layer_nonoutput_dpb_params_idx[i] may be present in the VPS.
[0095] For example, the syntax element vps_sublayer_dpb_params_present_flag may also be used to control the presence of the syntax elements max_dec_pic_buffering_minus1[], max_num_reorder_pics[], and max_latency_increase_plus1[] in the dpb_parameters() syntax structure of the VPS. Additionally, when the syntax element vps_sublayer_dpb_params_present_flag is not present, the value of the syntax element vps_sublayer_dpb_params_present_flag may be inferred to be equal to 0.
[0096] For example, the syntax element dpb_size_only_flag[i] may also indicate whether the syntax elements max_num_reorder_pics[] and max_latency_increase_plus1[] may be present in the i dpb_parameters() syntax structure of the VPS. For example, when the value of the syntax element dpb_size_only_flag[i] is 1, the syntax element dpb_size_only_flag[i] may indicate that the syntax elements max_num_reorder_pics[] and max_latency_increase_plus1[] are not present in the i dpb_parameters() syntax structure of the VPS, and when the value of the syntax element dpb_size_only_flag[i] is 0, the syntax element dpb_size_only_flag[i] may indicate that the syntax elements max_num_reorder_pics[] and max_latency_increase_plus1[] may be present in the i dpb_parameters() syntax structure of the VPS.
[0097] For example, the syntax element dpb_max_temporal_id[i] may also indicate the TemporalId of the highest sublayer representation of a DPB parameter for which the idpb_parameters() syntax structure in the VPS may be present. Additionally, the value of dpb_max_temporal_id[i] may be in the range of 0 to vps_max_sublayers_minus1. Additionally, for example, when the value of vps_max_sublayers_minus1 is 0, the value of dpb_max_temporal_id[i] may be inferred to be 0. Additionally, for example, when the value of vps_max_sublayers_minus1 is greater than 0 and vps_all_layers_same_num_sublayers_flag is 1, the value of dpb_max_temporal_id[i] may be inferred to be equal to vps_max_sublayers_minus1.
[0098] In addition, for example, the syntax element layer_output_dpb_params_idx[i] may specify an index into the dpb_parameters() syntax structure list of the VPS that applies to the i-th layer, which is the output layer of the OLS. When the syntax element layer_output_dpb_params_idx[i] is present, the value of the syntax element layer_output_dpb_params_idx[i] may be in the range of 0 to vps_num_dpb_params-1.
[0099] For example, when vps_independent_layer_flag[i] is 1, the dpb_parameters() syntax structure applied to the i-th layer as the output layer may be a dpb_parameters() syntax structure present in the SPS referenced by the layer.
[0100] Alternatively, for example, when vps_independent_layer_flag[i] is 0, the following may apply.
[0101] -When vps_num_dpb_params is 1, the value of layer_output_dpb_params_idx[i] can be inferred to be equal to 0.
[0102] - A bitstream conformance requirement may be that the value of layer_output_dpb_params_idx[i] is such that the value of dpb_size_only_flag[layer_output_dpb_params_idx[i]] is equal to 0.
[0103] Additionally, for example, the syntax element layer_nonoutput_dpb_params_idx[i] may specify an index into the dpb_parameters() syntax structure list of the VPS that applies to the i-th layer, which is a non-output layer of the OLS. When the syntax element layer_nonoutput_dpb_params_idx[i] is present, the value of the syntax element layer_nonoutput_dpb_params_idx[i] may be in the range of 0 to vps_num_dpb_params-1.
[0104] For example, when same_dpb_size_output_or_nonoutput_flag is 1, the following may apply.
[0105] - When vps_independent_layer_flag[i] is 1, the dpb_parameters() syntax structure applied to the i-th layer as a non-output layer may be the dpb_parameters() syntax structure present in the SPS referenced by the layer.
[0106] - When vps_independent_layer_flag[i] is 0, the value of layer_nonoutput_dpb_params_idx[i] can be inferred to be equal to layer_output_dpb_params_idx[i].
[0107] Alternatively, for example, when same_dpb_size_output_or_nonoutput_flag is 0, and when vps_num_dpb_params is 1, the value of layer_output_dpb_params_idx[i] may be inferred to be 0.
[0108] In addition, for example, the syntax structure of dpb_parameters() can be as follows:
[0109] [Table 3]
[0110]
[0111] Referring to Table 3, the dpb_parameters() syntax structure may provide information on the DPB size, maximum picture reordering number, and maximum delay of each CLVS of the CVS. The dpb_parameters() syntax structure may be represented as information on DPB parameters or DPB parameter information.
[0112] When the VPS includes a dpb_parameters() syntax structure, the OLS to which the dpb_parameters() syntax structure is applied can be specified by the VPS. In addition, when the dpb_parameters() syntax structure is included in the SPS, the dpb_parameters() syntax structure can be applied to the OLS including only the lowest layer among the layers of the reference SPS, where the lowest layer can be an independent layer.
[0113] The semantics of the syntax elements shown in Table 3 above may be as follows.
[0114] [Table 4]
[0115]
[0116]
[0117] For example, the syntax element max_dec_pic_buffering_minus1[i] plus 1 may specify the maximum required size of the DPB in units of the picture storage buffer for each CLVS of the CVS when Htid is equal to i. For example, max_dec_pic_buffering_minus1[i] may be information regarding the size of the DPB. For example, the value of the syntax element max_dec_pic_buffering_minus1[i] may be in the range of 0 to MaxDpbSize-1. Additionally, for example, when i is greater than 0, max_dec_pic_buffering_minus1[i] may be greater than or equal to max_dec_pic_buffering_minus1[i-1]. Additionally, for example, if max_dec_pic_buffering_minus1[i] does not exist for i in the range of 0 to maxSubLayersMinus1-1, then due to subLayerInfoFlag being equal to 0, the value of the syntax element max_dec_pic_buffering_minus1[i] may be inferred to be equal to max_dec_pic_buffering_minus1[maxSubLayersMinus1].
[0118] Additionally, for example, the syntax element max_num_reorder_pics[i] may specify, for each CLVS of a CVS, the maximum allowed number of CLVSs that may precede all pictures of the CLVS in decoding order and follow the corresponding picture in output order when Htid is equal to i. For example, max_num_reorder_pics[i] may be information about the maximum picture reordering number of the DPB. The value of max_num_reorder_pics[i] may be in the range of 0 to max_dec_pic_buffering_minus1[i]. Additionally, for example, when i is greater than 0, max_num_reorder_pics[i] may be greater than or equal to max_num_reorder_pics[i-1]. Additionally, for example, if max_num_reorder_pics[i] does not exist for i in the range of 0 to maxSubLayersMinus1-1, then since subLayerInfoFlag is equal to 0, the syntax element max_num_reorder_pics[i] may be inferred to be equal to max_num_reorder_pics[maxSubLayersMinus1].
[0119] In addition, for example, the syntax element max_latency_increase_plus1[i] whose value is not 0 can be used to calculate the value of MaxLatencyPictures[i]. MaxLatencyPictures[i] can specify the maximum number of pictures of the CLVS that can precede all pictures of the CLVS in output order and follow the corresponding picture in decoding order when Htid is equal to i for each CLVS of the CVS. For example, max_latency_increase_plus1[i] can be information about the maximum latency of the DPB.
[0120] For example, when max_latency_increase_plus1[i] is not 0, the value of MaxLatencyPictures[i] can be derived as follows.
[0121] [Formula 1]
[0122] MaxLatencyPictures[i]=max_num_reorder_pics[i]+max_latency_increase_plus1[i]-1
[0123] In addition, for example, if max_latency_increase_plus1[i] is 0, the corresponding limit may not be expressed. The value of max_latency_increase_plus1[i] can be between 0 and 2. 32 Additionally, for example, if max_latency_increase_plus1[i] does not exist for i in the range of 0 to maxSubLayersMinus1-1, then due to subLayerInfoFlag being equal to 0, the syntax element max_latency_increase_plus1[i] may be inferred to be equal to max_latency_increase_plus1[maxSubLayersMinus1].
[0124] In addition, DPB parameters can be used for output and removal of picture processing as shown in the following table.
[0125] [Table 5]
[0126]
[0127]
[0128] In addition, the DPB parameter signaling design in the traditional VVC standard may have at least the following problems.
[0129] First, the VVC draft considers the concept of sub-DPBs, but a physical decoding device may only have one DPB for decoding a multi-layer bitstream. Therefore, the decoding device needs to know the DPB size requirement before decoding the OLS in a given multi-layer bitstream, but the conventional VVC draft does not clearly disclose how to know this information.
[0130] For example, the DPB size required for the OLS in the bitstream cannot be simply derived from the sub-DPB size of each layer of the OLS. That is, the DPB size required for the OLS cannot be simply derived as the sum of max_dec_pic_buffering_minus1[]+1 values of the layers in the OLS. For example, the sum of max_dec_pic_buffering_minus1[]+1 of each layer in the OLS may be larger than the actual DPB size. For example, in a specific access unit, since each layer in the DPB may have a different reference picture list structure, the number of reconstructed pictures of each layer in the DPB may not be the maximum. Therefore, the DPB size required for the OLS cannot be simply derived from the sum of max_dec_pic_buffering_minus1[]+1 values of the layers in the OLS.
[0131] For example, the following table exemplarily shows the pictures required for each sub-DPB to exist for a bitstream with two spatial scalability layers, a group of pictures (GOP) size of 16, and no temporal sub-layer.
[0132] [Table 6]
[0133]
[0134]
[0135] Referring to Table 6, the base layer (ie, layer 0) may have a more complex RPL structure than layer 1, and considering the picture size between the two layers, the size of sub-DPB 0 may include more reference pictures than sub-DPB 1. In addition, for example, as shown in Table 6, the maximum number of reference pictures of the two layers (ie, 12) may be greater than the total number of actual pictures of the DPB (ie, 11).
[0136] Secondly, the bump processing may not be called when it is actually needed. When using the above example, the number of pictures of sub-DPB 1 does not reach the maximum sub-DPB size after the first slice header of the picture with the picture order count (POC) 37 of the second layer has been decoded, so the bump processing may not be called. That is, if the picture with POC 37 is included, the number of pictures of sub-DPB 1 can be 4, and the maximum number of pictures of sub-DPB 1 can be increased to 5. However, since the maximum number of pictures of DPB has been reached, the bump processing must be called at the corresponding time. This problem may occur because only the DPB parameters of the current layer are checked under the condition of calling the bump processing. Here, for example, the bump processing may refer to the process of deriving the desired pictures from the pictures in the DPB and removing the pictures that are not used as references from the DPB.
[0137] Therefore, this document proposes a solution to the above problems. The proposed embodiments can be applied individually or in combination.
[0138] As an example, a method of signaling DPB parameters mapped to the OLS in addition to signaling DPB parameters mapped to each layer is proposed.
[0139] In addition, as an example, a method may be proposed by which the value of max_dec_pic_buffering_minus1[i] is derived so that when there are no DPB parameters mapped to the OLS because signaling of the DPB parameters mapped to the OLS is optional, it is equal to the value obtained by subtracting 1 from the sum of the values obtained by adding 1 to max_dec_pic_buffering_minus1[i] for all layers in the OLS. The method proposed in this embodiment may be performed based on a flag indicating whether there are DPB parameters mapped to the OLS. For example, when the value of the flag is 1, the flag may indicate that there are DPB parameter indices for all OLSs including at least one or more layers. Otherwise, that is, when the value of the flag is 0, the flag may indicate that there are no DPB parameters mapped to the OLS (that is, the DPB parameter indices of the OLS). In addition, for example, the flag may be present for each OLS.
[0140] In addition, as an example, a method may be provided by which the value of max_dec_pic_buffering_minus1[highest temporal sublayer] of each OLS is not greater than the value obtained by subtracting 1 from the sum of the value obtained by subtracting 1 from MaxDpbSize and the value obtained by adding 1 to max_dec_pic_buffering_minus1[highest temporal sublayer] of the layer within the OLS, minus 1.
[0141] In addition, as an example, a method may be proposed in which the DPB parameters allocated to the OLS include only the DPB size.
[0142] In addition, as an example, a method of updating the condition for calling the bump process in consideration of the number of pictures in the DPB and the value of max_dec_pic_buffering_minus1[i] of the OLS being processed by the decoding device may be proposed.
[0143] Furthermore, for example, the embodiment can be applied according to the following process.
[0144] Figure 4 The encoding process according to the embodiment of this document is exemplified.
[0145] Reference Figure 4 , the encoding device can decode the (reconstructed) picture (S400). The encoding device can update the DPB based on the DPB parameters (S410). For example, the decoded picture can be basically inserted into the DPB, and the decoded picture can be used as a reference picture for inter-frame prediction. In addition, the decoded picture in the DPB can be deleted based on the DPB parameters. In addition, the encoding device can encode the image information including the DPB parameters (S420). In addition, although not shown, the encoding device can further decode the current picture based on the DPB updated after step S410. In addition, the decoded current picture can be inserted into the DPB, and the DPB including the decoded current picture can be further updated based on the DPB parameters before the next picture is decoded in the decoding order.
[0146] Figure 5 The decoding process according to the embodiment of this document is exemplified.
[0147] Reference Figure 5 , the decoding device can obtain image information including information about DPB parameters from the bitstream (S500). The decoding device can output a picture decoded from the DPB based on the information about the DPB parameters (S505). In addition, when the layer related to the DPB (or DPB parameters) is not an output layer but a reference layer, step S505 can be skipped.
[0148] In addition, the decoding device can update the DPB based on the information about the DPB parameters (S510). The decoded picture can be basically inserted into the DPB. Then, the DPB can be updated before decoding the current picture. For example, the decoded picture in the DPB can be deleted based on the information of the DPB parameters. Here, the DPB update can be referred to as DPB management.
[0149] The information about the DPB parameters may include the information / syntax elements disclosed in Tables 1 and 3 above. In addition, for example, different DPB parameters may be signaled depending on whether the current layer is an output layer or a reference layer, or different DPB parameters may be signaled depending on whether the DPB (or DPB parameters) is used for OLS (mapped to OLS) as in the embodiment proposed in this document.
[0150] In addition, the decoding device can decode the current picture based on the DPB (S520). For example, the decoding device can use the decoded picture of the DPB (before the current picture) as a reference picture to decode the current picture based on inter-frame prediction of the block / slice of the current picture.
[0151] In addition, although not shown, the encoding device can decode the current picture based on the DPB updated after the above step S410. In addition, the decoded current picture can be inserted into the DPB, and the DPB including the decoded current picture can be further updated based on the DPB parameters before decoding the next picture.
[0152] The syntax and DPB management process to which the embodiments proposed in this document are applied will be described below.
[0153] As an embodiment, the signaled video parameter set (VPS) syntax may be as follows.
[0154] [Table 7]
[0155]
[0156] Referring to Table 7, the VPS may include syntax elements vps_num_dpb_params, same_dpb_size_output_or_nonoutput_flag, vps_sublayer_dpb_params_present_flag, dpb_size_only_flag[i], dpb_max_temporal_id[i], layer_output_dpb_params_idx[i], and / or layer_nonoutput_dpb_params_idx[i].
[0157] In addition, referring to Table 7, the VPS may further include syntax elements vps_ols_dpb_params_present_flag and / or ols_dpb_params_idx[i].
[0158] For example, the syntax element vps_ols_dpb_params_present_flag may indicate whether ols_dpb_params_idx[] may be present. For example, when the value of vps_ols_dpb_params_present_flag is 1, vps_ols_dpb_params_present_flag may indicate that ols_dpb_params_idx[] may be present, and when the value of vps_ols_dpb_params_present_flag is 0, vps_ols_dpb_params_present_flag may indicate that ols_dpb_params_idx[] does not exist. Furthermore, when vps_ols_dpb_params_present_flag does not exist, the value of vps_ols_dpb_params_present_flag may be inferred to be 0.
[0159] Additionally, for example, when i is less than TotalNumOlss, and when vps_ols_dpb_params_present_flag is 1, and when vps_num_dpb_params is greater than 1, and if NumLayersInOls[i] is greater than 1, the syntax element ols_dpb_params_idx[i] may be signaled. ols_dpb_params_idx[i] may be expressed as vps_ols_dpb_params_idx[i].
[0160] For example, when NumLayersInOls[i] is greater than 1, the syntax element ols_dpb_params_idx[i] may specify the index of the dpb_parameters() syntax structure applied to the i-th OLS in the list of dpb_parameters() syntax structures of the VPS. That is, for example, the syntax element ols_dpb_params_idx[i] may indicate the dpb_parameters() syntax structure of the VPS for the target OLS (i.e., the i-th OLS). When ols_dpb_params_idx[i] exists, the value of ols_dpb_params_idx[i] may be in the range of 0 to vps_num_dpb_params-1.
[0161] In addition, for example, when NumLayersInOls[i] is equal to 1, a dpb_parameters( ) syntax structure applied to the i-th OLS may exist in an SPS referenced by a layer in the i-th OLS.
[0162] Furthermore, according to this embodiment, OlsMaxDecPicBufferingMinus1[Htid] may be defined as follows.
[0163] [Table 8]
[0164]
[0165]
[0166] For example, referring to Table 8, the value of OlsMaxDecPicBufferingMinus1[Htid] of the target OLS can be derived as follows.
[0167] For example, when the value of vps_ols_dpb_params_present_flag is 1, OlsMaxDecPicBufferingMinus1[Htid] may be derived such that it is equal to the value of max_dec_pic_buffering_minus1[Htid] in ols_dpb_params_idx[opOlsIdx].
[0168] In addition, for example, in other cases, that is, when the value of vps_ols_dpb_params_present_flag is 0, OlsMaxDecPicBufferingMinus1[Htid] may be derived such that it is equal to a value obtained by subtracting 1 from the sum of max_dec_pic_buffering_minus1[Htid]+1 of each layer in the target OLS.
[0169] Furthermore, according to the present embodiment, the screen output and removal process (ie, the DPB management process) can be defined as follows.
[0170] [Table 9]
[0171]
[0172]
[0173] For example, referring to Table 9, the number of pictures of the sub-DPB may be greater than or equal to max_dec_pic_buffering_minus1[Htid] + 1. In addition, for example, the number of pictures of the DPB may be greater than or equal to OlsMaxDecPicBufferingMinus1[Htid] + 1.
[0174] In addition, according to the present embodiment, the constraint on the maximum picture of the DPB (ie, the maximum number of pictures of the DPB) may be updated as follows: Here, the maximum number of pictures of the DPB may be expressed as a maximum DPB size.
[0175] [Table 10]
[0176]
[0177]
[0178] For example, referring to Table 10, when the level is not level 8.5, the value of OlsMaxDecPicBufferingMinus1[Htid]+1 may be less than or equal to MaxDpbSize.
[0179] Alternatively, as an embodiment, the signaled video parameter set (VPS) syntax may be as follows.
[0180] [Table 11]
[0181]
[0182] Referring to Table 11, the VPS may include syntax elements vps_num_dpb_params, same_dpb_size_output_or_nonoutput_flag, vps_sublayer_dpb_params_present_flag, dpb_size_only_flag[i], dpb_max_temporal_id[i], layer_output_dpb_params_idx[i], and / or layer_nonoutput_dpb_params_idx[i].
[0183] In addition, referring to Table 11, the VPS may further include syntax elements vps_ols_dpb_params_present_flag and / or ols_dpb_params_idx[i].
[0184] For example, when i is less than TotalNumOlss, and when vps_num_dpb_params is greater than 1, and if NumLayersInOls[i] is greater than 1, the syntax element vps_ols_dpb_params_present_flag may be signaled. Unlike the embodiment shown in Table 7 above where vps_ols_dpb_params_present_flag is signaled without a separate condition, vps_ols_dpb_params_present_flag may be signaled only when i is less than TotalNumOlss and vps_num_dpb_params is greater than 1.
[0185] For example, the syntax element vps_ols_dpb_params_present_flag may indicate whether ols_dpb_params_idx[] may be present. For example, when the value of vps_ols_dpb_params_present_flag is 1, vps_ols_dpb_params_present_flag may indicate that ols_dpb_params_idx[] may be present, and when the value of vps_ols_dpb_params_present_flag is 0, vps_ols_dpb_params_present_flag may indicate that ols_dpb_params_idx[] does not exist. Furthermore, when vps_ols_dpb_params_present_flag does not exist, the value of vps_ols_dpb_params_present_flag may be inferred to be 0.
[0186] Additionally, for example, when vps_ols_dpb_params_present_flag is 1, the syntax element ols_dpb_params_idx[i] may be signaled.
[0187] For example, when NumLayersInOls[i] is greater than 1, the syntax element ols_dpb_params_idx[i] may specify the index of the dpb_parameters() syntax structure applied to the i-th OLS in the list of dpb_parameters() syntax structures of the VPS. That is, for example, the syntax element ols_dpb_params_idx[i] may indicate the dpb_parameters() syntax structure of the VPS for the target OLS (i.e., the i-th OLS). When ols_dpb_params_idx[i] exists, the value of ols_dpb_params_idx[i] may be in the range of 0 to vps_num_dpb_params-1.
[0188] Figure 6 The image encoding method performed by the encoding device according to the present disclosure is briefly illustrated. Figure 6 The method disclosed in Figure 2 Specifically, for example, Figure 6 S600 to S620 may be performed by the entropy encoder of the encoding device. In addition, although not shown, the process of updating the DPB may be performed by the DPB of the encoding device, and the process of decoding the current picture may be performed by the predictor and residual processor of the encoding device.
[0189] The encoding device generates decoded picture buffer (DPB) parameter information (S600). The encoding device can generate and encode the decoded picture buffer (DPB) parameter information. The image information can include the decoded picture buffer (DPB) parameter information. For example, the video parameter set (VPS) syntax can include the DPB parameter information.
[0190] For example, the decoded picture buffer (DPB) parameter information may include DPB parameter information of a target output layer set (OLS). For example, the DPB parameter information of the target OLS may include information about the DPB size of the target OLS, information about the maximum picture reordering number of the DBP of the target OLS, and / or information about the maximum delay of the DBP of the target OLS. Here, the DPB size may indicate the maximum number of pictures that the DPB can include.
[0191] The syntax element for information on the DPB size of the target OLS may be the above-mentioned max_dec_pic_buffering_minus1[i], the syntax element for information on the maximum picture reordering number of the DBP of the target OLS may be the above-mentioned max_num_reorder_pics[i], and the syntax element for information on the maximum delay of the DBP of the target OLS may be the above-mentioned max_latency_increase_plus1[i].
[0192] Furthermore, for example, whether to perform bump processing on the pictures in the DPB may be determined based on the number of pictures in the DPB and the information about the DPB size of the target OLS. For example, when the number of pictures in the DPB is greater than or equal to a value derived based on the information about the DPB size, bump processing may be performed, whereas when the number of pictures in the DPB is less than the value derived based on the information about the DPB size, bump processing may not be performed. Here, for example, the value derived based on the information about the DPB size may be a value obtained by adding 1 to the value of the information about the DPB size.
[0193] The encoding device generates an OLSDPB parameter index for the DPB parameter information of the target output layer set (OLS) (S610). The encoding device may generate and encode the OLSDPB parameter index for the DPB parameter information of the target OLS. The image information may include the OLSDPB parameter index for the DPB parameter information of the target OLS. For example, the VPS syntax may include the OLSDPB parameter index.
[0194] For example, the OLSDPB parameter index of the target OLS may indicate the DPB parameter information of the target OLS. For example, the OLSDPB parameter index of the target OLS may indicate the DPB parameter information of the target OLS in the DPB parameter information. The syntax element of the OLSDPB parameter index may be the above-mentioned vps_ols_dpb_params_idx[i] or ols_dpb_params_idx[i].
[0195] Furthermore, for example, the encoding device may generate and encode an OLSDPB parameter flag to determine whether DPB parameter information for the target OLS exists. For example, the image information may include the OLSDPB parameter flag. Furthermore, for example, the VPS syntax may include the OLSDPB parameter flag. For example, the OLSDPB parameter flag may indicate whether DPB parameter information for the target OLS exists. For example, when the value of the OLSDPB parameter flag is 1, the OLSDPB parameter flag may indicate that DPB parameter information for the target OLS may exist, while when the value of the OLSDPB parameter flag is 0, the OLSDPB parameter flag may indicate that DPB parameter information for the target OLS does not exist. Furthermore, for example, an OLSDPB parameter index may be generated and encoded based on the OLSDPB parameter flag. For example, when the value of the OLSDPB parameter flag is 1, the OLSDPB parameter index may be generated / encoded / signaled, while when the value of the OLSDPB parameter flag is 0, the OLSDPB parameter index may not be generated / encoded / signaled. The syntax element of the OLSDPB parameter flag may be the aforementioned vps_ols_dpb_params_present_flag.
[0196] The encoding device encodes the image information including the OLS DPB parameter index and the DPB parameter information (S620). The encoding device may encode the DPB parameter information and the OLS DPB parameter index. The image information may include the DPB parameter information and the OLS DPB parameter index. In addition, the image information may include the above-mentioned OLS DPB parameter flag.
[0197] In addition, the encoding device can decode the picture of the target OLS. In addition, for example, the encoding device can update the DPB based on the DPB parameter information of the target OLS. For example, the encoding device can perform picture management processing on the decoded picture of the DPB based on the DPB parameter information. For example, the encoding device can add a decoded picture to the DPB, or can remove a decoded picture in the DPB. For example, a decoded picture in the DPB can be used as a reference picture for inter-frame prediction of the current picture, or a decoded picture in the DPB can be used as an output picture. The decoded picture may refer to a picture decoded before the current picture in the decoding order in the target OLS.
[0198] In addition, for example, the encoding device may decode the current picture of the target OLS based on the updated DPB. For example, the encoding device may derive prediction samples by performing inter-frame prediction on the blocks in the current picture based on the reference picture of the updated DPB, and may generate reconstructed samples and / or reconstructed pictures for the current picture based on the prediction samples. In addition, for example, the encoding device may derive residual samples for the blocks in the current picture, and may generate reconstructed samples and / or reconstructed pictures by adding the prediction samples and the residual samples.
[0199] Furthermore, for example, the encoding device may generate and encode prediction information for a block of the current picture. In this case, various prediction methods disclosed in this document, such as inter-frame prediction or intra-frame prediction, may be applied. For example, the encoding device may determine whether to perform inter-frame prediction or intra-frame prediction on the block, and may determine a specific inter-frame prediction mode or a specific intra-frame prediction mode based on the RD cost. Based on the determined mode, the encoding device may derive prediction samples for the current chroma block. The prediction information may include prediction mode information for the current chroma block. The image information may include the prediction information.
[0200] In addition, for example, the encoding apparatus may encode residual information of blocks of a picture.
[0201] For example, the encoding device may derive residual samples by subtracting original samples and predicted samples of the block.
[0202] Thereafter, for example, the encoding device may derive quantized residual samples by quantizing the residual samples, may derive transform coefficients based on the quantized residual samples, and may generate and encode residual information based on the transform coefficients. Alternatively, for example, the encoding device may derive quantized residual samples by quantizing the residual samples, may derive transform coefficients by transforming the quantized residual samples, and may generate and encode residual information based on the transform coefficients. The image information may include residual information. In addition, for example, the encoding device may encode the image information and output the encoded image information in the form of a bitstream.
[0203] The encoding device can generate a reconstructed sample and / or a reconstructed picture by adding the predicted sample and the residual sample. Thereafter, as described above, an in-loop filtering process such as ALF processing, SAO and / or deblocking filtering can be applied to the reconstructed sample as needed to improve the subjective / objective video quality.
[0204] In addition, the bit stream including the image information can be sent to the decoding device via a network or a (digital) storage medium. Here, the network may include a broadcast network, a communication network, etc., and the digital storage medium may include various storage media such as a universal serial bus (USB), a secure digital (SD), a compact disc (CD), a digital video disc (DVD), a Blu-ray, a hard disk drive (HDD), a solid-state drive (SSD), etc.
[0205] Figure 7 An encoding device for performing an image encoding method according to the present disclosure is briefly illustrated. Figure 6 The method disclosed in Figure 7 Specifically, for example, Figure 7The entropy encoder of the encoding device may perform S600 to S620. In addition, although not shown, the process of updating the DPB may be performed by the DPB of the encoding device, and the process of decoding the current picture may be performed by the predictor and residual processor of the encoding device.
[0206] Figure 8 The image decoding method performed by the decoding device according to the present disclosure is briefly illustrated. Figure 8 The method disclosed in Figure 3 Specifically, for example, Figure 8 S800 may be performed by an entropy decoder of a decoding device; Figure 8 S810 to S820 may be performed by the DPB of the decoding device; Figure 8 S830 may be performed by a predictor and a residual processor of the decoding device.
[0207] The decoding apparatus obtains image information including an OLS decoded picture buffer (DPB) parameter index and DPB parameter information of a target output layer set (OLS) (S800). The decoding apparatus may obtain image information including decoded picture buffer (DPB) parameter information of a target output layer set (OLS) and an OLS DPB parameter index of the target OLS.
[0208] For example, a decoding device may obtain a video parameter set (VPS) syntax from a bitstream. Image information may include the VPS syntax. Image information may be received in a bitstream. The VPS syntax may include an OLSDPB parameter index and DPB parameter information for a target OLS. That is, for example, the decoding device may obtain the OLSDPB parameter index and DPB parameter information for a target OLS using the VPS syntax.
[0209] For example, the OLSDPB parameter index of the target OLS may indicate the DPB parameter information of the target OLS. For example, the DPB parameter information may include the DPB parameter information of the target OLS, and the OLSDPB parameter index of the target OLS may indicate the DPB parameter information of the target OLS in the DPB parameter information. The syntax element of the OLSDPB parameter index may be the above-mentioned vps_ols_dpb_params_idx[i] or ols_dpb_params_idx[i].
[0210] Furthermore, for example, the decoding device may obtain an OLSDPB parameter flag to determine whether DPB parameter information for the target OLS exists. For example, the image information may include the OLSDPB parameter flag. Furthermore, for example, the VPS syntax may include the OLSDPB parameter flag. For example, the OLSDPB parameter flag may indicate whether DPB parameter information for the target OLS exists. For example, when the value of the OLSDPB parameter flag is 1, the OLSDPB parameter flag may indicate that DPB parameter information for the target OLS may exist, while when the value of the OLSDPB parameter flag is 0, the OLS DPB parameter flag may indicate that DPB parameter information for the target OLS does not exist. Furthermore, for example, the OLSDPB parameter index may be obtained based on the OLSDPB parameter flag. For example, when the value of the OLSDPB parameter flag is 1, the OLSDPB parameter index may be signaled / obtained, while when the value of the OLSDPB parameter flag is 0, the OLSDPB parameter index may not be signaled / obtained. The syntax element for the OLSDPB parameter flag may be the aforementioned vps_ols_dpb_params_present_flag.
[0211] The decoding device derives DPB parameter information of the target OLS based on the OLS DPB parameter index (S810). The decoding device may derive the DPB parameter information of the target OLS based on the OLS DPB parameter index. For example, the decoding device may derive the DPB parameter information of the target OLS indicated by the OLS DPB parameter index. The DPB parameter information may include the DPB parameter information of the target OLS.
[0212] For example, the DPB parameter information of the target OLS may include information about the DPB size of the target OLS, information about the maximum picture reordering number of the DBP of the target OLS, and / or information about the maximum delay of the DBP of the target OLS. Here, the DPB size may indicate the maximum number of pictures that the DPB may include.
[0213] The syntax element for information on the DPB size of the target OLS may be the above-mentioned max_dec_pic_buffering_minus1[i], the syntax element for information on the maximum picture reordering number of the DBP of the target OLS may be the above-mentioned max_num_reorder_pics[i], and the syntax element for information on the maximum delay of the DBP of the target OLS may be the above-mentioned max_latency_increase_plus1[i].
[0214] The decoding device updates the DPB based on the DPB parameter information of the target OLS (S820). The decoding device may update the DPB based on the DPB parameter information. For example, the decoding device may perform picture management processing on the decoded picture of the DPB based on the DPB parameter information. For example, the decoding device may add a decoded picture to the DPB, or may remove a decoded picture in the DPB. For example, a decoded picture in the DPB may be used as a reference picture for inter-frame prediction of the current picture, or a decoded picture in the DPB may be used as an output picture. The decoded picture may refer to a picture decoded before the current picture in the decoding order in the target OLS.
[0215] For example, the decoding device may determine whether to perform bump processing on the pictures in the DPB based on the information about the DPB size of the target OLS and the number of pictures in the DPB, and may perform bump processing on the pictures in the DPB based on the determination result. For example, when the number of pictures in the DPB is greater than or equal to the value derived based on the information about the DPB size, the bump processing may be performed, and when the number of pictures in the DPB is less than the value derived based on the information about the DPB size, the bump processing may not be performed. Here, for example, the value derived based on the information about the DPB size may be a value obtained by adding 1 to the value of the information about the DPB size.
[0216] The decoding device decodes the current picture based on the updated DPB (S830). The decoding device can decode the current picture based on the updated DPB. For example, the decoding device can derive prediction samples by performing inter-frame prediction on the blocks in the current picture based on the reference picture of the updated DPB, and can generate reconstructed samples and / or reconstructed pictures of the current picture based on the prediction samples. In addition, for example, the decoding device can derive residual samples of the blocks in the current picture based on the residual information received through the bitstream, and can generate reconstructed pictures and / or reconstructed samples by adding the prediction samples and the residual samples.
[0217] Thereafter, as described above, in-loop filtering processing such as ALF processing, SAO and / or deblocking filtering may be applied to the reconstructed samples as needed in order to improve the subjective / objective video quality.
[0218] Figure 9 A decoding device for performing an image decoding method according to the present disclosure is briefly illustrated. Figure 8 The method disclosed in Figure 9 Specifically, for example, Figure 9 The entropy decoder of the decoding device can perform Figure 8 S800 in Figure 9 The DPB of the decoding device can perform Figure 8 S810 to S820, Figure 9The predictor and residual processor of the decoding device can perform Figure 8 S830.
[0219] According to the above-mentioned document, the DPB parameters of the OLS can be signaled, and thus, the DPB can be adaptively updated for the OLS, and the overall coding efficiency can be improved.
[0220] In addition, according to this document, index information indicating the DPB parameters of the OLS can be signaled, and thus, the DPB parameters can be adaptively derived for the OLS, and the overall encoding efficiency can be improved by updating the DPB of the OLS based on the derived DPB parameters.
[0221] In the above-mentioned embodiment, method is described based on the flow chart with a series of steps or square frames. The present disclosure is not limited to the order of the above steps or square frames. Some steps or square frames can be performed in an order different from the above-mentioned other steps or square frames or performed simultaneously. In addition, it will be understood by those skilled in the art that the steps shown in the flow chart are not exclusive and may also include other steps, or may delete one or more steps in the flow chart without affecting the scope of the present disclosure.
[0222] The embodiments described in this specification can be implemented on a processor, microprocessor, controller, or chip. For example, the functional units shown in each figure can be implemented on a computer, processor, microprocessor, controller, or chip. In this case, information (e.g., information about instructions) or algorithms used for implementation can be stored in a digital storage medium.
[0223] In addition, the decoding device and encoding device to which the present disclosure is applied may be included in the following devices: multimedia broadcast transmission / reception devices, mobile communication terminals, home theater video devices, digital theater video devices, surveillance cameras, video chat devices, real-time communication devices such as video communication, mobile streaming devices, storage media, portable cameras, VoD service providing devices, over-the-top (OTT) video devices, Internet streaming service providing devices, three-dimensional (3D) video devices, teleconferencing video devices, transportation user devices (e.g., vehicle user devices, aircraft user devices, and ship user devices), and medical video devices; and the decoding device and encoding device to which the present disclosure is applied may be used to process video signals or data signals. For example, over-the-top (OTT) video devices may include game consoles, Blu-ray players, Internet access televisions, home theater systems, smart phones, tablet computers, digital video recorders (DVRs), etc.
[0224] In addition, the processing method of the present invention can be generated in the form of a program executed by a computer and can be stored in a computer-readable recording medium. The multimedia data with a data structure according to the present invention can also be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices in which computer-readable data is stored. The computer-readable recording medium may include, for example, a BD, a universal serial bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. In addition, the computer-readable recording medium includes a medium implemented in the form of a carrier wave (for example, via transmission over the Internet). In addition, the bit stream generated by the encoding method can be stored in a computer-readable recording medium or transmitted via a wired / wireless communication network.
[0225] In addition, the embodiments of the present disclosure may be implemented using a computer program product according to a program code, and the program code may be executed in a computer through the embodiments of the present disclosure. The program code may be stored on a computer-readable carrier.
[0226] Figure 10 A structural diagram of a content streaming system to which the present disclosure is applied is illustrated.
[0227] A content streaming system to which embodiments of the present disclosure are applied may mainly include an encoding server, a streaming server, a network server, a media storage, a user device, and a multimedia input device.
[0228] The encoding server compresses content input from a multimedia input device such as a smartphone, camera, or camcorder into digital data to generate a bitstream and transmits the bitstream to the streaming server. As another example, when the multimedia input device such as a smartphone, camera, or camcorder directly generates the bitstream, the encoding server can be omitted.
[0229] A bitstream may be generated by an encoding method or a bitstream generating method to which an embodiment of the present disclosure is applied, and a streaming server may temporarily store the bitstream in the process of transmitting or receiving the bitstream.
[0230] The streaming server transmits multimedia data to user devices via a network server based on user requests, and the network server serves as an intermediary for notifying users of services. When a user requests a desired service from the network server, the network server transmits the request to the streaming server, which then transmits the multimedia data to the user. In this case, the content streaming system may include a separate control server. In this case, the control server is used to control commands and responses between devices within the content streaming system.
[0231] The streaming server can receive content from a media storage and / or encoding server. For example, when receiving content from an encoding server, the content can be received in real time. In this case, in order to provide a stable streaming service, the streaming server can store the bitstream for a predetermined time.
[0232] Examples of user devices may include mobile phones, smart phones, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigators, touch-screen PCs, tablet PCs, ultrabooks, wearable devices (e.g., smart watches, smart glasses, and head-mounted displays), digital TVs, desktop computers, and digital signage, etc. Each server within the content streaming system may operate as a distributed server, in which case data received from each server may be distributed.
[0233] The claims described in this disclosure can be combined in various ways. For example, the technical features of the method claims of this disclosure can be combined to implement a device, and the technical features of the device claims of this disclosure can be combined to implement a method. Furthermore, the technical features of the method claims of this disclosure and the technical features of the device claims of this disclosure can be combined to implement a device, and the technical features of the method claims of this disclosure and the technical features of the device claims of this disclosure can be combined to implement a method.
Claims
1. An image decoding method performed by a decoding device, the image decoding method comprising the following steps: Obtaining image information, the image information including decoded picture buffer (DPB) parameter information for one or more OLSs including a target output layer set (OLS) and an OLSDPB parameter index for the target OLS; deriving DPB parameter information for the target OLS based on the OLSDPB parameter index; updating the DPB based on the DPB parameter information for the target OLS; as well as Decode the current picture based on the updated DPB, The OLSDPB parameter index specifies an index of the DPB parameter information applied to the target OLS, and the target OLS is a multi-layer OLS.
2. The image decoding method according to claim 1, wherein: The DPB parameter information and the OLS DPB parameter index are included in the video parameter set VPS syntax.
3. The image decoding method according to claim 1, wherein: An OLSDPB parameter flag is obtained to determine whether the DPB parameter information for the target OLS exists.
4. The image decoding method according to claim 3, wherein: The OLSDPB parameter index is obtained based on the OLSDPB parameter flag.
5. The image decoding method according to claim 4, wherein: When the value of the OLSDPB parameter flag is 1, the OLSDPB parameter index is obtained.
6. The image decoding method according to claim 4, wherein: The OLSDPB parameter flag is included in the video parameter set VPS syntax.
7. The image decoding method according to claim 1, wherein: The DPB parameter information for the target OLS includes information about a DPB size, information about a maximum picture reordering number of the DPB for the target OLS, and information about a maximum delay of the DPB.
8. An image encoding method performed by an encoding device, the image encoding method comprising the following steps: generating decoded picture buffer (DPB) parameter information for one or more OLSs including a target output layer set (OLS); Generate an OLSDPB parameter index for the DPB parameter information of the target OLS; as well as Encoding the image information including the OLSDPB parameter index and the DPB parameter information, The OLSDPB parameter index specifies an index of the DPB parameter information applied to the target OLS, and the target OLS is a multi-layer OLS.
9. The image encoding method according to claim 8, wherein: The DPB parameter information and the OLS DPB parameter index are included in the video parameter set VPS syntax.
10. The image encoding method according to claim 8, wherein: An OLSDPB parameter flag indicating whether the DPB parameter information for the target OLS exists is encoded.
11. The image encoding method according to claim 10, wherein: The OLSDPB parameter index is generated based on the OLSDPB parameter flag.
12. The image encoding method according to claim 11, wherein: When the value of the OLSDPB parameter flag is 1, the OLSDPB parameter index is generated.
13. The image encoding method according to claim 11, wherein: The OLSDPB parameter flag is included in the video parameter set VPS syntax.
14. A non-transitory computer-readable storage medium having a computer program stored therein, wherein the computer program is executed by a processor to implement the following steps of a method, the method comprising the following steps: generating decoded picture buffer (DPB) parameter information for one or more OLSs including a target output layer set (OLS); Generate an OLSDPB parameter index for the DPB parameter information of the target OLS; Encoding the image information including the OLSDPB parameter index and the DPB parameter information; as well as generating a bit stream including the image information, The OLSDPB parameter index specifies an index of the DPB parameter information applied to the target OLS, and the target OLS is a multi-layer OLS.
15. A method for transmitting image data, the method comprising the following steps: Obtaining a bitstream of image information, the image information including decoded picture buffer (DPB) parameter information for one or more OLSs including a target output layer set (OLS) and an OLSDPB parameter index of the DPB parameter information for the target OLS; as well as sending the data of the bitstream including the image information, wherein the image information includes the DPB parameter information and the OLSDPB parameter index, The OLSDPB parameter index specifies an index of the DPB parameter information applied to the target OLS, and the target OLS is a multi-layer OLS.
Citation Information
Patent Citations
Signaling for sub-decoded picture buffer (SUB-DPB) based DPB operations in video coding
CN105637878A