Image decoding method and apparatus therefor
By deriving variables to update the DPB state in the decoding device, the problem of high transmission and storage costs in high-resolution image encoding is solved, and more efficient image encoding and DPB management are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- LG ELECTRONICS INC
- Filing Date
- 2021-05-04
- Publication Date
- 2026-04-10
AI Technical Summary
Existing technologies incur high transmission and storage costs when processing high-resolution, high-quality images, necessitating improved image encoding efficiency and effective management of the Decoding Picture Buffer (DPB) to optimize the encoding process.
By deriving a variable in the decoding device to determine whether to clear the decoded image buffer (DPB) and updating the DPB state based on that variable, image decoding is optimized during the decoding process, avoiding changes to the DPB state of all layers for each image.
It improves image encoding efficiency, reduces the cost of image data transmission and storage, and optimizes the DPB management process.
Smart Images

Figure CN115668932B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The disclosure relates to an image encoding technique, and more particularly, to an image decoding method and apparatus for performing a DPB management process in an image encoding system. BACKGROUND
[0002] Recently, in various fields, the demand for high-resolution, high-quality images such as HD (High Definition) images and UHD (Ultra High Definition) images is increasing. Because image data has high resolution and high quality, the amount of information or bits to be transmitted increases relative to conventional image data. Therefore, when transmitting image data using a medium such as a conventional wired / wireless broadband line or storing image data using an existing storage medium, the transmission cost and storage cost thereof increase.
[0003] Therefore, there is a need for an efficient image compression technique for efficiently transmitting, storing, and reproducing information of high-resolution, high-quality images. SUMMARY
[0004] TECHNICAL PROBLEM
[0005] The disclosure provides a method and apparatus for improving image encoding efficiency.
[0006] Another technical challenge of the disclosure is to provide a method and apparatus for performing a DPB management process.
[0007] TECHNICAL SOLUTION
[0008] According to an embodiment of the disclosure, an image decoding method performed by a decoding device is provided. The method includes deriving a value of a variable based on whether a current picture is a first picture of a current access unit (AU) that is a coded video sequence start access unit (CVSS AU) other than the AU 0, updating the DPB based on the variable, and decoding the current picture based on the updated DPB. The variable indicates whether all picture storage buffers in a decoded picture buffer (DPB) are emptied without output.
[0009] According to another embodiment of the disclosure, a decoding device for performing image decoding is provided. The decoding device includes a DPB for deriving a value of a variable based on whether a current picture is a first picture of a current access unit (AU) that is a coded video sequence start access unit (CVSS AU) other than the AU 0 and updating the DPB based on the variable, and a predictor for decoding the current picture based on the updated DPB. The variable indicates whether all picture storage buffers in a decoded picture buffer (DPB) are emptied without output.
[0010] According to another embodiment of the disclosure, an image encoding method performed by an encoding device is provided. The method includes deriving a value of a variable based on whether a current picture is a first picture of a current access unit (AU) that is a coded video sequence start access unit (CVSS AU) other than AU 0, updating the DPB based on the variable, and encoding image information of the current picture. The variable indicates whether all picture storage buffers in a decoded picture buffer (DPB) are emptied without output.
[0011] According to another embodiment of the disclosure, a video encoding device is provided. The encoding device includes a DPB configured to derive a value of a variable based on whether a current picture is a first picture of a current access unit (AU) that is a coded video sequence start access unit (CVSS AU) other than AU 0 and update the DPB based on the variable, and an entropy encoder configured to encode image information of the current picture. The variable indicates whether all picture storage buffers in a decoded picture buffer (DPB) are emptied without output.
[0012] According to another embodiment of the disclosure, a computer-readable digital storage medium having stored therein a bitstream including image information causing an image decoding method to be performed. In the computer-readable digital storage medium, the image decoding method includes deriving a value of a variable based on whether a current picture is a first picture of a current access unit (AU) that is a coded video sequence start access unit (CVSS AU) other than AU 0, updating the DPB based on the variable, and decoding the current picture based on the updated DPB. The variable indicates whether all picture storage buffers in a decoded picture buffer (DPB) are emptied without output.
[0013] Technical Effects
[0014] According to the disclosure, whether to perform a process of removing pictures in the DPB without outputting them can be determined before only decoding a first picture of a CVSS AU other than AU 0, rather than before decoding all pictures of a CVSS AU other than AU 0. By doing so, the DPB state affecting all layers in a CVS can not be changed for each picture, and encoding efficiency can be improved.
[0015] According to the present disclosure, a variable indicating whether pictures in the DPB are removed without output can be determined before decoding the first picture of the CVSS AUs other than AU 0, rather than before decoding all pictures of the CVSS AUs other than AU 0. By doing so, the DPB status of all layers in the CVS can not be changed for each picture, and coding efficiency can be improved. BRIEF DESCRIPTION OF DRAWINGS
[0016] Figure 1 A video / image encoding apparatus to which an embodiment of the present disclosure is applicable is exemplarily illustrated.
[0017] Figure 2 is a schematic view illustrating a configuration of a video / image encoding apparatus to which an embodiment of the present disclosure can be applied.
[0018] Figure 3 is a schematic view illustrating a configuration of a video / image decoding apparatus to which an embodiment of the present disclosure can be applied.
[0019] Figure 4 An encoding process according to an embodiment of the present disclosure is exemplarily illustrated.
[0020] Figure 5 A decoding process according to an embodiment of the present disclosure is exemplarily illustrated.
[0021] Figure 6 An image encoding method by an encoding apparatus according to the present document is schematically shown.
[0022] Figure 7 An encoding apparatus for performing an image encoding method according to the present document is schematically shown.
[0023] Figure 8 An image decoding method by a decoding apparatus according to the present document is schematically shown.
[0024] Figure 9 A decoding apparatus for performing an image decoding method according to the present document is schematically shown.
[0025] Figure 10 A configuration diagram of a content streaming system to which the present disclosure is applied is exemplarily illustrated. DETAILED DESCRIPTION
[0026] The present disclosure can be modified in various forms, and specific embodiments thereof will be described and illustrated in the accompanying drawings. However, these embodiments are not intended to limit the present disclosure. The terms used in the following description are merely used to describe specific embodiments, and are not intended to limit the present disclosure. Singular expressions include plural expressions unless it is clearly different from the context. Terms such as "include" and "have" are intended to indicate that there is a feature, number, step, operation, element, component, or a combination thereof described in the following description, and it should be understood that the possibility of existence or addition of one or more different features, numbers, steps, operations, elements, components, or combinations thereof is not excluded.
[0027] In addition, the elements in the drawings described in the present disclosure are independently drawn for the purpose of conveniently explaining different specific functions, and this does not mean that the elements are implemented by independent hardware or independent software. For example, two or more of the elements can be combined to form a single element, or one element can be divided into a plurality of elements. Embodiments in which elements are combined and / or divided do not depart from the concept of the present disclosure.
[0028] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. In addition, throughout the drawings, like reference numerals are used to refer to like elements, and the same description will be omitted for similar elements.
[0029] Figure 1 An example of a video / image encoding apparatus to which embodiments of the present disclosure can be applied is briefly illustrated.
[0030] Referring to Figure 1 , a video / image encoding system can include a first apparatus (a source apparatus) and a second apparatus (a sink). The source apparatus can transmit encoded video / image information or data in the form of a file or a stream to the sink via a digital storage medium or a network.
[0031] The source apparatus can include a video source, an encoding device, and a transmitter. The sink can include a receiver, a decoding device, and a renderer. The encoding device can be referred to as a video / image encoding device, and the decoding device can be referred to as a video / image decoding device. The transmitter can be included in the encoding device. The receiver can be included in the decoding device. The renderer can include a display, and the display can be configured as a separate apparatus or an external component.
[0032] The video source can acquire a video / image through a process of capturing, synthesizing, or generating a video / image. The video source can include a video / image capturing device and / or a video / image generating device. The video / image capturing device can include, for example, one or more cameras, a video / image archive including previously captured videos / images, etc. The video / image generating device can include, for example, a computer, a tablet, and a smartphone, and can (electronically) generate a video / image. For example, a virtual video / image can be generated through a computer, etc. In this case, the video / image capturing process can be replaced by a process of generating related data.
[0033] The encoding device can encode an input video / image. The encoding device can perform a series of processes such as prediction, transformation, and quantization to achieve compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0034] The transmitter can transmit the encoded image / image information or data output in the form of a bitstream to a receiver of a receiving device in the form of a file or a stream through a digital storage medium or a network. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter can include an element for generating a media file through a predetermined file format, and can include an element for transmission through a broadcasting / communication network. The receiver can receive / extract a bitstream and transmit the received bitstream to a decoding device.
[0035] The decoding device can decode a video / image by performing a series of processes such as dequantization, inverse transformation, and prediction corresponding to the operations of the encoding device.
[0036] The renderer can render the decoded video / image. The rendered video / image can be displayed through a display.
[0037] The present disclosure relates to video / image encoding. For example, the methods / embodiments disclosed in the present disclosure can be applied to the methods disclosed in the Versatile Video Coding (VVC), EVC (Elementary Video Coding) standard, AOMedia Video 1 (AV1) standard, second generation Audio Video Coding standard (AVS2), or next generation video / image encoding standards (e.g., H.267 or H.268, etc.).
[0038] The present disclosure proposes various embodiments of video / image encoding, and unless otherwise mentioned, the embodiments can be performed in combination with each other.
[0039] In the disclosure, a video can refer to a series of images over time. A picture generally refers to a unit representing one image in a specific time region, and a sub-picture / slice / tile is a unit that constitutes a part of a picture at the time of encoding. A sub-picture / slice / tile can include one or more coding tree units (CTUs). One picture can consist of one or more sub-pictures / slices / tiles. One picture can consist of one or more tile groups. One tile group can include one or more tiles. A brick can refer to a rectangular region of CTU rows within a tile in a picture. A tile can be partitioned into multiple bricks, each of which consists of one or more CTU rows within the tile. A tile that is not partitioned into multiple bricks can also be referred to as a brick. Brick scanning is a specific order of partitioning CTUs of a picture in which the CTUs are ordered in a raster scan in the bricks, the bricks within a tile are consecutively ordered in a raster scan of the bricks of the tile, and the tiles in a picture are consecutively ordered in a raster scan of the tiles of the picture. In addition, a sub-picture can refer to a rectangular region of one or more slices within a picture. That is, a sub-picture contains one or more slices that collectively cover a rectangular region of a picture. A tile is a rectangular region of CTUs within a specific tile column and a specific tile row in a picture. A tile column is a rectangular region of CTUs whose height is equal to the height of a picture and whose width is specified by a syntax element in a picture parameter set. A tile row is a rectangular region of CTUs whose height is specified by a syntax element in a picture parameter set and whose width is equal to the picture width. Tile scanning is a specific order of partitioning CTUs of a picture in which the CTUs are consecutively ordered in a raster scan in the tiles, while the tiles in a picture are consecutively ordered in a raster scan of the tiles of the picture. A slice includes an integer number of bricks of a picture that can be exclusively contained in a single NAL unit. A slice can consist of either multiple complete tile groups or a consecutive sequence of complete bricks of only one tile. In the disclosure, a tile group can be used interchangeably with a slice. For example, in the disclosure, a tile group / tile group header can be referred to as a slice / slice header.
[0040] A pixel or pel can mean a minimum unit constituting one picture (or image). In addition, a "sample" can be used as a term corresponding to a pixel. A sample can generally represent a pixel or a value of a pixel, can represent only a pixel / value of a pixel of a luminance component, or can represent only a pixel / value of a pixel of a chrominance component.
[0041] A unit can represent a basic unit of image processing. A unit can include at least one of a specific area of a picture and information related to the area. One unit can include one luma block and two chroma (e.g., cb, cr) blocks. In some cases, a unit can be used interchangeably with terms such as a block or an area. In general, an MxN block can include a set (or an array) of M columns and N rows of samples (or sample array) or transform coefficients.
[0042] In the present specification, "A or B" can mean "only A", "only B", or "both A and B". In other words, in the present specification, "A or B" can be interpreted as "A and / or B". For example, "A, B, or C" in the present specification means "only A", "only B", "only C", or "any one of A, B, and C and any combination thereof".
[0043] In the present specification, a slash ( / ) or a comma (,) used in the present specification can mean "and / or". For example, "A / B" can mean "A and / or B". Accordingly, "A / B" can mean "only A", "only B", or "both A and B". For example, "A, B, C" can mean "A, B, or C".
[0044] In the present specification, "at least one of A and B" can mean "only A", "only B", or "both A and B". In addition, in the present specification, the expression "at least one of A or B" or "at least one of A and / or B" can be interpreted as the same as "at least one of A and B".
[0045] In addition, in the present specification, "at least one of A, B, and C" means "only A", "only B", "only C", or "any combination of A, B, and C". In addition, "at least one of A, B, or C" or "at least one of A, B, and / or C" can mean "at least one of A, B, and C".
[0046] In addition, the parentheses used in the present specification can mean "for example". Specifically, when indicating "prediction (intra prediction)", "intra prediction" can be proposed as an example of "prediction". In other words, "prediction" in the present specification is not limited to "intra prediction", and "intra prediction" can be proposed as an example of "prediction". In addition, even when indicating "prediction (i.e., intra prediction)", "intra prediction" can be proposed as an example of "prediction".
[0047] In the present specification, technical features described separately in one drawing can be implemented separately or can be implemented simultaneously.
[0048] The following drawings are created in order to explain specific examples of the present specification. Since names of specific apparatuses described in the drawings or names of specific signals / messages / fields are presented in an exemplary manner, technical features of the present specification are not limited to the specific names used in the following drawings.
[0049] Figure 2 is a schematic diagram illustrating a configuration of a video / image encoding apparatus to which embodiments of the present disclosure can be applied. Hereinafter, the video encoding apparatus can include an image encoding apparatus.
[0050] Referring to Figure 2 The encoding apparatus 200 includes an image partitioner 210, a predictor 220, a residual processor 230, and an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 can include an inter-predictor 221 and an intra-predictor 222. The residual processor 230 can include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 can further include a subtractor 231. The adder 250 can be referred to as a reconstructor or a reconstructed block generator. According to an embodiment, the image partitioner 210, the predictor 220, the residual processor 230, the entropy encoder 240, the adder 250, and the filter 260 can be constituted by at least one hardware component (e.g., an encoder chipset or a processor). In addition, the memory 270 can include a decoded picture buffer (DPB) or can be constituted by a digital storage medium. The hardware component can further include the memory 270 as an internal / external component.
[0051] The image partitioner 210 can partition an input image (or picture or frame) input to the encoding apparatus 200 into one or more processors. For example, the processor can be referred to as a coding unit (CU). In this case, the coding unit can be recursively partitioned from a coding tree unit (CTU) or a largest coding unit (LCU) according to a quad tree binary tree ternary (QTBTT) structure. For example, one coding unit can be partitioned into a plurality of coding units having a deeper depth based on a quad tree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quad tree structure can be applied first, and then the binary tree structure and / or the ternary structure can be applied. Alternatively, the binary tree structure can be applied first. The encoding process according to the disclosure can be performed based on the final coding unit that is no longer partitioned. In this case, the largest coding unit can be used as the final coding unit based on coding efficiency according to image characteristics, or if necessary, the coding unit can be recursively partitioned into a coding unit having a deeper depth and having an optimal size, and the coding unit can be used as the final coding unit. Here, the encoding process can include a process of prediction, transformation, and reconstruction, which will be described later. As another example, the processor can also include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit can be separated or partitioned from the final coding unit described above. The prediction unit can be a unit of sample prediction, and the transform unit can be a unit for deriving a transform coefficient and / or a unit for deriving a residual signal from the transform coefficient.
[0052] In some cases, the term unit can be used interchangeably with terms such as block or region. In general, an MxN block can represent a set of samples or transform coefficients consisting of M columns and N rows. The sample can generally represent a pixel or a pixel value, can represent only a pixel / pixel value of a luminance component, or can represent only a pixel / pixel value of a chrominance component. The sample can be used as a term corresponding to a picture (or image) of pixels or picture elements.
[0053] In the encoding device 200, a prediction signal (prediction block, prediction sample array) output from the inter-predictor 221 or the intra-predictor 222 is subtracted from an input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array) and the generated residual signal is sent to the transformer 232. In this case, as illustrated, a unit for subtracting a prediction signal (prediction block, prediction sample array) from an input image signal (original block, original sample array) in the encoding device 200 can be referred to as a subtractor 231. The predictor can perform prediction on a block to be processed (hereinafter, referred to as a current block) and generate a prediction block including predicted samples of the current block. The predictor can determine whether to apply intra-prediction or inter-prediction on a basis of the current block or CU. As described later in the description of each prediction mode, the predictor can generate various information related to prediction, such as prediction mode information, and send the generated information to the entropy encoder 240. The information about prediction can be encoded in the entropy encoder 240 and output in the form of a bitstream.
[0054] The intra-predictor 222 can predict the current block by referring to samples in the current picture. Depending on the prediction mode, the referred samples can be located in the vicinity of the current block, or can be far away from the current block. In intra-prediction, the prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes can include, for example, a DC mode and a planar mode. Depending on the level of detail of the prediction direction, the directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes. However, this is only an example, and depending on the settings, more or less directional prediction modes can be used. The intra-predictor 222 can determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring block.
[0055] The inter predictor 221 can derive a prediction block of a current block based on a reference block (a reference sample array) specified by a motion vector on a reference picture. Here, to reduce the amount of motion information transmitted in the inter prediction mode, motion information can be predicted in units of a block, a sub-block, or a sample based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the neighboring blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block can be the same or different. The temporal neighboring block can be referred to as a collocated reference block, a collocated CU (colCU), etc., and the reference picture including the temporal neighboring block can be referred to as a collocated picture (colPic). For example, the inter predictor 221 can configure a motion information candidate list based on the neighboring blocks, and generate information indicating which candidate to use to derive a motion vector and / or a reference picture index of the current block. Inter prediction can be performed based on various prediction modes. For example, in the case of a skip mode and a merge mode, the inter predictor 221 can use motion information of the neighboring blocks as motion information of the current block. In the skip mode, unlike the merge mode, a residual signal can not be transmitted. In the case of a motion vector prediction (MVP) mode, a motion vector of the neighboring block can be used as a motion vector predictor, and a motion vector of the current block can be indicated by signaling a motion vector difference.
[0056] The predictor 220 can generate a prediction signal based on various prediction methods described below. For example, the predictor can not only apply intra prediction or inter prediction to predict one block, but also can simultaneously apply both intra prediction and inter prediction. This can be referred to as combined inter-intra prediction (CIIP). In addition, the predictor can predict a block based on an intra block copy (IBC) prediction mode or a palette mode. The IBC prediction mode or the palette mode can be used for content image / video encoding of a game or the like, for example, screen content coding (SCC). The IBC basically performs prediction in the current picture, but can be performed similarly to inter prediction in that a reference block is derived in the current picture. That is, the IBC can use at least one of the inter prediction techniques described in the disclosure. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, sample values within a picture can be signaled based on information about a palette table and a palette index.
[0057] The prediction signal generated by the predictor (including the inter-predictor 221 and / or the intra-predictor 222) can be used to generate a reconstructed signal or to generate a residual signal. The transformer 232 can generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique can include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a karhunen-loève transform (KLT), a graph-based transform (GBT), or a conditional nonlinear transform (CNT). Here, the GBT denotes a transform obtained from a graph when relationship information between pixels is represented by a graph. The CNT refers to a transform generated based on a prediction signal generated using all previously reconstructed pixels. Also, the transform process can be applied to square pixel blocks having the same size, or can be applied to blocks having a variable size other than square.
[0058] The quantizer 233 can quantize the transform coefficients and transmit them to the entropy encoder 240, and the entropy encoder 240 can encode the quantized signal (information about the quantized transform coefficients) and output a bitstream. The information about the quantized transform coefficients can be referred to as residual information. The quantizer 233 can rearrange the block-type quantized transform coefficients into a one-dimensional vector form based on a coefficient scan order, and generate information about the quantized transform coefficients based on the one-dimensional vector-form quantized transform coefficients. Information about the transform coefficients can be generated. The entropy encoder 240 can perform various encoding methods such as, for example, exponential Golomb, context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), and the like. The entropy encoder 240 can encode information (e.g., values of syntax elements, etc.) required for video / image reconstruction, together with or separately from the quantized transform coefficients. The encoded information (e.g., encoded video / image information) can be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer). The video / image information can further include information about various parameter sets such as an adaptive parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information can further include general constraint information. In the disclosure, information and / or syntax elements transmitted / signaled from the encoding apparatus to the decoding apparatus can be included in the video / picture information. The video / picture information can be encoded through the above-described encoding process and included in the bitstream. The bitstream can be transmitted through a network, or can be stored in a digital storage medium. The network can include a broadcast network and / or a communication network, and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) that transmits a signal output from the entropy encoder 240 and / or a storage unit (not shown) that stores the signal can be included as an internal / external element of the encoding apparatus 200, and alternatively, the transmitter can be included in the entropy encoder 240.
[0059] The quantized transform coefficients output from the quantizer 233 can be used to generate a prediction signal. For example, a residual signal (a residual block or a residual sample) can be reconstructed by applying dequantization and inverse transform to the quantized transform coefficients using the dequantizer 234 and the inverse transformer 235. The adder 250 adds the reconstructed residual signal to the prediction signal output from the inter-predictor 221 or the intra-predictor 222 to generate a reconstructed signal (a reconstructed picture, a reconstructed block, a reconstructed sample array). If a block to be processed has no residual (such as a case where a skip mode is applied), a prediction block can be used as a reconstructed block. The adder 250 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-prediction of a next block to be processed in a current picture, and can be used for inter-prediction of a next picture through filtering as described below.
[0060] Further, during picture encoding and / or reconstruction, luma mapping with chroma scaling (LMCS) can be applied.
[0061] The filter 260 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 260 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and store the modified reconstructed picture in the memory 270 (specifically, a DPB of the memory 270). The various filtering methods can include, for example, a deblocking filter, a sample adaptive offset, an adaptive loop filter, a bilateral filter, etc. The filter 260 can generate various information related to filtering, and transmit the generated information to the entropy encoder 240, as described later in descriptions of the various filtering methods. The information related to filtering can be encoded by the entropy encoder 240 and output in the form of a bitstream.
[0062] The modified reconstructed picture transmitted to the memory 270 can be used as a reference picture in the inter-predictor 221. When inter-prediction is applied by the encoding apparatus, prediction mismatch between the encoding apparatus 200 and a decoding apparatus can be avoided, and encoding efficiency can be improved.
[0063] The DPB of the memory 270 can store the modified reconstructed picture used as a reference picture in the inter-predictor 221. The memory 270 can store motion information of a block from which motion information in a current picture is derived (or encoded), and / or motion information of a reconstructed block in a picture. The stored motion information can be transmitted to the inter-predictor 221, and used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory 270 can store reconstructed samples of a reconstructed block in a current picture, and can deliver the reconstructed samples to the intra-predictor 222.
[0064] Figure 3 is a schematic diagram illustrating a configuration of a video / image decoding apparatus to which embodiments of the present disclosure can be applied.
[0065] Referring to Figure 3 , the decoding apparatus 300 can include an entropy decoder 310, a residue processor 320, a predictor 330, an adder 340, a filter 350, a memory 360. The predictor 330 can include an inter-predictor 332 and an intra-predictor 331. The residue processor 320 can include a dequantizer 321 and an inverse transformer 322. According to an embodiment, the entropy decoder 310, the residue processor 320, the predictor 330, the adder 340, and the filter 350 can be constituted by hardware components (e.g., a decoder chipset or a processor). In addition, the memory 360 can include a decoded picture buffer (DPB), or can be constituted by a digital storage medium. The hardware components can further include the memory 360 as an internal / external component.
[0066] When a bitstream including video / image information is input, the decoding apparatus 300 can reconstruct an image corresponding to the processing of the video / image information processed in the encoding apparatus of Figure 2 . For example, the decoding apparatus 300 can derive a unit / block based on block partitioning-related information obtained from the bitstream. The decoding apparatus 300 can perform decoding using a processor applied in the encoding apparatus. Accordingly, the processor for decoding can be, for example, an encoding unit, and can partition an encoding unit from a coding tree unit or a largest coding unit according to a quad tree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units can be derived from the encoding unit. The reconstructed image signal decoded and output by the decoding apparatus 300 can be reproduced by a reproducing apparatus.
[0067] The decoding apparatus 300 can receive a bitstream from Figure 2The signal output from the encoding apparatus can be received and decoded by the entropy decoder 310. For example, the entropy decoder 310 can parse a bitstream to derive information (e.g., video / image information) required for image reconstruction (or picture reconstruction). The video / image information can further include information on various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information can further include general constraint information. The decoding apparatus can further decode a picture based on the information on the parameter sets and / or the general constraint information. The signaled / received information and / or syntax elements described later in the disclosure can be decoded through the decoding process and obtained from the bitstream. For example, the entropy decoder 310 decodes information in the bitstream based on an encoding method such as exponential Golomb coding, CAVLC, or CABAC, and outputs syntax elements and quantized values of transform coefficients of a residual required for image reconstruction. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element in the bitstream, determine a context model using information of a decoding target syntax element, decoding information of a decoding target block, or a symbol / bin decoded in a previous stage, and arithmetically decode the bins by predicting a probability of occurrence of the bins according to the determined context model, and generate a symbol corresponding to a value of each syntax element. In this case, after the context model is determined, the CABAC entropy decoding method can update the context model by using information of the decoded symbol / bin for the context model of the next symbol / bin. Information related to prediction among the information decoded by the entropy decoder 310 can be provided to the predictors (inter-predictor 332 and intra-predictor 331), and residual values (that is, quantized transform coefficients and related parameter information) on which entropy decoding is performed in the entropy decoder 310 can be input to the residual processor 320. The residual processor 320 can derive a residual signal (a residual block, a residual sample, a residual sample array). In addition, information on filtering among the information decoded by the entropy decoder 310 can be provided to the filter 350. Further, a receiver (not shown) for receiving a signal output from the encoding apparatus can be further configured as an internal / external element of the decoding apparatus 300, or the receiver can be a component of the entropy decoder 310. In addition, the decoding apparatus according to the disclosure can be referred to as a video / image / picture decoding apparatus, and the decoding apparatus can be classified into an information decoder (a video / image / picture information decoder) and a sample decoder (a video / image / picture sample decoder). The information decoder can include the entropy decoder 310, and the sample decoder can include at least one of the inverse quantizer 321, the inverse transformer 322, the adder 340, the filter 350, the memory 360, the inter-predictor 332, and the intra-predictor 331.
[0068] The dequantizer 321 can dequantize the quantized transform coefficients and output the transform coefficients. The dequantizer 321 can rearrange the quantized transform coefficients in the form of a two-dimensional block. In this case, the rearrangement can be performed based on a coefficient scanning order performed in the encoding apparatus. The dequantizer 321 can perform dequantization on the quantized transform coefficients by using a quantization parameter (e.g., quantization step length information), and obtain the transform coefficients.
[0069] The inverse transformer 322 inverse-transforms the transform coefficients to obtain a residual signal (a residual block, a residual sample array).
[0070] The predictor can perform prediction on the current block and generate a prediction block including predicted samples of the current block. The predictor can determine whether to apply intra prediction or inter prediction to the current block based on information about prediction output from the entropy decoder 310, and can determine a specific intra / inter prediction mode.
[0071] The predictor can generate a predicted signal based on various prediction methods described below. For example, the predictor can not only apply intra prediction or inter prediction to predict one block, but also can simultaneously apply intra prediction and inter prediction. This can be referred to as combined inter and intra prediction (CIIP). In addition, the predictor can predict a block based on an intra block copy (IBC) prediction mode or a palette mode. The IBC prediction mode or the palette mode can be used for content image / video encoding of games, etc., for example, screen content coding (SCC). The IBC basically performs prediction in the current picture, but can be performed similarly to inter prediction in that a reference block is derived in the current picture. That is, the IBC can use at least one of the inter prediction techniques described in the disclosure. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, sample values within a picture can be signaled based on information about a palette table and a palette index.
[0072] The intra predictor 331 can predict the current block by referring to samples in the current picture. Depending on the prediction mode, the referred samples can be located in the vicinity of the current block, or can be far from the current block. In intra prediction, the prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The intra predictor 331 can determine a prediction mode applied to the current block by using a prediction mode applied to a neighboring block.
[0073] The inter predictor 332 can derive a prediction block of the current block based on reference samples array of a reference block specified by a motion vector on a reference picture. In this case, to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of a block, a sub-block, or a sample based on the correlation of the motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the neighboring blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter predictor 332 can configure a motion information candidate list based on the neighboring blocks and derive a motion vector and / or a reference picture index of the current block based on received candidate selection information. The inter prediction can be performed based on various prediction modes, and information about the prediction can include information indicating a mode of the inter prediction for the current block.
[0074] The adder 340 can generate a reconstructed signal (a reconstructed picture, a reconstructed block, a reconstructed sample array) by adding the obtained residual signal to a prediction signal (a prediction block, a prediction sample array) output from the predictor (including the inter predictor 332 and / or the intra predictor 331). If there is no residual for a block to be processed (for example, when a skip mode is applied), the prediction block can be used as the reconstructed block.
[0075] The adder 340 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra prediction of a next block to be processed in the current picture, can be output through filtering as described below, or can be used for inter prediction of a next picture.
[0076] In addition, luminance mapping and chrominance scaling (LMCS) can be applied in the picture decoding process.
[0077] The filter 350 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 350 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture and store the modified reconstructed picture in the memory 360 (specifically, the DPB of the memory 360). The various filtering methods can include, for example, a deblocking filter, a sample adaptive offset, an adaptive loop filter, a bilateral filter, etc.
[0078] The reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter prediction 332. The memory 360 can store motion information of a block from which motion information in the current picture is derived (or decoded) and / or motion information of a reconstructed block in the picture. The stored motion information can be sent to the inter prediction 332 to be utilized as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory 360 can store reconstructed samples of a reconstructed block in the current picture and can transfer the reconstructed samples to the intra prediction 331.
[0079] In the disclosure, the embodiments described in the filter 260, the inter prediction 221, and the intra prediction 222 of the encoding apparatus 200 can be the same as or respectively applied to correspond to the filter 350, the inter prediction 332, and the intra prediction 331 of the decoding apparatus 300. The same content can also be applied to the inter prediction 332 and the intra prediction 331.
[0080] In the disclosure, at least one of quantization / dequantization and / or transform / inverse transform can be omitted. When the quantization / dequantization is omitted, the quantized transform coefficient can be referred to as a transform coefficient. When the transform / inverse transform is omitted, the transform coefficient can be referred to as a coefficient or a residual coefficient, or for the unity of expression, can still be referred to as a transform coefficient.
[0081] In the disclosure, the quantized transform coefficient and the transform coefficient can be referred to as a transform coefficient and a scaled transform coefficient, respectively. In this case, the residual information can include information on the transform coefficient, and the information on the transform coefficient can be signaled through a residual coding syntax. The transform coefficient can be derived based on the residual information (or the information on the transform coefficient), and the scaled transform coefficient can be derived by inverse transforming (scaling) the transform coefficient. The residual sample can be derived based on inverse transforming (transforming) the scaled transform coefficient. This can also be applied / expressed in other parts of the disclosure.
[0082] Further, a picture output and removal process in a decoded picture buffer (DPB) can be performed. The picture output and removal process in the decoded picture buffer (DPB) in the existing VVC standard of the video / image encoding system can be as shown in the following table.
[0083] [table 1]
[0084]
[0085]
[0086]
[0087]
[0088] For example, according to the VVC standard of the video / image encoding system, the picture output process can be called once per picture, as disclosed in the above table, before decoding the current picture (but after parsing the slice header of the first slice of the current picture).
[0089] In addition, for example, with reference to Table 1, when the current access unit (AU) is a coded video sequence start AU (CVSS AU) other than AU 0, the following sequence of steps can be applied.
[0090] - First, for the decoder under test, the variable NoOutputOfPriorPicsFlag can be derived as follows.
[0091] - When the values of PicWidthMaxInSamplesY, PicHeightMaxInSamplesY, MaxChromaFormat, MaxBitDepthMinus8 or max_dec_pic_buffering_minus1[Htid] derived for the current AU are different from the values of PicWidthMaxInSamplesY, PicHeightMaxInSamplesY, MaxChromaFormat, MaxBitDepthMinus8 or max_dec_pic_buffering_minus1[Htid] derived for the previous AU, NoOutputOfPriorPicsFlag can be set to equal to 1 by the decoder under test, regardless of the value of ph_no_output_of_prior_pics_flag for the current AU.
[0092] - Otherwise, NoOutputOfPriorPicsFlag can be set to equal to the value of ph_no_output_of_prior_pics_flag for the current AU.
[0093] - Second, the variable NoOutputOfPriorPicsFlag derived for the decoder under test can be applied to the hypothetical reference decoder (HRD). Thus, when the value of NoOutputOfPriorPicsFlag is 1, all picture storage buffers of the DPB can be emptied without outputting the pictures they contain, and the DPB fullness can be set to 0.
[0094] In addition, for example, with reference to Table 1, when all the following conditions are true for any picture k of the DPB, all such pictures k of the DPB can be removed from the DPB.
[0095] - picture k is marked as "unused for reference".
[0096] - picture k has a PictureOutputFlag equal to 0 or a DPB output time of picture k is less than or equal to the CPB removal time of the first decoding unit (DU), denoted as DU m, of the current picture n; i.e., DpbOutputTime[ k ] is less than or equal to DuCpbRemovalTime[ m ].
[0097] In addition, for example, with reference to Table 1, when the current access unit (AU) is a coded video sequence start AU (CVSS AU) other than AU 0, the following ordered steps can be applied.
[0098] - First, for the decoder under test, the variable NoOutputOfPriorPicsFlag can be derived as follows.
[0099] - When the values of PicWidthMaxInSamplesY, PicHeightMaxInSamplesY, MaxChromaFormat, MaxBitDepthMinus8, or max_dec_pic_buffering_minus1[ Htid ] derived for the current AU are different from the values of PicWidthMaxInSamplesY, PicHeightMaxInSamplesY, MaxChromaFormat, MaxBitDepthMinus8, or max_dec_pic_buffering_minus1[ Htid ] derived for the previous AU, NoOutputOfPriorPicsFlag can be set by the decoder under test to equal to 1, regardless of the value of ph_no_output_of_prior_pics_flag for the current AU.
[0100] - Otherwise, NoOutputOfPriorPicsFlag can be set to equal to the value of ph_no_output_of_prior_pics_flag for the current AU.
[0101] - Second, the variable NoOutputOfPriorPicsFlag derived for the decoder under test can be applied to the HRD (hypothetical reference decoder) as follows.
[0102] For example, when the value of NoOutputOfPriorPicsFlag is 1, all picture storage buffers of the DPB can be emptied without outputting the pictures they contain, and the DPB fullness can be set to 0.
[0103] Otherwise (i.e., when the value of NoOutputOfPriorPicsFlag is 0), all picture storage buffers of the DPB containing pictures that are marked as “not needed for output” and “unused for reference” can be emptied (without output), and all non-empty picture storage buffers in the DPB can be emptied by repeatedly invoking the “bumping” process specified in clause C.5.2.4 of the VVC standard, and the DPB fullness can be set to 0.
[0104] Further, the bumping process can consist of the following sequence of steps.
[0105] 1. Among all pictures of the DPB that are marked as “needed for output”, the first picture (or pictures) to be output can be selected as the picture(s) with the smallest PicOrderCntVal value.
[0106] 2. Each of the pictures in ascending order of nuh layer id can be cropped using the consistent cropping window for the picture, the cropped picture can be output, and the picture can be marked as “not needed for output”.
[0107] 3. Each picture storage buffer containing a picture that is marked as “unused for reference” and is one of the pictures that were cropped and output can be emptied, and the fullness of the DPB can be decremented by 1.
[0108] Further, for example, with reference to Table 1, when the current AU is not a CVSS AU, all picture storage buffers containing pictures that are marked as “not needed for output” and “unused for reference” can be emptied (without output). For each picture storage buffer, the DPB fullness can be decremented by 1. Further, the “bumping” process specified in clause C.5.2.4 of the VVC standard can be repeatedly invoked while further decrementing the DPB fullness by 1 for each additional picture storage buffer that is emptied until none of the following conditions are true.
[0109] The number of pictures in the DPB that are marked as “needed for output” is greater than max num reorder pics[Htid].
[0110] - max latency increase plusl [Htid] is not equal to 0, and there is at least one picture in the DPB that is marked as “needs output” and whose associated variable PicLatencyCount is greater than or equal to MaxLatencyPictures [Htid].
[0111] - the number of pictures in the DPB is greater than or equal to max dec pic buffering minusl [Htid] + 1.
[0112] On the other hand, the existing VVC standard for the above-mentioned picture output and removal process can have the following problems.
[0113] For example, first, a picture can be marked as “used for short-term reference” after all slices of the picture are decoded. Therefore, the picture cannot be in the clean-out state in the DPB when the picture is decoded. As a result, the picture storage number of the DPB can be affected.
[0114] Second, the assignment of the output status (i.e., needs output) of a picture can be performed during the bumping process. According to the existing VVC standard, for a picture of an AU that is the start of a coded video sequence AU, the process cannot be invoked. Accordingly, the value of PicLatencyCount related to the corresponding picture cannot be initialized.
[0115] As described above, the process of outputting and removing a picture in the DPB can be invoked once per picture, but the process can affect the state of the DPB (i.e., the state of the pictures stored in the DPB) shared by all layers of the CVS. In consideration of the above fact, the process of deriving NoOutputOfPriorPicsFlag from the DPB based on the value of NoOutputOfPriorPicsFlag and removing a picture can be problematic. According to the existing video / image standard, for all pictures of a CVS AU except AU 0, the process of deriving and removing a picture from the DPB based on the value of the flag (i.e., NoOutputOfPriorPicsFlag) can be invoked. As described above, the execution of the process can be possible only for the first picture. The process starting from the second picture can remove the previous picture from the DPB before outputting the previous picture (i.e., the picture in the previous order in the decoding order). This behavior can not be a correct decoder behavior.
[0116] Accordingly, the present disclosure proposes a solution to the above-mentioned problems. The proposed embodiments can be applied independently or in combination.
[0117] As an example, the process of deriving the value of the flag or variable indicating whether a reference picture is removed from the DPB without being output can be invoked only once per access unit (AU). That is, for example, a method can be proposed such that the process of deriving the value of the flag or variable indicating whether a reference picture is removed from the DPB without being output is invoked only once per access unit (AU). Here, the variable can be NoOutputOfPriorPicsFlag.
[0118] In addition, as an example, the process of deriving the value of NoOutputOfPriorPicsFlag can be invoked before the decoding process of the first picture in a coded video sequence start AU (CVSS AU) but after parsing the slice header of the first slice of the current picture. That is, for example, a method can be proposed that performs the process of deriving the value of NoOutputOfPriorPicsFlag before the decoding process of the first picture in a CVSS AU but after parsing the slice header of the first slice of the current picture.
[0119] In addition, as an example, the process of removing a picture stored in the DPB without outputting it can be invoked only once per AU when NoOutputOfPriorPicsFlag is 1. That is, for example, a method can be proposed that removes a picture stored in the DPB without outputting it only once per AU when NoOutputOfPriorPicsFlag is 1.
[0120] In addition, as an example, the process of removing a picture stored in the DPB without outputting it can be invoked before the decoding process of the first picture in a CVSS AU but after parsing the slice header of the first slice of the current picture when NoOutputOfPriorPicsFlag is 1. That is, for example, a method can be proposed that invokes the process of removing a picture stored in the DPB without outputting it before the decoding process of the first picture in a CVSS AU but after parsing the slice header of the first slice of the current picture when NoOutputOfPriorPicsFlag is 1.
[0121] In addition, as an example, the picture removal in the above-described embodiments can not include current picture removal in the DPB. That is, for example, a method can be proposed that the picture removal in the above-described embodiments does not include current picture removal in the DPB.
[0122] The above-described embodiments can be implemented as follows. For example, the above-described embodiments can be expressed based on the VVC standard specification, as described below.
[0123] [Table 2]
[0124]
[0125]
[0126]
[0127]
[0128] For example, referring to Table 2, when the current picture is the first picture and the current AU (i.e. the AU including the current picture) is a coded video sequence start AU (CVSS AU) except for AU 0, the following sequence of steps can be applied.
[0129] - First, for the decoder under test, the variable NoOutputOfPriorPicsFlag can be derived as follows.
[0130] - When the value of PicWidthMaxInSamplesY, PicHeightMaxInSamplesY, MaxChromaFormat, MaxBitDepthMinus8 or max_dec_pic_buffering_minus1 [Htid] derived for the current AU is different from the value of PicWidthMaxInSamplesY, PicHeightMaxInSamplesY, MaxChromaFormat, MaxBitDepthMinus8 or max_dec_pic_buffering_minus1 [Htid] derived for the previous AU, NoOutputOfPriorPicsFlag can be set to equal to 1 by the decoder under test, regardless of the value of ph_no_output_of_prior_pics_flag of the current AU.
[0131] - Otherwise, NoOutputOfPriorPicsFlag can be set to equal to the value of ph_no_output_of_prior_pics_flag of the current AU.
[0132] - Second, the variable NoOutputOfPriorPicsFlag derived for the decoder under test can be applied to the hypothetical reference decoder (HRD). Thus, when the value of NoOutputOfPriorPicsFlag is 1, all picture storage buffers of the DPB can be emptied without outputting the pictures they contain, and the DPB fullness can be set to 0.
[0133] In addition, for example, referring to Table 2, when the current AU is not a CVSS AU or the current AU is a CVSS AU other than AU 0 but the current picture is not the first picture in the current AU and when all the following conditions are true for any picture k in the DPB, all such pictures k in the DPB can be removed from the DPB.
[0134] - picture k is marked as "unused for reference".
[0135] - picture k has a PictureOutputFlag equal to 0 or a DPB output time of picture k is less than or equal to the CPB removal time of the first decoding unit (DU) of the current picture n, denoted as DU m; i.e., DpbOutputTime[ k ] is less than or equal to DuCpbRemovalTime[ m ].
[0136] In addition, for example, referring to Table 2, when the current picture is the first picture and the current AU (i.e., the AU that includes the current picture) is a coded video sequence start AU (CVSS AU) other than AU 0, the following steps in order can be applied.
[0137] - First, for the decoder under test, the variable NoOutputOfPriorPicsFlag can be derived as follows.
[0138] - When the value of PicWidthMaxInSamplesY, PicHeightMaxInSamplesY, MaxChromaFormat, MaxBitDepthMinus8, or max_dec_pic_buffering_minus1[ Htid ] derived for the current AU is different from the value of PicWidthMaxInSamplesY, PicHeightMaxInSamplesY, MaxChromaFormat, MaxBitDepthMinus8, or max_dec_pic_buffering_minus1[ Htid ] derived for the previous AU, NoOutputOfPriorPicsFlag can be set to equal to 1 by the decoder under test, regardless of the value of ph_no_output_of_prior_pics_flag for the current AU.
[0139] - Otherwise, NoOutputOfPriorPicsFlag can be set to equal to the value of ph_no_output_of_prior_pics_flag for the current AU.
[0140] - Secondly, the variable NoOutputOfPriorPicsFlag derived for the decoder under test can be applied to the hypothetical reference decoder (HRD) as follows.
[0141] - For example, when the value of NoOutputOfPriorPicsFlag is 1, all picture storage buffers of the DPB can be emptied without outputting the pictures they contain, and the DPB fullness can be set to 0.
[0142] - Otherwise (i.e., when the value of NoOutputOfPriorPicsFlag is 0), all picture storage buffers in the DPB containing pictures marked as “not needed for output” and “unused for reference” can be emptied (without output), and all non-empty picture storage buffers in the DPB can be emptied by repeatedly invoking the “hollowing” process specified in clause C.5.2.4 of the VVC standard, and the DPB fullness can be set to 0.
[0143] In addition, for example, with reference to Table 2, when the current AU is not a CVSS AU or the current AU is a CVSS AU other than AU 0 but the current picture is not the first picture of the current AU, all picture storage buffers containing pictures marked as “not needed for output” and “unused for reference” can be emptied (without output). For each picture storage buffer that is emptied, the DPB fullness can be decremented by 1. In addition, the “hollowing” process specified in clause C.5.2.4 of the VVC standard can be repeatedly invoked while further decrementing the DPB fullness by 1 for each additional picture storage buffer that is emptied until none of the following conditions are true.
[0144] - The number of pictures in the DPB marked as “needed for output” is greater than max num reorder pics [Htid].
[0145] - max latency increase plusl [Htid] is not equal to 0, and there is at least one picture in the DPB that is marked as “needed for output” and whose associated variable PicLatencyCount is greater than or equal to MaxLatencyPictures [Htid].
[0146] - The number of pictures in the DPB is greater than or equal to max dec pic buffering minusl [Htid] + 1.
[0147] Alternatively, the above-described embodiments can be implemented as follows. For example, the above-described embodiments can be expressed based on the VVC standard specification, as follows.
[0148] [Table 3]
[0149]
[0150]
[0151]
[0152] For example, referring to Table 3, when the current picture is the first picture of the current AU and the current AU is a coded video sequence start AU (CVSS AU) except for AU 0, the following sequence of steps can be applied.
[0153] - First, for the decoder under test, the variable NoOutputOfPriorPicsFlag can be derived as follows.
[0154] - When the value of PicWidthMaxInSamplesY, PicHeightMaxInSamplesY, MaxChromaFormat, MaxBitDepthMinus8 or max_dec_pic_buffering_minus1 [Htid] derived for the current AU is different from the value of PicWidthMaxInSamplesY, PicHeightMaxInSamplesY, MaxChromaFormat, MaxBitDepthMinus8 or max_dec_pic_buffering_minus1 [Htid] derived for the previous AU, NoOutputOfPriorPicsFlag can be set to equal to 1 by the decoder under test, regardless of the value of ph_no_output_of_prior_pics_flag for the current AU.
[0155] - Otherwise, NoOutputOfPriorPicsFlag can be set to equal to the value of ph_no_output_of_prior_pics_flag for the current AU.
[0156] - Second, the variable NoOutputOfPriorPicsFlag derived for the decoder under test can be applied to the hypothetical reference decoder (HRD). Thus, when the value of NoOutputOfPriorPicsFlag is 1, all picture storage buffers of the DPB can be emptied without outputting the pictures they contain, and the DPB fullness can be set to 0.
[0157] In addition, for example, referring to Table 3, when the current AU is not a CVSS AU or the current AU is a CVSS AU other than AU 0 but the current picture is not the first picture of the current AU and when all the following conditions are true for any picture k in the DPB, all such pictures k in the DPB can be removed from the DPB.
[0158] - picture k is marked as "unused for reference".
[0159] - picture k has a PictureOutputFlag equal to 0 or a DpbOutputTime of picture k is less than or equal to the CPB removal time of the first decoding unit (DU) of the current picture n, denoted as DU m; i.e., DpbOutputTime[ k ] is less than or equal to DuCpbRemovalTime[ m ].
[0160] In addition, for example, referring to Table 3, when the current picture is the first picture and the current access unit (AU) is a coded video sequence start AU (CVSS AU) other than AU 0, the following steps in order can be applied.
[0161] - First, for the decoder under test, the variable NoOutputOfPriorPicsFlag can be derived as follows.
[0162] - When the value of PicWidthMaxInSamplesY, PicHeightMaxInSamplesY, MaxChromaFormat, MaxBitDepthMinus8 or max_dec_pic_buffering_minus1[ Htid ] derived for the current AU is different from the value of PicWidthMaxInSamplesY, PicHeightMaxInSamplesY, MaxChromaFormat, MaxBitDepthMinus8 or max_dec_pic_buffering_minus1[ Htid ] derived for the previous AU, NoOutputOfPriorPicsFlag can be set to equal to 1 by the decoder under test, regardless of the value of ph_no_output_of_prior_pics_flag for the current AU.
[0163] - Otherwise, NoOutputOfPriorPicsFlag can be set to equal to the value of ph_no_output_of_prior_pics_flag for the current AU.
[0164] - Second, the variable NoOutputOfPriorPicsFlag derived for the decoder under test can be applied to the hypothetical reference decoder (HRD) as follows.
[0165] - For example, when the value of NoOutputOfPriorPicsFlag is 1, all picture storage buffers of the DPB can be emptied without outputting the pictures they contain, and the DPB fullness can be set to 0.
[0166] - Otherwise (i.e., when the value of NoOutputOfPriorPicsFlag is 0), all picture storage buffers in the DPB containing pictures that are marked as “not needed for output” and “unused for reference” can be emptied (without output), and all non-empty picture storage buffers in the DPB can be emptied by repeatedly invoking the “hollowing” process specified in clause C.5.2.4 of the VVC standard, and the DPB fullness can be set to 0.
[0167] In addition, for example, with reference to Table 3, when the current AU is not a CVSS AU or the current AU is a CVSS AU other than AU 0 but the current picture is not the first picture of the current AU, all picture storage buffers containing pictures that are marked as “not needed for output” and “unused for reference” can be emptied (without output). For each picture storage buffer that is emptied, the DPB fullness can be decremented by 1. In addition, the “hollowing” process specified in clause C.5.2.4 of the VVC standard can be repeatedly invoked while further decrementing the DPB fullness by 1 for each additional picture storage buffer that is emptied until none of the following conditions are true.
[0168] - The number of pictures in the DPB that are marked as “needed for output” is greater than max_num_reorder_pics[Htid].
[0169] - max_latency_increase_plus1[Htid] is not equal to 0, and there is at least one picture in the DPB that is marked as “needed for output” and whose associated variable PicLatencyCount is greater than or equal to MaxLatencyPictures[Htid].
[0170] - The number of pictures in the DPB is greater than or equal to max_dec_pic_buffering_minus1[Htid] + 1.
[0171] Furthermore, for example, the embodiments can be applied according to the following procedure. One or more steps of the procedure, which will be described later, can be omitted.
[0172] Figure 4 An example illustrates an encoding process according to an embodiment of the disclosure.
[0173] Referring to Figure 4 The encoding device decodes (retrieves) pictures (S400). The encoding device can decode pictures of a current AU.
[0174] The encoding device manages the DPB based on the DPB parameters (S410). Here, the DPB management can be referred to as DPB update. The DPB management process can include a process of marking and / or removing decoded pictures in the DPB. The decoded pictures can be used as a reference for inter prediction of subsequent pictures. That is, the decoded pictures can be used as reference pictures for inter prediction of pictures that are subsequent in decoding order. Each decoded picture can be substantially inserted into the DPB. In addition, the DPB can be updated generally before decoding a current picture. When a layer related to the DPB is not an output layer (or the DPB parameters are not related to the output layer) and is a reference layer, the decoded pictures in the DPB cannot be output. If a layer related to the DPB (or the DPB parameters) is an output layer, the decoded pictures in the DPB can be output based on the DPB and / or the DPB parameters. The DPB management can include outputting the decoded pictures from the DPB.
[0175] The encoding device encodes picture information including information related to the DPB parameters (S420). The information related to the DPB parameters can include the information / syntax elements disclosed in the above-described embodiments and / or the syntax elements disclosed in a table to be described later.
[0176] [Table 4]
[0177]
[0178] For example, the above-described Table 4 can represent a video parameter set (VPS) including syntax elements for the DPB parameters to be signaled.
[0179] The semantics for the syntax elements shown in the above Table 4 can be as follows.
[0180] [Table 5]
[0181]
[0182]
[0183]
[0184] For example, the syntax element vps num dpb params can indicate a number of dpb parameters ( ) syntax structures in the VPS. For example, a value of vps num dpb params can be in a range of 0 to 16. In addition, when the syntax element vps num dpb params is not present, a value of the syntax element vps num dpb params can be inferred to be equal to 0.
[0185] In addition, for example, the syntax element same dpb size output or nonoutput flag can indicate whether the syntax element layer nonoutput dpb params idx [ i ] can be present in the VPS. For example, when a value of the syntax element same dpb size output or nonoutput flag is 1, the syntax element same dpb size output or nonoutput flag can indicate that the syntax element layer nonoutput dpb params idx [ i ] is not present in the VPS, and when a value of the syntax element same dpb size output or nonoutput flag is 0, the syntax element same dpb size output or nonoutput flag can indicate that the syntax element layer nonoutput dpb params idx [ i ] can be present in the VPS.
[0186] In addition, for example, the syntax element vps sublayer dpb params present flag can be used to control the presence of the syntax elements max dec pic buffering minusl [ ], max num reorder pics [ ], and max latency increase plusl [ ] in the dpb parameters ( ) syntax structure of the VPS. In addition, when the syntax element vps sublayer dpb params present flag is not present, a value of the syntax element vps sublayer dpb params present flag can be inferred to be equal to 0.
[0187] Additionally, for example, the syntax element dpb_size_only_flag[ i ] can indicate whether the syntax elements max_num_reorder_pics[ ] and max_latency_increase_plusl[ ] can be present in the i-th dpb_parameters( ) syntax structure of the VPS. For example, when the value of the syntax element dpb_size_only_flag[ i ] is 1, the syntax element dpb_size_only_flag[ i ] can indicate that the syntax elements max_num_reorder_pics[ ] and max_latency_increase_plusl[ ] are not present in the i-th dpb_parameters( ) syntax structure of the VPS. When the value of the syntax element dpb_size_only_flag[ i ] is 0, the syntax element dpb_size_only_flag[ i ] can indicate that the syntax elements max_num_reorder_pics[ ] and max_latency_increase_plusl[ ] can be present in the i-th dpb_parameters( ) syntax structure of the VPS.
[0188] Additionally, for example, the syntax element dpb_max_temporal_id[ i ] can indicate the highest sublayer representation Temporalld in which DPB parameters can be present in the i-th dpb_parameters( ) syntax structure of the VPS. Additionally, the value of dpb_max_temporal_id[ i ] can be in the range of 0 to vps_max_sublayers_minusl. Additionally, for example, when the value of vps_max_sublayers_minusl is 0, the value of dpb_max_temporal_id[ i ] can be inferred to be 0. Additionally, for example, when the value of vps_max_sublayers_minusl is greater than 0 and vps_all_layers_same_num_sublayers_flag is 1, the value of dpb_max_temporal_id[ i ] can be inferred to be equal to vps_max_sublayers_minusl.
[0189] Additionally, for example, the syntax element layer_output_dpb_params_idx[ i ] can specify an index to a dpb_parameters( ) syntax structure in the list of dpb_parameters( ) syntax structures in the VPS that applies to the dpb_parameters( ) syntax structure of the i-th layer that is an output layer of an OLS. When the syntax element layer_output_dpb_params_idx[ i ] is present, the value of the syntax element layer_output_dpb_params_idx[ i ] can be in the range of 0 to vps_num_dpb_params - 1.
[0190] For example, when vps_independent_layer_flag[ i ] is 1, it can be a dpb_parameters( ) syntax structure in the SPS that the layer of dpb_parameters( ) syntax structure referring to is applied to the i-th layer that is an output layer.
[0191] Alternatively, for example, when vps_independent_layer_flag[ i ] is 0, the following can be applied.
[0192] - When vps_num_dpb_params is 1, the value of layer_output_dpb_params_idx[ i ] can be inferred to be 0.
[0193] - The requirement of bitstream conformance is that the value of layer_output_dpb_params_idx[ i ] makes the value of dpb_size_only_flag[ layer_output_dpb_params_idx[ i ] ] to be 0.
[0194] Additionally, for example, the syntax element layer_nonoutput_dpb_params_idx[ i ] can specify an index to a dpb_parameters( ) syntax structure in the list of dpb_parameters( ) syntax structures in the VPS that applies to the dpb_parameters( ) syntax structure of the i-th layer that is a non-output layer of an OLS. When the syntax element layer_nonoutput_dpb_params_idx[ i ] is present, the value of the syntax element layer_nonoutput_dpb_params_idx[ i ] can be in the range of 0 to vps_num_dpb_params - 1.
[0195] For example, when same_dpb_size_output_or_nonoutput_flag is 1, the following can be applied.
[0196] - When vps_independent_layer_flag[ i ] is equal to 1, it can be the dpb_parameters( ) syntax structure in the SPS referred to by the layer of the dpb_parameters( ) syntax structure applied to the i-th layer that is a non-output layer.
[0197] - When vps_independent_layer_flag[ i ] is equal to 0, the value of layer_nonoutput_dpb_params_idx[ i ] can be inferred to be equal to layer_output_dpb_params_idx[ i ].
[0198] Alternatively, for example, when same_dpb_size_output_or_nonoutput_flag is equal to 0, if vps_num_dpb_params is equal to 1, the value of layer_output_dpb_params_idx[ i ] can be inferred to be 0.
[0199] Further, for example, the dpb_parameters( ) syntax structure as the DPB parameters syntax structure disclosed in Table 4 above can be as follows.
[0200] [Table 6]
[0201]
[0202] Referring to Table 6, the dpb_parameters( ) syntax structure can provide information on a DPB size, a maximum picture reorder number, and a maximum latency for each CLVS of a CVS. The dpb_parameters( ) syntax structure can denote information on DPB parameters or DPB parameter information.
[0203] When the dpb_parameters( ) syntax structure is included in a VPS, the VPS can specify an OLS to which the dpb_parameters( ) syntax structure is applied. In addition, when the dpb_parameters( ) syntax structure is included in an SPS, the dpb_parameters( ) syntax structure can be applied to only an OLS including a lowest layer among layers referring to the SPS, wherein the lowest layer can be an independent layer.
[0204] Semantics for the syntax elements shown in Table 6 above can be as follows.
[0205] [Table 7]
[0206]
[0207]
[0208] For example, the value obtained by adding 1 to the syntax element max_dec_pic_buffering_minus1[i] can represent the maximum required size of the DPB in units of the image storage buffer for each CLVS of CVS when Htid equals i. For example, max_dec_pic_buffering_minus1[i] can be information about the DPB size. For example, the value of the syntax element max_dec_pic_buffering_minus1[i] can be in the range of 0 to MaxDpbSize-1. Additionally, for example, when i is greater than 0, max_dec_pic_buffering_minus1[i] may be greater than or equal to max_dec_pic_buffering_minus1[i-1]. Additionally, for example, when there is no max_dec_pic_buffering_minus1[i] for i in the range from 0 to maxSubLayersMinus1-1, the value of the syntax element max_dec_pic_buffering_minus1[i] can be inferred to be equal to max_dec_pic_buffer_minus1[maxSubLayersMinus1] because subLayerInfoFlag is 0.
[0209] Additionally, for example, the syntax element `max_num_reorder_pics[i]` can represent, for each CLVS of a CVS, the maximum allowed number of CLVS pictures that, when `Htid` equals `i`, can be in decoding order before all pictures in the CLVS and in output order after those pictures. For example, `max_num_reorder_pics[i]` can be information about the maximum picture reordering number of the DPB. The value of `max_num_reorder_pics[i]` can be in the range of 0 to `max_dec_pic_buffering_minus1[i]`. Additionally, for example, when `i` is greater than 0, `max_num_reorder_pics[i]` can be greater than or equal to `max_num_reorder_pics[i-1]`. Additionally, for example, when `max_num_reorder_pics[i]` does not exist for `i` in the range of 0 to `maxSubLayersMinus1-1`, because `subLayerInfoFlag` is 0, the syntax element `max_num_reorder_pics[i]` can be inferred to be equal to `max_num_reorder_pics[maxSubLayersMinus1]`.
[0210] In addition, for example, the value of MaxLatencyPictures[i] can be calculated using a syntax element max_latency_increase_plus1[i] whose value is not 0. MaxLatencyPictures[i] can represent, for each CLVS of the CVS, the maximum number of pictures of the CLVS that can precede all pictures of the CLVS in output order and follow these pictures in decoding order when Htid is equal to i. For example, max_latency_increase_plus1[i] can be information about the maximum latency of the DPB.
[0211] For example, when max_latency_increase_plus1[i] is not 0, the value of MaxLatencyPictures[i] can be derived as follows.
[0212] [Formula 1]
[0213] MaxLatencyPictures[i] = max_num_reorder_pics[i] + max_latency_increase_plus1[i] - 1
[0214] On the other hand, for example, if max_latency_increase_plus1[i] is 0, the corresponding limit can not be expressed. The value of max_latency_increase_plus1[i] can be in the range of 0 to 2 32 - 2. In addition, for example, when max_latency_increase_plus1[i] does not exist for i in the range of 0 to maxSubLayersMinus1 - 1, the syntax element max_latency_increase_plus1[i] can be inferred to be equal to max_latency_increase_plus1[maxSubLayersMinus1] because subLayerInfoFlag is 0.
[0215] The above-described DPB management can be performed based on information / syntax elements related to the above-described DPB parameters. Other DPB parameters can be signaled depending on whether the current layer is an output layer or a reference layer, or can be signaled depending on whether the DPB (or DPB parameters) is used for (mapped to) OLS.
[0216] Furthermore, although in the above description, the syntax element max_latency_increase_plus1[i] is used to calculate the value of MaxLatencyPictures[i], another syntax element can be used to calculate the value of MaxLatencyPictures[i]. For example, the value of MaxLatencyPictures[i] can be calculated using a syntax element max_latency_increase_minus1[i] whose value is not 0. Figure 4The decoded current picture can be inserted into the DPB, and the DPB including the decoded current picture can be updated based on the DPB parameters before a next picture of the current picture is decoded in the decoding order.
[0217] Figure 5 An exemplary illustrates a decoding process according to an embodiment of the disclosure.
[0218] The decoding device obtains picture information including information related to the DPB parameters from the bitstream (S500). The decoding device can obtain the picture information including the information related to the DPB parameters. The information / syntax elements related to the DPB parameters can be as described above.
[0219] The decoding device manages the DPB based on the DPB parameters (S510). Here, the DPB management can be referred to as DPB update. The DPB management process can include a process of marking and / or removing decoded pictures in the DPB. The decoding device can derive the DPB parameters based on the information related to the DPB parameters, and can perform the DPB management process based on the derived DPB parameters.
[0220] The decoding device decodes / outputs the current picture based on the DPB (S520). The decoding device can decode the current picture based on the updated / managed DPB. For example, blocks / slices in the current picture can be decoded based on inter prediction using (previously) decoded pictures in the DPB as reference pictures.
[0221] Figure 6 An exemplary illustrates a picture encoding method by an encoding device according to the disclosure. Figure 6 The method disclosed in the Figure 2 The encoding device illustrated in the can perform. Specifically, for example, Figure 6 S600 to S610 in the can be performed by a DPB of the encoding device, and S620 can be performed by an entropy encoder of the encoding device. In addition, although not shown, a process of decoding the current picture can be performed by a predictor and a residual processor of the encoding device.
[0222] The encoding device derives a value of a variable based on whether the current picture is a first picture of a current access unit (CVSS AU) that starts an access unit (AU) other than AU 0 (S600). The encoding device can derive the value of the variable to update the DPB before decoding the current picture and after generating / encoding a slice header of the current picture. The DPB can include pictures decoded before the current picture.
[0223] For example, the encoding device can derive a value of a variable based on whether the current picture is a first picture of a current access unit (AU) that is a CVSS AU other than the AU 0. Here, the variable can indicate whether all picture storage buffers in a decoded picture buffer (DPB) are emptied without output. The current AU can be an AU that includes the current picture. In addition, for example, the AU 0 can be a first AU of a bitstream in decoding order. That is, for example, the AU 0 can be a first AU of the bitstream to be decoded. On the other hand, the encoding device can generate / encode a slice header of the current picture, and then derive a value of the variable based on whether the current picture is a first picture of a current access unit (AU) that is a CVSS AU other than the AU 0.
[0224] For example, the encoding device can determine whether the current picture is a first picture of a current access unit (AU) that is a CVSS AU other than the AU 0. When the current AU is a CVSS AU other than the AU 0 and the current picture is a first picture of the current AU, the encoding device can derive a value of a variable.
[0225] For example, when the current AU is a CVSS AU other than the AU 0 and the current picture is a first picture of the current AU, the encoding device can determine whether at least one of parameters of the current AU is different from parameters of a previous AU of the current AU in decoding order. When the at least one of the parameters of the current AU is different from the parameters of the previous AU, the value of the variable can be set equal to 1, and when the parameters of the current AU are the same as the parameters of the previous AU, the value of the variable can be set equal to a value of a syntax element for the variable. The encoding device can generate / encode picture information of the current picture, and the picture information can include the syntax element. The syntax element can be the ph_no_output_of_prior_pics_flag described above. In addition, the parameters of the current AU include a parameter of a maximum picture width, a parameter of a maximum picture height, a parameter of available chroma formats, a parameter of a maximum bit depth, and a parameter of a maximum DPB size. The parameter of the maximum picture width, the parameter of the maximum picture height, the parameter of the available chroma formats, the parameter of the maximum bit depth, and the parameter of the maximum DPB size can be the PicWidthMaxInSamplesY, the PicHeightMaxInSamplesY, the MaxChromaFormat, the MaxBitDepthMinus8, and the max_dec_pic_buffering_minus1[Htid], respectively.
[0226] On the other hand, for example, when the current access unit (AU) is not a CVSS AU or the current picture is not a first picture of a current AU that is a CVSS AU other than the AU 0, the encoding device can not derive a value of a variable.
[0227] By so doing, the variable can be derived before decoding only the current picture that is the first picture of the current AU, instead of before decoding all pictures of the current AU that is a CVSS AU other than AU 0. The process of emptying all picture storage buffers in the decoded picture buffer (DPB) without output can be performed before decoding only the current picture that is the first picture of the current AU.
[0228] The encoding device updates the DPB based on the variable (S610). For example, the encoding device can update the DPB based on the variable.
[0229] For example, when the value of the variable is 1, all picture storage buffers in the DPB can be emptied without output, and the DPB fullness can be set equal to 0. Also, for example, when the value of the variable is 0, picture storage buffers in the DPB including a particular picture can be emptied without output, and a concave-convex process can be performed on non-empty picture storage buffers in the DPB. Also, the DPB fullness can be set to 0. Here, for example, the particular picture can be a picture that is marked as “no output needed” and “unused for reference”. The concave-convex process can be as described above.
[0230] Also, for example, when the current picture is not the first picture of the current AU that is a CVSS AU other than AU 0, the encoding device can remove from the DPB a particular picture in the DPB that satisfies a first condition and a second condition. Here, the first condition can be that the particular picture is a picture that is marked as “unused for reference”, and the second condition can be that the particular picture has a picture output flag equal to 0 or a DPB output time (DPB) of the particular picture is less than or equal to a CPB removal time of a first decoding unit (DU) of the current picture. Here, the picture output flag can be the PictureOutputFlag described above.
[0231] Also, for example, when the current picture is not the first picture of the current AU that is a CVSS AU other than AU 0, picture storage buffers in the DPB including a particular picture can be emptied without output. Here, the particular picture can be a picture that is marked as “no output needed” and “unused for reference”. Also, for the emptied picture storage buffers, the DPB fullness can be decremented by 1. That is, for example, the DPB fullness can be decremented by 1 each time a picture storage buffer is emptied. Furthermore, the concave-convex process described above can be repeatedly performed while further decrementing the DPB fullness by 1 for each additional picture storage buffer that is emptied, when one or more of the following conditions are true.
[0232] For example, the first condition can be that the number of pictures in the DPB that are marked as "output required" is greater than a syntax element max_num_reorder_pics[Htid] of the current AU. The second condition can be that a syntax element max_latency_increase_plus1[Htid] of the current AU is not equal to 0 and there is at least one picture in the DPB that is marked as "output required" and whose associated variable PicLatencyCount is greater than or equal to MaxLatencyPictures[Htid]. The third condition can be that the number of pictures in the DPB is greater than or equal to a value obtained by adding 1 to a syntax element max_dec_pic_buffering_minus1[Htid] of the current AU. The picture information can include the syntax elements of the current AU.
[0233] The encoding device encodes picture information of the current picture (S620). The encoding device can encode the picture information including syntax elements for updating the DPB. In addition, the picture information can include a slice header of the current picture.
[0234] Further, although not illustrated, the encoding device can decode the current picture based on the updated DPB. For example, the encoding device can derive prediction samples by performing inter prediction on blocks in the current picture based on reference pictures of the DPB, and can generate reconstructed samples and / or a reconstructed picture for the current picture based on the prediction samples. Further, for example, the encoding device can derive residual samples for blocks in the current picture, and can generate the reconstructed samples and / or the reconstructed picture by adding the prediction samples and the residual samples. As described above, in-loop filtering processes such as deblocking filtering, SAO, and / or ALF processes can be applied to the reconstructed samples in order to improve subjective / objective picture quality. The encoding device can generate / encode prediction-related information and / or residual information for the blocks, and the picture information can include the prediction-related information and / or the residual information. In addition, the encoding device can insert the decoded current picture into the DPB. In addition, for example, the encoding device can derive DPB parameters for the current AU, and can generate DBP-related information for the DPB parameters. The picture information can include the DBP-related information.
[0235] Further, the bitstream including the picture information can be transmitted to a decoding device through a network or a (digital) storage medium. Here, the network can include a broadcasting network and / or a communication network, and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD.
[0236] Figure 7 An encoding device performing an image encoding method according to the present disclosure is schematically illustrated. Figure 7 The method disclosed in the present disclosure can be performed byFigure 6 The encoding device disclosed in the foregoing can perform. Specifically, for example, Figure 7 The DPB of the encoding device of the foregoing can perform S600 to S610, and Figure 7 The entropy encoder of the encoding device of the foregoing can perform S620. In addition, although not shown, the process of decoding the current picture can be performed by the predictor and the residual processor of the encoding device.
[0237] Figure 8 An image decoding method performed by a decoding device according to the present disclosure is schematically illustrated. Figure 8 The method disclosed in the foregoing can be performed by Figure 3 The decoding device illustrated in the foregoing can perform. Specifically, for example, Figure 8 S800 to S810 in the foregoing can be performed by the DPB of the decoding device, and Figure 8 S820 in the foregoing can be performed by the predictor and the residual processor of the decoding device.
[0238] The decoding device derives a value of a variable based on whether the current picture is a first picture of a current access unit (AU) that is a coded video sequence start access unit (CVSS AU) other than access unit (AU) 0 (S800). The decoding device can derive the value of the variable based on whether the current picture is a first picture of a current AU that is a CVSS AU other than access unit (AU) 0. Here, the variable can indicate whether all picture storage buffers in a decoded picture buffer (DPB) are emptied without output. The current AU can be an AU that includes the current picture. In addition, for example, AU 0 can be a first AU of a bitstream in decoding order. That is, for example, AU 0 can be a first AU of a bitstream to be decoded by the decoding device. On the other hand, the decoding device can parse a slice header of the current picture, and then can derive the value of the variable based on whether the current picture is a first picture of a current AU that is a CVSS AU other than access unit (AU) 0.
[0239] For example, the decoding device can determine whether the current picture is a first picture of a current AU that is a CVSS AU other than access unit (AU) 0. When the current AU is a CVSS AU other than AU 0 and the current picture is a first picture of the current AU, the decoding device can derive the value of the variable.
[0240] For example, when the current AU is a CVSS AU other than AU 0 and the current picture is the first picture of the current AU, the decoding device can determine whether at least one of the parameters of the current AU is different from the parameters of a previous AU of the current AU in decoding order. When at least one of the parameters of the current AU is different from the parameters of the previous AU, the value of the variable can be set equal to 1. When the parameters of the current AU are the same as the parameters of the previous AU, the value of the variable can be set equal to the value of a syntax element signaled for the variable. The decoding device can obtain picture information of the current picture, and the picture information can include the syntax element. The syntax element can be the ph_no_output_of_prior_pics_flag described above. In addition, the parameters of the current AU include a parameter of a maximum picture width, a parameter of a maximum picture height, a parameter of available chroma formats, a parameter of a maximum bit depth, and a parameter of a maximum DPB size. The parameter of the maximum picture width, the parameter of the maximum picture height, the parameter of the available chroma formats, the parameter of the maximum bit depth, and the parameter of the maximum DPB size can be PicWidthMaxInSamplesY, PicHeightMaxInSamplesY, MaxChromaFormat, MaxBitDepthMinus8, and max_dec_pic_buffering_minusl[Htid], respectively.
[0241] On the other hand, for example, when the current access unit (AU) is not a CVSS AU or the current picture is not the first picture of the current AU that is a CVSS AU other than AU 0, the decoding device can not derive the value of the variable.
[0242] By doing so, the variable can be derived before decoding only the current picture that is the first picture of the current AU, instead of before decoding all pictures of the current AU that is a CVSS AU other than AU 0. The process of emptying all picture storage buffers in the decoded picture buffer (DPB) without output can be performed before decoding only the current picture that is the first picture of the current AU.
[0243] The decoding device updates the DPB based on the variable (S810). For example, the decoding device can update the DPB based on the variable. Before being updated, the DPB can include pictures decoded before the current picture.
[0244] For example, when the value of the variable is 1, all picture storage buffers in the DPB can be emptied without output, and the DPB fullness can be set equal to 0. Also, for example, when the value of the variable is 0, a picture storage buffer in the DPB including a particular picture can be emptied without output, and a concave-convex process can be performed on non-empty picture storage buffers in the DPB. Also, the DPB fullness can be set equal to 0. Here, for example, the particular picture can be a picture marked as "no output required" and "unused for reference". The concave-convex process can be as described above.
[0245] Also, for example, when the current picture is not the first picture of the current AU as a CVSS AU other than AU 0, the decoding device can remove a particular picture in the DPB that satisfies a first condition and a second condition from the DPB. Here, the first condition can be that the particular picture is a picture marked as "unused for reference". The second condition can be that the particular picture has a picture output flag equal to 0 or a DPB output time (DPB) of the particular picture is less than or equal to a CPB removal time of a first decoding unit (DU) of the current picture. Here, the picture output flag can be the PictureOutputFlag described above.
[0246] Also, for example, when the current picture is not the first picture of the current AU as a CVSS AU other than AU 0, a picture storage buffer in the DPB including a particular picture can be emptied without output. Here, the particular picture can be a picture marked as "no output required" and "unused for reference". Also, for the emptied picture storage buffer, the DPB fullness can be decremented by 1. That is, for example, the DPB fullness can be decremented by 1 each time a picture storage buffer is emptied. Furthermore, when at least one of the following conditions is true, the above-described concave-convex process can be repeatedly performed while further decrementing the DPB fullness by 1 for each additional picture storage buffer that is emptied until all conditions are not true.
[0247] For example, the first condition can be that the number of pictures in the DPB that are marked as "output required" is greater than a syntax element max num reorder pics [Htid] of the current AU. The second condition can be that a syntax element max latency increase plusl [Htid] of the current AU is not equal to 0 and there is at least one picture in the DPB that is marked as "output required" and whose associated variable PicLatencyCount is greater than or equal to MaxLatencyPictures [Htid]. The third condition can be that the number of pictures in the DPB is greater than or equal to a value obtained by adding 1 to a syntax element max dec pic buffering minusl [Htid] of the current AU. The picture information can include the syntax elements of the current AU.
[0248] The decoding device decodes the current picture based on the updated DPB (S820). For example, the decoding device can decode the current picture based on the updated DPB. For example, the decoding device can derive prediction samples by performing inter prediction on blocks in the current picture based on reference pictures of the DPB, and can generate reconstructed samples and / or a reconstructed picture for the current picture based on the prediction samples. Also, for example, the decoding device can derive residual samples for blocks in the current picture based on residual information of the current picture received through the bitstream, and can generate reconstructed samples and / or a reconstructed picture by adding the prediction samples and the residual samples. The picture information can include the residual information. In addition, the decoding device can insert the decoded current picture into the DPB.
[0249] As described above, in-loop filtering processes such as deblocking filtering, SAO, and / or ALF processes can be applied to the reconstructed samples in order to improve subjective / objective picture quality, if necessary, thereafter.
[0250] Figure 9 A decoding device that performs a picture decoding method according to the present disclosure is illustratively shown. Figure 8 The method disclosed in the Figure 9 The decoding device illustrated in the Figure 9 The DPB of the decoding device of Figure 8 S800 to S810 in the Figure 9 The predictor and the residual processor of the decoding device of Figure 8 S820 in the
[0251] According to the disclosure above, whether the process of removing pictures in the DPB without outputting them can be determined before decoding only the first picture of the CVSS AUs other than AU 0, rather than before decoding all pictures of the CVSS AUs other than AU 0. By doing so, the DPB status affecting all layers in the CVS can not be changed for each picture, and coding efficiency can be improved.
[0252] In addition, according to the disclosure, the variable indicating whether pictures in the DPB are removed without outputting them can be determined before decoding only the first picture of the CVSS AUs other than AU 0, rather than before decoding all pictures of the CVSS AUs other than AU 0. By doing so, the DPB status affecting all layers in the CVS can not be changed for each picture, and coding efficiency can be improved.
[0253] In the above embodiments, the methods are described based on flowcharts having a series of steps or blocks. The disclosure is not limited to the order of the above steps or blocks. Some steps or blocks can be performed in a different order from other steps or blocks described above or simultaneously. Furthermore, those skilled in the art will understand that the steps shown in the flowcharts are not exclusive and can further include other steps, or one or more steps of the flowcharts can be deleted without affecting the scope of the disclosure.
[0254] The embodiments described in the present specification can be executed by being implemented on a processor, a microprocessor, a controller, or a chip. For example, the functional units shown in each drawing can be executed by being implemented on a computer, a processor, a microprocessor, a controller, or a chip. In this case, information for implementation (for example, information on instructions) or an algorithm can be stored in a digital storage medium.
[0255] In addition, the decoding apparatus and the encoding apparatus applying the disclosure can be included in a multimedia broadcast transmitting / receiving apparatus, a mobile communication terminal, a home theater video apparatus, a digital theater video apparatus, a surveillance camera, a video chat apparatus, a real-time communication apparatus such as video communication, a mobile streaming apparatus, a storage medium, a camcorder, a VoD service providing apparatus, an over-the-top (OTT) video apparatus, an Internet streaming service providing apparatus, a three-dimensional (3D) video apparatus, a teleconference video apparatus, a transport user apparatus (for example, a vehicle user apparatus, an airplane user apparatus, and a ship user apparatus), and a medical video device, and the decoding apparatus and the encoding apparatus applying the disclosure can be used to process a video signal or a data signal. For example, the over-the-top (OTT) video apparatus can include a game console, a Blu-ray player, an Internet access television, a home theater system, a smart phone, a tablet, a digital video recorder (DVR), etc.
[0256] In addition, the processing method according to the present disclosure can be generated in the form of a program executed by a computer, and can be stored in a computer-readable recording medium. Multimedia data having a data structure according to the present disclosure can also be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices in which computer-readable data are stored. The computer-readable recording medium can include, for example, a BD, a universal serial bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. In addition, the computer-readable recording medium includes a medium implemented in the form of a carrier wave (for example, transmission via the Internet). In addition, a bitstream generated by the encoding method can be stored in a computer-readable recording medium or transmitted through a wired / wireless communication network.
[0257] In addition, the embodiments of the present disclosure can be implemented using a computer program product according to program codes, and the program codes can be executed in a computer by the embodiments of the present disclosure. The program codes can be stored on a computer-readable carrier.
[0258] Figure 10 A structure diagram of a content streaming system to which the present disclosure is applied is exemplified.
[0259] The content streaming system to which the embodiments of the present disclosure are applied can mainly include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.
[0260] The encoding server compresses content input from a multimedia input device such as a smartphone, a camera, or a camcorder, etc. into digital data to generate a bitstream and transmits the bitstream to the streaming server. As another example, when a multimedia input device such as a smartphone, a camera, or a camcorder, etc. directly generates a bitstream, the encoding server can be omitted.
[0261] The bitstream can be generated by applying the encoding method or the bitstream generation method according to the embodiments of the present disclosure, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.
[0262] The streaming server transmits multimedia data to a user device through a web server based on a user request, and the web server serves as a medium for informing a service to a user. When a user requests a desired service from the web server, the web server delivers the request to the streaming server, and the streaming server transmits multimedia data to the user. In this case, the content streaming system can include a separate control server. In this case, the control server is used to control commands / responses between devices within the content streaming system.
[0263] The streaming server can receive the content from the media storage and / or the encoding server. For example, when receiving the content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a predetermined period of time.
[0264] Examples of the user device can include a mobile phone, a smart phone, a laptop computer, a digital broadcasting terminal, a personal digital assistant (PDA), a portable multimedia player (PMP), a navigation, a tablet PC, a slate PC, an ultrabook, a wearable device (e.g., a smart watch, smart glasses, and a head-mounted display), a digital TV, a desktop computer, a digital signage, and the like. Each server within the content streaming system can operate as a distributed server, in which case the data received from each server can be distributed.
[0265] The technical features of the method claims of the present disclosure can be combined in various ways. For example, the technical features of the method claims of the present disclosure can be combined to be implemented as an apparatus, and the technical features of the apparatus claims of the present disclosure can be combined to be implemented as a method. Furthermore, the technical features of the method claims of the present disclosure and the technical features of the apparatus claims of the present disclosure can be combined to be implemented as an apparatus, and the technical features of the method claims of the present disclosure and the technical features of the apparatus claims of the present disclosure can be combined to be implemented as a method.
Claims
1. An image decoding method performed by a decoding device, the image decoding method comprising the steps of: deriving a value of a variable based on whether a current picture is a first picture of a current access unit (AU) other than an AU 0, i.e., a coded video sequence start (CVSS) AU, the variable indicating whether all picture storage buffers in a decoded picture buffer (DPB) are emptied without output; updating the DPB based on the variable; and decoding the current picture based on the updated DPB, wherein the AU 0 is a first AU in a bitstream.
2. The image decoding method of claim 1, wherein, When the value of the variable is 1, all picture storage buffers in the DPB are emptied without output, and a DPB fullness is set equal to 0.
3. The image decoding method of claim 2, wherein, When the value of the variable is 0, picture storage buffers in the DPB including a particular picture are emptied without output, and a hole-filling process is performed on non-empty picture storage buffers in the DPB, and wherein the particular picture is a picture that is marked as not needed for output and not used for reference.
4. The image decoding method of claim 1, wherein, Deriving the value of the variable comprises the steps of: when the current picture is the first picture of the current AU that is the CVSS AU other than the AU 0, determining whether at least one of parameters for the current AU is different in decoding order from parameters for a previous AU of the current AU, wherein when the at least one of parameters for the current AU is different from parameters for the previous AU, the value of the variable is set equal to 1, and wherein when parameters for the current AU are the same as parameters for the previous AU, the value of the variable is set equal to a value of a syntax element signaled for the variable.
5. The image decoding method of claim 4, wherein, The parameters for the current AU include a parameter for maximum picture width, a parameter for maximum picture height, a parameter for available chroma format, a parameter for maximum bit depth, and a parameter for maximum DPB size for the current AU.
6. The image decoding method of claim 1, wherein, Updating the DPB comprises the steps of: when the current picture is not the first picture of the current AU that is the CVSS AU other than the AU 0, removing a particular picture from the DPB that satisfies a first condition and a second condition, wherein the first condition is that the particular picture is a picture that is marked as not used for reference, wherein the second condition is that the particular picture has a picture output flag equal to 0 or a DPB output time of the particular picture is less than or equal to a CPB removal time of a first decoding unit (DU) of the current picture.
7. The image decoding method of claim 1, wherein, when the current picture is not the first picture of the current AU that is the CVSS AU other than the AU 0, picture storage buffers in the DPB including a particular picture are emptied without output, and wherein the particular picture is a picture that is marked as not needed for output and not used for reference.
8. An image encoding method performed by an encoding device, the image encoding method comprising the steps of: deriving a value of a variable based on whether the current picture is a first picture of a current access unit, AU, other than an AU0, i.e. a CVSS AU, the variable indicating whether all picture storage buffers in a decoded picture buffer, DPB, are emptied without output; updating the DPB based on the variable; and encoding picture information for the current picture, wherein the AU0 is a first AU in a bitstream.
9. The image coding method of claim 8, wherein, When the value of the variable is 1, all picture storage buffers in the DPB are emptied without output and a DPB fullness is set equal to 0.
10. The image coding method of claim 9, wherein, When the value of the variable is 0, picture storage buffers in the DPB including a particular picture are emptied without output and a hole filling process is performed on non-empty picture storage buffers in the DPB, and wherein the particular picture is a picture marked as not needed for output and not used for reference.
11. The image coding method of claim 8, wherein, The deriving the value of the variable comprises the steps of: when the current picture is the first picture of the current AU other than the AU0, i.e. a CVSS AU, determining whether at least one of parameters for the current AU is different in decoding order from parameters for a previous AU of the current AU, wherein the value of the variable is set equal to 1 when the at least one of the parameters for the current AU is different from the parameters for the previous AU, wherein the value of the variable is set equal to a value of a syntax element for the variable when parameters for the current AU are the same as parameters for the previous AU, and wherein the picture information comprises the syntax element.
12. The image coding method of claim 11, wherein, The parameters for the current AU comprise a parameter for maximum picture width, a parameter for maximum picture height, a parameter for available chroma format, a parameter for maximum bit depth, and a parameter for maximum DPB size for the current AU.
13. The image coding method of claim 8, wherein, The updating the DPB comprises the steps of: when the current picture is not the first picture of the current AU other than the AU0, i.e. a CVSS AU, removing a particular picture from the DPB that satisfies a first condition and a second condition, wherein the first condition is that the particular picture is a picture marked as not used for reference, wherein the second condition is that the particular picture has a picture output flag equal to 0 or a DPB output time of the particular picture is less than or equal to a CPB removal time of a first decoding unit, DU, of the current picture.
14. The image coding method of claim 8, wherein, When the current picture is not the first picture of the current AU other than the AU0, i.e. a CVSS AU, picture storage buffers in the DPB including a particular picture are emptied without output, and wherein the particular picture is a picture marked as not needed for output and not used for reference.
15. A method for transmitting data for picture information, the method comprising the steps of: deriving a value of a variable based on whether the current picture is a first picture of a current access unit, AU, other than an AU 0, which is a coded video sequence start AU, i.e., a CVSS AU, the variable indicating whether all picture storage buffers in a decoded picture buffer, DPB, are emptied without output, the AU 0 being a first AU in a bitstream; updating the DPB based on the variable; encoding picture information for the current picture to generate the bitstream; and transmitting the data including the bitstream.
Citation Information
Patent Citations
Video encoding and decoding
CN106416250A
Multi-view video codec supporting residual prediction
CN107439015A