Image encoding / decoding method and device for signaling HRD parameters, and computer-readable recording medium in which bitstream is stored
The image encoding/decoding method and apparatus address the challenge of efficiently compressing high-resolution images by effectively signaling HRD parameters, resulting in improved encoding/decoding efficiency and reduced costs.
Patent Information
- Application Number
- JP2025064903
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-03-30
- Filing Date
- 2025-04-10
- Publication Date
- 2025-06-26
AI Technical Summary
The increasing demand for high-resolution, high-quality images has led to a need for highly efficient image compression techniques to reduce transmission and storage costs.
An image encoding/decoding method and apparatus that improves encoding/decoding efficiency by efficiently signaling HRD (Hypothetical Reference Decoder) parameters, including obtaining and processing HRD parameter syntax structures and mapping them to multi-layer OLSs.
The proposed method and apparatus achieve improved encoding/decoding efficiency, enabling efficient transmission and storage of high-resolution images while reducing costs.
Smart Images

Figure 2025096503000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an image encoding / decoding method and apparatus, and more particularly, to an image encoding / decoding method and apparatus for signaling HRD (Hypothetical reference decoder) related parameters, and a computer-readable recording medium storing a bitstream generated by the image encoding method / apparatus of the present disclosure, and the like.
Background Art
[0002] Recently, the demand for high-resolution, high-quality images, such as HD (High Definition) images and UHD (Ultra High Definition) images, has been increasing in various fields. As the image data becomes higher in resolution and quality, the amount of information or bits to be transmitted relatively increases compared to conventional image data. The increase in the amount of information or bits to be transmitted leads to an increase in transmission costs and storage costs.
[0003] Accordingly, there is a need for a highly efficient image compression technique for effectively transmitting, storing, and reproducing information of high-resolution, high-quality images.
Summary of the Invention
Problems to be Solved by the Invention
[0004] An object of the present disclosure is to provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.
[0005] Another object of the present disclosure is to provide an image encoding / decoding method and apparatus that improves encoding / decoding efficiency by efficiently signaling HRD parameters.
[0006] Another object of the present disclosure is to provide a method for transmitting a bitstream generated by the image encoding method or apparatus according to the present disclosure.
[0007] Another object of the present disclosure is to provide a recording medium storing a bitstream generated by an image encoding method or apparatus according to the present disclosure.
[0008] Another object of the present disclosure is to provide a recording medium storing a bitstream received by an image decoding apparatus according to the present disclosure, decoded, and used for restoring an image.
[0009] The technical problems to be solved by the present disclosure are not limited to the above-described technical problems, and other technical problems not described above will be clearly understood by those of ordinary skill in the technical field to which the present disclosure pertains from the following description.
Means for Solving the Problems
[0010] An image decoding method performed by an image decoding apparatus according to an aspect of the present disclosure may include: obtaining first information indicating the number of one or more HRD (hypothetical reference decoder) parameter syntax structures in a VPS (Video Parameter Set); obtaining the one or more HRD parameter syntax structures from the VPS based on the first information; obtaining second information regarding mapping between one or more multi-layer OLS (output layer set) and the one or more HRD parameter syntax structures from the VPS based on the first information; selecting an HRD parameter syntax structure applied to a current OLS based on the second information; and processing the current OLS based on the selected HRD parameter syntax structure.
[0011] In the image decoding method of the present disclosure, the number of the one or more HRD parameter syntax structures in the VPS may not be greater than the number of the one or more multi-layer OLSs.
[0012] In the image decoding method of the present disclosure, each of the one or more HRD parameter syntax structures in the VPS can be mapped to at least one of the one or more multi-layer OLSs.
[0013] In the image decoding method of the present disclosure, based on the fact that the number of the one or more HRD parameter syntax structures in the VPS is greater than 1 and the number of the one or more HRD parameter syntax structures in the VPS is not equal to the number of the one or more multi-layer OLSs, the second information can be obtained from the VPS.
[0014] In the image decoding method of the present disclosure, based on the fact that the number of the one or more HRD parameter syntax structures in the VPS is 1, the second information is not obtained from the VPS, and the second information can be inferred to have a value of 0.
[0015] In the image decoding method of the present disclosure, based on the fact that the number of the one or more HRD parameter syntax structures in the VPS is greater than 1 and the number of the one or more HRD parameter syntax structures in the VPS is equal to the number of the one or more multi-layer OLSs, the second information is not obtained from the VPS, and the second information for the i-th multi-layer OLS can be inferred to have a value of i.
[0016] In the image decoding method of the present disclosure, based on the fact that the current OLS includes only a single layer, the HRD parameter syntax structure applied to the current OLS can be obtained from the SPS (Sequence Parameter Set).
[0017] An image decoding apparatus according to another aspect of the present disclosure includes a memory and at least one processor. The at least one processor obtains first information indicating the number of one or more HRD (hypothetical reference decoder) parameter syntax structures in a VPS (Video Parameter Set), obtains the one or more HRD parameter syntax structures from the VPS based on the first information, obtains second information regarding the mapping between one or more multi-layer OLS (output layer set) and the one or more HRD parameter syntax structures from the VPS based on the first information, selects an HRD parameter syntax structure to be applied to the current OLS based on the second information, and can process the current OLS based on the selected HRD parameter syntax structure.
[0018] An image encoding method performed by an image encoding apparatus according to another aspect of the present disclosure may include encoding first information indicating the number of one or more HRD (hypothetical reference decoder) parameter syntax structures in a VPS (Video Parameter Set), encoding the one or more HRD parameter syntax structures in the VPS based on the first information, encoding second information regarding the mapping between one or more multi-layer OLS (output layer set) and the one or more HRD parameter syntax structures in the VPS based on the first information, and processing the current OLS based on the HRD parameter syntax structure applied to the current OLS.
[0019] In the image encoding method of the present disclosure, the number of the one or more HRD parameter syntax structures in the VPS may not be greater than the number of the one or more multi-layer OLSs.
[0020] In the image encoding method of the present disclosure, each of the one or more HRD parameter syntax structures in the VPS can be mapped to at least one of the one or more multi-layer OLSs.
[0021] In the image encoding method of the present disclosure, based on the fact that the number of the one or more HRD parameter syntax structures in the VPS is greater than 1 and the number of the one or more HRD parameter syntax structures in the VPS is not equal to the number of the one or more multi-layer OLSs, the second information can be encoded in the VPS.
[0022] In the image encoding method of the present disclosure, based on the fact that the number of the one or more HRD parameter syntax structures in the VPS is 1, the second information is not encoded in the VPS, and the second information can be inferred to have a value of 0.
[0023] In the image encoding method of the present disclosure, based on the fact that the number of the one or more HRD parameter syntax structures in the VPS is greater than 1 and the number of the one or more HRD parameter syntax structures in the VPS is equal to the number of the one or more multi-layer OLSs, the second information is not encoded in the VPS, and the second information for the i-th multi-layer OLS can be inferred to have a value of i.
[0024] In the image encoding method of the present disclosure, based on the fact that the current OLS includes only a single layer, the HRD parameter syntax structure applied to the current OLS can be encoded in the SPS (Sequence Parameter Set).
[0025] A transmission method according to another aspect of the present disclosure can transmit a bitstream generated by the image encoding apparatus or the image encoding method of the present disclosure.
[0026] A computer-readable recording medium according to another aspect of the present disclosure can store a bitstream generated by the image encoding method or image encoding apparatus of the present disclosure.
[0027] The features briefly summarized and described above with respect to the present disclosure are merely exemplary aspects of the detailed description of the present disclosure to follow, and do not limit the scope of the present disclosure.
Advantages of the Invention
[0028] According to the present disclosure, it is possible to provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.
[0029] Also, according to the present disclosure, it is possible to provide an image encoding / decoding method and apparatus capable of improving the encoding / decoding efficiency by efficiently signaling HRD parameters.
[0030] Also, according to the present disclosure, it is possible to provide a method of transmitting a bitstream generated by the image encoding method or apparatus according to the present disclosure.
[0031] Also, according to the present disclosure, it is possible to provide a recording medium storing a bitstream generated by the image encoding method or apparatus according to the present disclosure.
[0032] Also, according to the present disclosure, it is possible to provide a recording medium storing a bitstream received by the image decoding apparatus according to the present disclosure, decoded, and used for restoring an image.
[0033] The effects obtained in the present disclosure are not limited to the effects described above, and other effects not described above will be clearly understood by those of ordinary skill in the technical field to which the present disclosure pertains from the following description.
Brief Description of the Drawings
[0034]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Embodiments for Carrying Out the Invention
[0035] Hereinafter, with reference to the accompanying drawings, embodiments of the present disclosure will be described in detail so that those having ordinary knowledge in the technical field to which the present disclosure pertains can easily implement them. However, the present disclosure can be realized in various different forms and is not limited to the embodiments described herein.
[0036] In describing the embodiments of the present disclosure, if it is determined that a specific description of a known configuration or function may obscure the gist of the present disclosure, the detailed description thereof will be omitted. And in the drawings, parts not related to the description of the present disclosure are omitted, and the same reference numerals are given to the same parts.
[0037] In the present disclosure, when a certain component is "connected", "coupled", or "connected" to another component, this can include not only a direct connection relationship but also an indirect connection relationship in which another component exists between them. Also, when a certain component "includes" or "has" another component, this means that, unless otherwise stated to the contrary, it does not exclude another component but can further include another component.
[0038] In the present disclosure, terms such as "first" and "second" are used only for the purpose of distinguishing one component from another and do not limit the order or importance, etc. between components unless otherwise specified. Therefore, within the scope of the present disclosure, the first component of one embodiment may be referred to as the second component in another embodiment, and similarly, the second component of one embodiment may be referred to as the first component in another embodiment.
[0039] In the present disclosure, the components distinguished from each other are for clearly explaining their respective features, and do not necessarily mean that the components are separated. That is, a plurality of components may be integrated and configured as one hardware or software unit, or one component may be distributed and configured as a plurality of hardware or software units. Therefore, even without separate mention, such integrated or distributed embodiments are also included in the scope of the present disclosure.
[0040] In the present disclosure, the components described in various embodiments do not necessarily mean essential components, and some may be optional components. Therefore, embodiments constituted by a subset of the components described in one embodiment are also included in the scope of the present disclosure. Further, embodiments that include additional components in addition to the components described in various embodiments are also included in the scope of the present disclosure.
[0041] The present disclosure relates to image encoding and decoding, and the terms used in the present disclosure can have their ordinary meanings in the technical field to which the present disclosure belongs, unless newly defined in the present disclosure.
[0042] In the present disclosure, "picture" generally means a unit indicating any one image in a specific time period, and a slice / tile is an encoding unit constituting a part of a picture, and one picture can be composed of one or more slices / tiles. Further, a slice / tile can include one or more CTUs (coding tree units).
[0043] In the present disclosure, "pixel" or "pel" can mean the smallest unit that constitutes a picture (or image). Also, the term "sample" can be used as a term corresponding to a pixel. A sample can generally indicate a pixel or a pixel value, and can also indicate only the pixel / pixel value of the luma component, or only the pixel / pixel value of the chroma component.
[0044] In the present disclosure, "unit" can indicate the basic unit of image processing. A unit can include at least one of a specific region of a picture and information related to the region. A unit can, in some cases, be used interchangeably with terms such as "sample array", "block", or "area". In general, an M×N block can include a set (or array) of samples (or sample arrays) or transform coefficients consisting of M columns and N rows.
[0045] In the present disclosure, "current block" can mean any one of "current coding block", "current coding unit", "block to be coded", "block to be decoded", or "block to be processed". When prediction is performed, "current block" can mean "current prediction block" or "block to be predicted". When transformation (inverse transformation) / quantization (inverse quantization) is performed, "current block" can mean "current transformation block" or "block to be transformed". When filtering is performed, "current block" can mean "block to be filtered".
[0046] In the present disclosure, unless explicitly stated as a chroma block, the "current block" can mean a block that includes all luma component blocks and chroma component blocks, or the "luma block of the current block". The luma component block of the current block can be explicitly expressed as including an explicit description of the luma component block, such as "luma block" or "current luma block". Also, the chroma component block of the current block can be explicitly expressed as including an explicit description of the chroma component block, such as "chroma block" or "current chroma block".
[0047] In the present disclosure, "A or B" can mean "only A", "only B", or "both A and B". In other words, in the present disclosure, "A or B" can be interpreted as "A and / or B". For example, in the present disclosure, "A, B or C" can mean "only A", "only B", "only C", or "any combination of A, B and C".
[0048] In the present disclosure, " / " and "," can mean "and / or". For example, "A / B" can mean "A and / or B". Thus, "A / B" can mean "only A", "only B", or "both A and B". For example, "A, B, C" can mean "A, B or C".
[0049] In the present disclosure, "at least one of A and B" can mean "only A", "only B", or "both A and B". Also, in the present disclosure, expressions such as "at least one of A or B" and "at least one of A and / or B" can be interpreted identically to "at least one of A and B".
[0050] Also, in the present disclosure, "at least one of A, B, and C" can mean "only A", "only B", "only C", or "any combination of A, B, and C". Also, "at least one of A, B, or C" and "at least one of A, B, and / or C" can mean "at least one of A, B, and C".
[0051] Also, the parentheses used in the present disclosure can mean "for example". Specifically, when it is indicated as "prediction (intra-prediction)", "intra-prediction" can be proposed as an example of "prediction". In other words, "prediction" in the present disclosure is not limited to "intra-prediction", and "intra-prediction" can be proposed as an example of "prediction". Also, when it is indicated as "prediction (i.e., intra-prediction)", "intra-prediction" can be proposed as an example of "prediction".
[0052] In the present disclosure, the technical features separately described within one drawing may be realized separately or simultaneously.
[0053] Overview of the video coding system
[0054] Figure 1 shows a video coding system according to the present disclosure.
[0055] A video coding system according to an embodiment may include an encoding device 10 and a decoding device 20. The encoding device 10 can transmit encoded video and / or image information or data to the decoding device 20 in a file or streaming format via a digital storage medium or a network.
[0056] An encoding device 10 according to an embodiment may include a video source generation unit 11, an encoding unit 12, and a transmission unit 13. A decoding device 20 according to an embodiment may include a reception unit 21, a decoding unit 22, and a rendering unit 23. The encoding unit 12 may be referred to as a video / image encoding unit, and the decoding unit 22 may be referred to as a video / image decoding unit. The transmission unit 13 may be included in the encoding unit 12. The reception unit 21 may be included in the decoding unit 22. The rendering unit 23 may also include a display unit, and the display unit may be configured as a separate device or an external component.
[0057] The video source generation unit 11 can obtain video / images through processes such as capture, synthesis, or generation of video / images. The video source generation unit 11 may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. The video / image generation device may include, for example, a computer, a tablet, and a smartphone, etc., and can (electronically) generate video / images. For example, virtual video / images can be generated via a computer or the like. In this case, the video / image capture process can be replaced by a process in which related data is generated.
[0058] The symbolization unit 12 can encode the input video / image. The symbolization unit 12 can perform a series of procedures such as prediction, transformation, quantization, etc. for compression and symbolization efficiency. The symbolization unit 12 can output the encoded data (encoded video / image information) in the form of a bitstream.
[0059] The transmission unit 13 can transmit the encoded video / image information or data output in the form of a bitstream to the receiving unit 21 of the decoding device 20 via a digital storage medium or a network in the form of a file or a streaming. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray (registered trademark), HDD, SSD, etc. The transmission unit 13 can include elements for generating a media file via a predetermined file format and can include elements for transmission via a broadcast / communication network. The receiving unit 21 can extract / receive the bitstream from the storage medium or the network and transmit it to the decoding unit 22.
[0060] The decoding unit 22 can perform a series of procedures such as inverse quantization, inverse transformation, prediction, etc. corresponding to the operation of the symbolization unit 12 to decode the video / image.
[0061] The rendering unit 23 can render the decoded video / image. The rendered video / image can be displayed via the display unit.
[0062] Overview of the image encoding device
[0063] FIG. 2 is a diagram schematically showing an image encoding device to which an embodiment according to the present disclosure can be applied.
[0064] As shown in FIG. 2, the image encoding apparatus 100 can include an image dividing unit 110, a subtraction unit 115, a conversion unit 120, a quantization unit 130, an inverse quantization unit 140, an inverse conversion unit 150, an addition unit 155, a filtering unit 160, a memory 170, an inter prediction unit 180, an intra prediction unit 185, and an entropy encoding unit 190. The inter prediction unit 180 and the intra prediction unit 185 can be collectively referred to as a “prediction unit”. The conversion unit 120, the quantization unit 130, the inverse quantization unit 140, and the inverse conversion unit 150 can be included in a residual processing unit. The residual processing unit can further include the subtraction unit 115.
[0065] All or at least a part of the plurality of components constituting the image encoding apparatus 100 can be realized by one hardware component (for example, an encoder or a processor) according to an embodiment. Further, the memory 170 can include a DPB (decoded picture buffer) and can be realized by a digital storage medium.
[0066] The image segmentation unit 110 can divide an input image (or picture, frame) input to the image encoding apparatus 100 into one or more processing units. As an example, the processing unit can be called a coding unit (CU). The coding unit can be obtained by recursively dividing a coding tree unit (CTU) or a largest coding unit (LCU) in a QT / BT / TT (Quad-tree / Binary-tree / Ternary-tree) structure. For example, one coding unit can be divided into a plurality of coding units with a deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. For the division of the coding unit, the quadtree structure can be applied first, and the binary tree structure and / or the ternary tree structure can be applied later. Based on the final coding unit that cannot be further divided, the coding procedure according to the present disclosure can be performed. The largest coding unit can be used as the final coding unit, and the coding units with a deeper depth obtained by dividing the largest coding unit can also be used as the final coding unit. Here, the coding procedure can include procedures such as prediction, conversion, and / or restoration, which will be described later. As another example, the processing unit of the coding procedure can be a prediction unit (PU: Prediction Unit) or a transform unit (TU: Transform Unit). The prediction unit and the transform unit can be divided or partitioned from the final coding unit, respectively. The prediction unit can be a unit of sample prediction, and the transform unit can be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.
[0067] The prediction unit (inter prediction unit 180 or intra prediction unit 185) can perform a prediction on a processing target block (current block) and generate a predicted block including prediction samples for the current block. The prediction unit can determine whether intra prediction is applied in units of the current block or CU, or whether inter prediction is applied. The prediction unit can generate various information regarding the prediction of the current block and transmit it to the entropy encoding unit 190. The information regarding the prediction can be encoded by the entropy encoding unit 190 and output in the form of a bitstream.
[0068] The intra prediction unit 185 can predict the current block by referring to samples within the current picture. The samples to be referred to can be located in the neighborhood of the current block or at a distance according to the intra prediction mode and / or intra prediction technique. The intra prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes can include, for example, the DC mode and the Planar mode. The directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes according to the degree of fineness of the prediction direction. However, this is only an example, and more or fewer directional prediction modes can be used based on the settings. The intra prediction unit 185 can also determine the prediction mode to be applied to the current block using the prediction mode applied to the neighboring blocks.
[0069] The inter prediction unit 180 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on the reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the peripheral block and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the peripheral blocks can include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different from each other. The temporal neighboring block can be called by names such as a collocated reference block, a collocated CU (colCU), etc. The reference picture including the temporal neighboring block can be called a collocated picture (colPic). For example, the inter prediction unit 180 can construct a motion information candidate list based on the peripheral blocks, and generate information indicating which candidate is used to derive the motion vector and / or the reference picture index of the current block. Inter prediction can be performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the inter prediction unit 180 can use the motion information of the peripheral blocks as the motion information of the current block. In the case of the skip mode, unlike the merge mode, the residual signal cannot be transmitted.In the case of the motion information prediction (MVP) mode, the motion vectors of neighboring blocks are used as motion vector predictors, and the motion vector difference and the indicator for the motion vector predictor are encoded to signal the motion vector of the current block. The motion vector difference can mean the difference between the motion vector of the current block and the motion vector predictor.
[0070] The prediction unit can generate a prediction signal based on various prediction methods and / or prediction techniques described below. For example, the prediction unit can apply intra prediction or inter prediction for predicting the current block, and can also apply intra prediction and inter prediction simultaneously. A prediction method that applies intra prediction and inter prediction simultaneously for predicting the current block can be called CIIP (combined inter and intra prediction). In addition, the prediction unit can also perform intra block copy (IBC) for predicting the current block. Intra block copy can be used for content image / video coding such as games, for example, like SCC (screen content coding). IBC is a method of predicting the current block using a restored reference block within the current picture at a position a predetermined distance away from the current block. When IBC is applied, the position of the reference block within the current picture can be encoded as a vector (block vector) corresponding to the predetermined distance. IBC basically performs prediction within the current picture, but can be performed in the same way as inter prediction in terms of deriving a reference block within the current picture. That is, IBC can use at least one of the inter prediction techniques described in the present disclosure.
[0071] The prediction signal generated by the prediction unit can be used to generate a restored signal or can be used to generate a residual signal. The subtraction unit 115 can subtract the prediction signal (predicted block, predicted sample array) output from the prediction unit from the input image signal (original block, original sample array) to generate a residual signal (residual signal, residual block, residual sample array). The generated residual signal can be transmitted to the conversion unit 120.
[0072] The conversion unit 120 can apply a conversion technique to the residual signal to generate transform coefficients. For example, the conversion technique can include at least one of DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), KLT (Karhunen-Loeve Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT means a transform obtained from a graph when representing the relationship information between pixels as a graph. CNT means a transform obtained based on generating a prediction signal using all previously reconstructed pixels. The conversion process can also be applied to a pixel block having the same size of a square, and can also be applied to a non-square, variable-size block.
[0073] The quantization unit 130 can quantize the transform coefficients and transmit them to the entropy encoding unit 190. The entropy encoding unit 190 can encode the quantized signal (information regarding the quantized transform coefficients) and output it in the form of a bit stream. The information regarding the quantized transform coefficients can be called residual information. The quantization unit 130 can reorder the quantized transform coefficients in block form into a one-dimensional vector form based on the coefficient scan order, and can also generate the information regarding the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form.
[0074] The entropy encoding unit 190 can perform various encoding methods such as, for example, exponential Golomb, CAVLC (context-adaptive variable length coding), CABAC (context-adaptive binary arithmetic coding), etc. The entropy encoding unit 190 can also encode, together or separately, information necessary for video / image restoration (such as the values of syntax elements, etc.) in addition to the quantized transform coefficients. The encoded information (such as the encoded video / image information) can be transmitted or stored in the form of a bit stream in units of NAL (network abstraction layer) units. The video / image information can further include information regarding various parameter sets such as an adaptive parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Also, the video / image information can further include general constraint information. The signaling information, the transmitted information, and / or the syntax elements mentioned in the present disclosure can be encoded through the above-described encoding procedure and included in the bit stream.
[0075] The bitstream can be transmitted via a network or stored in a digital storage medium. Here, the network can include a broadcast network and / or a communication network, etc., and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmission unit (not shown) for transmitting and / or a storage unit (not shown) for storing the signal output from the entropy encoding unit 190 can be provided as internal / external elements of the image encoding apparatus 100, or the transmission unit can also be provided as a component of the entropy encoding unit 190.
[0076] The quantized transform coefficients output from the quantization unit 130 can be used to generate a residual signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients via the inverse quantization unit 140 and the inverse transformation unit 150, a residual signal (residual block or residual sample) can be restored.
[0077] The addition unit 155 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the restored residual signal to the prediction signal output from the inter prediction unit 180 or the intra prediction unit 185. When there is no residual for the block to be processed as in the case where the skip mode is applied, the predicted block can be used as the reconstructed block. The addition unit 155 can be called a restoration unit or a reconstructed block generation unit. The generated reconstructed signal can be used for intra prediction of the next block to be processed within the current picture, and can also be used for inter prediction of the next picture after passing through filtering as described later.
[0078] The filtering unit 160 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 160 can apply various filtering methods to the restored picture to generate a modified restored picture, and can save the modified restored picture in the memory 170, specifically in the DPB of the memory 170. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit 160 can generate various information related to filtering as described later in the description of each filtering method and transmit it to the entropy encoding unit 190. The information related to filtering can be encoded by the entropy encoding unit 190 and output in the form of a bitstream.
[0079] The modified restored picture transmitted to the memory 170 can be used as a reference picture by the inter prediction unit 180. When inter prediction is applied through this, the image encoding apparatus 100 can avoid prediction mismatches between the image encoding apparatus 100 and the image decoding apparatus, and can also improve the encoding efficiency.
[0080] The DPB in the memory 170 can save the modified restored picture for use as a reference picture by the inter prediction unit 180. The memory 170 can save the motion information of the block where the motion information in the current picture was derived (or encoded) and / or the motion information of the block in the already restored picture. The saved motion information can be transmitted to the inter prediction unit 180 for utilization as the motion information of the spatial neighboring blocks or the motion information of the temporal neighboring blocks. The memory 170 can save the restored samples of the restored blocks in the current picture and transmit them to the intra prediction unit 185.
[0081] Overview of the image decoding device
[0082] FIG. 3 is a diagram schematically showing an image decoding apparatus to which an embodiment according to the present disclosure can be applied.
[0083] As shown in FIG. 3, the image decoding apparatus 200 can be configured to include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an addition unit 235, a filtering unit 240, a memory 250, an inter prediction unit 260, and an intra prediction unit 265. The inter prediction unit 260 and the intra prediction unit 265 can be collectively referred to as a "prediction unit". The inverse quantization unit 220 and the inverse transform unit 230 can be included in a residual processing unit.
[0084] All or at least a part of the plurality of components constituting the image decoding apparatus 200 can be realized by one hardware component (for example, a decoder or a processor) according to an embodiment. Further, the memory 170 can include a DPB and can be realized by a digital storage medium.
[0085] The image decoding apparatus 200 that has received a bitstream including video / image information can execute a process corresponding to the process performed by the image encoding apparatus 100 in FIG. 2 to restore an image. For example, the image decoding apparatus 200 can perform decoding using the processing unit applied in the image encoding apparatus. Therefore, the decoding processing unit can be, for example, a coding unit. The coding unit can be obtained by dividing a coding tree unit or a maximum coding unit. Then, the restored image signal decoded and output via the image decoding apparatus 200 can be played back via a playback apparatus (not shown).
[0086] The image decoding apparatus 200 can receive the signal output from the image encoding apparatus of FIG. 2 in the form of a bit stream. The received signal can be decoded via the entropy decoding unit 210. For example, the entropy decoding unit 210 can parse the bit stream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information can further include information regarding various parameter sets such as an Adaptive Parameter Set (APS), a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), or a Video Parameter Set (VPS). Also, the video / image information can further include general constraint information. The image decoding apparatus can further use the information regarding the parameter set and / or the general constraint information for decoding an image. The signaling information, the received information, and / or the syntax elements referred to in the present disclosure can be obtained from the bit stream by being decoded via the decoding procedure. For example, the entropy decoding unit 210 can decode the information in the bit stream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the value of the syntax element necessary for image restoration, the quantized value of the transform coefficient regarding the residual. More specifically, the CABAC entropy decoding method receives a bin corresponding to each syntax element from the bit stream, determines a context model using the syntax element information to be decoded, the decoding information of the surrounding blocks and the block to be decoded, or the information of the symbol / bin decoded in the previous step, predicts the occurrence probability of the bin based on the determined context model, and performs arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element. At this time, the CABAC entropy decoding method can update the context model using the information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model.Of the information decoded by the entropy decoding unit 210, the information related to prediction is provided to the prediction units (inter prediction unit 260 and intra prediction unit 265), and the residual values that have undergone entropy decoding in the entropy decoding unit 210, that is, the quantized transform coefficients and related parameter information, can be input to the inverse quantization unit 220. Also, of the information decoded by the entropy decoding unit 210, the information related to filtering can be provided to the filtering unit 240. On the other hand, a receiving unit (not shown) that receives the signal output from the image encoding device can be further provided as an internal / external element of the image decoding device 200, or the receiving unit can be provided as a component of the entropy decoding unit 210.
[0087] On the other hand, the image decoding device according to the present disclosure can be called a video / image / picture decoding device. The image decoding device can also include an information decoder (video / image / picture information decoder) and / or a sample decoder (video / image / picture sample decoder). The information decoder can include the entropy decoding unit 210, and the sample decoder can include at least one of the inverse quantization unit 220, the inverse transform unit 230, the addition unit 235, the filtering unit 240, the memory 250, the inter prediction unit 260, and the intra prediction unit 265.
[0088] In the inverse quantization unit 220, the quantized transform coefficients can be inverse quantized to output transform coefficients. The inverse quantization unit 220 can reorder the quantized transform coefficients in a two-dimensional block format. In this case, the reordering can be performed based on the coefficient scan order performed in the image encoding device. The inverse quantization unit 220 can perform inverse quantization on the quantized transform coefficients using a quantization parameter (for example, quantization step size information) to obtain transform coefficients.
[0089] In the inverse conversion unit 230, the conversion coefficients can be inversely converted to obtain a residual signal (residual block, residual sample array).
[0090] The prediction unit can perform prediction on the current block and generate a predicted block including prediction samples for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied to the current block based on the information regarding the prediction output from the entropy decoding unit 210, and can determine a specific intra / inter prediction mode (prediction technique).
[0091] The prediction unit can generate a prediction signal based on various prediction methods (techniques) described later, which is the same as described in the explanation of the prediction unit of the image encoding apparatus 100.
[0092] The intra prediction unit 265 can predict the current block by referring to samples within the current picture. The explanation of the intra prediction unit 185 can be similarly applied to the intra prediction unit 265.
[0093] The inter prediction unit 260 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the neighboring blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit 260 can construct a motion information candidate list based on the neighboring blocks, and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction can be performed based on various prediction modes (techniques), and the information regarding the prediction can include information indicating the mode (technique) of the inter prediction for the current block.
[0094] The addition unit 235 can generate a restored signal (restored picture, restored block, restored sample array) by adding the obtained residual signal to the prediction signal (predicted block, predicted sample array) output from a prediction unit (including the inter prediction unit 260 and / or the intra prediction unit 265). When there is no residual for the processing target block as in the case where the skip mode is applied, the predicted block can be used as the restored block. The description of the addition unit 155 can be similarly applied to the addition unit 235. The addition unit 235 may also be referred to as a restoration unit or a restored block generation unit. The generated restored signal can be used for intra prediction of the next processing target block in the current picture, and can also be used for inter prediction of the next picture through filtering as described later.
[0095] The filtering unit 240 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 240 can apply various filtering methods to the restored picture to generate a modified restored picture, and can store the modified restored picture in the memory 250, specifically, in the DPB of the memory 250. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, and the like.
[0096] The (modified) restored picture stored in the DPB of the memory 250 can be used as a reference picture in the inter prediction unit 260. The memory 250 can store the motion information of the block where the motion information in the current picture has been derived (or decoded) and / or the motion information of the block in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 260 for utilization as the motion information of the spatial neighboring blocks or the motion information of the temporal neighboring blocks. The memory 250 can store the restored samples of the restored blocks in the current picture and can transmit them to the intra prediction unit 265.
[0097] In this specification, the embodiments described in the filtering unit 160, inter prediction unit 180, and intra prediction unit 185 of the image encoding apparatus 100 can be similarly or correspondingly applied to the filtering unit 240, inter prediction unit 260, and intra prediction unit 265 of the image decoding apparatus 200, respectively.
[0098] General image / video coding procedure
[0099] In image / video coding, pictures constituting an image / video can be encoded / decoded according to a series of decoding orders. The picture order corresponding to the output order of the decoded pictures can be set to be different from the decoding order. Based on this, during inter prediction, not only forward prediction but also backward prediction can be performed.
[0100] FIG. 4 shows an example of a schematic picture decoding procedure to which the embodiments of the present disclosure can be applied.
[0101] Each procedure shown in FIG. 4 can be performed by the image encoding apparatus of FIG. 3. For example, step S410 can be performed by the entropy decoding unit 210, step S420 can be performed by a prediction unit including the intra prediction unit 265 and the inter prediction unit 260, step S430 can be performed by a residual processing unit including the inverse quantization unit 220 and the inverse transform unit 230, step S440 can be performed by the addition unit 235, and step S450 can be performed by the filtering unit 240. Step S410 can include the information decoding procedure described in the present disclosure, step S420 can include the inter / intra prediction procedure described in the present disclosure, step S430 can include the residual processing procedure described in the present disclosure, step S440 can include the block / picture restoration procedure described in the present disclosure, and step S450 can include the in-loop filtering procedure described in the present disclosure.
[0102] Referring to FIG. 4, the picture decoding procedure can generally include, as shown in the description of FIG. 3, an image / video information acquisition procedure (S410) from the bitstream (by decoding), a picture restoration procedure (S420 - S440), and an in-loop filtering procedure (S450) for the restored picture. The picture restoration procedure can be performed based on the predicted samples and residual samples obtained through the inter / intra prediction (S420) and residual processing (S430, inverse quantization and inverse transformation for the quantized transform coefficients) processes described in the present disclosure. Through the in-loop filtering procedure for the restored picture generated by the picture restoration procedure, a modified restored picture can be generated, and the modified restored picture can be output as the decoded picture, and can also be stored in the decoded picture buffer or memory 250 of the decoding device and used as a reference picture in the inter prediction procedure during the decoding of subsequent pictures. In some cases, the in-loop filtering procedure can be omitted. In this case, the restored picture can be output as the decoded picture, and can also be stored in the decoded picture buffer or memory 250 of the decoding device and used as a reference picture in the inter prediction procedure during the decoding of subsequent pictures. The in-loop filtering procedure (S450) can include, as described above, a deblocking filtering procedure, a SAO (sample adaptive offset) procedure, an ALF (adaptive loop filter) procedure, and / or a bilateral filter procedure, etc., and some or all of them can be omitted. Also, one or some of the deblocking filtering procedure, SAO (sample adaptive offset) procedure, ALF (adaptive loop filter) procedure, and bilateral filter procedure can be sequentially applied, or all of them can be sequentially applied. For example, after the deblocking filtering procedure is applied to the restored picture, the SAO procedure can be performed.Alternatively, for example, after a deblocking filtering procedure is applied to the reconstructed picture, the ALF procedure can be performed. This can also be done in the encoding apparatus in the same manner.
[0103] FIG. 5 shows an example of a schematic picture encoding procedure to which the embodiments of the present disclosure can be applied.
[0104] Each procedure shown in FIG. 5 can be performed by the image encoding apparatus of FIG. 2. For example, step S510 can be performed by a prediction unit including the intra prediction unit 185 or the inter prediction unit 180, step S520 can be performed by a residual processing unit including the conversion unit 120 and / or the quantization unit 130, and step S530 can be performed by the entropy encoding unit 190. Step S510 can include the inter / intra prediction procedure described in the present disclosure, step S520 can include the residual processing procedure described in the present disclosure, and step S530 can include the information encoding procedure described in the present disclosure.
[0105] Referring to FIG. 5, the picture encoding procedure is not only a procedure for roughly encoding information for picture restoration (e.g., prediction information, residual information, partitioning information, etc.) and outputting it in the form of a bitstream, as described for FIG. 2, but also a procedure for generating a restored picture for the current picture, and a procedure (optional) for applying in-loop filtering to the restored picture. The encoding device can derive (corrected) residual samples from the quantized transform coefficients via the inverse quantization unit 140 and the inverse transform unit 150, and generate a restored picture based on the prediction samples that are the output of step S510 and the (corrected) residual samples. The restored picture generated in this way can be the same as the restored picture generated by the decoding device described above. Through the in-loop filtering procedure for the restored picture, a corrected restored picture can be generated, which can be stored in the decoded picture buffer or memory 170, and can be used as a reference picture in the inter-prediction procedure when encoding subsequent pictures, similar to the case of the decoding device. As described above, in some cases, part or all of the in-loop filtering procedure can be omitted. When the in-loop filtering procedure is performed, (in-loop) filtering-related information (parameters) can be encoded by the entropy encoding unit 190 and output in the form of a bitstream, and the decoding device can perform the in-loop filtering procedure in the same way as the encoding device based on the filtering-related information.
[0106] Through such in-loop filtering procedures, noise generated during image / video coding, such as blocking artifacts and ringing artifacts, can be reduced, and subjective / objective visual quality can be improved. Also, by performing the in-loop filtering procedures in both the encoding device and the decoding device, the encoding device and the decoding device can derive the same prediction result, enhancing the reliability of picture coding and reducing the amount of data to be transmitted for picture coding.
[0107] As described above, picture restoration procedures can be performed not only in the decoding device but also in the encoding device. Restored blocks can be generated based on intra prediction / inter prediction for each block unit, and a restored picture including the restored blocks can be generated. When the current picture / slice / tile group is an I picture / slice / tile group, the blocks included in the current picture / slice / tile group can be restored based only on intra prediction. On the other hand, when the current picture / slice / tile group is a P or B picture / slice / tile group, the blocks included in the current picture / slice / tile group can be restored based on intra prediction or inter prediction. In this case, inter prediction can be applied to some blocks within the current picture / slice / tile group, and intra prediction can also be applied to some of the remaining blocks. The color components of a picture can include a luma component and a chroma component, and unless explicitly limited in the present disclosure, the methods and examples proposed in the present disclosure can be applied to the luma component and the chroma component.
[0108] Example of coding hierarchy and structure
[0109] The coded video / image according to the present disclosure can be processed, for example, according to the coding hierarchy and structure described later.
[0110] FIG. 6 is a diagram showing an example of a hierarchical structure for a coded image / video.
[0111] The coded image / video can be divided into a VCL (video coding layer) that performs decoding processing of the image / video and handles itself, a lower-level system that transmits and stores the encoded information, and a NAL (network abstraction layer) that exists between the VCL and the lower-level system and is responsible for network adaptation functions.
[0112] In the VCL, it is possible to generate VCL data including compressed image data (slice data), or to generate a parameter set including information such as a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), a Video Parameter Set (VPS), or a SEI (Supplemental Enhancement Information) message that is additionally required for the decoding processing of the image.
[0113] In the NAL, a NAL unit can be generated by adding header information (NAL unit header) to the RBSP (Raw Byte Sequence Payload) generated in the VCL. At this time, the RBSP refers to slice data, parameter sets, SEI messages, etc. generated in the VCL. The NAL unit header can include NAL unit type information specified by the RBSP data included in the corresponding NAL unit.
[0114] As shown in FIG. 6, the NAL unit can be classified into a VCL NAL unit and a Non-VCL NAL unit according to the type of RBSP generated by the VCL. The VCL NAL unit can mean a NAL unit containing information (slice data) for an image, and the Non-VCL NAL unit can mean a NAL unit containing information (parameter set or SEI message) necessary for decoding the image.
[0115] As described above, the above-mentioned VCL NAL unit and Non-VCL NAL unit can be transmitted via a network with header information according to the data standard of the lower system. For example, the NAL unit can be transformed into a data format of a predetermined standard such as the H.266 / VVC file format, RTP (Real-time Transport Protocol), TS (Transport Stream), and transmitted via various networks.
[0116] As described above, the NAL unit type can be specified according to the RBSP data structure (structure) included in the NAL unit, and information on such a NAL unit type can be stored and signaled in the NAL unit header. For example, it can be roughly classified into a VCL NAL unit type and a Non-VCL NAL unit type depending on whether the NAL unit contains information (slice data) for an image. The VCL NAL unit type can be classified according to the nature and type of the picture contained in the VCL NAL unit, and the Non-VCL NAL unit type can be classified according to the type of parameter set.
[0117] The following lists an example of the NAL unit type specified according to the type of parameter set / information included in the Non-VCL NAL unit type.
[0118] - DCI (Decoding capability information) NAL unit type (NUT): Type for NAL units containing DCI
[0119] - VPS (Video Parameter Set) NUT: Type for NAL units containing VPS
[0120] - SPS (Sequence Parameter Set) NUT: Type for NAL units containing SPS
[0121] - PPS (Picture Parameter Set) NUT: Type for NAL units containing PPS
[0122] - APS (Adaptation Parameter Set) NUT: Type for NAL units containing APS
[0123] - PH (Picture header) NUT: Type for NUL units containing a picture header
[0124] The above-mentioned NAL unit types have syntax information for the NAL unit type, and the syntax information can be stored in the NAL unit header and signaled. For example, the syntax information is nal_unit_type, and the NAL unit type can be specified using the value of nal_unit_type.
[0125] On one hand, a picture can include a plurality of slices, and one slice can include a slice header and slice data. In this case, one picture header can be further added for the plurality of slices (slice header and slice data set) within one picture. The picture header (picture header syntax) can include information / parameters that are commonly applicable to the picture. The slice header (slice header syntax) can include information / parameters that are commonly applicable to the slice. The APS (APS syntax) or PPS (PPS syntax) can include information / parameters that are commonly applicable to one or more slices or pictures. The SPS (SPS syntax) can include information / parameters that are commonly applicable to one or more sequences. The VPS (VPS syntax) can include information / parameters that are commonly applicable to multi-layers. The DCI can include information / parameters related to decoding capability.
[0126] In the present disclosure, the high level syntax (HLS) can include at least one of the APS syntax, PPS syntax, SPS syntax, VPS syntax, DCI syntax, picture header syntax, and slice header syntax. Also, in the present disclosure, the low level syntax (LLS) can include, for example, slice data syntax, CTU syntax, coding unit syntax, transform unit syntax, etc.
[0127] On the one hand, in the present disclosure, the image / video information encoded by an encoding device and signaled in the form of a bitstream to a decoding device not only includes information related to partitioning within a picture, intra / inter prediction information, residual information, in-loop filtering information, etc., but can also include the information of the slice header, the information of the picture header, the information of the APS, the information of the PPS, the information of the SPS, the information of the VPS, and / or the information of the DCI. Further, the image / video information can further include general constraint information and / or the information of the NAL unit header.
[0128] High level syntax signalling and semantics
[0129] As described above, the image / video information according to the present disclosure can include High Level Syntax (HLS). The image encoding method and / or the image decoding method can be performed based on the image / video information.
[0130] Video Parameter Set signalling
[0131] A Video Parameter Set (VPS) is a parameter set used for the transmission of hierarchical information. The hierarchical information can include, for example, information related to an output layer set (OLS), information related to a profile tier level, information related to the relationship between the OLS and a hypothetical reference decoder, information related to the relationship between the OLS and the DPB, etc. The VPS can be not essential for the decoding of the bitstream.
[0132] The VPS RBSP (raw byte sequence payload) must be made available for the decoding process by being included in at least one Access Unit (AU) with a TemporalID of 0 before being referenced, or by being provided via external means.
[0133] All VPS NAL units having a vps_video_parameter_set_id with a specific value within a CVS (coded video sequence) must have the same content.
[0134] FIG. 7 is a diagram exemplarily showing the syntax structure of a VPS according to an embodiment of the present disclosure.
[0135] The syntax structure of the VPS shown in FIG. 7 includes only the syntax elements related to the present disclosure, and various other syntax elements not shown in FIG. 7 can be included in the VPS.
[0136] In the example shown in FIG. 7, the vps_video_parameter_set_id provides an identifier for the VPS. Other syntax elements can refer to the VPS using the vps_video_parameter_set_id. The value of the vps_video_parameter_set_id must be greater than 0.
[0137] The value obtained by adding 1 to vps_max_layers_minus1 can indicate the maximum number of allowable layers within each CVS that refers to the VPS.
[0138] The value obtained by adding 1 to vps_max_sublayers_minus1 can indicate the maximum number of temporal sublayers that can exist in the layers within each CVS that refers to the VPS. vps_max_sublayers_minus1 can have a value from 0 to 6.
[0139] The vps_all_layers_same_num_sublayers_flag can be signaled when vps_max_layers_minus1 is greater than 0 and vps_max_sublayers_minus1 is greater than 0. A vps_all_layers_same_num_sublayers_flag with a first value (e.g., 1) can indicate that the number of temporal sublayers is the same for all layers within each CVS that refers to the VPS. A vps_all_layers_same_num_sublayers_flag with a second value (e.g., 0) can indicate that the layers within each CVS that refers to the VPS can have a different number of temporal sublayers. When the vps_all_layers_same_num_sublayers_flag does not exist, its value can be inferred as the first value (e.g., 1).
[0140] The vps_all_independent_layers_flag can be signaled when vps_max_layers_minus1 is greater than 0. A vps_all_independent_layers_flag with a first value (e.g., 1) can indicate that all layers within the CVS are independently encoded without using inter-layer prediction. A vps_all_independent_layers_flag with a second value (e.g., 0) can indicate that one or more layers within the CVS can use inter-layer prediction. When the vps_all_independent_layers_flag does not exist, its value can be inferred as the first value (e.g., 1).
[0141] each_layer_is_an_ols_flag can be signaled when vps_max_layers_minus1 is greater than 0. Also, each_layer_is_an_ols_flag can be signaled when vps_all_independent_layers_flag is the first value. The each_layer_is_an_ols_flag with the first value (e.g., 1) can indicate whether each OLS contains only one layer. Also, the each_layer_is_an_ols_flag with the first value (e.g., 1) can indicate that each layer itself within the CVS referring to the VPS is an OLS (i.e., the one layer contained in the OLS is the only output layer). Also, the each_layer_is_an_ols_flag with the second value (e.g., 0) can indicate that at least one OLS can contain more than one layer. When vps_max_layers_minus1 is 0, the value of each_layer_is_an_ols_flag can be inferred as 1. Otherwise, when vps_all_independent_layers_flag is 0, the value of each_layer_is_an_ols_flag can be inferred as 0.
[0142] When each_layer_is_an_ols_flag is the second value (e.g., 0) and vps_all_independent_layers_flag is the second value (e.g., 0), ols_mode_idc can be signaled.
[0143] The ols_mode_idc with the first value (e.g., 0) can indicate that the total number of OLSs specified by the VPS is equal to vps_max_layers_minus1 + 1. At this time, the i-th OLS can contain the layers with layer indices from 0 to i. Also, for each OLS, only the layer with the highest layer index within the OLS (the top layer) can be output.
[0144] The ols_mode_idc having the second value (e.g., 1) can indicate that the total number of OLSs specified by the VPS is equal to vps_max_layers_minus1 + 1. At this time, the i-th OLS can include the layers with layer indices from 0 to i. Further, for each OLS, all the layers within the OLS can be output.
[0145] The ols_mode_idc having the third value (e.g., 2) can indicate that the total number of OLSs specified by the VPS is explicitly signaled. Further, for each OLS, it can indicate that the output layer is explicitly signaled. The other layers that are not the output layer can be the direct reference layer or the indirect reference layer of the output layer of the OLS.
[0146] When vps_all_independent_layers_flag is 1 and each_layer_is_an_ols_flag is 0, the value of ols_mode_idc can be inferred to be the third value (e.g., 2).
[0147] When ols_mode_idc is 2, num_output_layer_sets_minus1 and ols_output_layer_flag[i][j] can be explicitly signaled.
[0148] The value obtained by adding 1 to num_output_layer_sets_minus1 can indicate the total number of OLSs specified by the VPS.
[0149] When ols_mode_idc is 2, ols_output_layer_flag[i][j] can indicate whether the j-th layer of the i-th OLS is an output layer. ols_output_layer_flag[i][j] with the first value (for example, 1) can indicate that the layer having the same layer identifier (nuh_layer_id) as vps_layer_id[j] is the output layer of the i-th OLS. ols_output_layer_flag[i][j] with the second value (for example, 0) can indicate that the layer having the same layer identifier (nuh_layer_id) as vps_layer_id[j] is not the output layer of the i-th OLS.
[0150] Hereinafter, the HRD parameters signaled in the VPS will be described.
[0151] When each_layer_is_an_ols_flag has the second value (for example, 0), vps_general_hrd_params_present_flag can be signaled. vps_general_hrd_params_present_flag with the first value (for example, 1) can indicate that there are HRD parameters different from the general_hrd_parameters() syntax structure in the VPS. vps_general_hrd_params_present_flag with the second value (for example, 0) can indicate that there are no HRD parameters different from the general_hrd_parameters() syntax structure in the VPS. When vps_general_hrd_params_present_flag does not exist, its value can be inferred as the second value (for example, 0).
[0152] When the i-th OLS contains one layer (NumLayersInOls[i] is equal to 1), the general_hrd_parameters() syntax structure applied to the i-th OLS may exist in the SPS (Sequence Parameter Set) referred to by the layer within the i-th OLS.
[0153] The vps_sublayer_cpb_params_present_flag can be signaled when vps_max_sublayers_minus1 is greater than 0. A vps_sublayer_cpb_params_present_flag with a first value (e.g., 1) can indicate that the ols_hrd_parameters() syntax structure within the VPS contains HRD parameters for sublayers where the TemporalId (Temporal Identifier) ranges from 0 to hrd_max_tid[i]. A vps_sublayer_cpb_params_present_flag with a second value (e.g., 0) can indicate that the ols_hrd_parameters() syntax structure within the VPS contains HRD parameters for sublayers where the TemporalId is only hrd_max_tid[i]. When vps_max_sublayers_minus1 is 0, the vps_sublayer_cpb_params_present_flag can be inferred to have the second value (e.g., 0).
[0154] When the vps_sublayer_cpb_params_present_flag has the second value (e.g., 0), it can be inferred that the HRD parameters for sublayers where the TemporalId ranges from 0 to hrd_max_tid[i] - 1 are the same as the HRD parameters for sublayers where the TemporalId is hrd_max_tid[i].
[0155] The value obtained by adding 1 to num_ols_hrd_params_minus1 can indicate the number of ols_hrd_parameters() syntax structures within the VPS. num_ols_hrd_params_minus1 can have a value from 0 to TotalNumOlss - 1. TotalNumOlss can indicate the total number of OLSs specified by the VPS. In the present disclosure, the HRD parameter can mean ols_hrd_parameters(). Therefore, the number of HRD parameter syntax structures can mean the number of ols_hrd_parameters() syntax structures.
[0156] hrd_max_tid[i] can be signaled when vps_max_sublayers_minus1 is greater than 0 and vps_all_layers_same_num_sublayers_flag is the second value (e.g., 0). hrd_max_tid[i] can indicate the TemporalId of the top sublayer in which the related HRD parameter is included in the i-th ols_hrd_parameters() syntax structure.
[0157] hrd_max_tid[i] can have a value from 0 to vps_max_sublayers_minus1. When vps_max_sublayers_minus1 is 0, the value of hrd_max_tid[i] can be inferred to be 0. When vps_max_sublayers_minus1 is greater than 0 and vps_all_layers_same_num_sublayers_flag is 1, the value of hrd_max_tid[i] can be inferred to be equal to vps_max_sublayers_minus1.
[0158] As shown in Figure 7, a variable firstSubLayer indicating the temporal identifier (TemporalId) of the first sublayer can be derived based on vps_sublayer_cpb_params_present_flag to be 0 or hrd_max_tid[i]. Specifically, when vps_sublayer_cpb_params_present_flag is 1, firstSubLayer is derived to be 0; otherwise, firstSubLayer can be derived to be hrd_max_tid[i]. Based on the derived firstSubLayer and hrd_max_tid[i], the ols_hrd_parameters() syntax structure can be signaled.
[0159] When the value obtained by adding 1 to num_ols_hrd_params_minus1 is not equal to TotalNumOlss and num_ols_hrd_params_minus1 is greater than 0, ols_hrd_idx[i] can be signaled. At this time, ols_hrd_idx[i] can be signaled for the i-th OLS when the number of layers (NumLayersInOls[i]) included in the i-th OLS is greater than 1. ols_hrd_idx[i] is an index for the list of ols_hrd_parameters() in the VPS and can be the index of the ols_hrd_parameters() applied to the i-th OLS. ols_hrd_idx[i] can have a value from 0 to num_ols_hrd_params_minus1. When the number of layers (NumLayersInOls[i]) included in the i-th OLS is 1, the ols_hrd_parameters() syntax structure applied to the i-th OLS can exist in the SPS referred to by the layer in the i-th OLS.
[0160] In the present disclosure, ols_hrd_idx[i] is the index of ols_hrd_parameters() applied to the i-th OLS or the i-th multi-layer OLS, and can be called the mapping information (information regarding the mapping) between the (multi-layer) OLS and the HRD parameter syntax structure (ols_hrd_parameters()).
[0161] If the value obtained by adding 1 to num_ols_hrd_param_minus1 is equal to TotalNumOlss, the value of ols_hrd_idx[i] can be inferred to be equal to i. Otherwise, if NumLayersInOls[i] is greater than 1 and num_ols_hrd_params_minus1 is 0, the value of ols_hrd_idx[i] can be inferred to be 0.
[0162] HRD signalling in VPS and SPS
[0163] Hereinafter, the signaling of HRD parameters according to the present disclosure will be described more specifically. The HRD parameters can be signaled for each output layer set (OLS). A hypothetical reference decoder (HRD) is a virtual decoder model that specifies the limitations on the variability of a NAL unit stream conforming to a standard specification or a byte stream conforming to a standard specification that can be generated during the encoding process.
[0164] As described with reference to FIG. 7, the HRD parameters can be included in the VPS and signaled. Alternatively, the HRD parameters can also be included in the SPS and signaled.
[0165] FIG. 8 is a diagram showing the syntax structure of the SPS for signaling HRD parameters according to an embodiment of the present disclosure.
[0166] In the example shown in FIG. 8, the sps_ptl_dpb_hrd_params_present_flag with the first value (e.g., 1) can indicate that the profile_tier_level() syntax structure and the dpb_parameters() syntax structure are present in the SPS. The profile_tier_level() is a syntax structure for transmitting parameters for the profile tier level, and the dpb_parameters() can be a syntax structure for transmitting DPB (decoded picture buffer) parameters. Also, the sps_ptl_dpb_hrd_params_present_flag with the first value (e.g., 1) can indicate that the general_hrd_parameters() syntax structure and the ols_hrd_parameters() syntax structure may be present in the SPS. The sps_ptl_dpb_hrd_params_present_flag with the second value (e.g., 0) can indicate that the four syntax structures are not present in the SPS. The value of the sps_ptl_dpb_hrd_params_present_flag can be the same as the value of vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]]. That is, the value of the sps_ptl_dpb_hrd_params_present_flag can be encoded as the value of vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]].
[0167] In the above, vps_independent_layer_flag[i] can be a syntax element included in the VPS and transmitted. A vps_independent_layer_flag[i] with a first value (e.g., 1) can indicate that the layer with index i is an independent layer that does not use inter-layer prediction. A vps_independent_layer_flag[i] with a second value (e.g., 0) can indicate that the layer with index i can use inter-layer prediction. If vps_independent_layer_flag[i] does not exist, its value can be inferred as the first value (e.g., 1).
[0168] If sps_ptl_dpb_hrd_params_present_flag is 1, sps_general_hrd_params_present_flag can be signaled.
[0169] A sps_general_hrd_params_present_flag with a first value (e.g., 1) can indicate that the SPS includes both the general_hrd_parameters() syntax structure and the ols_hrd_parameters() syntax structure. A sps_general_hrd_params_present_flag with a second value (e.g., 0) can indicate that the SPS does not include either the general_hrd_parameters() syntax structure or the ols_hrd_parameters() syntax structure.
[0170] According to what is shown in FIG. 8, when sps_max_sublayers_minus1 is greater than 0, sps_sublayer_cpb_params_present_flag can be signaled. At this time, the value obtained by adding 1 to sps_max_sublayers_minus1 can indicate the maximum number of temporal sublayers that may exist in each CLVS (coded layer video sequence) referring to the SPS. The sps_sublayer_cpb_params_present_flag with the first value (for example, 1) can indicate that the ols_hrd_parameters() syntax structure in the SPS includes HRD parameters for sublayers where the TemporalId (Temporal Identifier) is from 0 to sps_max_sublayers_minus1. The sps_sublayer_cpb_params_present_flag with the second value (for example, 0) can indicate that the ols_hrd_parameters() syntax structure in the SPS includes HRD parameters for the sublayer where the TemporalId is only sps_max_sublayers_minus1. When sps_max_sublayers_minus1 is 0, the value of sps_sublayer_cpb_params_present_flag can be inferred as the second value (for example, 0).
[0171] When sps_sublayer_cpb_params_present_flag is the second value (for example, 0), it can be inferred that the HRD parameters for sublayers where the TemporalId (Temporal Identifier) is from 0 to sps_max_sublayers_minus1 - 1 are equal to the HRD parameters for the sublayer where the TemporalId is sps_max_sublay_minus1.
[0172] Figure 9 is a diagram showing the general_hrd_parameters() syntax structure according to an embodiment of the present disclosure.
[0173] As shown in Figure 9, the general_hrd_parameters() syntax structure can include some of the sequence level HRD parameters used in the HRD operation. As a requirement for the integrity of the bitstream, the content of general_hrd_parameters() present in the VPS or SPS within the bitstream must be the same.
[0174] When the general_hrd_parameters() syntax structure is included in the VPS, the general_hrd_parameters() syntax structure can be applied to all OLSs specified by the VPS. When the general_hrd_parameters() syntax structure is included in the SPS, the general_hrd_parameters() syntax structure can be applied to the OLS that includes only the lowest layer among the hierarchies referring to the SPS. At this time, the lowest layer may be an independent layer.
[0175] As shown in Figure 9, the general_hrd_parameters() syntax structure can include syntax elements such as num_units_in_tick, time_scale, and general_nal_hrd_params_present_flag as HRD parameters. The HRD parameters shown in Figure 9 can have the same meaning as the conventional HRD parameters. Therefore, specific descriptions of HRD parameters that are less relevant to the present disclosure are omitted.
[0176] Figure 10 is a diagram showing the ols_hrd_parameters() syntax structure according to an embodiment of the present disclosure.
[0177] When the ols_hrd_parameters() syntax structure is included in the VPS, the OLSs to which the ols_hrd_parameters() syntax structure is applied can be specified by the VPS. When the ols_hrd_parameters() syntax structure is included in the SPS, the ols_hrd_parameters() syntax structure can be applied to the OLS that includes only the lowest layer among the layers referring to the SPS. At this time, the lowest layer may be an independent layer.
[0178] As shown in FIG. 10, the ols_hrd_parameters() syntax structure can include syntax elements such as fixed_pic_rate_general_flag, fixed_pic_rate_within_cvs_flag, and elemental_duration_in_tc_minus1 as HRD parameters. The HRD parameters shown in FIG. 10 can have the same meaning as the conventional HRD parameters. Therefore, a specific description of HRD parameters with little relevance to the present disclosure is omitted.
[0179] FIG. 11 is a diagram showing the sublayer_hrd_parameters() syntax structure according to an embodiment of the present disclosure.
[0180] The sublayer_hrd_parameters() syntax structure can be included in and signaled by the ols_hrd_parameters() syntax structure of FIG. 10.
[0181] As shown in FIG. 11, the syntax structure of sublayer_hrd_parameters() can include syntax elements such as bit_rate_value_minus1, cpb_size_value_minus1, and cpb_size_du_value_minus1 as HRD parameters. The HRD parameters shown in FIG. 11 can have the same meaning as the conventional HRD parameters. Therefore, a specific description of HRD parameters with little relevance to the present disclosure is omitted.
[0182] For reference, the output time can mean the time when the picture restored from the DPB is output. The output time can be specified by the HRD according to the output timing DPB operation.
[0183] Two sets of HRD parameters, namely NAL HRD parameters and VCL HRD parameters, can be used. The HRD parameters can be signaled via the general_hrd_parameters() syntax structure and the ols_hrd_parameters() syntax structure. The general_hrd_parameters() syntax structure and the ols_hrd_parameters() syntax structure may be included in the VPS for signaling or may be included in the SPS for signaling.
[0184] For example, DPB management can be performed based on the HRD parameters. For example, deletion of pictures from the DPB and / or output of (decoded) pictures before decoding of the current picture can be performed based on the HRD parameters.
[0185] According to the method for signaling HRD parameters described with reference to FIGS. 7 to 11, at least the following problems may occur.
[0186] - As described above, num_ols_hrd_params_minus1 can be restricted to have values from 0 to TotalNumOlss - 1. However, there is no restriction that each HRD parameter signaled in the VPS should be associated with at least one OLS. Therefore, since the VPS can include HRD parameters that are not used, there is a problem that the signaling efficiency decreases.
[0187] - As described above, num_ols_hrd_params_minus1 can be restricted to have values from 0 to TotalNumOlss - 1. However, since there can be one or more OLSs that include only one layer, the restriction allows the unused HRD parameter structure to be included and signaled in the VPS. An OLS that includes only one layer is not associated with the HRD parameter structure signaled in the VPS.
[0188] The signaling of HRD parameters related to OLSs can include other drawbacks not mentioned in the present disclosure by including the problems described above.
[0189] Examples according to the present disclosure for solving at least one of the above problems can include at least one of the following configurations. The following configurations can be applied individually or in combination with other configurations.
[0190] Configuration 1: Each HRD parameter structure signaled in the VPS can be restricted to be associated with at least one OLS.
[0191] Configuration 2: The number of HRD parameter structures signaled by the VPS (i.e., num_ols_hrd_params_minus1) shall not be greater than the number of OLSs that include more than one layer. That is, the number of HRD parameter structures signaled by the VPS can be restricted so as not to be greater than the number obtained by subtracting the number of OLSs that include only one layer from the total number of OLSs.
[0192] FIG. 12 is a diagram for explaining an example of an image encoding method to which an embodiment according to the present disclosure can be applied.
[0193] The image encoding apparatus can derive HRD parameters (S1210) and encode image / video information (S1220). At this time, the image / video information can include information related to the derived HRD parameters.
[0194] Although not shown in FIG. 12, the image encoding apparatus can perform DPB management based on the HRD parameters derived in step S1210.
[0195] FIG. 13 is a diagram for explaining an example of an image decoding method to which an embodiment according to the present disclosure can be applied.
[0196] The image decoding apparatus can obtain image / video information from a bitstream (S1310). At this time, the image / video information can include information related to HRD parameters.
[0197] The image decoding apparatus can decode a picture based on the obtained HRD parameters (S1320).
[0198] FIG. 14 is a diagram for explaining another example of an image decoding method to which an embodiment according to the present disclosure can be applied.
[0199] The image decoding device can obtain image / video information from a bitstream (S1410). At this time, the image / video information can include information related to HRD parameters.
[0200] The image decoding device can perform DPB management based on the obtained HRD parameters (S1420).
[0201] The image decoding device can decode a picture based on the DPB (S1430). For example, blocks / slices within the current picture can be decoded based on inter prediction that uses already restored pictures within the DPB as reference pictures.
[0202] In the example described with reference to FIGS. 12 to 14, the information related to the HRD parameters can include at least one of the information / syntax elements described in relation to at least one of the embodiments of the present disclosure. Also, as described above, DPB management can be performed based on the HRD parameters. For example, deletion of pictures from the DPB and / or output of (decoded) pictures before decoding of the current picture can be performed based on the HRD parameters.
[0203] According to an embodiment of the present disclosure for solving at least some of the above-described problems, each HRD parameter structure signaled by the VPS can be restricted to be associated with at least one OLS.
[0204] As described above, the ols_hrd_parameters() applied to the i-th OLS can be specified by ols_hrd_idx[i]. According to this embodiment, each of all the ols_hrd_parameters() signaled in the VPS can be restricted to be applied to at least one OLS. That is, each ols_hrd_parameters() in the VPS can be specified by at least one ols_hrd_idx[i].
[0205] According to this embodiment, each of the ols_hrd_parameters() in the VPS is used at least once. That is, the ols_hrd_parameters() that are not used are not signaled. Therefore, according to this embodiment, there is an effect that the signaling of the ols_hrd_parameters() can be performed efficiently.
[0206] According to another embodiment of the present disclosure for solving at least a part of the above-described problems, the number of HRD parameter structures signaled in the VPS (i.e., num_ols_hrd_params_minus1) can be restricted to be not greater than the number of OLSs including more than one hierarchy.
[0207] According to the example described with reference to FIG. 7, the value obtained by adding 1 to num_ols_hrd_params_minus1 indicates the number of ols_hrd_parameters() syntax structures in the VPS, and num_ols_hrd_params_minus1 can have a value from 0 to TotalNumOlss - 1.
[0208] However, as described above, an OLS containing only one layer can exist, and the OLS containing only one layer is not associated with the HRD parameter structure signaled in the VPS. The range of the value of num_ols_hrd_params_minus1 in the example of FIG. 7 may cause inaccurate signaling. Therefore, if the range of the number of ols_hrd_parameters() syntax structures is specified by the number of OLSs with multiple layers (multi-layer OLSs) (NumMultiLayerOlss) instead of the total number of OLSs (TotalNumOlss), accurate signaling can be performed.
[0209] According to the present disclosure, the value obtained by adding 1 to num_ols_hrd_params_minus1 indicates the number of ols_hrd_parameters() syntax structures in the VPS, and num_ols_hrd_params_minus1 can have a value from 0 to NumMultiLayerOlss - 1. At this time, the number of OLSs with multiple layers (multi-layer OLSs) (NumMultiLayerOlss) can be the same as the result of subtracting the number of OLSs containing only one layer (NumSingleLayerOlss) from the total number of OLSs (TotalNumOlss).
[0210] As described above, each HRD parameter structure signaled in the VPS can be restricted to be associated (mapped) with at least one OLS. Also, the number of HRD parameter structures signaled in the VPS (that is, num_ols_hrd_params_minus1) can be restricted to be not greater than the number of OLSs with multiple layers (multi-layer OLSs). The above two embodiments can be combined to form another embodiment as follows.
[0211] FIG. 15 is a diagram for explaining the process of encoding HRD parameters based on num_ols_hrd_params_minus1 according to another embodiment of the present disclosure.
[0212] The image encoding device can encode num_ols_hrd_params_minus1 into the VPS (S1510). The value obtained by adding 1 to num_ols_hrd_params_minus1 indicates the number of ols_hrd_parameters() syntax structures in the VPS, and num_ols_hrd_params_minus1 can have values from 0 to NumMultiLayerOlss - 1.
[0213] The image encoding device can encode num_ols_hrd_params_minus1 + 1 ols_hrd_parameters() syntax structures into the VPS (S1520).
[0214] The image encoding device can determine the following condition 1 (S1530).
[0215] Condition 1: (num_ols_hrd_params_minus1 + 1!= NumMultiLayerOlss && num_ols_hrd_params_minus1 > 0)?
[0216] The condition 1 is a condition for encoding ols_hrd_idx into the bitstream. When the condition 1 is satisfied (S1530 - Yes), the image encoding device can encode ols_hrd_idx into the VPS (S1540). According to the embodiment described with reference to FIG. 15, ols_hrd_idx[i] is an index for the list of ols_hrd_parameters() in the VPS, and can be an index of ols_hrd_parameters() applied to an OLS (multi - layer OLS) including the i - th plurality of layers. That is, ols_hrd_idx[i] in the VPS is signaled for the multi - layer OLSs and can have values from 0 to num_ols_hrd_params_minus1.
[0217] When the condition 1 is not satisfied (S1530-No), the image encoding device does not encode ols_hrd_idx into the VPS, and its value can be inferred (S1550). Specifically, when num_ols_hrd_params_minus1 is 0, the value of ols_hrd_idx[i] can be inferred to be 0. Otherwise, when the value obtained by adding 1 to num_ols_hrd_param_minus1 is equal to NumMultiLayerOlss, the value of ols_hrd_idx[i] can be inferred to be equal to i.
[0218] Also, the ols_hrd_parameters() syntax structure applied to the OLS containing only a single layer is not encoded into the VPS and can be encoded into the SPS referenced by the layer within the OLS.
[0219] Also, each ols_hrd_parameters() within the VPS can be referenced by at least one ols_hrd_idx[i] (where i has a value from 0 to NumMultiLayerOlss-1).
[0220] The image encoding device can encode a picture based on the HRD parameters (S1560). At this time, the HRD parameters can be those obtained in step S1540, or the ols_hrd_parameters() within the VPS referenced by the ols_hrd_idx[i] inferred in step S1550. Or, in the case of an OLS containing only a single layer, the HRD parameters can be the ols_hrd_parameters() within the SPS referenced by the single layer.
[0221] FIG. 16 is a diagram for explaining the process of decoding HRD parameters based on num_ols_hrd_params_minus1 according to another embodiment of the present disclosure.
[0222] The image decoding device can obtain num_ols_hrd_params_minus1 from the VPS (S1610). The value obtained by adding 1 to num_ols_hrd_params_minus1 indicates the number of ols_hrd_parameters() syntax structures in the VPS, and num_ols_hrd_params_minus1 can have values from 0 to NumMultiLayerOlss - 1.
[0223] The image decoding device can obtain num_ols_hrd_params_minus1 + 1 ols_hrd_parameters() syntax structures from the VPS (S1620).
[0224] The image decoding device can determine the following Condition 2 (S1630).
[0225] Condition 2: (num_ols_hrd_params_minus1 + 1!= NumMultiLayerOlss && num_ols_hrd_params_minus1 > 0)?
[0226] The said Condition 2 is a condition for obtaining ols_hrd_idx from the bitstream. When the said Condition 2 is satisfied (S1630 - Yes), the image decoding device can obtain ols_hrd_idx from the VPS (S1640). According to the embodiment described with reference to FIG. 16, ols_hrd_idx[i] can be an index for the list of ols_hrd_parameters() in the VPS, and can be an index of ols_hrd_parameters() applied to an OLS (multi - layer OLS) including the i - th plurality of layers.
[0227] That is, ols_hrd_idx[i] in the VPS is signaled for the multi - layer OLS and can have values from 0 to num_ols_hrd_params_minus1.
[0228] When the condition 2 is not satisfied (S1630-No), the image decoding apparatus does not acquire ols_hrd_idx from the VPS and can infer its value (S1650). Specifically, when num_ols_hrd_params_minus1 is 0, the value of ols_hrd_idx[i] can be inferred to be 0. Otherwise, when the value obtained by adding 1 to num_ols_hrd_param_minus1 is equal to NumMultiLayerOlss, the value of ols_hrd_idx[i] can be inferred to be equal to i.
[0229] Also, the ols_hrd_parameters() syntax structure applied to an OLS including only a single layer is not acquired from the VPS and can be acquired from the SPS referenced by the layer within the OLS.
[0230] Also, each ols_hrd_parameters() in the VPS can be referenced by at least one ols_hrd_idx[i] (where i has a value from 0 to NumMultiLayerOlss-1).
[0231] The image decoding apparatus can decode a picture based on the HRD parameters (S1660). At this time, the HRD parameters can be those acquired in step S1640, or the ols_hrd_parameters() in the VPS referenced by the ols_hrd_idx[i] inferred in step S1650. Alternatively, in the case of an OLS including only a single layer, the HRD parameters can be the ols_hrd_parameters() in the SPS referenced by the single layer.
[0232] According to the embodiments described with reference to FIGS. 15 and 16, unnecessary signaling of the ols_hrd_parameters() syntax structure within the non-referred VPS can be prevented, and for the OLS containing only a single layer, the ols_hrd_parameters() syntax structure can more accurately and efficiently signal the HRD parameters by signaling only via the SPS.
[0233] In the method described with reference to FIGS. 15 and 16, some steps may be omitted, and the order may be changed with other steps. Also, steps not shown in FIGS. 15 and 16 can be added at any position.
[0234] Although the exemplary method of the present disclosure is presented in a series of operations for clarity of explanation, this is not for limiting the order in which the steps are performed. If necessary, each step can also be performed simultaneously or in a different order. To implement the method according to the present disclosure, it can further include other steps in addition to the exemplified steps, or include the remaining steps except for some steps, or include additional other steps except for some steps.
[0235] In the present disclosure, an image encoding device or an image decoding device that performs a predetermined operation (step) can perform an operation (step) of checking the execution conditions and situations of the said operation (step). For example, when it is described that a predetermined operation is performed if a predetermined condition is satisfied, the image encoding device or the image decoding device can perform the said predetermined operation after performing an operation of checking whether the said predetermined condition is satisfied.
[0236] The various embodiments of the present disclosure do not list all possible combinations, but are for explaining representative aspects of the present disclosure. The matters described in the various embodiments may be applied independently or in combinations of two or more.
[0237] In addition, various embodiments of the present disclosure can be realized by hardware, firmware, software, or a combination thereof. In the case of realization by hardware, it can be realized by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), general processors, controllers, microcontrollers, microprocessors, etc.
[0238] In addition, the image decoding device and the image encoding device to which the embodiments of the present disclosure are applied can be included in a multimedia broadcast transceiver, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video conferencing device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camcorder, an on-demand video (VoD) service providing device, an over-the-top video (OTT) device, an Internet streaming service providing device, a three-dimensional (3D) video device, an image phone video device, and a medical video device, etc., and can be used to process video signals or data signals. For example, as the over-the-top video (OTT) device, it can include a game console, a Blu-ray player, an Internet-connected TV, a home theater system, a smartphone, a tablet PC, a Digital Video Recoder (DVR), etc.
[0239] FIG. 17 is a diagram illustrating a content streaming system to which an embodiment of the present disclosure can be applied.
[0240] As shown in FIG. 17, a content streaming system to which an embodiment of the present disclosure is applied can generally include an encoding server, a streaming server, a Web server, a media storage, a user device, and a multimedia input device.
[0241] The encoding server compresses content input from a multimedia input device such as a smartphone, a camera, or a camcorder into digital data to generate a bitstream and transmits the bitstream to the streaming server. As another example, when a multimedia input device such as a smartphone, a camera, or a video camera directly generates a bitstream, the encoding server can be omitted.
[0242] The bitstream can be generated by an image encoding method and / or an image encoding device to which an embodiment of the present disclosure is applied, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.
[0243] The streaming server transmits multimedia data to the user device based on a user's request via the Web server, and the Web server can serve as a medium to inform the user of what services are available. When the user requests a desired service from the Web server, the Web server transmits this to the streaming server, and the streaming server can transmit multimedia data to the user. At this time, the content streaming system can include a separate control server, and in this case, the control server can play a role in controlling commands / responses between each device in the content streaming system.
[0244] The streaming server can receive content from a media storage and / or an encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.
[0245] Examples of the user device may include a mobile phone, a smart phone, a laptop computer, a digital broadcast terminal, a PDA (personal digital assistants), a PMP (portable multimedia player), a navigation device, a slate PC, a tablet PC, an ultrabook, a wearable device, for example, a smartwatch, smart glass, an HMD (head mounted display), a digital TV, a desktop computer, a digital signage, and the like.
[0246] Each server in the content streaming system can be operated as a distributed server. In this case, the data received from each server can be processed distributively.
[0247] The scope of the present disclosure includes software or machine-executable commands (for example, an operating system, an application, firmware, a program, etc.) that enable operations according to methods of various embodiments to be executed on a device or a computer, and a non-transitory computer-readable medium on which such software or commands are stored and can be executed on a device or a computer.
Industrial Applicability
[0248] Examples according to the present disclosure can be used to encode / decode images.
Claims
1. An image decoding method performed by an image decoding device, the image decoding method comprising: obtaining first information indicating a number of one or more hypothetical reference decoder (HRD) parameter syntax structures in a video parameter set (VPS); obtaining the one or more HRD parameter syntax structures from the VPS based on the first information; obtaining second information regarding a mapping between one or more multi-layer output layer sets (OLS) and the one or more HRD parameter syntax structures from the VPS based on the first information; selecting an HRD parameter syntax structure to be applied to the current OLS based on the second information; and processing the current OLS based on the selected HRD parameter syntax structure; The first information is obtained based on whether the VPS includes the one or more HRD parameter syntax structures; Each of the one or more HRD parameter syntax structures in the VPS is mapped to at least one multi-layer OLS of the one or more multi-layer OLSs.
2. The image decoding method of claim 1 , wherein a number of the one or more HRD parameter syntax structures in the VPS is not greater than a number of the one or more multi-layer OLSs.
3. 2. The image decoding method of claim 1, wherein the second information is obtained from the VPS based on the number of the one or more HRD parameter syntax structures in the VPS being greater than one and the number of the one or more HRD parameter syntax structures in the VPS not being equal to the number of the one or more multi-layer OLSs.
4. 4. The image decoding method of claim 3, wherein, based on a number of the one or more HRD parameter syntax structures in the VPS being 1, the second information is not obtained from the VPS and the second information is inferred to be equal to a value of 0.
5. 4. The image decoding method of claim 3, wherein, based on the number of the one or more HRD parameter syntax structures in the VPS being greater than 1 and the number of the one or more HRD parameter syntax structures in the VPS being equal to the number of the one or more multi-layer OLSs, the second information is not obtained from the VPS, and it is inferred that the second information of an i-th multi-layer OLS is equal to the value of i.
6. The image decoding method of claim 1 , wherein the HRD parameter syntax structure applied to the current OLS is obtained from a Sequence Parameter Set (SPS) on the basis that the current OLS includes only a single layer.
7. An image coding method performed by an image coding device, the image coding method comprising: encoding first information indicating a number of one or more hypothetical reference decoder (HRD) parameter syntax structures in a video parameter set (VPS); encoding the one or more HRD parameter syntax structures in the VPS based on the first information; encoding second information regarding a mapping between one or more multi-layer output layer sets (OLS) in the VPS and the one or more HRD parameter syntax structures based on the first information; and processing the current OLS based on an HRD parameter syntax structure applied to the current OLS; the first information is encoded based on whether the VPS includes the one or more HRD parameter syntax structures; Each of the one or more HRD parameter syntax structures in the VPS is mapped to at least one multi-layer OLS of the one or more multi-layer OLSs.
8. The image coding method of claim 7 , wherein a number of the one or more HRD parameter syntax structures in the VPS is not greater than a number of the one or more multi-layer OLSs.
9. 8. The image encoding method of claim 7, wherein the second information is encoded within the VPS based on the number of the one or more HRD parameter syntax structures in the VPS being greater than one and the number of the one or more HRD parameter syntax structures in the VPS not being equal to the number of the one or more multi-layer OLSs.
10. 10. The image encoding method of claim 9, wherein the second information is not encoded in the VPS and the second information is inferred to be equal to a value of 0 based on the number of the one or more HRD parameter syntax structures in the VPS being 1.
11. 10. The image encoding method of claim 9, wherein, based on the number of the one or more HRD parameter syntax structures in the VPS being greater than one and the number of the one or more HRD parameter syntax structures in the VPS being equal to the number of multi-layer OLSs, the second information is not encoded in the VPS, and it is inferred that the second information of an i-th multi-layer OLS is equal to the value of i.
12. The image encoding method of claim 7 , wherein the HRD parameter syntax structure applied to the current OLS is encoded within a Sequence Parameter Set (SPS) based on the current OLS including only a single layer.
13. 1. A method for transmitting a bitstream, comprising: generating the bitstream; transmitting the bitstream; The bitstream comprises: encoding first information indicating a number of one or more hypothetical reference decoder (HRD) parameter syntax structures in a video parameter set (VPS); encoding the one or more HRD parameter syntax structures in the VPS based on the first information; encoding second information regarding a mapping between one or more multi-layer output layer sets (OLS) in the VPS and the one or more HRD parameter syntax structures based on the first information; processing the current OLS based on an HRD parameter syntax structure applied to the current OLS; the first information is encoded based on whether the VPS includes the one or more HRD parameter syntax structures; The method of claim 1, wherein each of the one or more HRD parameter syntax structures in the VPS is mapped to at least one multi-layer OLS of the one or more multi-layer OLSs.
Citation Information
Patent Citations
Image encoding / decoding method and apparatus for signaling HRD parameters, and computer-readable recording medium storing a bitstream - Patents.com
JP7667351B2