Image encoding / decoding method and apparatus for signaling HRD parameters, and computer-readable recording medium storing a bitstream
The image encoding/decoding method and apparatus address the challenge of high-resolution image data transmission costs by efficiently signaling HRD parameters, improving encoding/decoding efficiency and reducing storage costs.
Patent Information
- Application Number
- JP2025513270
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-09-07
- Filing Date
- 2023-09-01
- Publication Date
- 2025-08-22
AI Technical Summary
The increasing demand for high-resolution, high-quality images leads to a significant increase in transmission and storage costs due to the higher amount of information required, necessitating highly efficient image compression techniques.
An image encoding/decoding method and apparatus that efficiently signal HRD parameters, including operations for determining compliance with video coding standards and updating a Hypothetical Stream Scheduler, allowing for improved encoding/decoding efficiency.
The method and apparatus provide enhanced encoding/decoding efficiency by effectively signaling HRD parameters, enabling efficient transmission and storage of high-resolution images.
Smart Images

Figure 2025527904000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an image encoding / decoding method and apparatus, and more particularly to an image encoding / decoding method and apparatus that signal HRD (Hypothetical reference decoder)-related parameters, as well as a computer-readable recording medium storing a bitstream generated by the image encoding method / apparatus of the present disclosure. [Background technology]
[0002] Recently, demand for high-resolution, high-quality images, such as HD (High Definition) images and UHD (Ultra High Definition) images, has been increasing in various fields. As image data becomes higher in resolution and quality, the amount of information or bits to be transmitted increases relatively compared to conventional image data. The increase in the amount of information or bits to be transmitted results in an increase in transmission costs and storage costs.
[0003] This requires highly efficient image compression techniques for effectively transmitting, storing, and reproducing high-resolution, high-quality image information. Summary of the Invention [Problem to be solved by the invention]
[0004] An object of the present disclosure is to provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.
[0005] Another object of the present disclosure is to provide an image encoding / decoding method and apparatus that improves the efficiency of encoding / decoding by efficiently signaling HRD parameters.
[0006] Another object of the present disclosure is to provide a method for transmitting a bitstream generated by the image encoding method or apparatus according to the present disclosure.
[0007] Another object of the present disclosure is to provide a recording medium storing a bitstream generated by the image encoding method or apparatus according to the present disclosure.
[0008] Another object of the present disclosure is to provide a recording medium storing a bitstream that is received by an image decoding device according to the present disclosure, decoded, and used to restore an image.
[0009] The technical problems to be solved by the present disclosure are not limited to the above-mentioned technical problems, and other technical problems not mentioned above will be clearly understood by a person having ordinary skill in the technical field to which the present disclosure pertains from the following description. [Means for solving the problem]
[0010] An image decoding method performed by an image decoding device according to one aspect of the present disclosure includes a step of determining whether a Reference Sample Resampling (RPR) operation is allowed, and a step of performing a Hypothetical Reference Decoder (HRD) operation based on whether the RPR operation is allowed, wherein the HRD operation includes at least one of an operation of determining conformance of a bitstream including encoded video data to a video coding standard, an operation of determining conformance of a video decoder to the video coding standard, or an HRD picture output timing signaling operation, and the HRD operation may be performed by updating a schedule index for a Hypothetical Stream Scheduler (HSS).
[0011] An image encoding method performed by an image encoding device according to one aspect of the present disclosure includes the steps of determining whether a Reference Sample Resampling (RPR) operation is allowed and encoding reference sample resampling allowance information, and encoding HRD parameters for a Hypothetical Reference Decoder (HRD) operation based on whether the RPR operation is allowed, wherein the HRD operation includes at least one of an operation of determining compliance of a bitstream including encoded video data with a video coding standard, an operation of determining compliance of a video decoder with the video coding standard, or an HRD picture output timing signaling operation, and the HRD operation may be performed by updating a schedule index for a Hypothetical Stream Scheduler (HSS).
[0012] A transmission method according to another aspect of the present disclosure can transmit a bitstream generated by the image encoding device or image encoding method of the present disclosure.
[0013] A computer-readable recording medium according to another aspect of the present disclosure can store a bitstream generated by the image encoding method or image encoding device of the present disclosure.
[0014] The features described above in this brief summary of the present disclosure are merely exemplary aspects of the detailed description of the present disclosure that follows and are not intended to limit the scope of the present disclosure. [Effects of the Invention]
[0015] According to the present disclosure, an image encoding / decoding method and apparatus with improved encoding / decoding efficiency can be provided.
[0016] Furthermore, the present disclosure can provide an image encoding / decoding method and apparatus that can improve the efficiency of encoding / decoding by efficiently signaling HRD parameters.
[0017] The present disclosure also provides a method for transmitting a bitstream generated by the image encoding method or apparatus according to the present disclosure.
[0018] Furthermore, according to the present disclosure, a recording medium storing a bitstream generated by the image encoding method or apparatus according to the present disclosure can be provided.
[0019] The present disclosure may also provide a recording medium storing a bitstream that is received by an image decoding device according to the present disclosure, decoded, and used to restore an image.
[0020] The effects obtained by the present disclosure are not limited to the effects described above, and other effects not described above will be clearly understood by those having ordinary skill in the art to which the present disclosure pertains from the following description. [Brief explanation of the drawings]
[0021] [Figure 1] 1 is a diagram illustrating a video coding system to which embodiments of the present disclosure can be applied; [Figure 2] 1 is a diagram schematically illustrating an image encoding device to which an embodiment of the present disclosure can be applied. [Figure 3] FIG. 1 is a diagram schematically illustrating an image decoding device to which an embodiment of the present disclosure can be applied. [Figure 4] 1 illustrates an example of a general picture decoding procedure to which the embodiments of the present disclosure can be applied. [Figure 5] 1 shows an example of a general picture encoding procedure to which the embodiments of the present disclosure can be applied. [Figure 6] FIG. 1 shows an example of a hierarchical structure for coded images / video. [Figure 7]FIG. 2 is a diagram illustrating an exemplary syntax structure of a VPS according to one embodiment of the present disclosure. [Figure 8] A diagram showing the syntax structure of an SPS for signaling HRD parameters according to one embodiment of the present disclosure. [Figure 9] FIG. 10 illustrates a general_hrd_parameter() syntax structure according to one embodiment of the present disclosure. [Figure 10] FIG. 10 illustrates an ols_hrd_parameters() syntax structure according to one embodiment of the present disclosure. [Figure 11] A diagram illustrating a sublayer_hrd_parameters() syntax structure according to one embodiment of the present disclosure. [Figure 12] FIG. 1 is a diagram illustrating an example of an image encoding method to which an embodiment of the present disclosure can be applied. [Figure 13] FIG. 1 is a diagram illustrating an example of an image decoding method to which an embodiment of the present disclosure can be applied. [Figure 14] FIG. 10 is a diagram for explaining another example of an image decoding method to which an embodiment of the present disclosure can be applied. [Figure 15] FIG. 10 is a diagram for explaining another example of an image decoding method to which an embodiment of the present disclosure can be applied. [Figure 16] FIG. 1 is a diagram illustrating an example of an image encoding method to which an embodiment of the present disclosure can be applied. [Figure 17] FIG. 1 illustrates a content streaming system to which an embodiment of the present disclosure can be applied. DETAILED DESCRIPTION OF THE INVENTION
[0022] The present disclosure will be described in detail below with reference to the accompanying drawings, so that those skilled in the art can easily implement the present disclosure. However, the present disclosure may be embodied in various different forms and is not limited to the embodiments described herein.
[0023] In describing the embodiments of the present disclosure, if it is determined that a detailed description of a known configuration or function may obscure the gist of the present disclosure, the detailed description thereof will be omitted. In addition, in the drawings, parts that are not related to the description of the present disclosure will be omitted, and similar parts will be designated by similar reference numerals.
[0024] In this disclosure, when a component is referred to as being "coupled," "coupled," or "connected" to another component, this includes not only a direct connection, but also an indirect connection where another component exists between them. Furthermore, when a component is referred to as "including" or "having" another component, this does not mean that the other component is excluded, but that the component can further include the other component, unless otherwise specified.
[0025] In this disclosure, terms such as "first" and "second" are used only to distinguish one component from another component, and do not limit the order or importance of the components unless otherwise specified. Therefore, within the scope of this disclosure, a first component in one embodiment may be referred to as a second component in another embodiment, and similarly, a second component in one embodiment may be referred to as a first component in another embodiment.
[0026] In this disclosure, components that are distinguished from one another are used to clearly describe the characteristics of each component and do not necessarily mean that the components are separate. In other words, multiple components may be integrated into a single hardware or software unit, or a single component may be distributed into multiple hardware or software units. Therefore, even if not otherwise specified, such integrated or distributed embodiments are also included within the scope of this disclosure.
[0027] In this disclosure, the components described in various embodiments are not necessarily essential components, and some may be optional components. Therefore, an embodiment consisting of a subset of the components described in one embodiment is also within the scope of this disclosure. Furthermore, an embodiment including other components in addition to the components described in various embodiments is also within the scope of this disclosure.
[0028] The present disclosure relates to image encoding and decoding, and terms used in this disclosure may have their ordinary meaning in the technical field to which the present disclosure belongs unless they are newly defined in this disclosure.
[0029] In this disclosure, a "picture" generally refers to a unit representing any one image in a specific time period, and a slice / tile is a coding unit constituting a part of a picture, and one picture may be composed of one or more slices / tiles. Furthermore, a slice / tile may include one or more coding tree units (CTUs).
[0030] In this disclosure, "pixel" or "pel" may refer to the smallest unit constituting one picture (or image). Also, "sample" may be used as a term corresponding to pixel. A sample may generally indicate a pixel or a pixel value, may indicate only a pixel / pixel value of a luma component, or may indicate only a pixel / pixel value of a chroma component.
[0031] In this disclosure, the term "unit" may refer to a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to that region. The term "unit" may be used interchangeably with terms such as "sample array," "block," or "area," depending on the situation. In general, an M×N block may include a set (or array) of samples or transform coefficients consisting of M columns and N rows.
[0032] In the present disclosure, a "current block" may refer to any one of a "current coding block," a "current coding unit," a "block to be coded," a "block to be decoded," or a "block to be processed." When prediction is performed, a "current block" may refer to a "current predicted block" or a "block to be predicted." When transformation (inverse transformation) / quantization (inverse quantization) is performed, a "current block" may refer to a "current transformed block" or a "block to be transformed." When filtering is performed, a "current block" may refer to a "block to be filtered."
[0033] Furthermore, in this disclosure, unless explicitly stated as a chroma block, the term "current block" may refer to a block including both a luma component block and a chroma component block, or to the "luma block of the current block." The luma component block of the current block may be expressed by explicitly including the term "luma block" or "current luma block." The chroma component block of the current block may be expressed by explicitly including the term "chroma block" or "current chroma block."
[0034] In the present disclosure, "A or B" can mean "A only," "B only," or "both A and B." In other words, in the present disclosure, "A or B" can be interpreted as "A and / or B." For example, in the present disclosure, "A, B or C" can mean "A only," "B only," "C only," or "any combination of A, B, and C."
[0035] As used in this disclosure, " / " and "," (comma) can mean "and / or." For example, "A / B" can mean "A and / or B." Thus, "A / B" can mean "A only," "B only," or "both A and B." For example, "A, B, C" can mean "A, B, or C."
[0036] In the present disclosure, "at least one of A and B" can mean "A only," "B only," or "both A and B." Also, in the present disclosure, the expressions "at least one of A or B" and "at least one of A and / or B" can be interpreted as being the same as "at least one of A and B."
[0037] Additionally, in this disclosure, "at least one of A, B, and C" can mean "A only," "B only," "C only," or "any combination of A, B, and C." Additionally, "at least one of A, B, or C" or "at least one of A, B, and / or C" can mean "at least one of A, B, and C."
[0038] Furthermore, parentheses used in the present disclosure may mean "for example." Specifically, when "prediction (intra prediction)" is displayed, "intra prediction" may be proposed as an example of "prediction." In other words, "prediction" in the present disclosure is not limited to "intra prediction," and "intra prediction" may be proposed as an example of "prediction." Furthermore, when "prediction (i.e., intra prediction)" is displayed, "intra prediction" may be proposed as an example of "prediction."
[0039] In the present disclosure, technical features described separately in one drawing may be realized separately or simultaneously.
[0040] Video Coding System Overview
[0041] FIG. 1 illustrates a video coding system according to this disclosure.
[0042] A video coding system according to one embodiment may include an encoding device 10 and a decoding device 20. The encoding device 10 may transmit encoded video and / or image information or data to the decoding device 20 in a file or streaming format via a digital storage medium or a network.
[0043] An encoding device 10 according to an embodiment may include a video source generation unit 11, an encoding unit 12, and a transmission unit 13. A decoding device 20 according to an embodiment may include a reception unit 21, a decoding unit 22, and a rendering unit 23. The encoding unit 12 may be referred to as a video / image encoding unit, and the decoding unit 22 may be referred to as a video / image decoding unit. The transmission unit 13 may be included in the encoding unit 12. The reception unit 21 may be included in the decoding unit 22. The rendering unit 23 may include a display unit, which may be configured as a separate device or an external component.
[0044] The video source generation unit 11 can acquire video / images through a video / image capture, synthesis, or generation process. The video source generation unit 11 can include a video / image capture device and / or a video / image generation device. The video / image capture device can include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. The video / image generation device can include, for example, a computer, a tablet, a smartphone, etc., and can (electronically) generate video / images. For example, virtual video / images can be generated via a computer, etc., in which case the video / image capture process can be replaced with a process in which related data is generated.
[0045] The encoder 12 may encode the input video / image. The encoder 12 may perform a series of steps such as prediction, transformation, and quantization for compression and coding efficiency. The encoder 12 may output the encoded data (encoded video / image information) in a bitstream format.
[0046] The transmitter 13 may transmit the encoded video / image information or data output in a bitstream format to the receiver 21 of the decoding device 20 in a file or streaming format via a digital storage medium or a network. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray®, HDD, and SSD. The transmitter 13 may include elements for generating a media file in a predetermined file format and elements for transmitting via a broadcasting / communication network. The receiver 21 may extract / receive the bitstream from the storage medium or network and transmit it to the decoder 22.
[0047] The decoding unit 22 can decode the video / image by performing a series of steps such as inverse quantization, inverse transformation, and prediction corresponding to the operations of the encoding unit 12.
[0048] The rendering unit 23 can render the decoded video / images, and the rendered video / images can be displayed via the display unit.
[0049] Overview of the image encoding device
[0050] FIG. 2 is a diagram schematically illustrating an image encoding device to which an embodiment of the present disclosure can be applied.
[0051] 2, the image encoding device 100 may include an image division unit 110, a subtraction unit 115, a transform unit 120, a quantization unit 130, an inverse quantization unit 140, an inverse transform unit 150, an addition unit 155, a filtering unit 160, a memory 170, an inter prediction unit 180, an intra prediction unit 185, and an entropy encoding unit 190. The inter prediction unit 180 and the intra prediction unit 185 may be collectively referred to as a "prediction unit." The transform unit 120, the quantization unit 130, the inverse quantization unit 140, and the inverse transform unit 150 may be included in a residual processing unit. The residual processing unit may further include a subtraction unit 115.
[0052] Depending on the embodiment, all or at least some of the components constituting the image encoding device 100 may be realized by a single hardware component (e.g., an encoder or a processor). Also, the memory 170 may include a decoded picture buffer (DPB) and may be realized by a digital storage medium.
[0053] The image division unit 110 may divide an input image (or picture, frame) input to the image encoding device 100 into one or more processing units. As an example, the processing units may be called coding units (CUs). The coding units may be obtained by recursively dividing a coding tree unit (CTU) or a largest coding unit (LCU) using a QT / BT / TT (quad-tree / binary-tree / ternary-tree) structure. For example, one coding unit may be divided into multiple coding units at deeper depths based on a quad-tree structure, a binary-tree structure, and / or a ternary-tree structure. To divide the coding units, the quad-tree structure may be applied first, and then the binary-tree structure and / or the ternary-tree structure may be applied later. The coding procedure according to the present disclosure may be performed based on the final coding unit that is not further divided. The maximum coding unit may be used as the final coding unit, or a lower-depth coding unit obtained by dividing the maximum coding unit may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and / or reconstruction, which will be described later. As another example, a processing unit of the coding procedure may be a prediction unit (PU) or a transform unit (TU). The prediction unit and the transform unit may be divided or partitioned from the final coding unit, respectively. The prediction unit may be a unit of sample prediction, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.
[0054] The prediction unit (inter prediction unit 180 or intra prediction unit 185) may perform prediction on a current block (current block) to generate a predicted block including prediction samples for the current block. The prediction unit may determine whether intra prediction or inter prediction is applied to the current block or CU. The prediction unit may generate various information related to prediction of the current block and transmit it to the entropy coding unit 190. The prediction information may be coded by the entropy coding unit 190 and output in a bitstream format.
[0055] The intra prediction unit 185 may predict the current block by referring to samples in the current picture. The referenced samples may be located in the neighborhood of the current block or may be located far away from the current block according to the intra prediction mode and / or intra prediction technique. The intra prediction modes may include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes may include, for example, a DC mode and a planar mode. The directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the granularity of the prediction direction. However, this is merely an example, and more or less directional prediction modes may be used depending on the settings. The intra prediction unit 185 may also determine the prediction mode to be applied to the current block using the prediction modes applied to neighboring blocks.
[0056] The inter prediction unit 180 may derive a predicted block for a current block based on a reference block (reference sample array) identified by a motion vector on a reference picture. To reduce the amount of motion information transmitted in inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation between the motion information of neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on the inter prediction direction (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block may be called a collocated reference block, a collocated CU (colCU), etc. The reference picture including the temporal neighboring block may be called a collocated picture (colPic). For example, the inter predictor 180 may construct a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to derive a motion vector and / or a reference picture index for the current block. Inter prediction may be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the inter predictor 180 may use motion information of neighboring blocks as motion information for the current block. In the case of skip mode, unlike in merge mode, a residual signal may not be transmitted.In the case of a motion vector prediction (MVP) mode, the motion vector of a neighboring block is used as a motion vector predictor, and the motion vector of the current block can be signaled by encoding a motion vector difference and an indicator for the motion vector predictor. The motion vector difference may mean the difference between the motion vector of the current block and the motion vector predictor.
[0057] The predictor may generate a prediction signal based on various prediction methods and / or prediction techniques, which will be described later. For example, the predictor may apply intra prediction or inter prediction to predict the current block, or may simultaneously apply intra prediction and inter prediction. A prediction method that simultaneously applies intra prediction and inter prediction to predict the current block may be referred to as combined inter and intra prediction (CIIP). The predictor may also perform intra block copy (IBC) to predict the current block. Intra block copy can be used for content image / video coding, such as screen content coding (SCC), for games. IBC is a method of predicting a current block using an already reconstructed reference block in a current picture that is located a predetermined distance away from the current block. When IBC is applied, the position of the reference block in the current picture may be coded as a vector (block vector) corresponding to the predetermined distance. IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that a reference block is derived within the current picture. That is, the IBC may use at least one of the inter prediction techniques described in this disclosure.
[0058] The prediction signal generated by the prediction unit may be used to generate a restored signal or a residual signal. The subtraction unit 115 may subtract the prediction signal (predicted block, predicted sample array) output from the prediction unit from the input image signal (original block, original sample array) to generate a residual signal (residual signal, residual block, residual sample array). The generated residual signal may be transmitted to the conversion unit 120.
[0059] The transform unit 120 may generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loeve transform (KLT), a graph-based transform (GBT), or a conditionally non-linear transform (CNT). Here, the GBT refers to a transform obtained from a graph representing inter-pixel relationship information. The CNT refers to a transform obtained based on a predicted signal generated using all previously reconstructed pixels. The transform process may be applied to pixel blocks having the same square size or to non-square blocks of variable size.
[0060] The quantization unit 130 may quantize the transform coefficients and transmit the quantized transform coefficients to the entropy coding unit 190. The entropy coding unit 190 may encode the quantized signal (information about the quantized transform coefficients) and output the encoded signal in a bitstream format. The information about the quantized transform coefficients may be referred to as residual information. The quantization unit 130 may rearrange the quantized transform coefficients in a block format into a one-dimensional vector format based on a coefficient scan order, and may generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector format.
[0061] The entropy coding unit 190 may perform various coding methods, such as exponential Golomb, context-adaptive variable length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy coding unit 190 may also code information necessary for video / image restoration (e.g., values of syntax elements) together with or separately from the quantized transform coefficients. The coded information (e.g., coded video / image information) may be transmitted or stored in a bitstream format in network abstraction layer (NAL) units. The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / image information may also include general constraint information. The signaling information, transmitted information and / or syntax elements mentioned in this disclosure may be encoded through the above-described encoding procedure and included in the bitstream.
[0062] The bitstream may be transmitted via a network or stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as a USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitting unit (not shown) that transmits and / or a storing unit (not shown) that stores the signal output from the entropy encoding unit 190 may be provided as an internal / external element of the image encoding device 100, or the transmitting unit may be provided as a component of the entropy encoding unit 190.
[0063] The quantized transform coefficients output from the quantization unit 130 can be used to generate a residual signal. For example, the residual signal (residual block or residual sample) can be reconstructed by applying inverse quantization and inverse transform to the quantized transform coefficients via the inverse quantization unit 140 and the inverse transform unit 150.
[0064] The adder 155 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter prediction unit 180 or the intra prediction unit 185. When there is no residual for the current block to be processed, such as when a skip mode is applied, the predicted block may be used as the reconstructed block. The adder 155 may be referred to as a reconstruction unit or a reconstructed block generation unit. The generated reconstructed signal may be used for intra prediction of the next current block to be processed in the current picture, and may also be used for inter prediction of the next picture after filtering, as will be described later.
[0065] The filtering unit 160 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 160 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture and store the modified reconstructed picture in the memory 170, specifically, in the DPB of the memory 170. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, etc. The filtering unit 160 may generate various information related to filtering and transmit it to the entropy coding unit 190, as will be described later in connection with each filtering method. The filtering information may be coded by the entropy coding unit 190 and output in a bitstream format.
[0066] The modified reconstructed picture transmitted to the memory 170 can be used as a reference picture in the inter prediction unit 180. When inter prediction is applied through this, the image encoding device 100 can avoid a prediction mismatch between the image encoding device 100 and the image decoding device, and can also improve encoding efficiency.
[0067] The DPB in the memory 170 may store modified reconstructed pictures for use as reference pictures in the inter predictor 180. The memory 170 may store motion information of blocks from which motion information in the current picture is derived (or coded) and / or motion information of already reconstructed intra-picture blocks. The stored motion information may be transmitted to the inter predictor 180 to be used as motion information of spatially surrounding blocks or temporally surrounding blocks. The memory 170 may store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra predictor 185.
[0068] Overview of the image decoding device
[0069] FIG. 3 is a diagram schematically illustrating an image decoding device to which an embodiment of the present disclosure can be applied.
[0070] 3, the image decoding apparatus 200 may include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an adder 235, a filtering unit 240, a memory 250, an inter prediction unit 260, and an intra prediction unit 265. The inter prediction unit 260 and the intra prediction unit 265 may be collectively referred to as a "prediction unit." The inverse quantization unit 220 and the inverse transform unit 230 may be included in a residual processing unit.
[0071] Depending on the embodiment, all or at least some of the components constituting the image decoding device 200 may be realized by a single hardware component (e.g., a decoder or a processor). Also, the memory 170 may include a DPB and may be realized by a digital storage medium.
[0072] The image decoding device 200, which receives a bitstream including video / image information, can reconstruct an image by performing a process corresponding to the process performed by the image encoding device 100 of FIG. 2. For example, the image decoding device 200 can perform decoding using a processing unit applied in the image encoding device. Therefore, the decoding processing unit can be, for example, a coding unit. The coding unit can be obtained by dividing a coding tree unit or a maximum coding unit. The reconstructed image signal decoded and output by the image decoding device 200 can be reproduced by a reproduction device (not shown).
[0073] The image decoding apparatus 200 may receive a signal output from the image encoding apparatus of FIG. 2 in a bitstream format. The received signal may be decoded via an entropy decoding unit 210. For example, the entropy decoding unit 210 may parse the bitstream to derive information (e.g., video / image information) necessary for image reconstruction (or picture reconstruction). The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / image information may also include general constraint information. The image decoding apparatus may further use the information on the parameter sets and / or the general constraint information to decode an image. The signaling information, received information, and / or syntax elements referred to in the present disclosure may be obtained from the bitstream by being decoded via the decoding procedure. For example, the entropy decoding unit 210 may decode information in a bitstream based on a coding method such as Exponential-Golomb coding, CAVLC, or CABAC, and output values of syntax elements required for image restoration and quantized values of transform coefficients related to residuals. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element from the bitstream, determines a context model using information on the syntax element to be decoded and decoded information on neighboring blocks and the block to be decoded, or information on symbols / bins decoded in a previous step, predicts the occurrence probability of the bins based on the determined context model, and performs arithmetic decoding of the bins to generate symbols corresponding to the values of each syntax element. After determining the context model, the CABAC entropy decoding method may update the context model using information on the decoded symbol / bin for the context model of the next symbol / bin.Among the information decoded by the entropy decoding unit 210, information related to prediction is provided to the prediction units (inter prediction unit 260 and intra prediction unit 265), and residual values entropy decoded by the entropy decoding unit 210, i.e., quantized transform coefficients and related parameter information, may be input to the inverse quantization unit 220. Also, among the information decoded by the entropy decoding unit 210, information related to filtering may be provided to the filtering unit 240. Meanwhile, a receiving unit (not shown) for receiving a signal output from the image encoding device may be further provided as an internal / external element of the image decoding device 200, or the receiving unit may be provided as a component of the entropy decoding unit 210.
[0074] Meanwhile, the image decoding apparatus according to the present disclosure may be referred to as a video / image / picture decoding apparatus. The image decoding apparatus may include an information decoder (video / image / picture information decoder) and / or a sample decoder (video / image / picture sample decoder). The information decoder may include an entropy decoding unit 210, and the sample decoder may include at least one of an inverse quantization unit 220, an inverse transform unit 230, an adder 235, a filtering unit 240, a memory 250, an inter prediction unit 260, and an intra prediction unit 265.
[0075] The inverse quantization unit 220 may inverse quantize the quantized transform coefficients and output the transform coefficients. The inverse quantization unit 220 may rearrange the quantized transform coefficients in a two-dimensional block format. In this case, the rearrangement may be performed based on the coefficient scanning order performed in the image encoding device. The inverse quantization unit 220 may perform inverse quantization on the quantized transform coefficients using a quantization parameter (e.g., quantization step size information) to obtain transform coefficients.
[0076] The inverse transform unit 230 can inversely transform the transform coefficients to obtain a residual signal (residual block, residual sample array).
[0077] The prediction unit may perform prediction on a current block and generate a predicted block including prediction samples for the current block. The prediction unit may determine whether intra prediction or inter prediction is applied to the current block based on information about the prediction output from the entropy decoding unit 210, and may determine a specific intra / inter prediction mode (prediction technique).
[0078] The prediction unit can generate a prediction signal based on various prediction methods (techniques) described below, as described in the description of the prediction unit of the image encoding device 100.
[0079] The intra predictor 265 may predict the current block by referring to samples in the current picture. The description of the intra predictor 185 may also be applied to the intra predictor 265.
[0080] The inter prediction unit 260 may derive a predicted block for a current block based on a reference block (reference sample array) identified by a motion vector on a reference picture. To reduce the amount of motion information transmitted in inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on correlations between motion information of neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on an inter prediction direction (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. For example, the inter prediction unit 260 may construct a motion information candidate list based on the neighboring blocks and derive a motion vector and / or a reference picture index for the current block based on received candidate selection information. Inter prediction may be performed based on various prediction modes (techniques), and the prediction information may include information indicating the inter prediction mode (technique) for the current block.
[0081] The adder 235 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to a prediction signal (predicted block, predicted sample array) output from a prediction unit (including the inter prediction unit 260 and / or intra prediction unit 265). When there is no residual for the current block, such as when a skip mode is applied, the predicted block can be used as the reconstructed block. The description of the adder 155 also applies to the adder 235. The adder 235 may also be referred to as a reconstruction unit or a reconstructed block generator. The generated reconstructed signal may be used for intra prediction of the next current block in the current picture, and may also be used for inter prediction of the next picture via filtering, as described below.
[0082] The filtering unit 240 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 240 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture, and may store the modified reconstructed picture in the memory 250, specifically, in a DPB of the memory 250. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, etc.
[0083] The (modified) reconstructed picture stored in the DPB of the memory 250 can be used as a reference picture in the inter predictor 260. The memory 250 can store motion information of a block from which motion information in the current picture is derived (or decoded) and / or motion information of a block in an already reconstructed picture. The stored motion information can be transmitted to the inter predictor 260 to be used as motion information of a spatially surrounding block or a temporally surrounding block. The memory 250 can store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra predictor 265.
[0084] In this specification, the embodiments described for the filtering unit 160, inter prediction unit 180 and intra prediction unit 185 of the image encoding device 100 can also be applied in a similar or corresponding manner to the filtering unit 240, inter prediction unit 260 and intra prediction unit 265 of the image decoding device 200, respectively.
[0085] General image / video coding procedures
[0086] In image / video coding, pictures constituting an image / video can be coded / decoded according to a sequence of decoding orders. A picture order corresponding to an output order of decoded pictures can be set to be different from the decoding order. Based on this, not only forward prediction but also backward prediction can be performed during inter prediction.
[0087] FIG. 4 shows an example of a general picture decoding procedure to which the embodiments of the present disclosure can be applied.
[0088] Each procedure shown in Fig. 4 may be performed by the image encoding apparatus of Fig. 3. For example, step S410 may be performed by the entropy decoding unit 210, step S420 may be performed by a prediction unit including an intra prediction unit 265 and an inter prediction unit 260, step S430 may be performed by a residual processing unit including an inverse quantization unit 220 and an inverse transform unit 230, step S440 may be performed by the addition unit 235, and step S450 may be performed by the filtering unit 240. Step S410 may include an information decoding procedure described in this disclosure, step S420 may include an inter / intra prediction procedure described in this disclosure, step S430 may include a residual processing procedure described in this disclosure, step S440 may include a block / picture reconstruction procedure described in this disclosure, and step S450 may include an in-loop filtering procedure described in this disclosure.
[0089] Referring to Figure 4, as shown in the description of Figure 3, the picture decoding procedure may generally include an image / video information acquisition procedure (S410) from a bitstream (through decoding), a picture reconstruction procedure (S420-S440), and an in-loop filtering procedure for the reconstructed picture (S450). The picture reconstruction procedure may be performed based on prediction samples and residual samples obtained through the inter / intra prediction (S420) and residual processing (S430, inverse quantization and inverse transform of quantized transform coefficients) processes described in this disclosure. A modified reconstructed picture may be generated through an in-loop filtering procedure for the reconstructed picture generated by the picture reconstruction procedure. The modified reconstructed picture may be output as a decoded picture or may be stored in a decoded picture buffer or memory 250 of the decoding device and used as a reference picture in the inter prediction procedure when decoding a subsequent picture. In some cases, the in-loop filtering procedure may be omitted. In this case, the reconstructed picture may be output as a decoded picture or may be stored in a decoded picture buffer or memory 250 of the decoding device and used as a reference picture in an inter-prediction procedure when decoding a subsequent picture. As described above, the in-loop filtering procedure (S450) may include a deblocking filtering procedure, a sample adaptive offset (SAO) procedure, an adaptive loop filter (ALF) procedure, and / or a bilateral filter procedure, some or all of which may be omitted. Furthermore, one or some of the deblocking filtering procedure, the sample adaptive offset (SAO) procedure, the adaptive loop filter (ALF) procedure, and the bilateral filter procedure may be applied sequentially, or all of them may be applied sequentially. For example, the SAO procedure may be performed after the deblocking filtering procedure is applied to the reconstructed picture.Or, for example, the ALF procedure can be performed after a deblocking filtering procedure is applied to the reconstructed picture, which can also be performed in the encoding device.
[0090] FIG. 5 shows an example of a general picture encoding procedure to which the embodiments of the present disclosure can be applied.
[0091] Each procedure shown in Fig. 5 may be performed by the image encoding apparatus of Fig. 2. For example, step S510 may be performed by a prediction unit including an intra prediction unit 185 or an inter prediction unit 180, step S520 may be performed by a residual processing unit including a transform unit 120 and / or a quantization unit 130, and step S530 may be performed by an entropy encoding unit 190. Step S510 may include an inter / intra prediction procedure described in this disclosure, step S520 may include a residual processing procedure described in this disclosure, and step S530 may include an information encoding procedure described in this disclosure.
[0092] Referring to FIG. 5, the picture encoding procedure, as described with reference to FIG. 2, may include not only a procedure of encoding information for picture reconstruction (e.g., prediction information, residual information, partitioning information, etc.) and outputting it in a bitstream format, but also a procedure of generating a reconstructed picture for a current picture and an optional procedure of applying in-loop filtering to the reconstructed picture. The encoding apparatus may derive (modified) residual samples from quantized transform coefficients via the inverse quantization unit 140 and the inverse transform unit 150, and may generate a reconstructed picture based on the prediction samples output in step S510 and the (modified) residual samples. The reconstructed picture generated in this manner may be the same as the reconstructed picture generated by the decoding apparatus described above. A modified reconstructed picture may be generated through an in-loop filtering procedure on the reconstructed picture, which may be stored in the decoded picture buffer or memory 170 and used as a reference picture in the inter prediction procedure when encoding a subsequent picture, as in the decoding apparatus. As described above, in some cases, some or all of the in-loop filtering procedure may be omitted. When the in-loop filtering procedure is performed, (in-loop) filtering-related information (parameters) can be coded by the entropy coding unit 190 and output in bitstream format, and the decoding device can perform the in-loop filtering procedure in the same manner as the coding device based on the filtering-related information.
[0093] This in-loop filtering procedure can reduce noise that occurs during image / video coding, such as blocking artifacts and ringing artifacts, and improve subjective / objective visual quality. Also, by performing the in-loop filtering procedure in both the encoding device and the decoding device, the encoding device and the decoding device can derive the same prediction result, thereby improving the reliability of picture coding and reducing the amount of data to be transmitted for picture coding.
[0094] As described above, a picture reconstruction procedure may be performed not only in a decoding device but also in an encoding device. Reconstructed blocks may be generated based on intra prediction / inter prediction for each block, and a reconstructed picture including the reconstructed blocks may be generated. If a current picture / slice / tile group is an I picture / slice / tile group, blocks included in the current picture / slice / tile group may be reconstructed based only on intra prediction. On the other hand, if the current picture / slice / tile group is a P or B picture / slice / tile group, blocks included in the current picture / slice / tile group may be reconstructed based on intra prediction or inter prediction. In this case, inter prediction may be applied to some blocks in the current picture / slice / tile group, and intra prediction may be applied to the remaining blocks. Color components of a picture may include luma components and chroma components, and unless explicitly limited in this disclosure, methods and embodiments proposed in this disclosure may be applied to luma components and chroma components.
[0095] Example of coding hierarchy and structure
[0096] Video / images coded according to this disclosure may be processed, for example, according to the coding hierarchy and structure described below.
[0097] FIG. 6 shows an example of a hierarchical structure for coded images / video.
[0098] The coded image / video can be divided into the VCL (video coding layer), which handles the image / video decoding process and itself, the lower system, which transmits and stores the coded information, and the NAL (network abstraction layer), which exists between the VCL and the lower system and is responsible for network adaptation functions.
[0099] The VCL can generate VCL data containing compressed image data (slice data), or it can generate parameter sets containing information such as a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), and a Video Parameter Set (VPS), or an SEI (Supplemental Enhancement Information) message that is additionally required for image decoding processing.
[0100] In NAL, NAL units can be generated by adding header information (NAL unit header) to RBSP (Raw Byte Sequence Payload) generated by VCL. RBSP refers to slice data, parameter sets, SEI messages, etc. generated by VCL. The NAL unit header can include NAL unit type information identified by the RBSP data included in the corresponding NAL unit.
[0101] As shown in Figure 6, NAL units can be classified into VCL NAL units and non-VCL NAL units according to the type of RBSP generated in the VCL. A VCL NAL unit can refer to a NAL unit containing information about an image (slice data), and a non-VCL NAL unit can refer to a NAL unit containing information necessary for decoding an image (parameter set or SEI message).
[0102] The VCL NAL unit and non-VCL NAL unit described above can be transmitted over a network with header information attached according to the data standard of the lower system. For example, the NAL unit can be transformed into a data format of a predetermined standard such as the H.266 / VVC file format, the Real-time Transport Protocol (RTP), or the Transport Stream (TS) and then transmitted over various networks.
[0103] As described above, the NAL unit type of an NAL unit can be identified according to the RBSP data structure included in the NAL unit, and information about the NAL unit type can be stored and signaled in the NAL unit header. For example, NAL units can be broadly classified into VCL NAL unit types and non-VCL NAL unit types depending on whether the NAL unit includes information about an image (slice data). VCL NAL unit types can be classified according to the nature and type of pictures included in the VCL NAL unit, and non-VCL NAL unit types can be classified according to the type of parameter set.
[0104] Below is a list of examples of NAL unit types identified by the type of parameter set / information included in the non-VCL NAL unit type.
[0105] -DCI (Decoding capability information) NAL unit type (NUT): Type for NAL units including DCI
[0106] -VPS (Video Parameter Set) NUT: Type for NAL units containing VPS
[0107] -SPS (Sequence Parameter Set) NUT: Type for NAL units containing SPS
[0108] -PPS (Picture Parameter Set) NUT: Type for NAL units containing PPS
[0109] -APS (Adaptation Parameter Set) NUT: Type for NAL units containing APS
[0110] -PH(Picture header) NUT: Type for NUL units containing picture headers
[0111] The above-mentioned NAL unit type has syntax information for the NAL unit type, and the syntax information can be stored in a NAL unit header and signaled. For example, the syntax information is nal_unit_type, and the NAL unit type can be identified using the value of nal_unit_type.
[0112] Meanwhile, one picture may include multiple slices, and one slice may include a slice header and slice data. In this case, one picture header may be added to multiple slices (slice header and slice data set) in one picture. The picture header (picture header syntax) may include information / parameters commonly applicable to the picture. The slice header (slice header syntax) may include information / parameters commonly applicable to the slices. The APS (APS syntax) or PPS (PPS syntax) may include information / parameters commonly applicable to one or more slices or pictures. The SPS (SPS syntax) may include information / parameters commonly applicable to one or more sequences. The VPS (VPS syntax) may include information / parameters commonly applicable to multiple layers. The DCI may include information / parameters related to decoding capability.
[0113] In the present disclosure, the high level syntax (HLS) may include at least one of the APS syntax, PPS syntax, SPS syntax, VPS syntax, DCI syntax, picture header syntax, and slice header syntax. Also, in the present disclosure, the low level syntax (LLS) may include, for example, slice data syntax, CTU syntax, coding unit syntax, transform unit syntax, etc.
[0114] Meanwhile, in the present disclosure, image / video information encoded from an encoding device to a decoding device and signaled in a bitstream format may include not only intra-picture partitioning-related information, intra / inter prediction information, residual information, in-loop filtering information, etc., but also the slice header information, the picture header information, the APS information, the PPS information, the SPS information, the VPS information, and / or the DCI information. Also, the image / video information may further include general constraint information and / or NAL unit header information.
[0115] High level syntax signaling and semantics
[0116] As described above, the image / video information according to the present disclosure may include a High Level Syntax (HLS), and an image encoding method and / or an image decoding method may be performed based on the image / video information.
[0117] Video Parameter Set signaling
[0118] A video parameter set (VPS) is a parameter set used for transmitting layer information. The layer information may include, for example, information on an output layer set (OLS), information on a profile tier level, information on a relationship between an OLS and a hypothetical reference decoder, information on a relationship between an OLS and a DPB, etc. The VPS may not be essential for decoding a bitstream.
[0119] Before it can be referenced, the VPS raw byte sequence payload (RBSP) must be available to the decoding process, either by being included in at least one Access Unit (AU) with TemporalID equal to 0, or by being provided via external means.
[0120] All VPS NAL units with a particular value of vps_video_parameter_set_id in a coded video sequence (CVS) must have the same content.
[0121] FIG. 7 is a diagram illustrating an exemplary syntax structure of a VPS according to one embodiment of the present disclosure.
[0122] The syntax structure of the VPS shown in FIG. 7 includes only syntax elements relevant to the present disclosure, and various other syntax elements not shown in FIG. 7 may be included in the VPS.
[0123] In the example shown in Figure 7, vps_video_parameter_set_id provides an identifier for the VPS. Other syntax elements can reference the VPS using vps_video_parameter_set_id. The value of vps_video_parameter_set_id must be greater than 0.
[0124] The value of vps_max_layers_minus1 plus 1 may indicate the maximum number of layers allowable within each CVS that references the VPS.
[0125] The value of vps_max_sublayers_minus1 plus 1 indicates the maximum number of temporal sublayers that can exist in a hierarchy within each CVS that references the VPS. vps_max_sublayers_minus1 can have a value of 0 to 6.
[0126] vps_all_layers_same_num_sublayers_flag may be signaled if vps_max_layers_minus1 is greater than 0 and vps_max_sublayers_minus1 is greater than 0. A first value (e.g., 1) of vps_all_layers_same_num_sublayers_flag may indicate that the number of temporal sublayers is the same for all layers in each CVS that references the VPS. A second value (e.g., 0) of vps_all_layers_same_num_sublayers_flag may indicate that layers in each CVS that references the VPS may not have the same number of temporal sublayers. When vps_all_layers_same_num_sublayers_flag is not present, its value may be inferred as the first value (e.g., 1).
[0127] vps_all_independent_layers_flag may be signaled if vps_max_layers_minus1 is greater than 0. A first value (e.g., 1) of vps_all_independent_layers_flag may indicate that all layers in the CVS are coded independently without using inter-layer prediction. A second value (e.g., 0) of vps_all_independent_layers_flag may indicate that one or more layers in the CVS can use inter-layer prediction. If vps_all_independent_layers_flag is not present, its value may be inferred as the first value (e.g., 1).
[0128] "each_layer_is_an_ols_flag" may be signaled when "vps_max_layers_minus1" is greater than 0. "each_layer_is_an_ols_flag" may also be signaled when "vps_all_independent_layers_flag" is a first value. "each_layer_is_an_ols_flag" with a first value (e.g., 1) may indicate whether each OLS includes only one layer. "each_layer_is_an_ols_flag" with a first value (e.g., 1) may indicate that each layer in the CVS that references the VPS is itself an OLS (i.e., the single layer included in the OLS is the only output layer). "each_layer_is_an_ols_flag" with a second value (e.g., 0) may indicate that at least one OLS can include more than one layer. If vps_max_layers_minus1 is 0, the value of each_layer_is_an_ols_flag can be inferred to be 1. Otherwise, if vps_all_independent_layers_flag is 0, the value of each_layer_is_an_ols_flag can be inferred to be 0.
[0129] If each_layer_is_an_ols_flag is a second value (e.g., 0) and vps_all_independent_layers_flag is a second value (e.g., 0), ols_mode_idc can be signaled.
[0130] An ols_mode_idc having a first value (e.g., 0) may indicate that the total number of OLSs identified by the VPS is equal to vps_max_layers_minus1+1. In this case, the i-th OLS may include layers with layer indexes from 0 to i. Also, for each OLS, only the layer with the highest layer index (top layer) within the OLS may be output.
[0131] A second value (e.g., 1) of ols_mode_idc may indicate that the total number of OLSs specified by the VPS is equal to vps_max_layers_minus1+1. In this case, the i-th OLS may include layers with layer indexes 0 to i. Furthermore, for each OLS, all layers within the OLS may be output.
[0132] A third value (e.g., 2) of ols_mode_idc can indicate that the total number of OLSs specified by the VPS is explicitly signaled. Additionally, for each OLS, the output layer can be explicitly signaled. The other layer, which is not the output layer, can be a direct reference layer or an indirect reference layer of the OLS output layer.
[0133] If vps_all_independent_layers_flag is 1 and each_layer_is_an_ols_flag is 0, the value of ols_mode_idc can be inferred to be a third value (e.g., 2).
[0134] If ols_mode_idc is 2, num_output_layer_sets_minus1 and ols_output_layer_flag[i][j] can be explicitly signaled.
[0135] The value of num_output_layer_sets_minus1 plus 1 can indicate the total number of OLSs specified by the VPS.
[0136] When ols_mode_idc is 2, ols_output_layer_flag[i][j] can indicate whether the jth layer of the ith OLS is the output layer. A first value (e.g., 1) of ols_output_layer_flag[i][j] can indicate that the layer with the same layer identifier (nuh_layer_id) as vps_layer_id[j] is the output layer of the ith OLS. A second value (e.g., 0) of ols_output_layer_flag[i][j] can indicate that the layer with the same layer identifier (nuh_layer_id) as vps_layer_id[j] is not the output layer of the ith OLS.
[0137] The HRD parameters signaled by the VPS are described below.
[0138] If each_layer_is_an_ols_flag is a second value (e.g., 0), vps_general_hrd_params_present_flag can be signaled. A first value (e.g., 1) of vps_general_hrd_params_present_flag can indicate that HRD parameters different from the general_hrd_parameters() syntax structure are present in the VPS. A second value (e.g., 0) of vps_general_hrd_params_present_flag can indicate that HRD parameters different from the general_hrd_parameters() syntax structure are not present in the VPS. If vps_general_hrd_params_present_flag is not present, its value can be inferred as the second value (e.g., 0).
[0139] If the i-th OLS contains one stratum (NumLayersInOls[i] is equal to 1), the general_hrd_parameters() syntax structure that applies to the i-th OLS may be present in the Sequence Parameter Set (SPS) referenced by the stratum within the i-th OLS.
[0140] vps_sublayer_cpb_params_present_flag can be signaled if vps_max_sublayers_minus1 is greater than 0. A first value (e.g., 1) of vps_sublayer_cpb_params_present_flag can indicate that the i-th ols_hrd_parameters() syntax structure in the VPS includes HRD parameters for a sublayer whose temporal layer identifier (TemporalId) is 0 through hrd_max_tid[i]. A second value (e.g., 0) of vps_sublayer_cpb_params_present_flag can indicate that the i-th ols_hrd_parameters() syntax structure in the VPS includes HRD parameters for a sublayer whose temporal layer identifier (TemporalId) is only hrd_max_tid[i]. If vps_max_sublayers_minus1 is 0, vps_sublayer_cpb_params_present_flag can be inferred to be a second value (e.g., 0).
[0141] If vps_sublayer_cpb_params_present_flag is a second value (e.g., 0), it can be inferred that the HRD parameters for sublayers whose temporal layer identifier (TemporalId) is 0 to hrd_max_tid[i]-1 are identical to the HRD parameters for sublayers whose temporal layer identifier (TemporalId) is hrd_max_tid[i].
[0142] The value of num_ols_hrd_params_minus1 plus 1 may indicate the number of ols_hrd_parameters() syntax structures in the VPS. num_ols_hrd_params_minus1 may have a value from 0 to TotalNumOlss-1. TotalNumOlss may indicate the total number of OLSs identified by the VPS. In this disclosure, HRD parameters may refer to ols_hrd_parameters(). Therefore, the number of HRD parameter syntax structures may refer to the number of ols_hrd_parameters() syntax structures.
[0143] hrd_max_tid[i] may be signaled if vps_max_sublayers_minus1 is greater than 0 and vps_all_layers_same_num_sublayers_flag is a second value (e.g., 0). hrd_max_tid[i] may indicate the temporal layer identifier (TemporalId) of the highest sublayer whose associated HRD parameters are included in the i-th ols_hrd_parameters() syntax structure.
[0144] hrd_max_tid[i] can have a value between 0 and vps_max_sublayers_minus1. If vps_max_sublayers_minus1 is 0, the value of hrd_max_tid[i] can be inferred to be 0. If vps_max_sublayers_minus1 is greater than 0 and vps_all_layers_same_num_sublayers_flag is 1, the value of hrd_max_tid[i] can be inferred to be equal to vps_max_sublayers_minus1.
[0145] As shown in Figure 7, the variable firstSubLayer indicating the temporal layer identifier (TemporalId) of the first sublayer can be induced to 0 or hrd_max_tid[i] based on vps_sublayer_cpb_params_present_flag. Specifically, if vps_sublayer_cpb_params_present_flag is 1, firstSubLayer can be induced to 0; otherwise, firstSubLayer can be induced to hrd_max_tid[i]. Based on the induced firstSubLayer and hrd_max_tid[i], the ols_hrd_parameters() syntax structure can be signaled.
[0146] If the value of num_ols_hrd_params_minus1 plus 1 is not equal to TotalNumOlss and num_ols_hrd_params_minus1 is greater than 0, ols_hrd_idx[i] can be signaled. In this case, ols_hrd_idx[i] can be signaled for the i-th OLS if the number of layers included in the i-th OLS (NumLayersInOls[i]) is greater than 1. ols_hrd_idx[i] is an index to the list of ols_hrd_parameters() in the VPS and can be the index of ols_hrd_parameters() applied to the i-th OLS. ols_hrd_idx[i] can have a value from 0 to num_ols_hrd_params_minus1. If the number of strata in the i-th OLS (NumLayersInOls[i]) is 1, then the ols_hrd_parameters() syntax structure that applies to the i-th OLS can be present in an SPS referenced by a strata in the i-th OLS.
[0147] In the present disclosure, ols_hrd_idx[i] is the index of ols_hrd_parameters() applied to the i-th OLS or the i-th multi-layer OLS, and can be referred to as mapping information (information about the mapping) between the (multi-layer) OLS and the HRD parameter syntax structure (ols_hrd_parameters()).
[0148] If num_ols_hrd_param_minus1 plus 1 is equal to TotalNumOlss, the value of ols_hrd_idx[i] can be inferred to be equal to i. Otherwise, if NumLayersInOls[i] is greater than 1 and num_ols_hrd_params_minus1 is 0, the value of ols_hrd_idx[i] can be inferred to be 0.
[0149] HRD signaling in VPS and SPS
[0150] Signaling of HRD parameters according to the present disclosure will be described in more detail below. HRD parameters can be signaled for each output layer set (OLS). A hypothetical reference decoder (HRD) is a hypothetical decoder model that specifies limitations on the variability of a conforming NAL unit stream or a conforming byte stream that can be generated during the encoding process.
[0151] The HRD parameters may be signaled in the VPS as described with reference to Figure 7. Alternatively, the HRD parameters may be signaled in the SPS.
[0152] FIG. 8 is a diagram illustrating a syntax structure of an SPS for signaling HRD parameters according to one embodiment of the present disclosure.
[0153] In the example shown in Figure 8, a first value (e.g., 1) of sps_ptl_dpb_hrd_params_present_flag may indicate that the profile_tier_level() syntax structure and the dpb_parameters() syntax structure are present in the SPS. The profile_tier_level() may be a syntax structure for transmitting parameters for the profile tier level, and the dpb_parameters() may be a syntax structure for transmitting DPB (decoded picture buffer) parameters. Furthermore, a first value (e.g., 1) of sps_ptl_dpb_hrd_params_present_flag may indicate that the general_hrd_parameters() syntax structure and the ols_hrd_parameters() syntax structure may be present in the SPS. A second value (e.g., 0) of sps_ptl_dpb_hrd_params_present_flag may indicate that the four syntax structures are not present in the SPS. The value of sps_ptl_dpb_hrd_params_present_flag may be the same as the value of vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]], i.e., the value of sps_ptl_dpb_hrd_params_present_flag may be encoded as the value of vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]].
[0154] In the above, vps_independent_layer_flag[i] may be a syntax element transmitted in a VPS. A first value (e.g., 1) of vps_independent_layer_flag[i] may indicate that the layer with index i is an independent layer that does not use inter-layer prediction. A second value (e.g., 0) of vps_independent_layer_flag[i] may indicate that the layer with index i can use inter-layer prediction. If vps_independent_layer_flag[i] does not exist, its value may be inferred as the first value (e.g., 1).
[0155] If sps_ptl_dpb_hrd_params_present_flag is 1, sps_general_hrd_params_present_flag can be signaled.
[0156] A first value (e.g., 1) of sps_general_hrd_params_present_flag may indicate that the SPS includes the general_hrd_parameters() and ols_hrd_parameters() syntax structures. A second value (e.g., 0) of sps_general_hrd_params_present_flag may indicate that the SPS does not include the general_hrd_parameters() or ols_hrd_parameters() syntax structures.
[0157] 8, sps_sublayer_cpb_params_present_flag may be signaled if sps_max_sublayers_minus1 is greater than 0. In this case, a value of sps_max_sublayers_minus1 plus 1 may indicate the maximum number of temporal sublayers that may exist in each coded layer video sequence (CLVS) that references the SPS. A first value (e.g., 1) of sps_sublayer_cpb_params_present_flag may indicate that the ols_hrd_parameters() syntax structure in the SPS includes HRD parameters for sublayers whose temporal layer identifiers (TemporalId) are 0 to sps_max_sublayers_minus1. A second value (e.g., 0) of sps_sublayer_cpb_params_present_flag may indicate that the ols_hrd_parameters() syntax structure in the SPS includes HRD parameters for sublayers whose temporal layer identifier (TemporalId) is only sps_max_sublayers_minus1. If sps_max_sublayers_minus1 is 0, the value of sps_sublayer_cpb_params_present_flag may be inferred as the second value (e.g., 0).
[0158] When sps_sublayer_cpb_params_present_flag is a second value (e.g., 0), the HRD parameters for sublayers whose temporal layer identifier (TemporalId) is 0 to sps_max_sublayers_minus1-1 can be inferred to be equal to the HRD parameters for sublayers whose temporal layer identifier (TemporalId) is sps_max_sublay_minus1.
[0159] FIG. 9 is a diagram illustrating a general_hrd_parameters() syntax structure according to one embodiment of the present disclosure.
[0160] As shown in Figure 9, the general_hrd_parameters() syntax structure can contain some of the sequence-level HRD parameters used for HRD operations. As a requirement for bitstream consistency, the contents of general_hrd_parameters() present in the VPS or SPS in the bitstream must be the same.
[0161] When the general_hrd_parameters() syntax structure is included in a VPS, it can be applied to all OLSs specified by the VPS. When the general_hrd_parameters() syntax structure is included in an SPS, it can be applied to an OLS that includes only the lowest layer among the layers that reference the SPS. In this case, the lowest layer may be an independent layer.
[0162] As shown in Fig. 9, the general_hrd_parameters() syntax structure can include syntax elements such as num_units_in_tick, time_scale, and general_nal_hrd_params_present_flag as HRD parameters. The HRD parameters shown in Fig. 9 may have the same meaning as conventional HRD parameters. Therefore, detailed descriptions of HRD parameters that are less relevant to the present disclosure will be omitted.
[0163] FIG. 10 is a diagram illustrating an ols_hrd_parameters() syntax structure according to one embodiment of the present disclosure.
[0164] When the ols_hrd_parameters() syntax structure is included in a VPS, the OLSs to which the ols_hrd_parameters() syntax structure is applied can be identified by the VPS. When the ols_hrd_parameters() syntax structure is included in an SPS, the ols_hrd_parameters() syntax structure can be applied to an OLS that includes only the lowest layer among layers that reference the SPS. In this case, the lowest layer may be an independent layer.
[0165] As shown in Figure 10, the ols_hrd_parameters() syntax structure can include syntax elements such as fixed_pic_rate_general_flag, fixed_pic_rate_within_cvs_flag, and elemental_duration_in_tc_minus1 as HRD parameters. The HRD parameters shown in Figure 10 may have the same meaning as conventional HRD parameters. Therefore, detailed descriptions of HRD parameters that are less relevant to the present disclosure will be omitted.
[0166] FIG. 11 is a diagram illustrating a sublayer_hrd_parameters() syntax structure according to one embodiment of the present disclosure.
[0167] The sublayer_hrd_parameters() syntax structure can be signaled by being included in the ols_hrd_parameters() syntax structure of FIG.
[0168] As shown in Fig. 11, the sublayer_hrd_parameters() syntax structure may include syntax elements such as bit_rate_value_minus1, cpb_size_value_minus1, and cpb_size_du_value_minus1 as HRD parameters. The HRD parameters shown in Fig. 11 may have the same meaning as conventional HRD parameters. Therefore, detailed descriptions of HRD parameters that are less relevant to the present disclosure will be omitted.
[0169] For reference, the output time may refer to the time at which a picture restored from a DPB is output, and may be specified by the HRD according to the output timing DPB operation.
[0170] Two sets of HRD parameters are available: NAL HRD parameters and VCL HRD parameters. The HRD parameters can be signaled via the general_hrd_parameters() syntax structure and the ols_hrd_parameters() syntax structure. The general_hrd_parameters() syntax structure and the ols_hrd_parameters() syntax structure can be signaled by being included in the VPS or the SPS.
[0171] For example, DPB management can be performed based on the HRD parameters, such as removing pictures from the DPB before decoding the current picture and / or outputting (decoded) pictures based on the HRD parameters.
[0172] In conventional image coding technology, for a hypothetical reference decoder (HRD) operation and a conformance test, a hypothetical stream scheduler (HSS), which is an entity that manages the HRD and conformance test, selects a specific ScIdx (schedule index) at the beginning of the operation / test to determine the HRD structure to be used for the operation. This schedule index is used from the beginning of the bitstream to the end for some operations. In the prior art, the HSS can change the schedule index when: if the contents of the selected previous general_timing_hrd_parameters() syntax structure for the AU containing AU m are different from those of the previous AU, the HSS selects the scIdx1 value of scIdx from the value of ScIdx provided in the selected general_timing_hrd_parameters() syntax structure for the AU containing AU m that results in BitRate[Htid][ScIdx1] or CpbSize[Htid][ScIdx1] for the AU containing AU m. The value of BitRate[Htid][ScIdx1] or CpbSize[Htid][ScIdx1] may be different from the value of BitRate[Htid][ScIdx0] or CpbSize[Htid][ScIdx0] for the ScIdx0 value of ScIdx used for the previous AU. Otherwise, the HSS continues to operate with the previous values of ScIdx, BitRate[Htid][ScIdx], and CpbSize[Htid][ScIdx]. This case can be applied to splicing cases. When two bitstreams are spliced and merged together, the HRD and general timing parameters may need to be updated because the bitstreams may undergo major changes starting from the splicing point. When splicing occurs, conditions related to changing the picture resolution may change.Similarly, it would be preferable to enable the HSS to change the selected schedule index when the picture resolution is changed not by splicing but by using the RPR function. However, in the prior art, such a function (e.g., the function to change the selected schedule index when the resolution is changed) is not defined / allowed, which may cause problems.
[0173] Furthermore, HRD operations and timing signaling have various problems. First, when reference picture resampling (RPR) is applicable, timing signaling and operations can become problematic. When coded media / video is transmitted / delivered from an encoder to a decoder, currently, HRD-related operations and timing signaling are performed under the assumption that certain conditions (e.g., network bandwidth) do not change within a coded video sequence (CVS) or coded layered video sequence (CLVS). That is, timing signaling exists at the beginning of a CVS, i.e., an access unit (AU) carrying an IRAP picture or a GDR picture, and conformance testing is initialized only at the beginning of the CVS. While this assumption (i.e., that transmission conditions do not change within a CVS) is reasonable in most cases, when the RPR function is enabled, the size of the coded picture may change in the middle of a CVS due to changes in certain conditions (e.g., changes in network bandwidth). For example, a picture coded at 1920x1080 resolution at the beginning of the CVS may be reduced to 1280x720 resolution at the CVS intermediate due to reduced network bandwidth. In such a situation, it may be necessary to allow re-initialization of timing operation calculations at the CVS intermediate when the picture size is changed (e.g., conformance test re-initialization for CPB timing, including arrival and removal timing, and DPB timing, including removal timing). Alternatively, it may be necessary to provide updated parameter / timing information instead of re-initialization.
[0174] Moreover, according to the prior art, problems may arise in calculating the number of conformance tests.
[0175] In order to solve the problems of the prior art, the present application proposes the following embodiments:
[0176] Configuration 1: Once DU-based conformance is permitted, DU-based tests can be applied to all test points of two types (i.e., Type 1 and Type 2) of tests starting at each CRA, each IDR, each GDR, and each DRAP. The number of bitstream conformance test runs can be: n0*n1*(2*n2+n3+n4)*n5. As an example, n0, n1, n2, n3, n4, and n5 can be the same as those described in the prior art, including VVC standard documents. As an example, n0 can be 2, n1 can be hrd_cpb_cnt_minus1+1, and n2 can be the number of BitstreamToDecode AUs associated with the BP SEI message applicable to each TargetOlsIdx, where the following conditions are all true:
[0177] -nal_unit_type is CRA_NUT.
[0178] - The bp_alt_cpb_params_present_flag of the associated BP SEI message is 1.
[0179] - There is one or more RASL pictures with pps_mixed_nalu_types_in_pic_flag set to 0 associated with the AU.
[0180] n3 may be the number of BitstreamToDecode IRAP or GDR AUs associated with the BP SEI message applicable to TargetOlsIdx, respectively, for which one or more of the following conditions are false:
[0181] -nal_unit_type is CRA_NUT.
[0182] - The bp_alt_cpb_params_present_flag of the associated BP SEI message is 1.
[0183] - There is one or more RASL pictures with pps_mixed_nalu_types_in_pic_flag set to 0 associated with the AU.
[0184] n4 may be the number of BitstreamToDecode AUs that are respectively associated with a DRAP indication SEI message applicable to TargetOlsIdx and whose associated PT SEI message has pt_cpb_alt_timing_info_present_flag equal to 1.
[0185] n5 can be derived as follows:
[0186] - If General_du_hrd_params_present_flag of the selected General_timing_hrd_parameters() syntax structure is 0, then n5 can be 1.
[0187] Otherwise, n5 may be 2.
[0188] Configuration 2: When RPR is applicable, the parameters and timing information updated in the conformance test / timing operation can be used to update the timing guidance of the decoding operation (e.g., arrival of a picture in the CPB, removal of a picture from the CPB, removal of a picture from the DPB, etc.).
[0189] Configuration 3: Updating of parameters and timing information, timing operations, and reinitialization of conformance tests can be done mid-CVS at pictures that satisfy the following:
[0190] - Picture with a.temporalId equal to 0
[0191] b. A picture that references one or more reference pictures with different resolutions / picture sizes
[0192] Configuration 4: For HRD operation and conformance testing, the selected schedule index can be updated or changed in the middle of CVS or CLVS.
[0193] Configuration 5: For HRD operation and conformance testing, the selected schedule index can be changed or updated in the AU where a major bit rate change occurs (i.e., the coded data of the current AU is larger or smaller than that of the previous AU).
[0194] Configuration 6: In pictures with temporalId=0, it is possible to further restrict the allowable schedule index change to a picture in the middle of a CVS / CLVS where a major bitrate change occurs.
[0195] Configuration 7: It can be further restricted so that a change in schedule index in the middle of a CVS / CLVS is allowed only if a change in picture resolution (i.e., RPR) within the CLVS is allowed.
[0196] Configuration 8: Additional HRD parameters may be present in association with some pictures and may be used to update or reinitialize conformance testing / timing operations.
[0197] Configuration 9: Additional HRD parameters can be present in a PPS (picture parameter set) referenced by a picture.
[0198] Configuration 10: Additional HRD parameters may be present in the picture header of a picture.
[0199] Configuration 11: There may be a flag that specifies whether additional HRD parameters are present in the PPS or picture header.
[0200] However, the above configurations are merely described separately for clarity of explanation, and each configuration can be used in combination.
[0201] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.
[0202] Each conformance test can consist of a combination of one option selected for each step. If there is more than one option for a step, only one option may be selected for a particular conformance test. All possible combinations of all steps can constitute the total conformance test set, and the number of bitstream conformance tests performed for each operating point under test can be n0*n1*(2*n2+n3+n4)*n5, where the values of n0, n1, n2, n3, n4, and n5 can be specified as follows:
[0203] -n0 can be 2.
[0204] -n1 can be hrd_cpb_cnt_minus1+1.
[0205] -n2 may be the number of AUs in BitstreamToDecode that are associated with the BP SEI message applicable to TargetOlsIdx, respectively, and for which all of the following conditions are true:
[0206] -nal_unit_type can be CRA_NUT.
[0207] The bp_alt_cpb_params_present_flag of the associated BP SEI message may be 1.
[0208] There may be one or more RASL pictures with pps_mixed_nalu_types_in_pic_flag equal to 0 associated with the AU.
[0209] - n3 may be the number of IRAP or GDR AUs in BitstreamToDecode, each associated with a BP SEI message applicable to TargetOlsIdx, for which one or more of the following conditions are false:
[0210] 1. nal_unit_type can be CRA_NUT.
[0211] 2. The bp_alt_cpb_params_present_flag of the associated BP SEI message may be 1.
[0212] 3. There can be one or more RASL pictures with pps_mixed_nalu_types_in_pic_flag=0 associated with the AU.
[0213] -n4 may be the number of BitstreamToDecode AUs that are associated with the DRAP indication SEI message applicable to TargetOlsIdx and whose associated PT SEI message has pt_cpb_alt_timing_info_present_flag equal to 1.
[0214] -n5 can be derived as follows:
[0215] 1. If the General_du_hrd_params_present_flag of the selected General_timing_hrd_parameters() syntax structure is 0, then n5 can be 1.
[0216] 2. Otherwise, n5 can be 2.
[0217] Meanwhile, n0 may correspond to a conformance test for Type I bitstream conformance and Type II bitstream conformance. n1 may correspond to a conformance test for each CPB delivery schedule. n2 may correspond to a conformance test for a bitstream starting from each CRA picture with an associated RASL picture and an alternative initial CPB removal delay and delay offset. This test may be performed twice: once for maintaining the bitstream and once for the bitstream that removes the RASL picture associated with the CRA. n3 may correspond to a conformance test for a bitstream starting from each IRAP or GDR AU, rather than a CRA with an associated RASL picture and an alternative initial CPB removal delay and delay offset. n4 corresponds to a conformance test for a bitstream starting from each IRAP having an associated DRAP picture with alternative timing information, and may correspond to the result of removing all AUs between the DRAP picture with alternative timing information and the previous IRAP. When General_du_hrd_params_present_flag is 1, n5 can correspond to AU-based conformance and DU-based conformance tests.
[0218] For example, the following steps can be applied to each test (e.g., conformance test), although some steps may be omitted, other steps may be added, or the order in which the steps are performed may be changed.
[0219] First, the operation point to be tested, represented by targetOp, can be selected by selecting a list for the OLS index opOlsIdx, the highest TemporalId value opTid, and the target sub-picture index value opSubpicIdxList[j] in the range from 0 to NumLayersIdOls[OlsIdx]-1. The opOlsIdx value can be in the range from 0 to TotalNumOlss-1, and the opTid value can be in the range from 0 to vps-max-sublayers-minus1. As an example, if opSubpicIdxList[] is not present, targetOp consists of a picture, but each pair of selected opOlsIdx and opTid values can invoke a sub-bitstream extraction process with enitreBitstream, opOlsIdx, and opTid as inputs, so that the output sub-bitstream BitstreamToDecode satisfies the following condition:
[0220] There is at least one VCL NAL unit in BitstreamToDecode whose TempiralId is equal to opTid. Otherwise (if opSubpicIdxList[] is present), targetOp consists of a subpicture, and the set of selected values of opOlsIdx, opTid, and opSubpicIdxList[j], for j in the range from 0 to numLayersInOls[opOlsIdx]-1, must satisfy the following condition: The subordinate bitstream output by the subpicture sub-bitstream extraction process BitStreamToDecode, invoked with inputs entireBitstream, opOlsIdx, opTid, and opSubpicIdxList[j], for j in the range from 0 to numLayersInOls[opOlsIdx]-1, must satisfy the following condition:
[0221] -BitstreamToDecode has one or more VCL NAL units with TemporalId equal to opTid.
[0222] - nuh_layer_id is the same as LayerIdInOls[opOlsIdx][j] and there is one or more VCL units whose sh_subpic_id is the same as the SubpicIdVal[opSubpicIdxList[j] value for each j in the range NumLayerInOls[opOlsIdx]-1.
[0223] Here, regardless of whether opSubpicIdxList[] exists or not, there can be at least one VCL NAL unit with nuh_layer_id equal to layerIdInOls[j] for each j in the range from 0 to NumLayerInOls[OpOlIdx]-1, in order to complete the bitstream conformance requirements of each IRAP or GDR UA.
[0224] Also, if opSubpicIdxList[] is not present, the following applies:
[0225] - If the layer of targetOp includes all layers of the entire bitstream and opTid is the same as the highest TemporalId value of all NAL units of the entire bitstream, BitsretamToDecode can be set to the same as the entireBitstream.
[0226] - Otherwise, the sub-bitstream extraction process is called with entireBitstream, opOlsIdx and opTid as inputs, and BitstreamToDecode can be output.
[0227] Otherwise (if opSubpicIdxList[] exists), the subpicture sub-bitstream extraction process is called with opSubpicIdxList[j], entireBitstream, opOlsIdx, and opTid as inputs, for j in the range from 0 to NumLayerSinOls[opOlsIdx]-1, and BitstreamToDecode can be output.
[0228] Furthermore, the values of TargetOlsIdx and Htid can be set to be the same as the opOlsIdx and opTid of targetOp, respectively.
[0229] The general_timing_hrd_parameters() syntax structure, the ols_timing_hrd_parameters() syntax structure, and the sublayer_hrd_parameters() syntax structure applicable to BitstreamToDecode can be selected as follows:
[0230] If NumLayersInOls[TargetOlsIdx] is 1, the general_timing_hrd_parameters() syntax structure of SPS and the ols_timing_hrd_parameters() syntax structure can be selected. Otherwise, the general_timing_hrd_parameters() syntax structure and the vps_ols_timing_hrd_idx[MultiLayerOlsIdx[TargetOlsIdx]-th ols_timing_hrd_parameters() syntax structure can be selected.
[0231] -To test a type I bitstream conformance point within a selected ols_timing_hrd_parameters() syntax structure, the sublayer_hrd_parameters(Htid) syntax structure that immediately follows the condition 'if(general_vcl_hrd_params_present_flag) (determine whether general_vcl_hrd_params_present_flag is 1)' is selected and the variable NalHrdModeFlag is set to 0; to test a type II bitstream conformance point, the sublayer_hrd_parameters(Htid) syntax structure that immediately follows the condition 'if(general_nal_hrd_params_present_flag) (determine whether general_nal_hrd_params_present_flag is 1)' is selected and the variable NalHrdModeFlag is set to 1. If BitstreamToDecode is a type II bitstream and NalHrdModeFlag is 0, all non-VCL NAL units except PH and filler data NAL units, and all leading_zero_8bits, zero_byte, start_code_prefix_one_3bytes, and trailing_zero_8bits syntax elements, if present, are discarded in BitstreamToDecode, and the remaining bitstream can be allocated to BitstreamTodecode.
[0232] The AU associated with the BP SEI message applicable to the TargetOp (which may be present in BitstreamToDecode or available via external means) may be selected as the HRD initialization point and named AU0.
[0233] When General_du_hrd_params_present_flag is 1 in the selected General_timing_hrd_parameters() syntax structure, the CPB can be scheduled to operate at the AU level (in this case, the DecodingUnitHrdFlag variable is set to 0) or at the DU level (in this case, the DecodingUnitHrdFlag variable is set equal to 1). Otherwise, DecodingUnitHrdFlag is set to 0 and the CPB can be scheduled to operate at the AU level.
[0234] For each AU in BitstreamToDecode starting from AU0, a BP SEI message (present in BitstreamToDecode or available via external means) associated with the AU and applied to TargetOlsIdx is selected, a PT SEI message (present in BitstreamToDecode or available via external means) associated with the AU and applied to TargetOlsIdx is selected, and if DecodingUnitHrdFlag is 1 and bp_du_cpb_params_in_pic_timing_sei_flag is 0, a DUI SEI message (present in BitstreamToDecode or available via external means not specified in this specification) associated with the DU of the AU and applied to TargetOlsIdx can be selected.
[0235] The ScIdx value can be selected. The selected ScIdx must be in the range of 0 to hrd_cpb_cnt_minus1.
[0236] When the BP SEI message associated with AU0 has bp_alt_cpb_params_present_flag equal to 0, the variable DefaultInitCpbParamsFlag is set equal to 1. When the BP SEI message associated with AU0 has bp_alt_cpb_params_present_flag equal to 1, for the initial CPB removal delay and delay offset selection, one of the following is applicable for selection:
[0237] - If NalHrdModeFlag is 1, the default initial CPB removal delay and delay offset, represented by bp_nal_initial_cpb_removal_delay[Htid][ScIdx] and bp_nal_initial_cpb_removal_offset[Htid][ScIdx] respectively in the selected BP SEI message, can be selected. Otherwise, the default initial CPB removal delay and delay offset, represented by bp_vcl_initial_cpb_removal_delay[Htid][ScIdx] and bp_vcl_initial_cpb_removal_offset[Htid][ScIdx] respectively in the selected BP SEI message, can be selected. The DefaultInitCpbParamsFlag variable can be set to 1.
[0238] -If NalHrdModeFlag is 1, the alternative initial CPB removal delay and delay offset _delta[Htid][ScIdx] and pt_nal_cpb_alt_initial_removal_offset_delta[Htid][ScIdx] indicated by bp_nal_initial_cpb_removal_delay[Htid][ScIdx] and bp_nal_initial_cpb_removal_offset[Htid][scIdx] respectively in the selected BP SEI message and pt_nal_cpb_alt_initial_removal_delay can be selected from the PT SEI message associated with the AU next to AU0 in decoding order. Otherwise, the alternative initial CPB removal delay and delay offset, indicated by bp_vcl_initial_cpb_removal_delay[Htid][ScIdx] and bp_vcl_initial_cpb_removal_offset[Htid][ScIdx], respectively, in the selected BP SEI message and pt_vcl_cpb_alt_initial_removal_delay_delta[Htid][ScIdx] and pt_vcl_cpb_alt_initial_removal_offset_delta[Htid][scIdx], respectively, can be selected from the PT SEI message associated with the AU next to AU0 in decoding order. The variable DefaultInitCpbParamsFlag is set to 0 and one of the following can apply:
[0239] -RASL AUs associated with CRA pictures contained in AU0, including RASL pictures with pps_mixed_nalu_types_in_pic_flag set to 0, can be removed from BitstreamToDecode, and the remaining bitstream can be allocated to BitstreamToDecode.
[0240] All AUs from AU0 onwards in decoding order up to the AU associated with the -DRAP indication SEI message can be discarded from BitstreamToDecode, and the remaining bitstream can be allocated to BitstreamToDecode.
[0241] If a current block (e.g., an AU) with temporalId=0 is associated with a BP SEI message, the picture of the block references at least one reference picture with a different resolution, the referenced PPS includes HRD parameters (i.e., the pps_hrd_parameters_update_present_flag value in the referenced PPS is 1), and the timing for CPB and DPB operations from the current block and the AU following the current CU in decoding order must utilize the HRD parameters signaled in the BP SEI associated with the referenced PPS and the current AU.
[0242] As another example, if a current block with temporalId=0 is associated with a BP SEI message, the picture of the block references at least one reference picture with a different resolution, and the selected ScIdx can be updated. When ScIdx is updated, the timing for CPB and DPB operations on the current block and blocks following the current block in decoding order must use the HRD parameters, BP SEI, PT SEI, and DUI SEI according to the new / updated ScIdx.
[0243] Buffering Period SEI Message Semantics
[0244] The BP SEI message may provide initial CPB removal delay and initial CPB removal delay offset information for initialization of the HRD at the associated AU position according to the decoding order.
[0245] When a BP SEI message is present, if an AU with TemporalId=0 is not a RASL or RADL picture and has at least one picture with ph_non_ref_pic_flag=0, the AU may be referred to as a notDiscardableAu.
[0246] If the current AU is not the first AU of the bitstream in decoding order, the AU prevNonDiscardableAu can be a previous AU in decoding order that is not a RASL or RADL picture, has at least one picture with ph_non_ref_pic_flag set to 0, and has TemporalId set to 0. Meanwhile, the presence or absence of a BP SEI message can be specified as follows:
[0247] - If NalHrdBpPresentFlag is 1 or VclHrdBpPresentFlag is 1, then for each AU in the CVS the following applies:
[0248] -If the AU is an IRAP or GDR AU, the BP SEI message applicable to the operation point must be associated with the AU.
[0249] Otherwise, if the AU is a notDiscardableAu that references at least one reference picture with a different resolution, the BP SEI messages applicable to the operation point must be associated with the AU.
[0250] Otherwise, if the AU is notDiscardableAu, the BP SEI messages applicable to the operation point may or may not be associated with the AU.
[0251] Otherwise, the AU must not be associated with a BP SEI message applicable to the operation point.
[0252] Otherwise (NalHrdBpPresentFlag and VclHrdBpPresentFlag are both 0), the AU of the CVS may not be associated with the BP SEI message.
[0253] However, for some applications, it is preferable to have more frequent BP SEI messages (e.g., for random access or bitstream splicing in IRAP AUs or non-IRAP AUs).
[0254] Table 1 below shows an example of a sequence parameter set.
[0255] [Table 1]
[0256] The syntax sps_extension_flag may indicate the presence of sps_extension_data_flag in the SPS RBSP syntax structure. A first value (e.g., 0) may indicate that sps_extension_data_flag is not present, and a second value (e.g., 1) may indicate that it is present.
[0257] The syntax sps_hrd_parameters_update_present_flag can indicate the possible presence of additional general timing and HRD parameters. A first value (e.g., 0) can indicate that additional general timing and HRD parameters are not present in the PPS, and a second value (e.g., 1) can indicate that additional general timing and HRD parameters can be present in the PPS. If the current syntax is not present in the bitstream, the value of this flag can be inferred to be 0.
[0258] There are no restrictions on the values that the syntax sps_extension_data_flag can have, and it can have any value.
[0259] Table 2 below shows an example of a picture parameter set.
[0260] [Table 2]
[0261] The syntax pps_extension_flag may indicate the presence of pps_extension_data_flag in the PPS RBSP syntax structure. A first value (e.g., 0) may indicate that pps_extension_data_flag is not present, and a second value (e.g., 1) may indicate that it is present. The syntax pps_hrd_parameters_update_present_flag may indicate the presence of general timing and HRD parameters. A first value (e.g., 0) may indicate that general timing and HRD parameters are not present, and a second value (e.g., 1) may indicate that general timing and HRD parameters are present. If the current syntax is not present in the bitstream, the value of this flag may be inferred to be 0.
[0262] The syntax structure general_timing_hrd_parameters() can provide some of the sequence level HRD parameters used in HRD operations.
[0263] When the syntax ols_timing_hrd_parameters() is included in a VPS, the OLS to which the syntax applies can be specified by the VPS. When the syntax ols_timing_hrd_parameters() is included in an SPS, the syntax may be applied to an OLS that includes only the lowest layer among the layers that reference the SPS, and this lowest layer may be an independent layer.
[0264] There is no restriction on the values that the syntax pps_extension_data_flag can have, and it can have any value.
[0265] As an example, if an RPR is applicable, updated parameters and timing information can be provided for conformance testing / timing operations and can be used to update the timing guidance of decoding operations (e.g., arrival of a picture in the CPB, removal of a picture from the CPB, removal of a picture from the DPB, etc.). Updated parameters and timing information / timing operations / reinitialization of conformance testing can be performed on pictures with temporalID=0 in the CVS intermediate and / or pictures that reference at least one reference picture with a different resolution / picture size. Additionally, additional HRD parameters can exist associated with a particular picture and can be used to update or reinitialize conformance testing / timing operations. The additional HRD parameters can exist in a picture parameter set (PPS) referenced by the picture. Alternatively, the additional HRD parameters can exist in the picture header of the picture. A flag specifying whether additional HRD parameters exist can exist in the PPS or picture header.
[0266] As an example, for HRD operations and conformance tests, the selected schedule index can be updated or changed in the middle of a CVS or CLVS. For HRD operations and conformance tests, the selected schedule index can be changed or updated in blocks where a major bitrate change occurs (i.e., the coded data of the current block (e.g., AU) is larger or smaller than that of the previous block). A schedule index change can be further restricted to be allowed in pictures in the middle of a CVS / CLVS where a major bitrate change occurs, where the temporal identifier (e.g., temporalId) is 0. A schedule index change can be further restricted to be allowed only if a picture resolution change (i.e., RPR) within the CLVS is allowed.
[0267] As an example, the values of ScIdx, BitRate[Htid][ScIdx] and CpbSize[Htid][ScIdx] can be constrained as follows:
[0268] -If at least one of the following conditions is met, the HSS may generate a BitRate[HtId] or CpBSize[HtId][ScIdx1] for the AU containing AUm by selecting the ScIdx1 value of ScIdx from among the ScIdx values provided in the selected general_timing_hrd_parameters() syntax structure containing AUm: The value of BitRate[HtId][ScIdx1] or CpBSize[HtId][ScIdx1] may be different from the value of BitRate[HtId][ScIdx0] or CPpBSize[HtId][ScIdx0] previously used for the AU:
[0269] 1. The contents of the syntax structure of general_timing_hrd_parameters() selected for the AU including AUm and the previous AU may differ.
[0270] 2. The picture sizes of the coded pictures of the current AU and the previous AU may be different, where the temporalId of the current AU may be 0.
[0271] Alternatively, RPR may be applicable (i.e., sps_res_change_in_clvs_allowed_flag value is 1), and the picture sizes of the coded pictures of the current AU and the previous AU may be different. Here, the temporalId of the current AU may be 0.
[0272] Otherwise, the HSS can continue to operate with the previous values of ScIdx, BitRate[Htid][ScIdx] and CpbSize[Htid][ScIdx].
[0273] When the HSS selects a value for BitRate[Htid][ScIdx] or CpbSize[Htid][ScIdx] that differs from the value of the previous AU, the following applies:
[0274] The variables BitRate[Htid][ScIdx] are applicable to the initial CPB arrival time of the current AU.
[0275] - The variables CpbSize[Htid][ScIdx] can be applied as follows:
[0276] 1. If the new value of CpbSize[Htid][ScIdx] is larger than the previous CPB size, it can be applied to the initial CPB arrival time of the current AU.
[0277] 2. Otherwise, the new value of CpbSize[Htid][ScIdx] is applicable to the CPB removal time of the current AU.
[0278] FIG. 12 is a diagram illustrating an example of an image encoding method to which an embodiment of the present disclosure can be applied.
[0279] The image encoding apparatus may derive HRD parameters (S1210) and encode image / video information (S1220). At this time, the image / video information may include information related to the induced HRD parameters.
[0280] Although not shown in FIG. 12, the image coding apparatus can perform DPB management based on the HRD parameters derived in step S1210.
[0281] FIG. 13 is a diagram illustrating an example of an image decoding method to which an embodiment of the present disclosure can be applied.
[0282] The image decoding apparatus may obtain image / video information from the bitstream (S1310). At this time, the image / video information may include information related to the HRD parameters.
[0283] The image decoding apparatus can decode the picture based on the obtained HRD parameters (S1320).
[0284] FIG. 14 is a diagram for explaining another example of an image decoding method to which an embodiment of the present disclosure can be applied.
[0285] The image decoding apparatus may obtain image / video information from the bitstream (S1410). At this time, the image / video information may include information related to the HRD parameters.
[0286] The image decoding apparatus can perform DPB management based on the acquired HRD parameters (S1420).
[0287] The image decoding apparatus may decode a picture based on the DPB (S1430). For example, a block / slice in the current picture may be decoded based on inter prediction using an already reconstructed picture in the DPB as a reference picture.
[0288] 12 to 14, the information related to the HRD parameters may include at least one of the information / syntax elements described in connection with at least one of the embodiments of the present disclosure. Also, as described above, DPB management may be performed based on the HRD parameters. For example, deletion of a picture from the DPB before decoding of the current picture and / or output of a (decoded) picture may be performed based on the HRD parameters.
[0289] 15 is a diagram illustrating an image processing method according to an embodiment of the present disclosure. The image processing method includes an image coding method and may be an image encoding or decoding method. The image processing method is performed by an image processing device, and the image processing device may include an image coding device. The image coding device may include both an image encoding device and an image decoding device.
[0290] For example, the image decoding method may include the above-described embodiments and configurations, where some of the configurations may be combined, the order may be changed, some configurations may be deleted, and other configurations may be added.
[0291] For example, it may be determined whether a reference sample resampling (RPR) operation is permitted by the image decoding method (S1510). The reference sample resampling operation may include upscaling or downscaling when the resolutions of the reference picture and the current picture are different. This is the same as described above, so a repeated description will be omitted. Furthermore, based on whether the RPR operation is permitted, a hypothetical reference decoder (HRD) operation may be performed (S1520). The HRD operation may include at least one of an operation of determining the conformance of a bitstream including encoded video data with a video coding standard, an operation of determining the conformance of a video decoder with the video coding standard, or an HRD picture output timing signaling operation. The HRD operation may be performed by updating a schedule index for a hypothetical stream scheduler (HSS). For example, the HRD operation may be performed using additional specific HRD parameters, which may be included in a picture header or a picture parameter set. Additional specific HRD parameters may be present only when the additional parameter presence flag has a specific value. Meanwhile, as an example, an HRD operation (e.g., an operation for determining the conformance of a bitstream to a video coding standard) may be determined based on at least one of the number of access units associated with a Buffering Period (BP) Supplemental Enhancement Information (SEI) message, the number of Intra Random Access Point (IRAP) or Gradual Decoding Refresh (GDR) access units, or the number of access units associated with a Dependent Random Access Point (DRAP) Indication SEI message.For example, the schedule index may be updated in the middle of a Coded Video Sequence (CVS) or a Coded Layered Video Sequence (CLVS) and may be updated based on a temporal identifier value. It may also be updated at an AU where a major bitrate change occurs. For example, the HRD operation may include an operation of updating output timing guidance based on updated HRD parameters and output timing information. Meanwhile, the output timing guidance described above may include guiding at least one of the arrival timing of a picture in a CPB, the removal timing of a picture from a CPB, or the removal timing of a picture from a DPB. For example, the HRD operation (e.g., an operation of determining the conformance of a bitstream to a video coding standard) may be performed a number of times derived based on the product of a first variable indicating the number of AUs in BitstreamToDecode associated with a BP SEI message applicable to TargetOlsIdx and a second variable related to the value of general_du_hrd_params_present_flag. In addition, an HRD operation (e.g., an operation of determining the conformance of a bitstream to a video coding standard) may be performed a number of times derived based on the product of a first variable indicating the number of IRAP or GDR AUs of BitstreamToDecode associated with a BP SEI message applicable to TargetOlsIdx and a second variable related to the value of general_du_hrd_params_present_flag. In addition, an HRD operation (e.g., an operation of determining the conformance of a bitstream to a video coding standard) may be performed a number of times derived based on the product of a first variable indicating the number of AUs of BitstreamToDecode associated with a DRAP indication SEI message applicable to TargetOlsIdx and whose associated PT SEI message has pt_cpb_alt_timing_info_present_flag equal to 1 and a second variable related to the value of general_du_hrd_params_present_flag.
[0292] On the other hand, since FIG. 15 corresponds to one embodiment of the present disclosure, some components not shown may be added, some components may be deleted, or the order of the embodiments described above may be changed.
[0293] 16 is a diagram illustrating an image encoding method according to an embodiment of the present disclosure. The image encoding method of FIG. 16 can be performed by an image encoding device and can include the embodiments and configurations described above. Here, some configurations can be combined, their order can be changed, some configurations can be deleted, and other configurations can be added.
[0294] For example, according to an image coding method, it may be determined whether a reference sample resampling (RPR) operation is permitted, and reference sample resampling permission information may be coded (S1610). A redundant description of the reference sample resampling will be omitted. Then, HRD parameters for a hypothetical reference decoder (HRD) operation may be coded (S1620). The HRD parameters may be coded based on whether the RPR operation is permitted. The HRD operation may include at least one of an operation of determining conformance of a bitstream including encoded video data with a video coding standard, an operation of determining conformance of a video decoder with the video coding standard, or an HRD picture output timing signaling operation. The HRD operation may be performed by updating a schedule index for a hypothetical stream scheduler (HSS). A redundant description of this operation will be omitted.
[0295] On the other hand, since FIG. 15 corresponds to one embodiment of the present disclosure, some components not shown may be added, some components may be deleted, or the order of the embodiments described above may be changed.
[0296] Meanwhile, although not shown, a bitstream transmission method and a bitstream recording medium may be proposed according to an embodiment of the present disclosure. The bitstream transmission method includes a configuration for transmitting a bitstream generated by an image coding method, and the image coding method may include the image coding method of FIG. 16. Also, the bitstream recording medium may be a recording medium that stores the bitstream generated by the image coding method. Similarly, the image coding method may include the image coding method of FIG. 16.
[0297] Although the exemplary method of the present disclosure is expressed as a series of operations for clarity of explanation, this is not intended to limit the order in which the steps are performed, and the steps may be performed simultaneously or in a different order if necessary. To achieve the method according to the present disclosure, the steps illustrated may include other steps, or some steps may be omitted and the remaining steps may be included, or some steps may be omitted and additional other steps may be included.
[0298] In the present disclosure, an image encoding device or an image decoding device that performs a predetermined operation (step) can perform the operation (step) to check the execution conditions and circumstances of the operation (step). For example, if it is described that a predetermined operation is performed when a predetermined condition is satisfied, the image encoding device or the image decoding device can perform the predetermined operation after performing an operation to check whether the predetermined condition is satisfied.
[0299] The various embodiments of the present disclosure are not intended to enumerate all possible combinations, but are intended to describe representative aspects of the present disclosure, and the matters described in the various embodiments may be applied independently or in combination of two or more.
[0300] Additionally, various embodiments of the present disclosure may be implemented using hardware, firmware, software, or a combination thereof, etc. In the case of a hardware implementation, the implementation may be using one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), general processors, controllers, microcontrollers, microprocessors, etc.
[0301] In addition, an image decoding apparatus and an image encoding apparatus to which an embodiment of the present disclosure is applied may be included in a multimedia broadcast transmitting / receiving apparatus, a mobile communication terminal, a home cinema video apparatus, a digital cinema video apparatus, a surveillance camera, a video conversation apparatus, a real-time communication apparatus such as video communication, a mobile streaming apparatus, a storage medium, a camcorder, a video on demand (VoD) service providing apparatus, an over-the-top (OTT) video apparatus, an internet streaming service providing apparatus, a three-dimensional (3D) video apparatus, an image telephone video apparatus, a medical video apparatus, etc., and may be used to process a video signal or a data signal. For example, an over-the-top (OTT) video apparatus may include a game console, a Blu-ray player, an internet-connected TV, a home theater system, a smartphone, a tablet PC, a digital video recorder (DVR), etc.
[0302] FIG. 17 is a diagram illustrating a content streaming system to which an embodiment of the present disclosure can be applied.
[0303] As shown in FIG. 17, a content streaming system to which an embodiment of the present disclosure is applied can broadly include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.
[0304] The encoding server compresses content input from a multimedia input device such as a smartphone, camera, or camcorder into digital data to generate a bitstream and transmits the bitstream to the streaming server. As another example, if a multimedia input device such as a smartphone, camera, or video camera directly generates a bitstream, the encoding server can be omitted.
[0305] The bitstream can be generated by an image encoding method and / or image encoding device to which an embodiment of the present disclosure is applied, and the streaming server can temporarily store the bitstream during the process of transmitting or receiving the bitstream.
[0306] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server serves as an intermediary for informing the user of available services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, which then transmits the multimedia data to the user. In this case, the content streaming system may include a separate control server, which may control commands and responses between devices in the content streaming system.
[0307] The streaming server may receive content from a media storage and / or an encoding server. For example, when receiving content from the encoding server, the content may be received in real time. In this case, the streaming server may store the bitstream for a certain period of time to provide a smooth streaming service.
[0308] Examples of the user device include a mobile phone, a smartphone, a laptop computer, a digital broadcasting terminal, a personal digital assistant (PDA), a portable multimedia player (PMP), a navigation system, a slate PC, a tablet PC, an ultrabook, a wearable device such as a smartwatch, smart glass, a head mounted display (HMD), a digital TV, a desktop computer, and digital signage.
[0309] Each server in the content streaming system can be operated as a distributed server, in which case data received from each server can be processed in a distributed manner.
[0310] The scope of the present disclosure includes software or machine-executable commands (e.g., operating systems, applications, firmware, programs, etc.) that cause operations according to the methods of various embodiments to be performed on a device or computer, and non-transitory computer-readable medium on which such software or commands can be stored and executed on a device or computer. [Industrial Applicability]
[0311] The embodiments of the present disclosure can be used to encode / decode images.
Claims
1. An image decoding method, comprising: determining whether Reference Sample Resampling (RPR) operation is allowed; and performing a Hypothetical Reference Decoder (HRD) operation based on the RPR operation being permitted; the HRD operations include at least one of an operation of determining conformance of a bitstream including encoded video data with a video coding standard, an operation of determining conformance of a video decoder with the video coding standard, or an HRD picture output timing signaling operation; The image decoding method, wherein the HRD operation is performed by updating a selected schedule index for a hypothetical stream scheduler (HSS).
2. The image decoding method according to claim 1 , wherein the HRD operation is performed by additional specific HRD parameters, the HRD parameters being included in a picture header or a picture parameter set.
3. The image decoding method according to claim 2 , wherein the additional specific HRD parameter is present only if an additional parameter present flag has a specific value.
4. 2. The image decoding method of claim 1, wherein the HRD operation is determined based on at least one of a number of access units associated with a Buffering Period (BP) Supplemental Enhancement Information (SEI) message, a number of Intra Random Access Point (IRAP) or Gradual Decoding Refresh (GDR) access units, or a number of access units associated with a Dependent Random Access Point (DRAP) indication SEI message.
5. 2. The image decoding method according to claim 1, wherein the schedule index is updated midway through a coded video sequence (CVS) or a coded layered video sequence (CLVS).
6. The image decoding method of claim 5 , wherein the schedule index is updated based on a temporal identifier value.
7. The image decoding method according to claim 1 , wherein the schedule index is updated in an access unit in which a change in bit rate occurs.
8. The image decoding method of claim 1 , wherein the HRD operations include updating output timing guidance based on updated HRD parameters and output timing information.
9. 9. The image decoding method of claim 8, wherein the output timing derivation includes derivation of at least one of arrival timing of a picture in a CPB, removal timing of a picture from a CPB, or removal timing of a picture from a DPB.
10. 2. The image decoding method of claim 1, wherein the HRD operation is performed a number of times derived based on the product of a first variable indicating the number of AUs of BitstreamToDecode associated with the BP SEI message applicable to TargetOlsIdx and a second variable related to the value of general_du_hrd_params_present_flag.
11. 2. The image decoding method of claim 1, wherein the HRD operation is performed a number of times derived based on the product of a first variable indicating the number of IRAPs or GDR AUs in BitstreamToDecode associated with the BP SEI message applicable to TargetOlsIdx and a second variable related to the value of general_du_hrd_params_present_flag.
12. 2. The image decoding method of claim 1, wherein the HRD operation is performed a number of times derived based on the product of a first variable indicating the number of AUs of BitstreamToDecode that are associated with a DRAP indication SEI message applicable to TargetOlsIdx and whose associated PT SEI message has a pt_cpb_alt_timing_info_present_flag equal to 1, and a second variable related to the value of general_du_hrd_params_present_flag.
13. 1. An image encoding method, comprising: determining whether a Reference Sample Resampling (RPR) operation is allowed and encoding reference sample resampling allowance information; encoding a Hypothetical Reference Decoder (HRD) parameter for an HRD operation based on whether the RPR operation is allowed; the HRD operations include at least one of an operation of determining conformance of a bitstream including encoded video data with a video coding standard, an operation of determining conformance of a video decoder with the video coding standard, or an HRD picture output timing signaling operation; The image coding method, wherein the HRD operation is performed by updating a selected schedule index for a Hypothetical Stream Scheduler (HSS).
14. A bitstream transmission method, comprising: transmitting a bitstream generated by the image encoding method, The image encoding method includes: determining whether a Reference Sample Resampling (RPR) operation is allowed and encoding reference sample resampling allowance information; encoding a Hypothetical Reference Decoder (HRD) parameter for an HRD operation based on whether the RPR operation is allowed; the HRD operations include at least one of an operation of determining conformance of a bitstream including encoded video data with a video coding standard, an operation of determining conformance of a video decoder with the video coding standard, or an HRD picture output timing signaling operation; The HRD operation is performed by updating a selected schedule index for a Hypothetical Stream Scheduler (HSS).
15. A non-transitory computer-readable recording medium storing a bitstream generated by the method of claim 13.