Image encoding / decoding method and device for signaling DPB-related information and PTL-related information, and computer-readable recording medium storing a bitstream
The image encoding/decoding method optimizes DPB and PTL information signaling within the VPS to enhance efficiency and reduce costs in high-resolution image transmission and storage.
Patent Information
- Application Number
- JP2024076796
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-04-02
- Filing Date
- 2024-05-09
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2041-03-26
AI Technical Summary
The increasing demand for high-resolution, high-quality images leads to a significant increase in transmission and storage costs due to the higher amount of information, necessitating highly efficient image compression techniques.
An image encoding/decoding method that efficiently signals DPB-related and PTL-related information by managing DPB parameter syntax structures and PTL syntax structures within the VPS, ensuring they do not exceed the number of multi-layer OLSs, and allowing for efficient mapping and selection of these structures for current OLSs.
This approach enhances encoding/decoding efficiency and reduces transmission and storage costs by optimizing the signaling of DPB and PTL information, enabling effective image restoration from a bitstream.
Smart Images

Figure 0007766134000001 
Figure 0007766134000002 
Figure 0007766134000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an image encoding / decoding method and apparatus, and more particularly to an image encoding / decoding method and apparatus that signal DPB (Decoded Picture Buffer) related information and PTL (Profile Tier Level) related information, as well as a computer-readable recording medium that stores a bitstream generated by the image encoding method / apparatus of the present disclosure. [Background technology]
[0002] Recently, demand for high-resolution, high-quality images, such as HD (High Definition) images and UHD (Ultra High Definition) images, has been increasing in various fields. As image data becomes higher in resolution and quality, the amount of information or bits to be transmitted increases relatively compared to conventional image data. The increase in the amount of information or bits to be transmitted results in an increase in transmission costs and storage costs.
[0003] This requires highly efficient image compression techniques for effectively transmitting, storing, and reproducing high-resolution, high-quality image information. Summary of the Invention [Problem to be solved by the invention]
[0004] An object of the present disclosure is to provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.
[0005] Another object of the present disclosure is to provide an image encoding / decoding method and device that improves the efficiency of encoding / decoding by efficiently signaling DPB-related information and PTL-related information.
[0006] Another object of the present disclosure is to provide a method for transmitting a bitstream generated by the image encoding method or apparatus according to the present disclosure.
[0007] Another object of the present disclosure is to provide a recording medium storing a bitstream generated by the image encoding method or apparatus according to the present disclosure.
[0008] Another object of the present disclosure is to provide a recording medium storing a bitstream that is received by an image decoding device according to the present disclosure, decoded, and used to restore an image.
[0009] The technical problems to be solved by the present disclosure are not limited to the above-mentioned technical problems, and other technical problems not mentioned above will be clearly understood by a person having ordinary skill in the technical field to which the present disclosure pertains from the following description. [Means for solving the problem]
[0010] An image decoding method performed by an image decoding device according to one embodiment of the present disclosure may include the steps of: acquiring first information indicating the number of one or more decoded picture buffer (DPB) parameter syntax structures in a VPS (Video Parameter Set); acquiring the one or more DPB parameter syntax structures from the VPS based on the first information; acquiring second information regarding mapping between one or more multi-layer output layer sets (OLSs) and the one or more DPB parameter syntax structures from the VPS based on the first information; selecting a DPB parameter syntax structure to be applied to a current OLS based on the second information; and processing the current OLS based on the selected DPB parameter syntax structure.
[0011] In the image decoding method of the present disclosure, the number of the one or more DPB parameter syntax structures in the VPS may not be greater than the number of the one or more multi-layer OLSs.
[0012] In the image decoding method of the present disclosure, each of the one or more DPB parameter syntax structures in the VPS can be mapped to at least one multi-layer OLS of the one or more multi-layer OLSs.
[0013] In the image decoding method of the present disclosure, the second information can be obtained from the VPS based on the number of the one or more DPB parameter syntax structures in the VPS being greater than one.
[0014] In the image decoding method of the present disclosure, based on the number of the one or more DPB parameter syntax structures in the VPS being not greater than 1, the second information is not obtained from the VPS and the second information can be inferred to be a value of 0.
[0015] In the image decoding method of the present disclosure, based on the fact that the current OLS includes a single layer, the DPB parameter syntax structure applied to the current OLS can be obtained from an SPS (Sequence Parameter Set).
[0016] The image decoding method of the present disclosure may further include the steps of: obtaining third information indicating the number of one or more PTL (profile tier level) syntax structures in a VPS; obtaining the one or more PTL syntax structures from the VPS based on the third information; obtaining fourth information regarding mapping between one or more OLSs and the one or more PTL syntax structures from the VPS based on the third information; and selecting a PTL syntax structure to be applied to the current OLS based on the fourth information.
[0017] In the image decoding method of the present disclosure, the number of the one or more PTL syntax structures in the VPS may not be greater than the total number of the one or more OLSs.
[0018] In the image decoding method of the present disclosure, each of the one or more PTL syntax structures in the VPS can be mapped to at least one OLS of the one or more OLSs.
[0019] An image decoding device according to another aspect of the present disclosure includes a memory and at least one processor, wherein the at least one processor acquires first information indicating the number of one or more decoded picture buffer (DPB) parameter syntax structures in a VPS (Video Parameter Set), acquires the one or more DPB parameter syntax structures from the VPS based on the first information, acquires second information regarding mapping between one or more multi-layer output layer sets (OLSs) and the one or more DPB parameter syntax structures from the VPS based on the first information, selects a DPB parameter syntax structure to be applied to a current OLS based on the second information, and processes the current OLS based on the selected DPB parameter syntax structure.
[0020] An image encoding method performed by an image encoding device according to another aspect of the present disclosure may include the steps of: encoding first information indicating the number of one or more decoded picture buffer (DPB) parameter syntax structures in a VPS (Video Parameter Set); encoding the one or more DPB parameter syntax structures into the VPS based on the first information; encoding second information regarding mapping between one or more multi-layer output layer sets (OLSs) and the one or more DPB parameter syntax structures into the VPS based on the first information; selecting a DPB parameter syntax structure to be applied to a current OLS based on the second information; and processing the current OLS based on the selected DPB parameter syntax structure.
[0021] In the image coding method of the present disclosure, the number of the one or more DPB parameter syntax structures in the VPS may not be greater than the number of the one or more multi-layer OLSs.
[0022] In the image coding method of the present disclosure, each of the one or more DPB parameter syntax structures in the VPS can be mapped to at least one multi-layer OLS of the one or more multi-layer OLSs.
[0023] In the image encoding method of the present disclosure, the second information may be encoded in the VPS based on the number of the one or more DPB parameter syntax structures in the VPS being greater than one.
[0024] A transmission method according to another aspect of the present disclosure can transmit a bitstream generated by the image encoding device or image encoding method of the present disclosure.
[0025] A computer-readable recording medium according to another aspect of the present disclosure can store a bitstream generated by the image encoding method or image encoding device of the present disclosure.
[0026] The features described above in this brief summary of the present disclosure are merely exemplary embodiments of the detailed description of the present disclosure that follows and are not intended to limit the scope of the present disclosure. [Effects of the Invention]
[0027] According to the present disclosure, an image encoding / decoding method and apparatus with improved encoding / decoding efficiency can be provided.
[0028] Furthermore, the present disclosure can provide an image encoding / decoding method and device that can improve the efficiency of encoding / decoding by efficiently signaling DPB-related information and PTL-related information.
[0029] The present disclosure also provides a method for transmitting a bitstream generated by the image encoding method or apparatus according to the present disclosure.
[0030] Furthermore, according to the present disclosure, a recording medium storing a bitstream generated by the image encoding method or apparatus according to the present disclosure can be provided.
[0031] The present disclosure may also provide a recording medium storing a bitstream that is received by an image decoding device according to the present disclosure, decoded, and used to restore an image.
[0032] The effects obtained by the present disclosure are not limited to the effects described above, and other effects not described above will be clearly understood by those having ordinary skill in the art to which the present disclosure pertains from the following description. [Brief explanation of the drawings]
[0033] [Figure 1] 1 is a diagram illustrating a video coding system to which embodiments of the present disclosure can be applied; [Figure 2] 1 is a diagram schematically illustrating an image encoding device to which an embodiment of the present disclosure can be applied. [Figure 3] FIG. 1 is a diagram schematically illustrating an image decoding device to which an embodiment of the present disclosure can be applied. [Figure 4] 1 illustrates an example of a general picture decoding procedure to which the embodiments of the present disclosure can be applied. [Figure 5] 1 shows an example of a general picture encoding procedure to which the embodiments of the present disclosure can be applied. [Figure 6] FIG. 1 shows an example of a hierarchical structure for coded images / video. [Figure 7] FIG. 2 is a diagram illustrating an exemplary syntax structure of a VPS according to one embodiment of the present disclosure. [Figure 8] A diagram illustrating a syntax structure for signaling DPB parameters according to the present disclosure. [Figure 9]FIG. 10 is a diagram illustrating an exemplary syntax structure of a VPS according to another embodiment of the present disclosure. [Figure 10] FIG. 1 is a diagram illustrating an example of an image encoding method to which an embodiment of the present disclosure can be applied. [Figure 11] FIG. 1 is a diagram illustrating an example of an image decoding method to which an embodiment of the present disclosure can be applied. [Figure 12] FIG. 10 is a diagram illustrating another example of an image decoding method to which an embodiment of the present disclosure can be applied. [Figure 13] 10 is a diagram illustrating a process of encoding DPB parameters based on information on the number of DPB parameters according to another embodiment of the present disclosure. [Figure 14] A diagram for explaining the process of decoding DPB parameters based on the number information of DPB parameters according to another embodiment of the present disclosure. [Figure 15] FIG. 1 illustrates a content streaming system to which an embodiment of the present disclosure can be applied. DETAILED DESCRIPTION OF THE INVENTION
[0034] The present disclosure will be described in detail below with reference to the accompanying drawings, so that those skilled in the art can easily implement the present disclosure. However, the present disclosure may be embodied in various different forms and is not limited to the embodiments described herein.
[0035] In describing the embodiments of the present disclosure, if it is determined that a detailed description of a known configuration or function may obscure the gist of the present disclosure, the detailed description thereof will be omitted. In addition, in the drawings, parts that are not related to the description of the present disclosure will be omitted, and similar parts will be designated by similar reference numerals.
[0036] In this disclosure, when a component is referred to as being "coupled," "coupled," or "connected" to another component, this includes not only a direct connection, but also an indirect connection where another component exists between them. Furthermore, when a component is referred to as "including" or "having" another component, this does not mean that the other component is excluded, but that the component can further include the other component, unless otherwise specified.
[0037] In this disclosure, terms such as "first" and "second" are used only to distinguish one component from another component, and do not limit the order or importance of the components unless otherwise specified. Therefore, within the scope of this disclosure, a first component in one embodiment may be referred to as a second component in another embodiment, and similarly, a second component in one embodiment may be referred to as a first component in another embodiment.
[0038] In this disclosure, components that are distinguished from one another are used to clearly describe the characteristics of each component and do not necessarily mean that the components are separate. In other words, multiple components may be integrated into a single hardware or software unit, or a single component may be distributed into multiple hardware or software units. Therefore, even if not otherwise specified, such integrated or distributed embodiments are also included within the scope of this disclosure.
[0039] In this disclosure, the components described in various embodiments are not necessarily essential components, and some may be optional components. Therefore, an embodiment consisting of a subset of the components described in one embodiment is also within the scope of this disclosure. Furthermore, an embodiment including other components in addition to the components described in various embodiments is also within the scope of this disclosure.
[0040] The present disclosure relates to image encoding and decoding, and terms used in this disclosure may have their ordinary meaning in the technical field to which the present disclosure belongs unless they are newly defined in this disclosure.
[0041] In this disclosure, a "picture" generally refers to a unit representing any one image in a specific time period, and a slice / tile is a coding unit constituting a part of a picture, and one picture may be composed of one or more slices / tiles. Furthermore, a slice / tile may include one or more coding tree units (CTUs).
[0042] In this disclosure, "pixel" or "pel" may refer to the smallest unit constituting one picture (or image). Also, "sample" may be used as a term corresponding to pixel. A sample may generally indicate a pixel or a pixel value, may indicate only a pixel / pixel value of a luma component, or may indicate only a pixel / pixel value of a chroma component.
[0043] In this disclosure, the term "unit" may refer to a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to that region. The term "unit" may be used interchangeably with terms such as "sample array," "block," or "area," depending on the situation. In general, an M×N block may include a set (or array) of samples or transform coefficients consisting of M columns and N rows.
[0044] In the present disclosure, a "current block" may refer to any one of a "current coding block," a "current coding unit," a "block to be coded," a "block to be decoded," or a "block to be processed." When prediction is performed, a "current block" may refer to a "current predicted block" or a "block to be predicted." When transformation (inverse transformation) / quantization (inverse quantization) is performed, a "current block" may refer to a "current transformed block" or a "block to be transformed." When filtering is performed, a "current block" may refer to a "block to be filtered."
[0045] Furthermore, in this disclosure, unless explicitly stated as a chroma block, the term "current block" may refer to a block including both a luma component block and a chroma component block, or to the "luma block of the current block." The luma component block of the current block may be expressed by explicitly including the term "luma block" or "current luma block." The chroma component block of the current block may be expressed by explicitly including the term "chroma block" or "current chroma block."
[0046] In the present disclosure, "A or B" can mean "A only," "B only," or "both A and B." In other words, in the present disclosure, "A or B" can be interpreted as "A and / or B." For example, in the present disclosure, "A, B or C" can mean "A only," "B only," "C only," or "any combination of A, B, and C."
[0047] As used in this disclosure, " / " and "," (comma) can mean "and / or." For example, "A / B" can mean "A and / or B." Thus, "A / B" can mean "A only," "B only," or "both A and B." For example, "A, B, C" can mean "A, B, or C."
[0048] In the present disclosure, "at least one of A and B" can mean "A only," "B only," or "both A and B." Also, in the present disclosure, the expressions "at least one of A or B" and "at least one of A and / or B" can be interpreted as being the same as "at least one of A and B."
[0049] Additionally, in this disclosure, "at least one of A, B, and C" can mean "A only," "B only," "C only," or "any combination of A, B, and C." Additionally, "at least one of A, B, or C" or "at least one of A, B, and / or C" can mean "at least one of A, B, and C."
[0050] Furthermore, parentheses used in the present disclosure may mean "for example." Specifically, when "prediction (intra prediction)" is displayed, "intra prediction" may be proposed as an example of "prediction." In other words, "prediction" in the present disclosure is not limited to "intra prediction," and "intra prediction" may be proposed as an example of "prediction." Furthermore, when "prediction (i.e., intra prediction)" is displayed, "intra prediction" may be proposed as an example of "prediction."
[0051] In the present disclosure, technical features described separately in one drawing may be realized separately or simultaneously.
[0052] Video Coding System Overview
[0053] FIG. 1 illustrates a video coding system according to this disclosure.
[0054] A video coding system according to one embodiment may include an encoding device 10 and a decoding device 20. The encoding device 10 may transmit encoded video and / or image information or data to the decoding device 20 in a file or streaming format via a digital storage medium or a network.
[0055] An encoding device 10 according to an embodiment may include a video source generation unit 11, an encoding unit 12, and a transmission unit 13. A decoding device 20 according to an embodiment may include a reception unit 21, a decoding unit 22, and a rendering unit 23. The encoding unit 12 may be referred to as a video / image encoding unit, and the decoding unit 22 may be referred to as a video / image decoding unit. The transmission unit 13 may be included in the encoding unit 12. The reception unit 21 may be included in the decoding unit 22. The rendering unit 23 may include a display unit, which may be configured as a separate device or an external component.
[0056] The video source generation unit 11 can acquire video / images through a video / image capture, synthesis, or generation process. The video source generation unit 11 can include a video / image capture device and / or a video / image generation device. The video / image capture device can include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. The video / image generation device can include, for example, a computer, a tablet, a smartphone, etc., and can (electronically) generate video / images. For example, virtual video / images can be generated via a computer, etc., in which case the video / image capture process can be replaced with a process in which related data is generated.
[0057] The encoder 12 may encode the input video / image. The encoder 12 may perform a series of steps such as prediction, transformation, and quantization for compression and coding efficiency. The encoder 12 may output the encoded data (encoded video / image information) in a bitstream format.
[0058] The transmitter 13 may transmit the encoded video / image information or data output in a bitstream format to the receiver 21 of the decoding device 20 in a file or streaming format via a digital storage medium or a network. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray®, HDD, and SSD. The transmitter 13 may include elements for generating a media file in a predetermined file format and elements for transmitting via a broadcasting / communication network. The receiver 21 may extract / receive the bitstream from the storage medium or network and transmit it to the decoder 22.
[0059] The decoding unit 22 can decode the video / image by performing a series of steps such as inverse quantization, inverse transformation, and prediction corresponding to the operations of the encoding unit 12.
[0060] The rendering unit 23 can render the decoded video / images, and the rendered video / images can be displayed via the display unit.
[0061] Overview of the image encoding device
[0062] FIG. 2 is a diagram schematically illustrating an image encoding device to which an embodiment of the present disclosure can be applied.
[0063] 2, the image encoding device 100 may include an image division unit 110, a subtraction unit 115, a transform unit 120, a quantization unit 130, an inverse quantization unit 140, an inverse transform unit 150, an addition unit 155, a filtering unit 160, a memory 170, an inter prediction unit 180, an intra prediction unit 185, and an entropy encoding unit 190. The inter prediction unit 180 and the intra prediction unit 185 may be collectively referred to as a "prediction unit." The transform unit 120, the quantization unit 130, the inverse quantization unit 140, and the inverse transform unit 150 may be included in a residual processing unit. The residual processing unit may further include a subtraction unit 115.
[0064] Depending on the embodiment, all or at least some of the components constituting the image encoding device 100 may be realized by a single hardware component (e.g., an encoder or a processor). Also, the memory 170 may include a decoded picture buffer (DPB) and may be realized by a digital storage medium.
[0065] The image division unit 110 may divide an input image (or picture, frame) input to the image encoding device 100 into one or more processing units. As an example, the processing units may be called coding units (CUs). The coding units may be obtained by recursively dividing a coding tree unit (CTU) or a largest coding unit (LCU) using a QT / BT / TT (quad-tree / binary-tree / ternary-tree) structure. For example, one coding unit may be divided into multiple coding units at deeper depths based on a quad-tree structure, a binary-tree structure, and / or a ternary-tree structure. To divide the coding units, the quad-tree structure may be applied first, and then the binary-tree structure and / or the ternary-tree structure may be applied later. The coding procedure according to the present disclosure may be performed based on the final coding unit that is not further divided. The maximum coding unit may be used as the final coding unit, or a lower-depth coding unit obtained by dividing the maximum coding unit may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and / or reconstruction, which will be described later. As another example, a processing unit of the coding procedure may be a prediction unit (PU) or a transform unit (TU). The prediction unit and the transform unit may be divided or partitioned from the final coding unit, respectively. The prediction unit may be a unit of sample prediction, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.
[0066] The prediction unit (inter prediction unit 180 or intra prediction unit 185) may perform prediction on a current block (current block) to generate a predicted block including prediction samples for the current block. The prediction unit may determine whether intra prediction or inter prediction is applied to the current block or CU. The prediction unit may generate various information related to prediction of the current block and transmit it to the entropy coding unit 190. The prediction information may be coded by the entropy coding unit 190 and output in a bitstream format.
[0067] The intra prediction unit 185 may predict the current block by referring to samples in the current picture. The referenced samples may be located in the neighborhood of the current block or may be located far away from the current block according to the intra prediction mode and / or intra prediction technique. The intra prediction modes may include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes may include, for example, a DC mode and a planar mode. The directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the granularity of the prediction direction. However, this is merely an example, and more or less directional prediction modes may be used depending on the settings. The intra prediction unit 185 may also determine the prediction mode to be applied to the current block using the prediction modes applied to neighboring blocks.
[0068] The inter prediction unit 180 may derive a predicted block for a current block based on a reference block (reference sample array) identified by a motion vector on a reference picture. To reduce the amount of motion information transmitted in inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation between the motion information of neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on the inter prediction direction (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block may be called a collocated reference block, a collocated CU (colCU), etc. The reference picture including the temporal neighboring block may be called a collocated picture (colPic). For example, the inter predictor 180 may construct a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to derive a motion vector and / or a reference picture index for the current block. Inter prediction may be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the inter predictor 180 may use motion information of neighboring blocks as motion information for the current block. In the case of skip mode, unlike in merge mode, a residual signal may not be transmitted.In the case of a motion vector prediction (MVP) mode, the motion vector of a neighboring block is used as a motion vector predictor, and the motion vector of the current block can be signaled by encoding a motion vector difference and an indicator for the motion vector predictor. The motion vector difference may mean the difference between the motion vector of the current block and the motion vector predictor.
[0069] The predictor may generate a prediction signal based on various prediction methods and / or prediction techniques, which will be described later. For example, the predictor may apply intra prediction or inter prediction to predict the current block, or may simultaneously apply intra prediction and inter prediction. A prediction method that simultaneously applies intra prediction and inter prediction to predict the current block may be referred to as combined inter and intra prediction (CIIP). The predictor may also perform intra block copy (IBC) to predict the current block. Intra block copy can be used for content image / video coding, such as screen content coding (SCC), for games. IBC is a method of predicting a current block using an already reconstructed reference block in a current picture that is located a predetermined distance away from the current block. When IBC is applied, the position of the reference block in the current picture may be coded as a vector (block vector) corresponding to the predetermined distance. IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that a reference block is derived within the current picture. That is, the IBC may use at least one of the inter prediction techniques described in this disclosure.
[0070] The prediction signal generated by the prediction unit may be used to generate a restored signal or a residual signal. The subtraction unit 115 may subtract the prediction signal (predicted block, predicted sample array) output from the prediction unit from the input image signal (original block, original sample array) to generate a residual signal (residual signal, residual block, residual sample array). The generated residual signal may be transmitted to the conversion unit 120.
[0071] The transform unit 120 may generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loeve transform (KLT), a graph-based transform (GBT), or a conditionally non-linear transform (CNT). Here, the GBT refers to a transform obtained from a graph representing inter-pixel relationship information. The CNT refers to a transform obtained based on a predicted signal generated using all previously reconstructed pixels. The transform process may be applied to pixel blocks having the same square size or to non-square blocks of variable size.
[0072] The quantization unit 130 may quantize the transform coefficients and transmit the quantized transform coefficients to the entropy coding unit 190. The entropy coding unit 190 may encode the quantized signal (information about the quantized transform coefficients) and output the encoded signal in a bitstream format. The information about the quantized transform coefficients may be referred to as residual information. The quantization unit 130 may rearrange the quantized transform coefficients in a block format into a one-dimensional vector format based on a coefficient scan order, and may generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector format.
[0073] The entropy coding unit 190 may perform various coding methods, such as exponential Golomb, context-adaptive variable length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy coding unit 190 may also code information necessary for video / image restoration (e.g., values of syntax elements) together with or separately from the quantized transform coefficients. The coded information (e.g., coded video / image information) may be transmitted or stored in a bitstream format in network abstraction layer (NAL) units. The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / image information may also include general constraint information. The signaling information, transmitted information and / or syntax elements mentioned in this disclosure may be encoded through the above-described encoding procedure and included in the bitstream.
[0074] The bitstream may be transmitted via a network or stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as a USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitting unit (not shown) that transmits and / or a storing unit (not shown) that stores the signal output from the entropy encoding unit 190 may be provided as an internal / external element of the image encoding device 100, or the transmitting unit may be provided as a component of the entropy encoding unit 190.
[0075] The quantized transform coefficients output from the quantization unit 130 can be used to generate a residual signal. For example, the residual signal (residual block or residual sample) can be reconstructed by applying inverse quantization and inverse transform to the quantized transform coefficients via the inverse quantization unit 140 and the inverse transform unit 150.
[0076] The adder 155 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter prediction unit 180 or the intra prediction unit 185. When there is no residual for the current block to be processed, such as when a skip mode is applied, the predicted block may be used as the reconstructed block. The adder 155 may be referred to as a reconstruction unit or a reconstructed block generation unit. The generated reconstructed signal may be used for intra prediction of the next current block to be processed in the current picture, and may also be used for inter prediction of the next picture after filtering, as will be described later.
[0077] The filtering unit 160 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 160 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture and store the modified reconstructed picture in the memory 170, specifically, in the DPB of the memory 170. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, etc. The filtering unit 160 may generate various information related to filtering and transmit it to the entropy coding unit 190, as will be described later in connection with each filtering method. The filtering information may be coded by the entropy coding unit 190 and output in a bitstream format.
[0078] The modified reconstructed picture transmitted to the memory 170 can be used as a reference picture in the inter prediction unit 180. When inter prediction is applied through this, the image encoding device 100 can avoid a prediction mismatch between the image encoding device 100 and the image decoding device, and can also improve encoding efficiency.
[0079] The DPB in the memory 170 may store modified reconstructed pictures for use as reference pictures in the inter predictor 180. The memory 170 may store motion information of blocks from which motion information in the current picture is derived (or coded) and / or motion information of already reconstructed intra-picture blocks. The stored motion information may be transmitted to the inter predictor 180 to be used as motion information of spatially surrounding blocks or temporally surrounding blocks. The memory 170 may store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra predictor 185.
[0080] Overview of the image decoding device
[0081] FIG. 3 is a diagram schematically illustrating an image decoding device to which an embodiment of the present disclosure can be applied.
[0082] 3, the image decoding apparatus 200 may include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an adder 235, a filtering unit 240, a memory 250, an inter prediction unit 260, and an intra prediction unit 265. The inter prediction unit 260 and the intra prediction unit 265 may be collectively referred to as a "prediction unit." The inverse quantization unit 220 and the inverse transform unit 230 may be included in a residual processing unit.
[0083] Depending on the embodiment, all or at least some of the components constituting the image decoding device 200 may be realized by a single hardware component (e.g., a decoder or a processor). Also, the memory 170 may include a DPB and may be realized by a digital storage medium.
[0084] The image decoding device 200, which receives a bitstream including video / image information, can reconstruct an image by performing a process corresponding to the process performed by the image encoding device 100 of FIG. 2. For example, the image decoding device 200 can perform decoding using a processing unit applied in the image encoding device. Therefore, the decoding processing unit can be, for example, a coding unit. The coding unit can be obtained by dividing a coding tree unit or a maximum coding unit. The reconstructed image signal decoded and output by the image decoding device 200 can be reproduced by a reproduction device (not shown).
[0085] The image decoding apparatus 200 may receive a signal output from the image encoding apparatus of FIG. 2 in a bitstream format. The received signal may be decoded via an entropy decoding unit 210. For example, the entropy decoding unit 210 may parse the bitstream to derive information (e.g., video / image information) necessary for image reconstruction (or picture reconstruction). The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / image information may also include general constraint information. The image decoding apparatus may further use the information on the parameter sets and / or the general constraint information to decode an image. The signaling information, received information, and / or syntax elements referred to in the present disclosure may be obtained from the bitstream by being decoded via the decoding procedure. For example, the entropy decoding unit 210 may decode information in a bitstream based on a coding method such as Exponential-Golomb coding, CAVLC, or CABAC, and output values of syntax elements required for image restoration and quantized values of transform coefficients related to residuals. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element from the bitstream, determines a context model using information on the syntax element to be decoded and decoded information on neighboring blocks and the block to be decoded, or information on symbols / bins decoded in a previous step, predicts the occurrence probability of the bins based on the determined context model, and performs arithmetic decoding of the bins to generate symbols corresponding to the values of each syntax element. After determining the context model, the CABAC entropy decoding method may update the context model using information on the decoded symbol / bin for the context model of the next symbol / bin.Among the information decoded by the entropy decoding unit 210, information related to prediction is provided to the prediction units (inter prediction unit 260 and intra prediction unit 265), and residual values entropy decoded by the entropy decoding unit 210, i.e., quantized transform coefficients and related parameter information, may be input to the inverse quantization unit 220. Also, among the information decoded by the entropy decoding unit 210, information related to filtering may be provided to the filtering unit 240. Meanwhile, a receiving unit (not shown) for receiving a signal output from the image encoding device may be further provided as an internal / external element of the image decoding device 200, or the receiving unit may be provided as a component of the entropy decoding unit 210.
[0086] Meanwhile, the image decoding apparatus according to the present disclosure may be referred to as a video / image / picture decoding apparatus. The image decoding apparatus may include an information decoder (video / image / picture information decoder) and / or a sample decoder (video / image / picture sample decoder). The information decoder may include an entropy decoding unit 210, and the sample decoder may include at least one of an inverse quantization unit 220, an inverse transform unit 230, an adder 235, a filtering unit 240, a memory 250, an inter prediction unit 260, and an intra prediction unit 265.
[0087] The inverse quantization unit 220 may inverse quantize the quantized transform coefficients and output the transform coefficients. The inverse quantization unit 220 may rearrange the quantized transform coefficients in a two-dimensional block format. In this case, the rearrangement may be performed based on the coefficient scanning order performed in the image encoding device. The inverse quantization unit 220 may perform inverse quantization on the quantized transform coefficients using a quantization parameter (e.g., quantization step size information) to obtain transform coefficients.
[0088] The inverse transform unit 230 can inversely transform the transform coefficients to obtain a residual signal (residual block, residual sample array).
[0089] The prediction unit may perform prediction on a current block and generate a predicted block including prediction samples for the current block. The prediction unit may determine whether intra prediction or inter prediction is applied to the current block based on information about the prediction output from the entropy decoding unit 210, and may determine a specific intra / inter prediction mode (prediction technique).
[0090] The prediction unit can generate a prediction signal based on various prediction methods (techniques) described below, as described in the description of the prediction unit of the image encoding device 100.
[0091] The intra predictor 265 may predict the current block by referring to samples in the current picture. The description of the intra predictor 185 may also be applied to the intra predictor 265.
[0092] The inter prediction unit 260 may derive a predicted block for a current block based on a reference block (reference sample array) identified by a motion vector on a reference picture. To reduce the amount of motion information transmitted in inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on correlations between motion information of neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on an inter prediction direction (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. For example, the inter prediction unit 260 may construct a motion information candidate list based on the neighboring blocks and derive a motion vector and / or a reference picture index for the current block based on received candidate selection information. Inter prediction may be performed based on various prediction modes (techniques), and the prediction information may include information indicating the inter prediction mode (technique) for the current block.
[0093] The adder 235 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to a prediction signal (predicted block, predicted sample array) output from a prediction unit (including the inter prediction unit 260 and / or intra prediction unit 265). When there is no residual for the current block, such as when a skip mode is applied, the predicted block can be used as the reconstructed block. The description of the adder 155 also applies to the adder 235. The adder 235 may also be referred to as a reconstruction unit or a reconstructed block generator. The generated reconstructed signal may be used for intra prediction of the next current block in the current picture, and may also be used for inter prediction of the next picture via filtering, as described below.
[0094] The filtering unit 240 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 240 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture, and may store the modified reconstructed picture in the memory 250, specifically, in a DPB of the memory 250. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, etc.
[0095] The (modified) reconstructed picture stored in the DPB of the memory 250 can be used as a reference picture in the inter predictor 260. The memory 250 can store motion information of a block from which motion information in the current picture is derived (or decoded) and / or motion information of a block in an already reconstructed picture. The stored motion information can be transmitted to the inter predictor 260 to be used as motion information of a spatially surrounding block or a temporally surrounding block. The memory 250 can store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra predictor 265.
[0096] In this specification, the embodiments described for the filtering unit 160, inter prediction unit 180 and intra prediction unit 185 of the image encoding device 100 can also be applied in a similar or corresponding manner to the filtering unit 240, inter prediction unit 260 and intra prediction unit 265 of the image decoding device 200, respectively.
[0097] General image / video coding procedures
[0098] In image / video coding, pictures constituting an image / video can be coded / decoded according to a sequence of decoding orders. A picture order corresponding to an output order of decoded pictures can be set to be different from the decoding order. Based on this, not only forward prediction but also backward prediction can be performed during inter prediction.
[0099] FIG. 4 shows an example of a general picture decoding procedure to which the embodiments of the present disclosure can be applied.
[0100] Each procedure shown in Fig. 4 may be performed by the image encoding apparatus of Fig. 3. For example, step S410 may be performed by the entropy decoding unit 210, step S420 may be performed by a prediction unit including an intra prediction unit 265 and an inter prediction unit 260, step S430 may be performed by a residual processing unit including an inverse quantization unit 220 and an inverse transform unit 230, step S440 may be performed by the addition unit 235, and step S450 may be performed by the filtering unit 240. Step S410 may include an information decoding procedure described in this disclosure, step S420 may include an inter / intra prediction procedure described in this disclosure, step S430 may include a residual processing procedure described in this disclosure, step S440 may include a block / picture reconstruction procedure described in this disclosure, and step S450 may include an in-loop filtering procedure described in this disclosure.
[0101] Referring to Figure 4, as shown in the description of Figure 3, the picture decoding procedure may generally include an image / video information acquisition procedure (S410) from a bitstream (through decoding), a picture reconstruction procedure (S420-S440), and an in-loop filtering procedure for the reconstructed picture (S450). The picture reconstruction procedure may be performed based on prediction samples and residual samples obtained through the inter / intra prediction (S420) and residual processing (S430, inverse quantization and inverse transform of quantized transform coefficients) processes described in this disclosure. A modified reconstructed picture may be generated through an in-loop filtering procedure for the reconstructed picture generated by the picture reconstruction procedure. The modified reconstructed picture may be output as a decoded picture or may be stored in a decoded picture buffer or memory 250 of the decoding device and used as a reference picture in the inter prediction procedure when decoding a subsequent picture. In some cases, the in-loop filtering procedure may be omitted. In this case, the reconstructed picture may be output as a decoded picture or may be stored in a decoded picture buffer or memory 250 of the decoding device and used as a reference picture in an inter-prediction procedure when decoding a subsequent picture. As described above, the in-loop filtering procedure (S450) may include a deblocking filtering procedure, a sample adaptive offset (SAO) procedure, an adaptive loop filter (ALF) procedure, and / or a bilateral filter procedure, some or all of which may be omitted. Furthermore, one or some of the deblocking filtering procedure, the sample adaptive offset (SAO) procedure, the adaptive loop filter (ALF) procedure, and the bilateral filter procedure may be applied sequentially, or all of them may be applied sequentially. For example, the SAO procedure may be performed after the deblocking filtering procedure is applied to the reconstructed picture.Or, for example, the ALF procedure can be performed after a deblocking filtering procedure is applied to the reconstructed picture, which can also be performed in the encoding device.
[0102] FIG. 5 shows an example of a general picture encoding procedure to which the embodiments of the present disclosure can be applied.
[0103] Each procedure shown in Fig. 5 may be performed by the image encoding apparatus of Fig. 2. For example, step S510 may be performed by a prediction unit including an intra prediction unit 185 or an inter prediction unit 180, step S520 may be performed by a residual processing unit including a transform unit 120 and / or a quantization unit 130, and step S530 may be performed by an entropy encoding unit 190. Step S510 may include an inter / intra prediction procedure described in this disclosure, step S520 may include a residual processing procedure described in this disclosure, and step S530 may include an information encoding procedure described in this disclosure.
[0104] Referring to FIG. 5, the picture encoding procedure, as described with reference to FIG. 2, may include not only a procedure of encoding information for picture reconstruction (e.g., prediction information, residual information, partitioning information, etc.) and outputting it in a bitstream format, but also a procedure of generating a reconstructed picture for a current picture and an optional procedure of applying in-loop filtering to the reconstructed picture. The encoding apparatus may derive (modified) residual samples from quantized transform coefficients via the inverse quantization unit 140 and the inverse transform unit 150, and may generate a reconstructed picture based on the prediction samples output in step S510 and the (modified) residual samples. The reconstructed picture generated in this manner may be the same as the reconstructed picture generated by the decoding apparatus described above. A modified reconstructed picture may be generated through an in-loop filtering procedure on the reconstructed picture, which may be stored in the decoded picture buffer or memory 170 and used as a reference picture in the inter prediction procedure when encoding a subsequent picture, as in the decoding apparatus. As described above, in some cases, some or all of the in-loop filtering procedure may be omitted. When the in-loop filtering procedure is performed, (in-loop) filtering-related information (parameters) can be coded by the entropy coding unit 190 and output in bitstream format, and the decoding device can perform the in-loop filtering procedure in the same manner as the coding device based on the filtering-related information.
[0105] This in-loop filtering procedure can reduce noise that occurs during image / video coding, such as blocking artifacts and ringing artifacts, and improve subjective / objective visual quality. Also, by performing the in-loop filtering procedure in both the encoding device and the decoding device, the encoding device and the decoding device can derive the same prediction result, thereby improving the reliability of picture coding and reducing the amount of data to be transmitted for picture coding.
[0106] As described above, a picture reconstruction procedure may be performed not only in a decoding device but also in an encoding device. Reconstructed blocks may be generated based on intra prediction / inter prediction for each block, and a reconstructed picture including the reconstructed blocks may be generated. If a current picture / slice / tile group is an I picture / slice / tile group, blocks included in the current picture / slice / tile group may be reconstructed based only on intra prediction. On the other hand, if the current picture / slice / tile group is a P or B picture / slice / tile group, blocks included in the current picture / slice / tile group may be reconstructed based on intra prediction or inter prediction. In this case, inter prediction may be applied to some blocks in the current picture / slice / tile group, and intra prediction may be applied to the remaining blocks. Color components of a picture may include luma components and chroma components, and unless explicitly limited in this disclosure, methods and embodiments proposed in this disclosure may be applied to luma components and chroma components.
[0107] Example of coding hierarchy and structure
[0108] Video / images coded according to this disclosure may be processed, for example, according to the coding hierarchy and structure described below.
[0109] FIG. 6 shows an example of a hierarchical structure for coded images / video.
[0110] The coded image / video can be divided into the VCL (video coding layer), which handles the image / video decoding process and itself, the lower system, which transmits and stores the coded information, and the NAL (network abstraction layer), which exists between the VCL and the lower system and is responsible for network adaptation functions.
[0111] The VCL can generate VCL data containing compressed image data (slice data), or it can generate parameter sets containing information such as a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), and a Video Parameter Set (VPS), or an SEI (Supplemental Enhancement Information) message that is additionally required for image decoding processing.
[0112] In NAL, NAL units can be generated by adding header information (NAL unit header) to RBSP (Raw Byte Sequence Payload) generated by VCL. RBSP refers to slice data, parameter sets, SEI messages, etc. generated by VCL. The NAL unit header can include NAL unit type information identified by the RBSP data included in the corresponding NAL unit.
[0113] As shown in Figure 6, NAL units can be classified into VCL NAL units and non-VCL NAL units according to the type of RBSP generated in the VCL. A VCL NAL unit can refer to a NAL unit containing information about an image (slice data), and a non-VCL NAL unit can refer to a NAL unit containing information necessary for decoding an image (parameter set or SEI message).
[0114] The VCL NAL unit and non-VCL NAL unit described above can be transmitted over a network with header information attached according to the data standard of the lower system. For example, the NAL unit can be transformed into a data format of a predetermined standard such as the H.266 / VVC file format, the Real-time Transport Protocol (RTP), or the Transport Stream (TS) and then transmitted over various networks.
[0115] As described above, the NAL unit type of an NAL unit can be identified according to the RBSP data structure included in the NAL unit, and information about the NAL unit type can be stored and signaled in the NAL unit header. For example, NAL units can be broadly classified into VCL NAL unit types and non-VCL NAL unit types depending on whether the NAL unit includes information about an image (slice data). VCL NAL unit types can be classified according to the nature and type of pictures included in the VCL NAL unit, and non-VCL NAL unit types can be classified according to the type of parameter set.
[0116] Below is a list of examples of NAL unit types identified by the type of parameter set / information included in the non-VCL NAL unit type.
[0117] -DCI (Decoding capability information) NAL unit type (NUT): Type for NAL units including DCI
[0118] -VPS (Video Parameter Set) NUT: Type for NAL units containing VPS
[0119] -SPS (Sequence Parameter Set) NUT: Type for NAL units containing SPS
[0120] -PPS (Picture Parameter Set) NUT: Type for NAL units containing PPS
[0121] -APS (Adaptation Parameter Set) NUT: Type for NAL units containing APS
[0122] -PH(Picture header) NUT: Type for NUL units containing picture headers
[0123] The above-mentioned NAL unit type has syntax information for the NAL unit type, and the syntax information can be stored in a NAL unit header and signaled. For example, the syntax information is nal_unit_type, and the NAL unit type can be identified using the value of nal_unit_type.
[0124] Meanwhile, one picture may include multiple slices, and one slice may include a slice header and slice data. In this case, one picture header may be added to multiple slices (slice header and slice data set) in one picture. The picture header (picture header syntax) may include information / parameters commonly applicable to the picture. The slice header (slice header syntax) may include information / parameters commonly applicable to the slices. The APS (APS syntax) or PPS (PPS syntax) may include information / parameters commonly applicable to one or more slices or pictures. The SPS (SPS syntax) may include information / parameters commonly applicable to one or more sequences. The VPS (VPS syntax) may include information / parameters commonly applicable to multiple layers. The DCI may include information / parameters related to decoding capability.
[0125] In the present disclosure, the high level syntax (HLS) may include at least one of the APS syntax, PPS syntax, SPS syntax, VPS syntax, DCI syntax, picture header syntax, and slice header syntax. Also, in the present disclosure, the low level syntax (LLS) may include, for example, slice data syntax, CTU syntax, coding unit syntax, transform unit syntax, etc.
[0126] Meanwhile, in the present disclosure, image / video information encoded from an encoding device to a decoding device and signaled in a bitstream format may include not only intra-picture partitioning-related information, intra / inter prediction information, residual information, in-loop filtering information, etc., but also the slice header information, the picture header information, the APS information, the PPS information, the SPS information, the VPS information, and / or the DCI information. Also, the image / video information may further include general constraint information and / or NAL unit header information.
[0127] High level syntax signaling and semantics
[0128] As described above, the image / video information according to the present disclosure may include a High Level Syntax (HLS), and an image encoding method and / or an image decoding method may be performed based on the image / video information.
[0129] DPB parameter signaling
[0130] A Decoded Picture Buffer (DPB) can conceptually be composed of sub-DPBs. Each sub-DPB can include a picture storage buffer for storing decoded pictures of one layer. Each picture storage buffer can contain decoded pictures marked as "used for reference" or decoded pictures being held for future output.
[0131] In a multilayer bitstream, DPB parameters are not assigned per output layer set (OLS), but per layer. Up to two DPB parameters can be assigned to each layer. One of the two DPB parameters is a DPB parameter for when the layer is an output layer, and the other is a DPB parameter for when the layer is used as a reference layer rather than an output layer. When the layer is an output layer, the layer may be used as a reference layer and for future output. When the layer is used as a reference layer rather than an output layer, the layer can only be used to reference pictures / slices / blocks in the output layer unless there is layer switching. According to the prior art, DPB parameters are signaled for each layer in an OLS. Signaling the DPB parameters can simplify the signaling of conventional DPB parameters.
[0132] FIG. 7 is a diagram illustrating an exemplary syntax structure of a VPS according to one embodiment of the present disclosure.
[0133] According to the example shown in FIG. 7, when vps_all_independent_layers_flag is 0, vps_num_dpb_params may be signaled. As will be described later, a first value (e.g., 0) of vps_all_independent_layers_flag may indicate that one or more layers in a coded video sequence (CVS) can use inter-layer prediction. Furthermore, a second value (e.g., 1) of vps_all_independent_layers_flag may indicate that all layers in a coded video sequence (CVS) are coded independently without using inter-layer prediction. In the above, a CVS may be understood as a bitstream or image / video information including a sequence of coded pictures for multiple layers. vps_num_dpb_params may indicate the number of dpb_parameters() syntax structures included in a video parameter set (VPS). For example, vps_num_dpb_params can have a value between 0 and 16, and if not present, the value of vps_num_dpb_params can be inferred (set) to 0.
[0134] If the value of vps_num_dpb_params is greater than 0, i.e., if the number of dpb_parameters() syntax structures included in the VPS is greater than 0, one or more DPB parameters may be signaled. For example, the one or more DPB parameters may include same_dpb_size_output_or_nonoutput_flag, vps_sublayer_dpb_params_present_flag, dpb_size_only_flag[i], dpb_max_temporal_id[i], layer_output_dpb_params_idx[i], layer_nonoutput_dpb_params_idx[i], and / or dpb_parameters().
[0135] A first value (e.g., 1) of same_dpb_size_output_or_nonoutput_flag can indicate that layer_nonoutput_dpb_params_idx[i] is not present in the VPS. A second value (e.g., 0) of same_dpb_size_output_or_nonoutput_flag can indicate that layer_nonoutput_dpb_params_idx[i] may be present in the VPS.
[0136] vps_sublayer_dpb_params_present_flag can be used to control whether max_dec_pic_buffering_minus1[], max_num_reorder_pics[], and / or max_latency_increase_plus1[] are present in dpb_parameters() in the VPS. If vps_sublayer_dpb_params_present_flag is not present, its value can be inferred as 0.
[0137] A first value (e.g., 1) of dpb_size_only_flag[i] may indicate that max_num_reorder_pics[] and / or max_latency_increase_plus1[] are not present in the i-th dpb_parameters() in the VPS. A second value (e.g., 0) of dpb_size_only_flag[i] may indicate that max_num_reorder_pics[] and / or max_latency_increase_plus1[] may be present in the i-th dpb_parameters() in the VPS.
[0138] dpb_max_temporal_id[i] may represent the temporal layer identifier (e.g., TemporalId) of the highest sublayer in which a DPB parameter may exist in the i-th dpb_parameters() in the VPS. dpb_max_temporal_id[i] may have a value between 0 and vps_max_sublayers_minus1. If vps_max_sublayers_minus1 is 0, the value of dpb_max_temporal_id[i] is not signaled and may be inferred to be 0. If vps_max_sublayers_minus1 is greater than 0 and vps_all_layers_same_num_sublayers_flag is 1, the value of dpb_max_temporal_id[i] may be inferred to be the same value as vps_max_sublayers_minus1.
[0139] layer_output_dpb_params_idx[i] is an index into the list of dpb_parameters() in the VPS, and can be an index indicating the dpb_parameters() that apply to the i-th layer when the i-th layer is an output layer included in the OLS. layer_output_dpb_params_idx[i] can have a value from 0 to vps_num_dpb_params-1.
[0140] If vps_independent_layer_flag[i] is 1, the dpb_parameters() applied to the i-th layer, which is the output layer, may be the dpb_parameters() present in the SPS referenced by that layer.
[0141] Otherwise, if vps_independent_layer_flag[i] is 0, the following applies:
[0142] -If vps_num_dpb_params is 1, the value of layer_output_dpb_params_idx[i] can be inferred to be 0.
[0143] -As a requirement for bitstream integrity, the value of layer_output_dpb_params_idx[i] must be such that dpb_size_only_flag[layer_output_dpb_params_idx[i]] is 0.
[0144] layer_nonoutput_dpb_params_idx[i] is an index into the list of dpb_parameters() in the VPS, and can be an index indicating the dpb_parameters() that apply to the i-th layer when the i-th layer is a non-output layer included in the OLS. layer_nonoutput_dpb_params_idx[i] can have a value from 0 to vps_num_dpb_params-1.
[0145] If same_dpb_size_output_or_nonoutput_flag is 1, the following applies:
[0146] - If vps_independent_layer_flag[i] is 1, the dpb_parameters() applied to the i-th layer, which is a non-output layer, may be the dpb_parameters() present in the SPS referenced by that layer.
[0147] - Otherwise, if vps_independent_layer_flag[i] is 0, the value of layer_nonoutput_dpb_params_idx[i] can be inferred to be the same value as layer_output_dpb_params_idx[i].
[0148] Otherwise, if same_dpb_size_output_or_nonoutput_flag is 0, then if vps_num_dpb_params is 1, the value of layer_output_dpb_params_idx[i] can be inferred to be 0.
[0149] FIG. 8 is a diagram illustrating a syntax structure for signaling DPB parameters according to the present disclosure.
[0150] As shown in Figure 8, the dpb_parameters() syntax structure may include information about the DPB size, the maximum number of picture reorders, and / or the maximum latency for each CLVS of a CVS. In the above, a CLVS (Coder Layer Video Sequence) may be understood as a bitstream or image / video information including a sequence of coded pictures belonging to the same layer.
[0151] When the dpb_parameters() syntax structure is included in a VPS, the OLS to which the dpb_parameters() syntax structure is applied can be specified by the VPS. When the dpb_parameters() syntax structure is included in an SPS, it can be applied to an OLS that includes only the lowest layer among layers that reference the SPS. In this case, the lowest layer may be an independent layer.
[0152] max_dec_pic_buffering_minus1[i]+1 can indicate the maximum required size of the DPB for each CLVS of the CVS. max_dec_pic_buffering_minus1[i] can have a value between 0 and MaxDpbSize-1.
[0153] max_num_reorder_pics[i] may indicate, for each CLVS of a CVS, the maximum allowable number of pictures in the CLVS that precede any picture in the CLVS in decoding order and follow it in output order. max_num_reorder_pics[i] may have values from 0 to max_dec_pic_buffering_minus1[i]. If i is greater than 0, max_num_reorder_pics[i] must have a value greater than or equal to max_num_reorder_pics[i1]. If max_num_reorder_pics[i] is not present, its value may be inferred to be the same as max_num_reorder_pics[maxSubLayersMinus1].
[0154] A non-zero value of max_latency_increase_plus1[i] can be used to calculate MaxLatencyPictures[i], which can indicate, for each CLVS in a CVS, the maximum number of pictures in the CLVS that precede any picture in the CLVS in output order and follow it in decoding order.
[0155] If max_latency_increase_plus1[i] is not 0, MaxLatencyPictures[i] can be calculated as follows:
[0156] MaxLatencyPictures[i]=max_num_reorder_pics[i]+max_latency_increase_plus1[i]-1
[0157] max_latency_increase_plus1[i] is between 0 and 2 32 It can have a value of -2. If max_latency_increase_plus1[i] is not present, its value can be inferred the same as max_latency_increase_plus1[maxSubLayersMinus1].
[0158] The DPB parameters can be used in the process of outputting or removing a decoded image from the DPB.
[0159] Video Parameter Set signaling
[0160] A video parameter set (VPS) is a parameter set used for transmitting layer information. The layer information may include, for example, information on an output layer set (OLS), information on a profile tier level, information on a relationship between an OLS and a hypothetical reference decoder, information on a relationship between an OLS and a DPB, etc. The VPS may not be essential for decoding a bitstream.
[0161] Before it can be referenced, the VPS raw byte sequence payload (RBSP) must be available to the decoding process, either by being included in at least one Access Unit (AU) with TemporalID equal to 0, or by being provided via external means.
[0162] All VPS NAL units with a particular value of vps_video_parameter_set_id in a coded video sequence (CVS) must have the same content.
[0163] FIG. 9 is a diagram illustrating an exemplary syntax structure of a VPS according to another embodiment of the present disclosure.
[0164] The syntax structure of the VPS shown in FIG. 9 includes only syntax elements relevant to the present disclosure, and various other syntax elements not shown in FIG. 9 may be included in the VPS.
[0165] In the example shown in Figure 9, vps_video_parameter_set_id provides an identifier for the VPS. Other syntax elements can reference the VPS using vps_video_parameter_set_id. The value of vps_video_parameter_set_id must be greater than 0.
[0166] The value of vps_max_layers_minus1 plus 1 may indicate the maximum number of layers allowable within each CVS that references the VPS.
[0167] vps_all_independent_layers_flag may be signaled if vps_max_layers_minus1 is greater than 0. A first value (e.g., 1) of vps_all_independent_layers_flag may indicate that all layers in the CVS are coded independently without using inter-layer prediction. A second value (e.g., 0) of vps_all_independent_layers_flag may indicate that one or more layers in the CVS can utilize inter-layer prediction. If vps_all_independent_layers_flag is not present, its value may be inferred as the first value (e.g., 1).
[0168] vps_num_ptls_minus1+1 may indicate the number of profile_tier_level() syntax structures in the VPS. The value of vps_num_ptls_minus1 must be less than TotalNumOlss. TotalNumOlss may indicate the total number of OLSs specified by the VPS. When vps_max_layers_minus1 is 0, TotalNumOlss may be derived to 1. Otherwise, if each_layer_is_an_ols_flag is 1 or ols_mode_idc is 0 or 1, TotalNumOlss may be derived to vps_max_layers_minus1+1. Otherwise, if ols_mode_idc is 2, TotalNumOlss may be derived to num_output_layer_sets_minus1+1. ols_mode_idc may be an indicator indicating the mode for deriving the total number of OLSs specified by the VPS. In the above, each_layer_is_an_ols_flag can indicate whether each OLS contains only one layer.
[0169] ols_ptl_idx[i] is an index into the list of profile_tier_level() in the VPS and may be the index of the profile_tier_level() that applies to the i-th OLS. When ols_ptl_idx[i] exists, it may have a value from 0 to vps_num_ptls_minus1. If vps_num_ptls_minus1 is 0, the value of ols_ptl_idx[i] may be inferred to be 0.
[0170] NumLayersInOls[i] may indicate the number of layers in the i-th OLS. When NumLayersInOls[i] is 1, the profile_tier_level() syntax structure that applies to the i-th OLS may also exist in the SPS referenced by a layer in the i-th OLS. As a requirement for bitstream consistency, when NumLayersInOls[i] is 1, the profile_tier_level() syntax structure in the VPS for the i-th OLS and the profile_tier_level() syntax structure in the SPS must be identical.
[0171] As shown in FIG. 9, when vps_all_independent_layers_flag is 0, vps_num_dpb_params can be signaled.
[0172] vps_num_dpb_params may indicate the number of DPB parameters (dpb_parameters()) syntax structures in the VPS. vps_num_dpb_params may have a value from 0 to 16, and if not present in the bitstream, its value may be inferred as 0.
[0173] Based on vps_num_dpb_params, the dpb_parameter() syntax structure can be signaled.
[0174] Also, if the number of layers in the i-th OLS is multiple (NumLayersInOls[i]>1) and the number of DPB parameters in the VPS is multiple (vps_num_dpb_params>1), ols_dpb_params_idx[i] can be signaled.
[0175] ols_dpb_params_idx[i] is an index into the list of dpb_parameters() in the VPS and may be the index of the dpb_parameters() that applies to the i-th OLS. When ols_dpb_params_idx[i] is present, it may have a value from 0 to vps_num_dpb_params-1. When ols_dpb_params_idx[i] is not present in the bitstream, its value may be inferred to be 0.
[0176] To improve the signaling of DPB-related information and PTL-related information within the VPS described with reference to Figures 7 to 9, an embodiment according to the present disclosure may include at least one of the following configurations, which may be applied individually or in combination with other configurations.
[0177] Configuration 1: Each of the DPB parameter syntax structures signaled in the VPS can be restricted to be associated with at least one output layer set (OLS).
[0178] Configuration 2: Alternatively, each DPB parameter syntax structure signaled in the VPS can be restricted to be associated with at least one multi-layer OLS, where the multi-layer OLS can refer to an output layer set including more than one layer.
[0179] Configuration 3: The number of DPB parameter syntax structures signaled in the VPS (i.e., vps_num_dpb_params) can be limited to not be greater than the number of output layer sets (multi-layer OLSs) containing more than one layer. In this case, the number of multi-layer OLSs can be derived by subtracting the number of OLSs containing only one layer from the total number of OLSs.
[0180] Configuration 4: Each PTL syntax structure signaled in the VPS can be restricted to be associated with at least one output layer set (OLS).
[0181] FIG. 10 is a diagram illustrating an example of an image encoding method to which an embodiment of the present disclosure can be applied.
[0182] The image encoding apparatus may induce DPB-related information and / or PTL-related information (S1010) and encode image / video information (S1020). In this case, the image / video information may include the induced DPB-related information and / or PTL-related information.
[0183] 10, the image coding apparatus may perform DPB management based on the DPB-related information derived in step S1010. Also, the image coding apparatus may process (encode) the current picture based on the DPB-related information and / or the PTL-related information. Alternatively, the image coding apparatus may process the current OLS based on the DPB-related information and / or the PTL-related information.
[0184] FIG. 11 is a diagram illustrating an example of an image decoding method to which an embodiment of the present disclosure can be applied.
[0185] The image decoding apparatus may obtain image / video information from the bitstream (S1110). At this time, the image / video information may include DPB-related information and / or PTL-related information.
[0186] The image decoding apparatus may process (decode) the picture based on the acquired DPB-related information and / or PTL-related information (S1120), or may process the current OLS based on the DPB-related information and / or PTL-related information.
[0187] FIG. 12 is a diagram for explaining another example of an image decoding method to which an embodiment of the present disclosure can be applied.
[0188] The image decoding apparatus may obtain image / video information from the bitstream (S1210). At this time, the image / video information may include DPB-related information and / or PTL-related information.
[0189] The image decoding apparatus can perform DPB management based on the acquired DPB-related information (S1220).
[0190] The image decoding apparatus may decode a picture based on the DPB (S1230). For example, a block / slice in a current picture may be decoded based on inter prediction using an already reconstructed picture in the DPB as a reference picture. Also, although not shown in FIG. 12, the image coding apparatus may decode the current picture based on the PTL-related information obtained in step S1210. Alternatively, the image decoding apparatus may process the current OLS based on the DPB-related information and / or the PTL-related information.
[0191] In the examples described with reference to FIGS. 10 to 12, the DPB-related information and / or PTL-related information may include at least one of the information / syntax elements described in connection with at least one of the embodiments of the present disclosure. Specifically, the DPB-related information may include the information / syntax elements described with reference to FIGS. 7 and 8. Furthermore, the PTL-related information may include the information / syntax elements described with reference to FIG. 9. The image / video information of the present disclosure may further include information described in connection with an output layer set (OLS) in the present disclosure. Furthermore, as described above, DPB management may be performed based on the DPB-related information. For example, storing a picture in the DPB before decoding the current picture, deleting a picture from the DPB, and / or outputting a (decoded) picture may be performed based on the DPB-related information. The DPB management may be further performed based on the OLS described in the present disclosure.
[0192] Various embodiments of the present disclosure for efficiently performing VPS signaling will be described below. The various embodiments of the present disclosure described below may be implemented alone or in combination with other embodiments.
[0193] According to one embodiment of the present disclosure, the number of multi-layer OLSs (NumMultiLayeredOlss) including multiple layers can be derived using NumLayersInOls[i]. As described above, NumLayersInOls[i] may indicate the number of layers in the i-th OLS. Specifically, to derive the number of multi-layer OLSs (NumMultiLayeredOlss), the number of multi-layer OLSs (NumMultiLayeredOlss) can be first set (initialized) to 0. Then, for all OLSs identified by the VPS (i = 1 to TotalNumOlss), it is checked whether NumLayersInOls[i] is greater than 1. If NumLayersInOls[i] is greater than 1, the value of NumMultiLayeredOlss is incremented by 1 to derive the number of multi-layer OLSs. The number of multi-layer OLSs derived as described above can be used in other embodiments of the present disclosure.
[0194] According to another embodiment of the present disclosure, each DPB parameter structure signaled in the VPS can be restricted to be associated with at least one OLS.
[0195] As described above, if the i-th OLS includes more than one layer, the dpb_parameters() applied to the i-th OLS can be identified by ols_dpb_params_idx[i]. If ols_dpb_params_idx[i] is present in the bitstream, its value can have a value between 0 and vps_num_dpb_params-1. If ols_dpb_params_idx[i] is not present in the bitstream, its value can be inferred as 0. If the i-th OLS includes only one layer (when NumLayersInOls[i] is equal to 1), the dpb_parameters() applied to the i-th OLS can be present in the SPS referenced by the layer in the i-th OLS. In this case, the dpb_parameters() applied to the i-th OLS can not be signaled in the VPS.
[0196] According to this embodiment, all dpb_parameters() signaled in a VPS can be restricted to be applied to at least one OLS. That is, each dpb_parameter() in a VPS can be restricted to be identified (referenced) by at least one ols_dpb_params_idx[i]. According to this embodiment, each dpb_parameter() in a VPS can be used at least once. That is, according to this embodiment, by not signaling unused dpb_parameters(), dpb_parameters() in a VPS can be signaled efficiently.
[0197] As described above, ols_dpb_params_idx[i] is an index into the list of dpb_parameters() in the VPS, and is the index of the dpb_parameters() that applies to the i-th OLS when the i-th OLS is a multi-layered OLS. That is, according to another embodiment of the present disclosure, ols_dpb_params_idx[i] is an index into the list of dpb_parameters() in the VPS, and can indicate the index of the dpb_parameters() that applies to the i-th multi-layered OLS. In this case, each of all dpb_parameters() signaled in the VPS can be restricted to be applied to at least one multi-layered OLS. That is, each dpb_parameters() in the VPS can be restricted to be referenced by at least one ols_dpb_params_idx[i] (i is 0 to NumMultiLayeredOlss-1).
[0198] According to another embodiment of the present disclosure, the number of DPB parameter structures signaled in the VPS (i.e., vps_num_dpb_params) can be limited to not be greater than the number of multi-layer OLSs (NumMultiLayeredOlss) that include more than one layer.
[0199] According to the example described with reference to FIG. 9, vps_num_dpb_params indicates the number of DPB parameters (dpb_parameters()) syntax structures in the VPS, and vps_num_dpb_params can have a value of 0-16.
[0200] However, as mentioned above, an OLS containing only one layer can exist, and an OLS containing only one layer is not associated with a DPB parameter structure signaled in the VPS. The range of values for vps_num_dpb_params in the example of Figure 9 may result in inaccurate signaling. Therefore, the range of the number of dpb_parameters() syntax structures must be specified by the number of multi-layer OLSs containing multiple layers (NumMultiLayeredOlss) to ensure accurate signaling.
[0201] In this case, the number of multi-layer OLSs (NumMultiLayeredOlss) can be derived by subtracting the number of OLSs including only one layer (NumSingleLayerOlss) from the total number of OLSs (TotalNumOlss). Alternatively, as described above, the number of multi-layer OLSs (NumMultiLayeredOlss) can be derived using NumLayersInOls[i].
[0202] As described above, each DPB parameter syntax structure signaled in the VPS may be restricted to be associated (mapped) with at least one multi-layer OLS. Also, the number of DPB parameter structures signaled in the VPS (i.e., vps_num_dpb_params) may be restricted to not be greater than the number of multi-layer OLSs. The above two embodiments may be combined to form another embodiment as follows.
[0203] FIG. 13 is a diagram illustrating a process of encoding DPB parameters based on information on the number of DPB parameters according to another embodiment of the present disclosure.
[0204] The image coding apparatus may encode information regarding the number of DPB parameters (e.g., vps_num_dpb_params) into the VPS (S1310). vps_num_dpb_params may indicate the number of dpb_parameters() syntax structures included in the VPS. As described above, the number of DPB parameters may be limited to not be greater than the number of multi-layered OLSs. Therefore, vps_num_dpb_params may have a value between 0 and NumMultiLayeredOlss. If there are no DPB parameters in the VPS, i.e., if the number of DPB parameters in the VPS is 0, the image coding apparatus may omit encoding vps_num_dpb_params. If encoding vps_num_dpb_params is omitted, its value may be inferred (set) to 0.
[0205] The image coding apparatus can encode the DPB parameters in the VPS (S1320). The image coding apparatus can encode vps_num_dpb_params dpb_parameters() syntax structures into the VPS based on the number of DPB parameters (vps_num_dpb_params) in the VPS (S1320).
[0206] The image coding apparatus may determine a condition for signaling OLS-DPB mapping information (S1330). In this case, the OLS-DPB mapping information may correspond to the above-mentioned ols_dpb_params_idx[i]. In the present disclosure, the OLS-DPB mapping information may be information regarding mapping between one or more multi-layer OLSs and one or more DPB parameter syntax structures.
[0207] The image coding apparatus may determine whether multiple parameters exist in the VPS. For example, the image coding apparatus may determine whether the number of DPB parameters (vps_num_dpb_params) in the VPS is greater than 1. If multiple DPB parameters exist in the VPS (S1330-Yes), the image coding apparatus may encode OLS-DPB mapping information into the VPS (S1340). The image coding apparatus may encode ols_dpb_params_idx[i] as the OLS-DPB mapping information. As described above, ols_dpb_params_idx[i] is an index to the list of dpb_parameters() in the VPS and may indicate the index of the dpb_parameters() that applies to the i-th multi-layer OLS. In this case, each of all dpb_parameters() signaled in the VPS may be restricted to be applied to at least one multi-layer OLS. That is, each dpb_parameters() in a VPS can be restricted to be referenced by at least one ols_dpb_params_idx[i] (i is 0 to NumMultiLayeredOlss-1). In addition, the dpb_parameters() syntax structure applied to an OLS that includes only a single layer can be coded in an SPS referenced by the layer in the OLS, rather than coded in the VPS.
[0208] If there are no DPB parameters in the VPS (S1330-No), the image coding apparatus may not code the OLS-DPB mapping information into the VPS (S1350). If the OLS-DPB mapping information is not coded, its value is inferred to be 0, as described above.
[0209] 13, the signaling condition for the OLS-DPB mapping information is determined based on whether the number of DPB parameters in the VPS is greater than 1. However, the signaling condition for the OLS-DPB mapping information is not limited thereto, and various other conditions not described in the present disclosure may also be determined.
[0210] The image coding apparatus may code the picture based on the DPB parameters (S1360). In this case, the DPB parameters may be dpb_parameters() in the VPS referenced by ols_dpb_params_idx[i] coded in step S1340 or inferred in step S1350. Alternatively, in the case of an OLS including only a single layer, the DPB parameters may be dpb_parameters() in the SPS referenced by the single layer.
[0211] FIG. 14 is a diagram illustrating a process of decoding DPB parameters based on information on the number of DPB parameters according to another embodiment of the present disclosure.
[0212] The image decoding apparatus may obtain information regarding the number of DPB parameters (e.g., vps_num_dpb_params) from the VPS (S1410). vps_num_dpb_params may indicate the number of dpb_parameters() syntax structures included in the VPS. As described above, the number of DPB parameters may be limited to not be greater than the number of multi-layered OLSs. Therefore, vps_num_dpb_params may have a value between 0 and NumMultiLayeredOlss. If vps_num_dpb_params does not exist, its value may be inferred (set) to 0.
[0213] The image decoding apparatus can acquire the DPB parameters in the VPS (S1420). The image decoding apparatus can acquire vps_num_dpb_params dpb_parameters() syntax structures from the VPS based on the number of DPB parameters in the VPS (vps_num_dpb_params) (S1420).
[0214] The image decoding apparatus may determine a condition for signaling OLS-DPB mapping information (S1430). In this case, the OLS-DPB mapping information may correspond to the above-mentioned ols_dpb_params_idx[i].
[0215] The image decoding apparatus may determine whether there are multiple parameters in the VPS. For example, the image decoding apparatus may determine whether the number of DPB parameters in the VPS (vps_num_dpb_params) is greater than 1. If there are multiple DPB parameters in the VPS (S1430-Yes), the image decoding apparatus may acquire OLS-DPB mapping information from the VPS (S1440). The image decoding apparatus may acquire ols_dpb_params_idx[i] as the OLS-DPB mapping information. As described above, ols_dpb_params_idx[i] is an index to the list of dpb_parameters() in the VPS and may indicate the index of dpb_parameters() applied to the i-th multi-layer OLS. In this case, each of all dpb_parameters() signaled in the VPS may be restricted to be applied to at least one multi-layer OLS. That is, each dpb_parameters() in a VPS can be restricted to be referenced by at least one ols_dpb_params_idx[i] (i is 0 to NumMultiLayeredOlss-1). Also, the dpb_parameters() syntax structure applied to an OLS that includes only a single layer can be obtained not from the VPS but from the SPS referenced by the layer in the OLS.
[0216] If there are no DPB parameters in the VPS (S1430-No), the image decoding apparatus may not acquire OLS-DPB mapping information from the VPS (S1450). If the OLS-DPB mapping information is not present in the bitstream, its value may be inferred to be 0.
[0217] 14, the signaling condition for the OLS-DPB mapping information is determined based on whether the number of DPB parameters in the VPS is greater than 1. However, the signaling condition for the OLS-DPB mapping information is not limited thereto, and various other conditions not described in the present disclosure may also be determined.
[0218] The image decoding apparatus may decode the picture based on the DPB parameters (S1460). At this time, the DPB parameters may be dpb_parameters() in the VPS referenced by ols_dpb_params_idx[i] obtained in step S1440 or inferred in step S1450. Alternatively, in the case of an OLS including only a single layer, the DPB parameters may be dpb_parameters() in the SPS referenced by the single layer.
[0219] According to the embodiments described with reference to Figures 13 and 14, unnecessary signaling of the dpb_parameters() syntax structure in an unreferenced VPS can be prevented, and the dpb_parameters() syntax structure for an OLS that includes only a single layer is signaled only via the SPS, thereby achieving the effect of more accurate and efficient signaling of DPB parameters.
[0220] In the methods described with reference to Figures 13 and 14, some steps may be omitted or the order of other steps may be changed, and steps not shown in Figures 13 and 14 may be added at any position.
[0221] According to another embodiment of the present disclosure, each PTL syntax structure signaled in a VPS can be restricted to be associated with at least one OLS.
[0222] As described above, ols_ptl_idx[i] is an index into the list of profile_tier_level() in the VPS and may be the index of the profile_tier_level() that applies to the i-th OLS. When ols_ptl_idx[i] exists, ols_ptl_idx[i] may have a value from 0 to vps_num_ptls_minus1. If vps_num_ptls_minus1 is 0, the value of ols_ptl_idx[i] may be inferred to be 0. In the present disclosure, the ols_ptl_idx may be information regarding the mapping between one or more OLSs and one or more PTL syntax structures.
[0223] NumLayersInOls[i] may indicate the number of layers in the i-th OLS. When NumLayersInOls[i] is 1, the profile_tier_level() syntax structure applied to the i-th OLS may also exist in the SPS referenced by a layer in the i-th OLS. As a requirement for bitstream consistency, when NumLayersInOls[i] is 1, the profile_tier_level() syntax structure in the VPS for the i-th OLS and the profile_tier_level() syntax structure in the SPS must be the same.
[0224] According to this embodiment, each of all profile_tier_level() signaled in a VPS can be restricted to be applied to at least one OLS. That is, each profile_tier_level() in a VPS can be identified by at least one ols_ptl_idx[i]. According to this embodiment, each profile_tier_level() in a VPS is used at least once. That is, unused profile_tier_level() is not signaled. Therefore, according to this embodiment, profile_tier_level() can be signaled efficiently.
[0225] Although the exemplary method of the present disclosure is expressed as a series of operations for clarity of explanation, this is not intended to limit the order in which the steps are performed, and the steps may be performed simultaneously or in a different order if necessary. To achieve the method according to the present disclosure, the steps illustrated may include other steps, or some steps may be omitted and the remaining steps may be included, or some steps may be omitted and additional other steps may be included.
[0226] In the present disclosure, an image encoding device or an image decoding device that performs a predetermined operation (step) can perform the operation (step) to check the execution conditions and circumstances of the operation (step). For example, if it is described that a predetermined operation is performed when a predetermined condition is satisfied, the image encoding device or the image decoding device can perform the predetermined operation after performing an operation to check whether the predetermined condition is satisfied.
[0227] The various embodiments of the present disclosure are not intended to enumerate all possible combinations, but are intended to describe representative aspects of the present disclosure, and the matters described in the various embodiments may be applied independently or in combination of two or more.
[0228] Additionally, various embodiments of the present disclosure may be implemented using hardware, firmware, software, or a combination thereof, etc. In the case of a hardware implementation, the implementation may be using one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), general processors, controllers, microcontrollers, microprocessors, etc.
[0229] In addition, an image decoding apparatus and an image encoding apparatus to which an embodiment of the present disclosure is applied may be included in a multimedia broadcast transmitting / receiving apparatus, a mobile communication terminal, a home cinema video apparatus, a digital cinema video apparatus, a surveillance camera, a video conversation apparatus, a real-time communication apparatus such as video communication, a mobile streaming apparatus, a storage medium, a camcorder, a video on demand (VoD) service providing apparatus, an over-the-top (OTT) video apparatus, an internet streaming service providing apparatus, a three-dimensional (3D) video apparatus, an image telephone video apparatus, a medical video apparatus, etc., and may be used to process a video signal or a data signal. For example, an over-the-top (OTT) video apparatus may include a game console, a Blu-ray player, an internet-connected TV, a home theater system, a smartphone, a tablet PC, a digital video recorder (DVR), etc.
[0230] FIG. 15 is a diagram illustrating a content streaming system to which the embodiments of the present disclosure can be applied.
[0231] As shown in FIG. 15, a content streaming system to which an embodiment of the present disclosure is applied can broadly include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.
[0232] The encoding server compresses content input from a multimedia input device such as a smartphone, camera, or camcorder into digital data to generate a bitstream and transmits the bitstream to the streaming server. As another example, if a multimedia input device such as a smartphone, camera, or video camera directly generates a bitstream, the encoding server can be omitted.
[0233] The bitstream can be generated by an image encoding method and / or image encoding device to which an embodiment of the present disclosure is applied, and the streaming server can temporarily store the bitstream during the process of transmitting or receiving the bitstream.
[0234] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server serves as an intermediary for informing the user of available services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, which then transmits the multimedia data to the user. In this case, the content streaming system may include a separate control server, which may control commands and responses between devices in the content streaming system.
[0235] The streaming server may receive content from a media storage and / or an encoding server. For example, when receiving content from the encoding server, the content may be received in real time. In this case, the streaming server may store the bitstream for a certain period of time to provide a smooth streaming service.
[0236] Examples of the user device include a mobile phone, a smartphone, a laptop computer, a digital broadcasting terminal, a personal digital assistant (PDA), a portable multimedia player (PMP), a navigation system, a slate PC, a tablet PC, an ultrabook, a wearable device such as a smartwatch, smart glass, a head mounted display (HMD), a digital TV, a desktop computer, and digital signage.
[0237] Each server in the content streaming system can be operated as a distributed server, in which case data received from each server can be processed in a distributed manner.
[0238] The scope of the present disclosure includes software or machine-executable commands (e.g., operating systems, applications, firmware, programs, etc.) that cause operations according to the methods of various embodiments to be performed on a device or computer, and non-transitory computer-readable medium on which such software or commands can be stored and executed on a device or computer. [Industrial Applicability]
[0239] The embodiments of the present disclosure can be used to encode / decode images.
Claims
1. An image decoding device, Memory and at least one processor coupled to the memory; The at least one processor Obtain first information indicating a number of one or more decoded picture buffer (DPB) parameter syntax structures in a video parameter set (VPS); obtaining the one or more DPB parameter syntax structures from the VPS based on the first information; obtaining second information regarding a mapping between one or more multi-layer output layer sets (OLSs) and the one or more DPB parameter syntax structures from the VPS based on the first information; Selecting a DPB parameter syntax structure to be applied to the current OLS based on the second information; processing the current OLS based on the selected DPB parameter syntax structure; obtaining third information indicating the number of one or more profile tier level (PTL) syntax structures in the VPS; obtaining the one or more PTL syntax structures from the VPS based on the third information; obtaining fourth information regarding a mapping between one or more OLSs and the one or more PTL syntax structures from the VPS based on the third information; and selecting a PTL syntax structure to be applied to the current OLS based on the fourth information; each of the one or more PTL syntax structures in the VPS is mapped to at least one OLS of the one or more OLSs; An image decoding device, wherein the second information is acquired further based on the number of OLSs.
2. The image decoding apparatus according to claim 1 , wherein the number of the one or more DPB parameter syntax structures in the VPS is not greater than the number of the one or more multi-layer OLSs.
3. The image decoding apparatus according to claim 1 , wherein each of the one or more DPB parameter syntax structures in the VPS is mapped to at least one multi-layer OLS of the one or more multi-layer OLSs.
4. The image decoding apparatus according to claim 1 , wherein the second information is obtained from the VPS based on the number of the one or more DPB parameter syntax structures in the VPS being greater than one.
5. 2. The image decoding device of claim 1, wherein the second information is not obtained from the VPS and the second information is inferred to be a value of 0 based on the number of the one or more DPB parameter syntax structures in the VPS being not greater than 1.
6. The image decoding apparatus according to claim 1 , wherein the DPB parameter syntax structure applied to the current OLS is obtained from a Sequence Parameter Set (SPS) based on the fact that the current OLS includes a single layer.
7. The image decoding apparatus according to claim 1 , wherein the number of the one or more PTL syntax structures in the VPS is not greater than a total number of the one or more OLSs.
8. An image encoding device, Memory and at least one processor coupled to the memory; The at least one processor encoding one or more decoded picture buffer (DPB) parameter syntax structures within a Video Parameter Set (VPS); encoding first information for indicating the number of one or more DPB parameter syntax structures in the VPS; determining a DPB parameter syntax structure to be applied to a current output layer set (OLS), the DPB parameter syntax structure to be applied to the current OLS being included in the one or more DPB parameter syntax structures; encoding second information regarding a mapping between one or more multi-layer OLSs and the one or more DPB parameter syntax structures in the VPS based on the determined DPB parameter syntax structure; encoding one or more profile tier level (PTL) syntax structures within the VPS; encoding third information for indicating the number of one or more PTL syntax structures in the VPS; determining a PTL syntax structure to be applied to the current OLS, the PTL syntax structure to be applied to the current OLS being included in the one or more PTL syntax structures; based on the determined PTL syntax structure, encoding fourth information regarding a mapping between one or more OLSs and the one or more PTL syntax structures in the VPS; each of the one or more PTL syntax structures in the VPS is mapped to at least one OLS of the one or more OLSs; An image encoding device, wherein the second information is encoded further based on the number of OLSs.
9. The image encoding device of claim 8 , wherein the number of the one or more DPB parameter syntax structures in the VPS is not greater than the number of the one or more multi-layer OLSs.
10. The image encoding device of claim 8 , wherein each of the one or more DPB parameter syntax structures in the VPS is mapped to at least one multi-layer OLS of the one or more multi-layer OLSs.
11. 1. An apparatus for transmitting a bitstream relating to an image, said apparatus comprising: at least one processor configured to generate the bitstream; a transmission unit configured to transmit the bitstream; The bitstream comprises: encoding one or more decoded picture buffer (DPB) parameter syntax structures in a Video Parameter Set (VPS); encoding first information to indicate a number of one or more DPB parameter syntax structures in the VPS; determining a DPB parameter syntax structure to be applied to a current output layer set (OLS), wherein the DPB parameter syntax structure to be applied to the current OLS is included in the one or more DPB parameter syntax structures; encoding second information regarding a mapping between one or more multi-layer OLSs and the one or more DPB parameter syntax structures in the VPS based on the determined DPB parameter syntax structure; encoding one or more profile tier level (PTL) syntax structures within the VPS; encoding third information to indicate the number of one or more profile tier level (PTL) syntax structures in the VPS; determining a PTL syntax structure to be applied to the current OLS, wherein the PTL syntax structure to be applied to the current OLS is included in the one or more PTL syntax structures; and encoding fourth information regarding a mapping between one or more OLSs and the one or more PTL syntax structures in the VPS based on the determined PTL syntax structure; each of the one or more PTL syntax structures in the VPS is mapped to at least one OLS of the one or more OLSs; The apparatus, wherein the second information is encoded further based on the number of OLSs.
Citation Information
Patent Citations
JPP7488360B
Coding device, decoding device, coding method, and decoding method
WO2021162016A1