Image encoding / decoding method and device based on SEI message containing layer identifier information, and method for transmitting bitstream
By introducing SEI messages into the image bitstream and including layer identifier information, the problem of low encoding and decoding efficiency of high-resolution image is solved, and more efficient image transmission and storage is achieved.
Patent Information
- Application Number
- JP2023565261
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-04-23
- Filing Date
- 2022-04-22
- Publication Date
- 2025-05-08
- Estimated Expiration
- 2042-04-22
AI Technical Summary
The prior art is difficult to effectively solve the problems of efficient encoding and decoding of high-resolution and high-quality images, resulting in increased transmission and storage costs.
By introducing supplementary enhancement information (SEI) messages into the bitstream, including layer identifier information, to support multi-layer image encoding and decoding, improving encoding and decoding efficiency.
It realizes the improvement of image encoding and decoding efficiency, reduces transmission and storage costs, and supports image recovery based on layer identifier information.
Smart Images

Figure 0007673242000008 
Figure 0007673242000009 
Figure 0007673242000010
Abstract
Description
[Technical field]
[0001] The present disclosure relates to an image encoding / decoding method and apparatus, and more particularly to an image encoding / decoding method and apparatus based on a supplemental enhancement information (SEI) message that includes layer identifier information for one or more layers included in a bitstream, and a method for transmitting a bitstream generated by the image encoding method / apparatus of the present disclosure. [Background technology]
[0002] Recently, the demand for high-resolution, high-quality images, for example, HD (High Definition) images and UHD (Ultra High Definition) images, is increasing in various fields. As the resolution and quality of image data increases, the amount of information or bits transmitted increases relatively compared to conventional image data. The increase in the amount of information or bits transmitted leads to an increase in transmission costs and storage costs.
[0003] This requires a highly efficient image compression technique for effectively transmitting, storing, and reproducing high-resolution, high-quality image information. Summary of the Invention [Problem to be solved by the invention]
[0004] An object of the present disclosure is to provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.
[0005] Another object of the present disclosure is to provide an image encoding / decoding method and device that performs image restoration based on layer identifier information in an SEI message.
[0006] Another object of the present disclosure is to provide a method and apparatus for obtaining information on one or more layers included in a bitstream based on an SEI message and encoding / decoding an image.
[0007] Another object of the present disclosure is to provide a method for transmitting a bitstream generated by the image encoding method or apparatus according to the present disclosure.
[0008] Another object of the present disclosure is to provide a recording medium storing a bitstream generated by the image encoding method or apparatus according to the present disclosure.
[0009] Another object of the present disclosure is to provide a recording medium storing a bitstream that is received by the image decoding device according to the present disclosure, decoded, and used to restore an image.
[0010] The technical problems to be solved by the present disclosure are not limited to the above-mentioned technical problems, and other technical problems not described above will be clearly understood by a person having ordinary skill in the art to which the present disclosure pertains from the following description. [Means for solving the problem]
[0011] According to one embodiment of the present disclosure, an image decoding method performed by an image decoding device includes the steps of receiving an SEI (supplemental enhancement information) message as a bitstream, the SEI message including information on one or more layers included in the bitstream, acquiring information on the one or more layers based on the received SEI message, and restoring an image in the bitstream based on the acquired information on the one or more layers, wherein the received SEI message may include layer identifier information for the one or more layers included in the bitstream.
[0012] According to one embodiment of the present disclosure, the layer identifier information can be obtained for the maximum number of layers in the bitstream.
[0013] According to one embodiment of the present disclosure, information regarding the maximum number of layers in the bitstream can be obtained from the SEI message.
[0014] According to an embodiment of the present disclosure, the SEI message may include information on layers other than one or more layers included in the bitstream.
[0015] According to one embodiment of the present disclosure, the maximum number of layers in the bitstream may not be limited to the number of one or more layers in the bitstream.
[0016] According to one embodiment of the present disclosure, the maximum number of layers in the bitstream may be limited to have a value not smaller than the number of one or more layers in the bitstream.
[0017] According to one embodiment of the present disclosure, the layer identifier information may be signaled for each of view id information or auxiliary id information.
[0018] According to one embodiment of the present disclosure, the layer identifier information may be included in the SEI message in ascending order of layer identifier values such that the layer identifier of the i-th layer has a greater value than the layer identifier of the i-1-th layer.
[0019] According to one embodiment of the present disclosure, the layer identifiers may be included in the SEI message in descending order of layer identifier values such that the layer identifier of the i-th layer has a smaller value than the layer identifier of the i-1-th layer.
[0020] According to one embodiment of the present disclosure, an image encoding method performed by an image encoding device includes a step of encoding an image in a bitstream based on information for one or more layers in the bitstream, and a step of encoding an SEI (supplemental enhancement information) message into the bitstream, the SEI message including information for one or more layers included in the bitstream, wherein the SEI message may include layer identifier information for the one or more layers included in the bitstream.
[0021] According to an embodiment of the present disclosure, the layer identifier information can be included in the SEI message for the maximum number of layers in the bitstream.
[0022] According to one embodiment of the present disclosure, information regarding the maximum number of layers in the bitstream may be included in the SEI message.
[0023] According to an embodiment of the present disclosure, a bitstream generated by an image encoding device or an image encoding method can be transmitted.
[0024] According to one embodiment of the present disclosure, the bitstream generated by the image coding method can be stored or recorded on a computer-readable medium.
[0025] The features described above in the brief summary of the present disclosure are merely exemplary embodiments of the detailed description of the present disclosure that follows and are not intended to limit the scope of the present disclosure. Effect of the Invention
[0026] According to the present disclosure, an image encoding / decoding method and apparatus with improved encoding / decoding efficiency can be provided.
[0027] According to the present disclosure, a method and apparatus for image encoding / decoding based on an SEI message containing information for one or more layers can be provided.
[0028] According to the present disclosure, a method and apparatus may be provided for encoding / decoding an image that includes layer identifier information for one or more layers.
[0029] According to the present disclosure, a method for transmitting a bitstream generated by an image encoding method or apparatus according to the present disclosure can be provided.
[0030] According to the present disclosure, a recording medium storing a bitstream generated by the image encoding method or apparatus according to the present disclosure can be provided.
[0031] Furthermore, according to the present disclosure, it is possible to provide a recording medium storing a bitstream that is received by the image decoding device according to the present disclosure, decoded, and used to restore an image.
[0032] The effects obtained by the present disclosure are not limited to the effects described above, and other effects not described above will be clearly understood by those having ordinary skill in the art to which the present disclosure pertains from the following description. [Brief description of the drawings]
[0033] [Figure 1] FIG. 1 is a schematic diagram illustrating a video coding system to which an embodiment of the present disclosure can be applied.
[0034] [Diagram 2] 1 is a diagram illustrating an image encoding device to which an embodiment of the present disclosure can be applied;
[0035] [Diagram 3] 1 is a diagram illustrating an image decoding device to which an embodiment of the present disclosure can be applied;
[0036] [Figure 4] FIG. 2 is a diagram illustrating an image decoding procedure.
[0037] [Diagram 5] FIG. 1 illustrates a schematic diagram of an image encoding procedure.
[0038] [Figure 6] FIG. 2 is a schematic diagram illustrating a coding hierarchy and structure according to the present disclosure.
[0039] [Figure 7-14] A diagram to explain syntax related to multiple layer information in a VPS according to the present disclosure.
[0040] [Figure 15] FIG. 2 illustrates an SDI SEI message syntax according to one embodiment of the present disclosure.
[0041] [Figure 16] FIG. 2 is a diagram illustrating an image decoding method according to an embodiment of the present disclosure.
[0042] [Figure 17] FIG. 1 illustrates an image encoding method according to an embodiment of the present disclosure.
[0043] [Figure 18] FIG. 1 is a diagram illustrating an image encoding / decoding device according to an embodiment of the present disclosure.
[0044] [Figure 19] FIG. 1 illustrates a content streaming system to which an embodiment of the present disclosure can be applied. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0045] According to one embodiment of the present disclosure, an image decoding method performed by an image decoding device includes the steps of receiving an SEI (supplemental enhancement information) message as a bitstream, the SEI message including information on one or more layers included in the bitstream, acquiring information on the one or more layers based on the received SEI message, and restoring an image in the bitstream based on the acquired information on the one or more layers, wherein the received SEI message may include layer identifier information for the one or more layers included in the bitstream. EXAMPLES
[0046] The present disclosure will be described in detail below with reference to the accompanying drawings so that those skilled in the art can easily implement the present disclosure. However, the present disclosure may be embodied in various different forms and is not limited to the embodiments described herein.
[0047] In describing the embodiments of the present disclosure, if it is determined that a specific description of a known configuration or function may make the gist of the present disclosure unclear, the detailed description thereof will be omitted. In addition, in the drawings, parts that are not related to the description of the present disclosure are omitted, and similar parts are denoted by similar reference numerals.
[0048] In the present disclosure, when a certain component is "connected," "coupled," or "connected" to another component, this includes not only a direct connection relationship, but also an indirect connection relationship in which another component exists between them. Furthermore, when a certain component is described as "including" or "having" another component, this does not mean that the other component is excluded, but that the other component can be further included, unless otherwise specified.
[0049] In this disclosure, terms such as "first" and "second" are used only for the purpose of distinguishing one component from another component, and do not limit the order or importance of the components unless otherwise specified. Therefore, within the scope of this disclosure, a first component in one embodiment may be called a second component in another embodiment, and similarly, a second component in one embodiment may be called a first component in another embodiment.
[0050] In this disclosure, components that are distinguished from one another are used to clearly describe the characteristics of each component, and do not necessarily mean that the components are separate. In other words, multiple components may be integrated and configured as a single hardware or software unit, or one component may be distributed and configured as multiple hardware or software units. Thus, even if not otherwise stated, such integrated or distributed embodiments are also included in the scope of the present disclosure.
[0051] In the present disclosure, the components described in the various embodiments are not necessarily essential components, and some may be optional components. Therefore, an embodiment consisting of a subset of the components described in one embodiment is also included in the scope of the present disclosure. In addition, an embodiment including other components in addition to the components described in the various embodiments is also included in the scope of the present disclosure.
[0052] The present disclosure relates to image encoding and decoding, and terms used in this disclosure may have ordinary meanings in the technical field to which the present disclosure belongs, unless they are newly defined in this disclosure.
[0053] In this disclosure, a "picture" generally means a unit indicating any one image in a particular time period, a slice / tile is a coding unit constituting a part of a picture, and one picture may be composed of one or more slices / tiles. Also, a slice / tile may include one or more coding tree units (CTUs).
[0054] In this disclosure, a "pixel" or a "pel" may refer to the smallest unit constituting one picture (or image). A term corresponding to a pixel may be a "sample." A sample may generally indicate a pixel or a pixel value, may indicate only a pixel / pixel value of a luma component, or may indicate only a pixel / pixel value of a chroma component.
[0055] In this disclosure, a "unit" may refer to a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to the region. A unit may be used interchangeably with terms such as "sample array," "block," or "area," depending on the case. In a general case, an M×N block may include a set (or array) of samples or transform coefficients consisting of M columns and N rows.
[0056] In the present disclosure, a "current block" may refer to any one of a "current coding block," a "current coding unit," a "block to be coded," a "block to be decoded," or a "block to be processed." If prediction is performed, a "current block" may refer to a "current predicted block" or a "block to be predicted." If transformation (inverse transformation) / quantization (inverse quantization) is performed, a "current block" may refer to a "current transformed block" or a "block to be transformed." If filtering is performed, a "current block" may refer to a "block to be filtered."
[0057] Furthermore, in this disclosure, unless explicitly stated as a chroma block, a "current block" may refer to a block including both a luma component block and a chroma component block, or a "luma block of the current block." A chroma block of the current block may be explicitly expressed by including an explicit description of a chroma block, such as a "chroma block" or a "current chroma block."
[0058] In the present disclosure, " / " and "," can be interpreted as "and / or." For example, "A / B" and "A, B" can be interpreted as "A and / or B." Also, "A / B / C" and "A, B, C" can mean "at least one of A, B, and / or C."
[0059] In this disclosure, "or" can be interpreted as "and / or." For example, "A or B" can mean 1) only "A," 2) only "B," or 3) "A and B." Alternatively, in this disclosure, "or" can mean "additionally or alternatively."
[0060] Video Coding System Overview
[0061] FIG. 1 is a diagram illustrating a video coding system in accordance with this disclosure.
[0062] A video coding system according to an embodiment may include an encoding device 10 and a decoding device 20. The encoding device 10 may transmit encoded video and / or image information or data to the decoding device 20 in a file or streaming format via a digital storage medium or a network.
[0063] The encoding device 10 according to an embodiment may include a video source generating unit 11, an encoding unit 12, and a transmitting unit 13. The decoding device 20 according to an embodiment may include a receiving unit 21, a decoding unit 22, and a rendering unit 23. The encoding unit 12 may be referred to as a video / image encoding unit, and the decoding unit 22 may be referred to as a video / image decoding unit. The transmitting unit 13 may be included in the encoding unit 12. The receiving unit 21 may be included in the decoding unit 22. The rendering unit 23 may include a display unit, which may be configured as a separate device or an external component.
[0064] The video source generating unit 11 can obtain videos / images through a process of capturing, synthesizing, or generating videos / images. The video source generating unit 11 can include a video / image capture device and / or a video / image generation device. The video / image capture device can include, for example, one or more cameras, a video / image archive containing previously captured videos / images, etc. The video / image generation device can include, for example, a computer, a tablet, a smartphone, etc., and can (electronically) generate videos / images. For example, a virtual video / image can be generated through a computer, etc., in which case the video / image capture process can be replaced by a process in which related data is generated.
[0065] The encoder 12 may encode the input video / image. The encoder 12 may perform a series of steps such as prediction, transformation, quantization, etc. for compression and encoding efficiency. The encoder 12 may output the encoded data (encoded video / image information) in a bitstream format.
[0066] The transmitting unit 13 may transmit the encoded video / image information or data output in a bitstream format to the receiving unit 21 of the decoding device 20 via a digital storage medium or a network in a file or streaming format. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit 13 may include elements for generating a media file via a predetermined file format and may include elements for transmitting via a broadcasting / communication network. The receiving unit 21 may extract / receive the bitstream from the storage medium or network and transmit it to the decoding unit 22.
[0067] The decoding unit 22 can decode the video / image by performing a series of steps such as inverse quantization, inverse transformation, and prediction corresponding to the operations of the encoding unit 12.
[0068] The rendering unit 23 can render the decoded video / images. The rendered video / images can be displayed via a display unit.
[0069] Overview of the image encoding device
[0070] FIG. 2 is a diagram illustrating an image encoding device to which an embodiment of the present disclosure can be applied.
[0071] As shown in Fig. 2, the image coding device 100 may include an image division unit 110, a subtraction unit 115, a transformation unit 120, a quantization unit 130, an inverse quantization unit 140, an inverse transformation unit 150, an addition unit 155, a filtering unit 160, a memory 170, an inter prediction unit 180, an intra prediction unit 185, and an entropy coding unit 190. The inter prediction unit 180 and the intra prediction unit 185 may be collectively referred to as a "prediction unit." The transformation unit 120, the quantization unit 130, the inverse quantization unit 140, and the inverse transformation unit 150 may be included in a residual processing unit. The residual processing unit may further include a subtraction unit 115.
[0072] All or at least some of the components constituting the image encoding device 100 may be realized by a single hardware component (e.g., an encoder or a processor) depending on the embodiment. Also, the memory 170 may include a decoded picture buffer (DPB) and may be realized by a digital storage medium.
[0073] The image division unit 110 may divide an input image (or picture, frame) input to the image encoding device 100 into one or more processing units. As an example, the processing units may be called coding units (CUs). The coding units may be obtained by recursively dividing a coding tree unit (CTU) or a largest coding unit (LCU) according to a QT / BT / TT (Quad-tree / Binary-tree / Ternal-tree) structure. For example, one coding unit may be divided into a plurality of coding units of a deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary-tree structure. For dividing the coding units, a quad-tree structure may be applied first, and a binary-tree structure and / or a ternary-tree structure may be applied later. A coding procedure according to the present disclosure may be performed based on a final coding unit that is not further divided. The maximum coding unit may be used as the final coding unit, and a lower depth coding unit obtained by dividing the maximum coding unit may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and / or restoration, which will be described later. As another example, a processing unit of the coding procedure may be a prediction unit (PU) or a transform unit (TU). The prediction unit and the transform unit may be divided or partitioned from the final coding unit, respectively. The prediction unit may be a unit of sample prediction, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.
[0074] The prediction unit (inter prediction unit 180 or intra prediction unit 185) may perform prediction on a block to be processed (current block) and generate a predicted block including prediction samples for the current block. The prediction unit may determine whether intra prediction or inter prediction is applied in units of a current block or a CU. The prediction unit may generate various information related to prediction of the current block and transmit it to the entropy encoding unit 190. The information related to prediction may be encoded by the entropy encoding unit 190 and output in a bitstream format.
[0075] The intra prediction unit 185 may predict the current block by referring to samples in the current picture. The referenced samples may be located in the neighborhood of the current block or may be located away from the current block according to an intra prediction mode and / or an intra prediction technique. The intra prediction mode may include a plurality of non-directional modes and a plurality of directional modes. The non-directional mode may include, for example, a DC mode and a planar mode. The directional mode may include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the fineness of the prediction direction. However, this is merely an example, and more or less directional prediction modes may be used depending on the setting. The intra prediction unit 185 may also determine a prediction mode to be applied to the current block using prediction modes applied to neighboring blocks.
[0076] The inter prediction unit 180 may derive a predicted block for a current block based on a reference block (reference sample array) identified by a motion vector on a reference picture. In this case, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information may be predicted in units of a block, a sub-block, or a sample based on the correlation of motion information between a neighboring block and a current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the neighboring block may include a spatial neighboring block present in the current picture and a temporal neighboring block present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different from each other. The temporal neighboring block may be called a collocated reference block, a collocated CU (colCU), etc. The reference picture including the temporal neighboring block may be called a collocated picture (colPic). For example, the inter prediction unit 180 may generate information indicating which candidate is used to derive a motion vector and / or a reference picture index of the current block by forming a motion information candidate list based on neighboring blocks. Inter prediction may be performed based on various prediction modes, and for example, in the case of a skip mode and a merge mode, the inter prediction unit 180 may use motion information of neighboring blocks as motion information of the current block. In the case of the skip mode, unlike the merge mode, a residual signal may not be transmitted.In the case of a motion vector prediction (MVP) mode, the motion vector of the current block can be signaled by using the motion vector of a neighboring block as a motion vector predictor and encoding a motion vector difference and an indicator for the motion vector predictor. The motion vector difference can mean the difference between the motion vector of the current block and the motion vector predictor.
[0077] The prediction unit may generate a prediction signal based on various prediction methods and / or prediction techniques, which will be described later. For example, the prediction unit may apply intra prediction or inter prediction for prediction of the current block, and may simultaneously apply intra prediction and inter prediction. A prediction method that simultaneously applies intra prediction and inter prediction for prediction of the current block may be called combined inter and intra prediction (CIIP). The prediction unit may also perform intra block copy (IBC) for prediction of the current block. Intra block copy can be used for content image / video coding such as games, for example, as in screen content coding (SCC). IBC is a method of predicting a current block using an already restored reference block in a current picture that is located a predetermined distance away from the current block. When IBC is applied, the position of the reference block in the current picture may be coded as a vector (block vector) corresponding to the predetermined distance. IBC is basically performed in the current picture, but may be performed similarly to inter prediction in that a reference block is derived in the current picture. That is, the IBC may use at least one of the inter prediction techniques described in this disclosure.
[0078] The prediction signal generated by the prediction unit may be used to generate a restored signal or a residual signal. The subtraction unit 115 may subtract the prediction signal (predicted block, prediction sample array) output from the prediction unit from the input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array). The generated residual signal may be transmitted to the conversion unit 120.
[0079] The transform unit 120 may generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loeve transform (KLT), a graph-based transform (GBT), or a conditionally non-linear transform (CNT). Here, the GBT refers to a transform obtained from a graph when the relationship information between pixels is expressed as a graph. The CNT refers to a transform obtained based on a predicted signal generated using all previously reconstructed pixels. The transform process may be applied to pixel blocks having the same square size, or may be applied to non-square, variable-sized blocks.
[0080] The quantization unit 130 may quantize the transform coefficients and transmit the quantized transform coefficients to the entropy coding unit 190. The entropy coding unit 190 may code the quantized signal (information on the quantized transform coefficients) and output the coded signal in a bitstream format. The information on the quantized transform coefficients may be referred to as residual information. The quantization unit 130 may rearrange the quantized transform coefficients in a block format into a one-dimensional vector format based on a coefficient scan order, and may generate information on the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector format.
[0081] The entropy coding unit 190 may perform various coding methods, such as exponential Golomb, context-adaptive variable length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy coding unit 190 may also code information required for video / image restoration (e.g., values of syntax elements, etc.) together or separately in addition to the quantized transform coefficients. The coded information (e.g., coded video / image information) may be transmitted or stored in a network abstraction layer (NAL) unit unit in a bitstream format. The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / image information may further include general constraint information. The signaling information, transmitted information and / or syntax elements referred to in this disclosure may be encoded through the above-mentioned encoding procedures and included in the bitstream.
[0082] The bitstream may be transmitted via a network or may be stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as a USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitting unit (not shown) for transmitting the signal output from the entropy encoding unit 190 and / or a storing unit (not shown) for storing the signal may be provided as an internal / external element of the image encoding device 100, or the transmitting unit may be provided as a component of the entropy encoding unit 190.
[0083] The quantized transform coefficients output from the quantization unit 130 can be used to generate a residual signal. For example, the quantized transform coefficients are subjected to inverse quantization and inverse transformation via the inverse quantization unit 140 and the inverse transformation unit 150, so that a residual signal (residual block or residual sample) can be restored.
[0084] The adder 155 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to a prediction signal output from the inter prediction unit 180 or the intra prediction unit 185. When there is no residual for the current block to be processed, such as when a skip mode is applied, a predicted block may be used as a reconstructed block. The adder 155 may be referred to as a reconstruction unit or a reconstructed block generation unit. The generated reconstructed signal may be used for intra prediction of the next current block to be processed in the current picture, and may also be used for inter prediction of the next picture after filtering as described below.
[0085] The filtering unit 160 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 160 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture, and may store the modified reconstructed picture in the memory 170, specifically, in the DPB of the memory 170. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, and the like. The filtering unit 160 may generate various information related to filtering, as will be described later in the description of each filtering method, and transmit the information to the entropy coding unit 190. The information related to filtering may be coded by the entropy coding unit 190 and output in a bitstream format.
[0086] The modified reconstructed picture transmitted to the memory 170 may be used as a reference picture in the inter prediction unit 180. When inter prediction is applied through this, the image encoding device 100 may avoid a prediction mismatch between the image encoding device 100 and the image decoding device, and may also improve encoding efficiency.
[0087] The DPB in the memory 170 may store modified reconstructed pictures to be used as reference pictures in the inter prediction unit 180. The memory 170 may store motion information of blocks from which motion information in the current picture is derived (or coded) and / or motion information of already reconstructed intra-picture blocks. The stored motion information may be transmitted to the inter prediction unit 180 to be used as motion information of spatial surrounding blocks or motion information of temporal surrounding blocks. The memory 170 may store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra prediction unit 185.
[0088] Overview of the image decoding device
[0089] FIG. 3 is a diagram illustrating an image decoding device to which an embodiment of the present disclosure can be applied.
[0090] 3, the image decoding device 200 may include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an adder 235, a filtering unit 240, a memory 250, an inter prediction unit 260, and an intra prediction unit 265. The inter prediction unit 260 and the intra prediction unit 265 may be collectively referred to as a "prediction unit." The inverse quantization unit 220 and the inverse transform unit 230 may be included in a residual processing unit.
[0091] All or at least some of the components constituting the image decoding device 200 may be realized by one hardware component (e.g., a decoder or a processor) depending on the embodiment. Also, the memory 170 may include a DPB and may be realized by a digital storage medium.
[0092] The image decoding device 200, which receives a bitstream including video / image information, can reconstruct an image by executing a process corresponding to the process performed by the image encoding device 100 of Fig. 2. For example, the image decoding device 200 can perform decoding using a processing unit applied in the image encoding device. Thus, the processing unit for decoding can be, for example, a coding unit. The coding unit can be obtained by dividing a coding tree unit or a maximum coding unit. Then, the reconstructed image signal decoded and output by the image decoding device 200 can be reproduced by a reproduction device (not shown).
[0093] The image decoding apparatus 200 may receive a signal output from the image encoding apparatus of FIG. 2 in the form of a bitstream. The received signal may be decoded via the entropy decoding unit 210. For example, the entropy decoding unit 210 may derive information (e.g., video / image information) required for image restoration (or picture restoration) by parsing the bitstream. The video / image information may further include information on various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. The image decoding apparatus may further use information on the parameter set and / or the general constraint information to decode an image. The signaling information, received information, and / or syntax elements referred to in the present disclosure may be obtained from the bitstream by being decoded via the decoding procedure. For example, the entropy decoding unit 210 may decode information in a bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output values of syntax elements required for image restoration and quantized values of transform coefficients related to the residual. More specifically, the CABAC entropy decoding method may receive bins corresponding to each syntax element from the bitstream, determine a context model using syntax element information to be decoded and decoded information of neighboring blocks and a block to be decoded, or information of a symbol / bin decoded in a previous step, predict the occurrence probability of the bin based on the determined context model, and perform arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element. In this case, the CABAC entropy decoding method may update the context model using information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model.Among the information decoded by the entropy decoding unit 210, information related to prediction is provided to a prediction unit (inter prediction unit 260 and intra prediction unit 265), and residual values entropy-decoded by the entropy decoding unit 210, i.e., quantized transform coefficients and related parameter information, may be input to the inverse quantization unit 220. Also, among the information decoded by the entropy decoding unit 210, information related to filtering may be provided to the filtering unit 240. Meanwhile, a receiving unit (not shown) for receiving a signal output from the image encoding device may be further provided as an internal / external element of the image decoding device 200, or the receiving unit may be provided as a component of the entropy decoding unit 210.
[0094] Meanwhile, the image decoding device according to the present disclosure may be called a video / image / picture decoding device. The image decoding device may include an information decoder (video / image / picture information decoder) and / or a sample decoder (video / image / picture sample decoder). The information decoder may include an entropy decoding unit 210, and the sample decoder may include at least one of an inverse quantization unit 220, an inverse transform unit 230, an adder 235, a filtering unit 240, a memory 250, an inter prediction unit 260, and an intra prediction unit 265.
[0095] The inverse quantization unit 220 may inverse quantize the quantized transform coefficients to output transform coefficients. The inverse quantization unit 220 may rearrange the quantized transform coefficients in a two-dimensional block format. In this case, the rearrangement may be performed based on a coefficient scan order performed in the image encoding device. The inverse quantization unit 220 may perform inverse quantization on the quantized transform coefficients using a quantization parameter (e.g., quantization step size information) to obtain transform coefficients.
[0096] The inverse transform unit 230 can inversely transform the transform coefficients to obtain a residual signal (residual block, residual sample array).
[0097] The prediction unit may perform prediction on a current block and generate a predicted block including a prediction sample for the current block. The prediction unit may determine whether intra prediction or inter prediction is applied to the current block based on the prediction information output from the entropy decoding unit 210, and may determine a specific intra / inter prediction mode (prediction technique).
[0098] The prediction unit can generate a prediction signal based on various prediction methods (techniques) described below, as has been described in the explanation of the prediction unit of the image encoding device 100.
[0099] The intra prediction unit 265 may predict the current block by referring to samples in the current picture. The description of the intra prediction unit 185 may be similarly applied to the intra prediction unit 265.
[0100] The inter prediction unit 260 may derive a predicted block for the current block based on a reference block (reference sample array) identified by a motion vector on a reference picture. In this case, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. For example, the inter prediction unit 260 may configure a motion information candidate list based on the neighboring blocks, and derive a motion vector and / or a reference picture index for the current block based on the received candidate selection information. Inter prediction may be performed based on various prediction modes (techniques), and the prediction information may include information indicating a mode (technique) of inter prediction for the current block.
[0101] The adder 235 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the acquired residual signal to a prediction signal (predicted block, predicted sample array) output from a prediction unit (including the inter prediction unit 260 and / or the intra prediction unit 265). When there is no residual for the current block, such as when a skip mode is applied, the predicted block can be used as the reconstructed block. The description of the adder 155 may also be applied to the adder 235. The adder 235 may also be referred to as a reconstruction unit or a reconstructed block generator. The generated reconstructed signal may be used for intra prediction of the next current block in the current picture, and may also be used for inter prediction of the next picture via filtering, as described below.
[0102] The filtering unit 240 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 240 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture, and may store the modified reconstructed picture in the memory 250, specifically, in the DPB of the memory 250. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, etc.
[0103] The (modified) reconstructed picture stored in the DPB of the memory 250 may be used as a reference picture in the inter prediction unit 260. The memory 250 may store motion information of a block from which motion information in the current picture is derived (or decoded) and / or motion information of a block in an already reconstructed picture. The stored motion information may be transmitted to the inter prediction unit 260 to be used as motion information of a spatial surrounding block or motion information of a temporal surrounding block. The memory 250 may store a reconstructed sample of a reconstructed block in the current picture and transmit it to the intra prediction unit 265.
[0104] In this specification, the embodiments described for the filtering unit 160, inter prediction unit 180 and intra prediction unit 185 of the image encoding device 100 can also be applied in a similar or corresponding manner to the filtering unit 240, inter prediction unit 260 and intra prediction unit 265 of the image decoding device 200, respectively.
[0105] Overview of image encoding / decoding procedures
[0106] FIG. 4 is a diagram illustrating an outline of an image decoding procedure.
[0107] In image / video coding, pictures constituting an image / video may be coded / decoded according to a series of decoding orders. A picture order corresponding to an output order of decoded pictures may be set to be different from the decoding order. Based on this, not only forward prediction but also backward prediction may be performed during inter prediction.
[0108] 4, S401 may be performed in the entropy decoding unit of the above-mentioned decoding device, S402 may be performed in the prediction unit, S403 may be performed in the residual processing unit, S404 may be performed in the addition unit, and S405 may be performed in the filtering unit. S401 may include the information decoding procedure described herein, S402 may include the inter / intra prediction procedure described herein, S403 may include the residual processing procedure described herein, S404 may include the block / picture reconstruction procedure described herein, and S405 may include the in-loop filtering procedure described herein.
[0109] As described above, the decoding procedure of FIG. 4 may include, in outline, an image / video information acquisition procedure (S401) from a bitstream (by decoding), a picture reconstruction procedure (S402 to S404), and an in-loop filtering procedure (S405) for the reconstructed picture. The picture reconstruction procedure may be performed based on a prediction sample and a residual sample obtained through the inter / intra prediction (S402) and residual processing (S403, inverse quantization and inverse transform for quantized transform coefficients) processes described herein. A modified reconstructed picture may be generated through an in-loop filtering procedure for the reconstructed picture generated by the picture reconstruction procedure, and the modified reconstructed picture may be output as a decoded picture or may be stored in a decoded picture buffer or memory of a decoding device and used as a reference picture in an inter prediction procedure when decoding a picture thereafter. In some cases, the in-loop filtering procedure may be omitted, in which case the reconstructed picture may be output as a decoded picture, or may be stored in a decoded picture buffer or memory of the decoding device and used as a reference picture in an inter-prediction procedure when decoding a picture thereafter. The in-loop filtering procedure (S405) may include a deblocking filtering procedure, a sample adaptive offset (SAO) procedure, an adaptive loop filter (ALF) procedure, and / or a bi-lateral filter procedure, as described above, and some or all of them may be omitted. In addition, one or some of the deblocking filtering procedure, the sample adaptive offset (SAO) procedure, the adaptive loop filter (ALF) procedure, and the bi-lateral filter procedure may be applied sequentially, or all of them may be applied sequentially. For example, the SAO procedure may be performed after the deblocking filtering procedure is applied to the reconstructed picture.Or, for example, the ALF procedure can be performed after a deblocking filtering procedure is applied to the reconstructed picture, which can be done in the encoding device as well.
[0110] FIG. 5 is a diagram illustrating an outline of an image encoding procedure.
[0111] 5, S501 may be performed in the prediction unit of the encoding device described above, S502 may be performed in the residual processing unit, and S503 may be performed in the entropy encoding unit 190. S501 may include the inter / intra prediction procedure described herein, S502 may include the residual processing procedure described herein, and S503 may include the information encoding procedure described herein.
[0112] As described above, the picture encoding procedure may include not only a procedure of encoding information for picture reconstruction (e.g., prediction information, residual information, partitioning information, etc.) and outputting it in a bitstream format, but also a procedure of generating a reconstructed picture for a current picture and a procedure of applying in-loop filtering to the reconstructed picture (optional). The encoding apparatus may derive a (modified) residual sample from the quantized transform coefficient through the inverse quantization unit and the inverse transform unit, and may generate a reconstructed picture based on the prediction sample output from S501 and the (modified) residual sample. The reconstructed picture generated in this manner may be the same as the reconstructed picture generated in the above-mentioned decoding apparatus. A modified reconstructed picture may be generated through an in-loop filtering procedure on the reconstructed picture, which may be stored in a decoded picture buffer or memory, and may be used as a reference picture in an inter prediction procedure when encoding a picture in the future, as in the case of the decoding apparatus. As described above, in some cases, some or all of the in-loop filtering procedure may be omitted. When the in-loop filtering procedure is performed, (in-loop) filtering related information (parameters) can be coded in an entropy coding unit and output in a bitstream format, and the decoding device can perform the in-loop filtering procedure in the same manner as the coding device based on the filtering related information.
[0113] Through such an in-loop filtering procedure, noises generated during image / video coding, such as blocking artifacts and ringing artifacts, can be reduced, and subjective / objective visual quality can be improved. In addition, by performing the in-loop filtering procedure in both the encoding device and the decoding device, the encoding device and the decoding device can derive the same prediction result, thereby improving the reliability of picture coding and reducing the amount of data to be transmitted for picture coding.
[0114] As described above, a picture reconstruction procedure may be performed not only in a decoding apparatus but also in an encoding apparatus. A reconstruction block may be generated based on intra prediction / inter prediction for each block, and a reconstruction picture including the reconstruction block may be generated. If a current picture / slice / tile group is an I picture / slice / tile group, blocks included in the current picture / slice / tile group may be reconstructed based only on intra prediction. Meanwhile, if a current picture / slice / tile group is a P or B picture / slice / tile group, blocks included in the current picture / slice / tile group may be reconstructed based on intra prediction or inter prediction. In this case, inter prediction may be applied to some blocks in the current picture / slice / tile group, and intra prediction may be applied to the remaining blocks. Color components of a picture may include luma components and chroma components, and unless explicitly limited herein, the methods and embodiments proposed herein may be applicable to luma components and chroma components.
[0115] Coding hierarchy and structure overview
[0116] FIG. 6 is a schematic diagram illustrating the coding hierarchy and structure.
[0117] Video / images coded according to this specification may be processed, for example, according to the coding hierarchy and structure described below.
[0118] The coded image is divided into the VCL (video coding layer), which handles the image decoding process and the image itself, a lower system that transmits and stores the coded information, and the NAL (network abstraction layer), which exists between the VCL and the lower system and is responsible for network adaptation functions.
[0119] In VCL, it is possible to generate VCL data including compressed image data (slice data), or to generate a parameter set including information such as a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), and a Video Parameter Set (VPS), or a Supplemental Enhancement Information (SEI) message that is additionally required for the image decoding process.
[0120] In NAL, a NAL unit can be generated by adding header information (NAL unit header) to RBSP (Raw Byte Sequence Payload) generated in VCL. In this case, RBSP refers to slice data, parameter set, SEI message, etc. generated in VCL. The NAL unit header can include NAL unit type information identified by the RBSP data included in the NAL unit.
[0121] As shown, NAL units can be divided into VCL NAL units and non-VCL NAL units according to the RBSP generated by the VCL. The VCL NAL unit can refer to a NAL unit that contains information about an image (slice data), and the non-VCL NAL unit can refer to a NAL unit that contains information required for decoding an image (parameter set or SEI message).
[0122] The above-mentioned VCL NAL unit and non-VCL NAL unit can be transmitted over a network with header information according to the data standard of the lower system. For example, the NAL unit can be transformed into a data format of a predetermined standard such as H.266 / VVC file format, RTP (Real-time Transport Protocol), TS (Transport Stream), etc., and transmitted over various networks.
[0123] As described above, the NAL unit type of a NAL unit can be identified according to the RBSP data structure included in the NAL unit, and information about such NAL unit type can be stored and signaled in the NAL unit header.
[0124] For example, NAL units can be largely classified into VCL NAL unit types and non-VCL NAL unit types depending on whether the NAL unit contains information about an image (slice data). The VCL NAL unit types can be classified according to the nature and type of pictures contained in the VCL NAL unit, and the non-VCL NAL unit types can be classified according to the type of parameter set.
[0125] Next, examples of NAL unit types identified by the types of parameter sets included in the Non-VCL NAL unit types are listed.
[0126] -APS (Adaptation Parameter Set) NAL unit: Type for NAL units containing APS
[0127] -DPS (Decoding Parameter Set) NAL unit: Type for NAL units containing DPS
[0128] -VPS (Video Parameter Set) NAL unit: Type for NAL units that contain VPS
[0129] -SPS (Sequence Parameter Set) NAL unit: Type for NAL units containing SPS
[0130] -PPS (Picture Parameter Set) NAL unit: Type for NAL units that contain PPS
[0131] The above-mentioned NAL unit type has syntax information for the NAL unit type, and the syntax information can be stored in a NAL unit header and signaled. For example, the syntax information can be nal_unit_type, and the NAL unit type can be specified by a value of nal_unit_type.
[0132] The slice header (slice header syntax) may include information / parameters commonly applicable to the slices. The APS (APS syntax) or PPS (PPS syntax) may include information / parameters commonly applicable to one or more slices or pictures. The SPS (SPS syntax) may include information / parameters commonly applicable to one or more sequences. The VPS (VPS syntax) may include information / parameters commonly applicable to multiple layers. The DPS (DPS syntax) may include information / parameters commonly applicable to videos in general. The DPS may include information / parameters related to concatenation of coded video sequences (CVSs). In this specification, a high level syntax (HLS) may include at least one of the APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, and slice header syntax.
[0133] In this specification, the image / video information encoded from the encoding device to the decoding device and signaled in bitstream format may include not only partitioning-related information within a picture, intra / inter prediction information, residual information, in-loop filtering information, etc., but also information contained in the slice header, information contained in the APS, information contained in the PPS, information contained in the SPS, and / or information contained in the VPS.
[0134] Signaling multilayer information within a VPS
[0135] The signaling of multiple layer information within the VPS is described in detail below.
[0136] For a multi-layer bitstream, the output layer set (OLS), which is the available set of decodable layers, such as dependencies between layers, profile, tier and level (PTL) information for the OLS, DPB information, and HRD information can be signaled as shown in the table below.
[0137] [Table 1-1] [Table 1-2] [Table 1-3] [Table 1-4]
[0138] Before being referenced, the VPS RBSP must be available to the decoding process, either by being contained in at least one AU with TemporalId equal to 0 or by being provided via external means.
[0139] All VPS NAL units with a particular vps_video_parameter_set_id value in the CVS must have the same content.
[0140] In Table 1, vps_video_parameter_set_id provides an identifier for the VPS for reference by other syntax elements. The value of vps_video_parameter_set_id must be greater than 0.
[0141] vps_max_layers_minus1+1 specifies the number of layers specified by the VPS, which is the maximum number of layers allowed in each CVS that references the VPS.
[0142] vps_max_sublayers_minus1+1 specifies the maximum number of temporal sublayers that may be present in a layer specified by the VPS. The value of vps_max_sublayers_minus1 must be in the range 0 to 6.
[0143] A vps_default_ptl_dpb_hrd_max_tid_flag of 1 specifies that the syntax elements vps_ptl_max_tid[i], vps_dpb_max_tid[i], and vps_ptl_max_tid[i] are not present and are inferred to be the same as the default value vps_max_sublayers_minus1. A vps_default_ptl_dpb_hrd_max_tid_flag of 0 specifies that the syntax elements vps_ptl_max_tid[i], vps_dpb_max_tid[i], and vps_ptl_max_tid[i] are present. Conversely, if the syntax elements are not present, then vps_default_ptl_dpb_hrd_max_tid_flag is inferred to be 1.
[0144] A vps_all_independent_layers_flag of 1 specifies that all layers specified by the VPS are coded independently without inter-layer prediction. A vps_all_independent_layers_flag of 0 specifies that one or more layers specified by the VPS can use inter-layer prediction. If the syntax element is not present, the value of vps_all_independent_layers_flag is inferred.
[0145] vps_layer_id[i] specifies the nuh_layer_id value of the i-th layer. If m and n are two non-negative integers, and m is less than n, then the value of vps_layer_id[m] must be less than vps_layer_id[n].
[0146] vps_independent_layer_flag[i] equal to 1 specifies that the layer with index i does not use inter-layer prediction. vps_independent_layer_flag[i] equal to 0 specifies that the layer with index i can use inter-layer prediction and that the syntax element vps_direct_ref_layer_flag[i][j], where j is in the range of 0 to i-1, is present in the VPS. If the syntax element is not present, the value of vps_independent_layer_flag[i] is inferred to be 1.
[0147] A vps_max_tid_ref_present_flag[i] equal to 1 specifies that the syntax element vps_max_tid_il_ref_pics_plus1[i][j] is present. A vps_max_tid_ref_present_flag[i] equal to 0 specifies that the syntax element vps_max_tid_il_ref_pics_plus1[i][j] is not present.
[0148] vps_direct_ref_layer_flag[i][j] equal to 0 specifies that the layer with index j is not a direct reference layer for the layer with index i. vps_direct_ref_layer_flag[i][j] equal to 1 specifies that the layer with index j is a direct reference layer for the layer with index i. If vps_direct_ref_layer_flag[i][j] is not present for i and j in the range 0 to vps_max_layers_minus1, then 0 is inferred. If vps_independent_layer_flag[i] is 0, then for j in the range 0 to i-1, there must be at least one j with vps_direct_ref_layer_flag[i][j] value 1.
[0149] The variables NumDirectRefLayers[i], DirectRefLayerIdx[i][d], NumRefLayers[i], RefLayerIdx[i][r], and LayerUsedAsRefLayerFlag[j] are derived as in FIG. 7.
[0150] The variable GeneralLayerIdx[i], which specifies the layer index of the layer whose nuh_layer_id is the same as vps_layer_id[i], is derived as shown in Figure 8.
[0151] For i and j, which may have any value different from each other but both in the range of 0 to vps_max_layers_minus1, if dependencyFlag[i][j] is 1, the values of tsps_chroma_format_idc and sps_bitdepth_minus8 applied to the i-th layer shall be identical to the values of sps_chroma_format_idc and sps_bitdepth_minus8 applied to the j-th layer, respectively. This is a bitstream consistency requirement.
[0152] vps_max_tid_il_ref_pics_plus1[i][j] equal to 0 specifies that a picture of the jth layer that is neither a GDR picture with ph_recovery_poc_cnt equal to 0 nor an IRAP picture shall not be used as an ILRP for picture decoding of the ith layer. vps_max_tid_il_ref_pics_plus1[i][j] greater than 0 specifies that any picture of the jth layer with TemporalId greater than vps_max_tid_il_ref_pics_plus1[i][j]-1 shall not be used as an ILRP for picture decoding of the ith layer. If this syntax element is not present, the value of vps_max_tid_il_ref_pics_plus1[i][j] shall be inferred to be equal to vps_max_sublayers_minus1+1.
[0153] vps_each_layer_is_an_ols_flag equal to 1 specifies that each OLS contains only one layer and each layer specified by the VPS is an OLS where the single-containing layer is the only output layer. vps_each_layer_is_an_ols_flag equal to 0 specifies that at least one OLS contains two or more layers. If vps_max_layers_minus1 is 0, then a value of 1 is inferred for vps_each_layer_is_an_ols_flag. Otherwise, if vps_all_independent_layers_flag is 0, then a value of 0 is inferred for vps_each_layer_is_an_ols_flag.
[0154] A vps_ols_mode_idc of 0 specifies that the number of total OLSs specified by the VPS is equal to vps_max_layers_minus1+1, the i-th OLS contains layers with layer indices in the range 0 to i, and for each OLS, the topmost layer of the OLS is the output layer.
[0155] A vps_ols_mode_idc of 1 specifies that the total number of OLSs specified by the VPS is equal to vps_max_layers_minus1+1, the i-th OLS contains layers with layer indices in the range 0 to i, and for each OLS, all layers of the OLS are output layers.
[0156] A vps_ols_mode_idc of 2 specifies that the total number of OLSs specified by the VPS is explicitly signaled, the output layer for each OLS is explicitly signaled, and other layers are direct or indirect reference layers to the output layer of the OLS.
[0157] The value of vps_ols_mode_idc must be in the range of 0 to 2. The value of vps_ols_mode_idc is 3, which is reserved for future ITU-T | ISO / IEC use.
[0158] If vps_all_independent_layers_flag is 1 and vps_each_layer_is_an_ols_flag is 0, then the value of vps_ols_mode_idc is inferred to be 2.
[0159] vps_num_output_layer_sets_minus1+1 specifies the number of total OLSs specified by the VPS when vps_ols_mode_idc is 2.
[0160] The variable olsModeIdc is derived as shown in Figure 9.
[0161] The variable TotalNumOlss, which specifies the total number of OLSs specified by the VPS, is derived as shown in Figure 10.
[0162] A vps_ols_output_layer_flag[i][j] of 1 specifies that the layer with nuh_layer_id equal to vps_layer_id[j] is the output layer of the ith OLS when vps_ols_mode_idc is 2. A vps_ols_output_layer_flag[i][j] of 0 specifies that the layer with nuh_layer_id equal to vps_layer_id[j] is not the output layer of the ith OLS when vps_ols_mode_idc is 2.
[0163] The variable NumOutputLayersInOls[i] that specifies the number of output layers of the i-th OLS, the variable NumSubLayersInLayerInOLS[i][j] that specifies the number of sublayers of the j-th layer of the i-th OLS, the variable OutputLayerIdInOls[i][j] that specifies the nuh_layer_id value of the j-th output layer of the i-th OLS, and the variable LayerUsedAsOutputLayerFlag[k] that specifies whether the k-th layer is used as an output layer in at least one OLS are derived as shown in Figure 11.
[0164] For each value in the range from 0 to vps_max_layers_minus1, the values of LayerUsedAsRefLayerFlag[i] and LayerUsedAsOutputLayerFlag[i] must not both be 0. In other words, there cannot be a layer that is neither an output layer of at least one OLS nor a direct reference layer of another layer.
[0165] In each OLS, at least one layer must be an output layer, in other words, for any value between 0 and TotalNumOlss-1, the value of NumOutputLayersInOls[i] must be 1 or greater.
[0166] The variable NumLayersInOls[i] that specifies the number of layers in the i-th OLS, the variable LayerIdInOls[i][j] that specifies the nuh_layer_id value of the j-th layer in the i-th OLS, the variable NumMultiLayerOlss that specifies the number of multi-layer OLSs (i.e., OLSs including two or more layers), and the variable MultiLayerOlsIdx[i] that specifies an index into the list of multi-layer OLSs in the i-th OLS when NumLayersInOls[i] is greater than 0 are derived as shown in FIG. 12.
[0167] Note 1 - The 0th OLS only includes the lowest layer (i.e. the layer with nuh_layer_id equal to vps_layer_id[0]), and only the included layers are output in the case of the 0th OLS.
[0168] The variable OlsLayerIdx[i][j], which specifies the OLS layer index of the layer whose nuh_layer_id is the same as LayerIdInOls[i][j], is derived as shown in Figure 13.
[0169] The lowest layer of each OLS must be an independent layer, i.e. for every i in the range 0 to TotalNumOlss-1, the value of vps_independent_layer_flag[GeneralLayerIdx[LayerIdInOls[i][0]]] must be 1.
[0170] Each layer must be included in at least one OLS specified by the VPS. In other words, there must be at least one pair for the values of i and j in each layer with a particular nuhLayerId value whose nuh_layer_id is the same as one of vps_layer_id[k], for k in the range 0 to vps_max_layers_minus1, where i is in the range 0 to TotalNumOlss-1 and j is in the range NumLayersInOls[i]-1, so that the value of LayerIdInOls[i][j] is the same as nuhLayerId.
[0171] vps_num_ptls_minus1+1 specifies the number of profile_tier_level() syntax structures in the VPS. The value of vps_num_ptls_minus1 must be less than TotalNumOlss. If this element is not present, the value of vps_num_ptls_minus1 is inferred to be 0.
[0172] A vps_pt_present_flag[i] of 1 specifies that profile, tier, and general constraint information is present in the i-th profile_tier_level() syntax structure of the VPS. A vps_pt_present_flag[i] of 0 specifies that profile, tier, and general constraint information is not present in the i-th profile_tier_level() syntax structure of the VPS. If vps_pt_present_flag[i] is 0, it is inferred that the profile, tier, and general constraint information for the i-th profile_tier_level() syntax structure of the VPS is the same as that of the (i-1)-th profile_tier_level() syntax structure of the VPS.
[0173] vps_ptl_max_tid[i] specifies the TemporalId of the highest sublayer representation whose level information is present in the i-th profile_tier_level() syntax structure in the VPS. The value of vps_ptl_max_tid[i] must be in the range of 0 to vps_max_sublayers_minus1. If this element is not present, the value of vps_ptl_max_tid[i] is inferred to be the same as vps_max_sublayers_minus1.
[0174] vps_ptl_alignment_zero_bit must be equal to 0.
[0175] vps_ols_ptl_idx[i] specifies the index into the list of profile_tier_level() syntax structures in the VPS of the profile_tier_level() syntax structure that applies to the i-th OLS. If this element is present, the value of vps_ols_ptl_idx[i] must be in the range of 0 to vps_num_ptls_minus1.
[0176] If this element is not present, the value of vps_ols_ptl_idx[i] is inferred as follows:
[0177] If -vps_num_ptls_minus1 is 0, the value of vps_ols_ptl_idx[i] is inferred to be 0.
[0178] - Otherwise (vps_num_ptls_minus1 is greater than 0 and vps_num_ptls_minus1+1 is equal to TotalNumOlss), the value of vps_ols_ptl_idx[i] is inferred to be i.
[0179] If NumLayersInOls[i] is 1, the profile_tier_level() syntax structure that applies to the i-th OLS also exists in the SPS referenced by the layer of the i-th OLS. If NumLayersInOls[i] is 1, it is a bitstream consistency requirement that the profile_tier_level() syntax structures signaled in the VPS and the SPS for the i-th OLS must be identical.
[0180] Each profile_tier_level() syntax structure of the VPS must be referenced as at least one vps_ols_ptl_idx[i] value, for i in the range of 0 to TotalNumOlss-1.
[0181] vps_num_dpb_params_minus1+1, if present, specifies the number of dpb_parameters() syntax structures in the VPS. The value of vps_num_dpb_params_minus1 must be in the range of 0 to NumMultiLayerOlss-1.
[0182] The variable VpsNumDpbParams, which specifies the number of dpb_parameters() syntax structures in the VPS, is derived as shown in Figure 14.
[0183] The vps_sublayer_dpb_params_present_flag is used to control the presence of the max_dec_pic_buffering_minus1[], max_num_reorder_pics[], and max_latency_increase_plus1[] syntax elements in the dpb_parameters() syntax structure in the VPS. If this element is not present, vps_sub_dpb_params_info_present_flag is inferred as 0.
[0184] vps_dpb_max_tid[i] specifies the TemporalId of the highest sublayer representation that a DPR parameter may be present in the i-th dpb_parameters() syntax structure in the VPS. The value of vps_dpb_max_tid[i] must be in the range of 0 to vps_max_sublayers_minus1. If this element is not present, the value of vps_dpb_max_tid[i] is inferred to be the same as vps_max_sublayers_minus1.
[0185] vps_ols_dpb_pic_width[i] specifies the width of each picture storage buffer for the i-th multi-layer OLS in luma samples.
[0186] vps_ols_dpb_pic_height[i] specifies the height of each picture storage buffer for the i-th multi-layer OLS, in luma samples.
[0187] vps_ols_dpb_chroma_format[i] specifies the maximum allowed value of sps_chroma_format_idc for all SPS referenced by the CLVS in the CVS for the i-th multi-layer OLS.
[0188] vps_ols_dpb_bitdepth_minus8[i] specifies the maximum allowed value of sps_bitdepth_minus8 for all SPSs referenced by the CLVS in the CVS for the i-th multi-layer OLS.
[0189] Note 2 - When decoding the i-th multi-layer OLS, the decoder can safely allocate memory for the DPB depending on the values of the vps_ols_dpb_pic_width[i], vps_ols_dpb_pic_height[i], vps_ols_dpb_chroma_format[i], and vps_ols_dpb_bitdepth_minus8[i] syntax elements.
[0190] vps_ols_dpb_params_idx[i] specifies the index, relative to the dpb_parameters() syntax structure in the VPS, of the dpb_parameters() syntax structure that applies to the i-th multi-layer OLS. If this element is present, the value of vps_ols_dpb_params_idx[i] must be in the range of 0 to VpsNumDebParams - 1.
[0191] If vps_ols_dpb_params_idx[i] is not present, the following is inferred:
[0192] -If VpsNumDpbParams is 1, the value of vps_ols_dpb_params_idx[i] is inferred to be 0.
[0193] - Otherwise (VpsNumDpbParams is greater than 1 and equal to NumMultiLayerOlss), the value of vps_ols_dpb_params_idx[i] is inferred to be i.
[0194] For single-layer OLS, the applicable dpb_parameters() syntax structures exist in the SPS that references the OLS layer.
[0195] Each dpb_parameters() syntax structure in a VPS must be referenced by at least one vps_ols_dpb_params_idx[i] value, for i in the range 0 to NumMultiLayerOlss-1.
[0196] A vps_general_hrd_params_present_flag of 1 specifies that the VPS includes the general_hrd_parameters() syntax structure and other HRD parameters. A vps_general_hrd_params_present_flag of 0 specifies that the VPS does not include the general_hrd_parameters() syntax structure or other HRD parameters.
[0197] If NumLayersInOls[i] is 1, the general_hrd_parameters() and ols_hrd_parameters() syntax structures that apply to the i-th OLS layer are in the SPS referenced by the i-th OLS layer.
[0198] A vps_sublayer_cpb_params_present_flag of 1 specifies that the i-th OLS_hrd_parameters() syntax structure of the VPS contains HRD parameters for sublayer representations with TemporalId in the range of 0 to vps_hrd_max_tid[i]. A vps_sublayer_cpb_params_present_flag of 0 specifies that the i-th ols_hrd_parameters() syntax structure of the VPS contains HRD parameters only for sublayer representations with TemporalId equal to vps_hrd_max_tid[i]. If vps_max_sublayers_minus1 is 0, the value of vps_sublayer_cpb_params_present_flag is inferred to be 0.
[0199] If vps_sublayer_cpb_params_present_flag is 0, the HRD parameters for sublayer representations with TemporalId in the range 0 to vps_hrd_max_tid[i]-1 are inferred to be identical to the HRD parameters for sublayer representations with TemporalId equal to vps_hrd_max_tid[i].
[0200] This includes the HRD parameters starting from the fixed_pic_rate_general_flag[i] syntax element up to the sublayer_hrd_parameters(i) syntax structure immediately under the “if(general_vcl_hrd_params_present_flag)” condition of the sublayer_hrd_parameters(i) syntax structure.
[0201] vps_num_ols_hrd_params_minus1+1 specifies the number of ols_hrd_parameters() syntax structures present in the VPS when vps_general_hrd_params_present_flag is 1. The value of vps_num_ols_hrd_params_minus1 must be in the range of 0 to NumMultiLayerOlss-1.
[0202] vps_hrd_max_tid[i] specifies the TemporalId of the top-level sublayer representation whose HRD parameters are included in the i-th ols_hrd_parameters() syntax structure. The value of vps_hrd_max_tid[i] must be in the range of 0 to vps_max_sublayers_minus1. If this element is not present, the value of vps_hrd_max_tid[i] is inferred to be vps_max_sublayers_minus1.
[0203] vps_ols_hrd_idx[i] specifies the index into the list of ols_hrd_parameters() syntax structures of the VPS for the ols_hrd_parameters() syntax structure that applies to the i-th multi-layer OLS. The value of vps_ols_hrd_idx[i] must be in the range of 0 to vps_num_ols_hrd_params_minus1.
[0204] If vps_ols_hrd_idx[i] does not exist, it is inferred as follows:
[0205] -If vps_num_ols_hrd_params_minus1 is 0, the value of vps_ols_hrd_idx[i] is inferred to be 0.
[0206] - Otherwise (vps_num_ols_hrd_params_minus1+1 is greater than 1 and is equal to NumMultiLayerOlss), the value of vps_ols_hrd_idx[i] is inferred to be i.
[0207] In the case of a single layer OLS, the applicable ols_hrd_parameters() syntax structure is present in the SPS that references the OLS layer.
[0208] Each ols_hrd_parameters() syntax structure of the VPS must be referenced as at least one vps_ols_hrd_idx[i] value, for i in the range of 1 to NumMultiLayerOlss-1.
[0209] A vps_extension_flag of 0 specifies that the vps_extension_data_flag syntax element is not present in the VPS RBSP syntax structure. A vps_extension_flag of 1 specifies that the vps_extension_data_flag syntax element may be present in the VPS RBSP syntax structure.
[0210] vps_extension_data_flag MAY have any value. The presence and value of this element have no effect on decoder conformance to the profile specified in this version of this specification. Decoders that conform to this version of this specification MUST ignore all vps_extension_data_flag syntax elements.
[0211] VPS and SPS
[0212] Signaling of VPS (video parameter set) and SPS (sequence parameter set) will be explained below.
[0213] The presence of a VPS is optional in the case of a single-layer bitstream. If a VPS is not present, the value of the syntax element sps_video_parameter_set_id is 0, and the values of some variables are inferred as described below.
[0214] If sps_video_parameter_set_id is 0, the following applies:
[0215] -The SPS does not reference a VPS, and when decoding each CLVS that references an SPS, the VPS is not referenced.
[0216] A value of -vps_max_layers_minus1 is inferred to be 0.
[0217] The value of -vps_max_sublayers_minus1 is inferred to be 6.
[0218] - A CVS must contain only one layer (i.e., all VCL NAL units in a CVS must have the same nuh_layer_id value).
[0219] -The value of GeneralLayerIdx[nuh_layer_id] is inferred to be 0.
[0220] The value of -vps_independent_layer_flag[GernalLayerIdx[nuh_layer_id]] is inferred to be 1.
[0221] Overview of SEI messages
[0222] The scalability dimension information (SDI) SEI message syntax will be described below with reference to Table 2.
[0223] [Table 2]
[0224] The following describes SDI (Scalability dimension information) SEI message semantics.
[0225] The SDI (scalability dimension information) SEI message provides scalability dimension information for each layer of bitstreamInScope (defined below), including, for example, 1) the view ID for each layer if bitstreamInScope is a multiview bitstream, and 2) the auxiliary ID for each layer if there is auxiliary information (e.g., depth or alpha) carried by one or more layers in bitstreamInScope.
[0226] The bitstreamInScope may be a sequence of AUs consisting of an AU including a current scalability dimension information SEI message and zero or more subsequent AUs in decoding order. In this case, the zero or more subsequent AUs may include all subsequent AUs up to the subsequent AU including the scalability dimension information SEI message. In this case, the subsequent AU including the scalability dimension information SEI message may not be included.
[0227] sdi_max_layers_minus1+1 can indicate the maximum number of layers in bitstreamInScope.
[0228] sdi_multiview_info_flag equal to 1 may indicate that bitstreamInScope may be a multiview bitstream and that the sdi_view_id_val[] syntax element is present in the scalability dimension information (SDI) SEI message. sdi_multiview_flag equal to 0 may indicate that bitstreamInScope is not a multiview bitstream and that the sdi_view_id_val[] syntax element is not present in the SDI SEI message.
[0229] sdi_auxiliary_info_flag equal to 1 may indicate that there may be auxiliary information conveyed by one or more layers in bitstreamInScope and that the sdi_aux_id[] syntax element is present in the SDI SEI message. sdi_auxiliary_info_flag equal to 0 may indicate that there may not be auxiliary information conveyed by one or more layers in bitstreamInScope and that the sdi_aux_id[] syntax element is not present in the SDI SEI message.
[0230] sdi_view_id_len may specify the length in bits of the sdi_view_id_val[i] syntax element.
[0231] sdi_view_id_val[i] may specify the view ID of the ith layer in bitstreamInScope. The length of the sdi_view_id_val[i] syntax element may be sdi_view_id_len bits. If this element is not present, the value of sdi_view_id_val[i] may be inferred to be 0.
[0232] An sdi_aux_id[i] of 0 may indicate that the i-th layer in bitstreamInScope does not contain an auxiliary picture. An sdi_aux_id[i] greater than 0 may indicate the type of auxiliary picture for the i-th layer in bitstreamInScope, as specified in Table 3 below.
[0233] Mapping of sdi_aux_id[i] to auxiliary picture type
[0234] [Table 3]
[0235] In Table 3 above, the resolution of auxiliary pictures associated with sdi_aux_id in the range of 128 to 159 may be specified via other means than the sdi_aux_id value.
[0236] For bitstreams according to the present disclosure, sdi_aux_id[i] can be in the range of 0 to 2 or in the range of 128 to 159. According to the present disclosure, even if the value of sdi_aux_id[i] is in the range of 0 to 2 or in the range of 128 to 159, the decoder must accept values of sdi_aux_id[i] in the range of 0 to 255.
[0237] As mentioned above, the SDI SEI message can provide additional information about layers present in the bitstream. To provide information about layers, a loop for the layer list in the bitstream can be included. However, the SDI SEI message does not include information about which layer each index i refers to. That is, the relationship between the i-th layer of the SDI SEI message and the actual layer itself can be defined by the layer information present in the VPS of the bitstream. This can cause undesirable results because the VPS may change and the entity that changes the VPS cannot know that the SEI message needs to be changed as well.
[0238] To solve this problem, the present disclosure proposes an SEI message including layer identifier information. That is, the present disclosure proposes an image encoding / decoding technique based on an SEI message including layer identifier information for layers included in a bitstream. This can improve coding efficiency by removing the dependency of layer identifier information on the VPS.
[0239] According to the present disclosure, as configurations for solving the above-mentioned problems, the following configurations 1 to 6 can be provided. The following configurations 1 to 6 may be realized independently, or two or more of them may be combined to be realized.
[0240] Configuration 1: A Scalable Dimension Information (SDI) SEI message can convey scalable dimensional information for a greater number of layers than the actual layers present in the associated CVS or bitstream.
[0241] Configuration 2: The value of sdi_max_layers_minus1 may not be limited to the number of layers in the corresponding CVS or bitstream, i.e., it may be greater than, equal to, or less than the number of layers in the bitstream.
[0242] Configuration 3: Alternatively, the value of sdi_max_layers_minus1+1 may not be smaller than the number of actual layers in the corresponding CVS or bitstream.
[0243] Configuration 4: For each view identifier and / or auxiliary identifier, layer id information can be signaled. In other words, the layer id information can be signaled using a loop for the layer in the SEI message.
[0244] Configuration 5: The layer identifier information in the iterative statement for the layers may be sorted in ascending order of layer identifier values or may be signaled in ascending order.
[0245] Configuration 6: Alternatively, the layer identifier information in the repeat statement for the layers can be sorted or signaled in descending order of layer identifier values.
[0246] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.
[0247] Example 1: SDI SEI Message Syntax and Semantics
[0248] Hereinafter, with reference to FIG. 15, an embodiment of the present disclosure will be described based on the SDI SEI message syntax and semantics according to the present disclosure.
[0249] As an example, an SDI SEI message may include scalability dimension information for each layer in a bitstream (e.g., bitstreamInScope (defined below)). For example, the SEI message may include 1) view ID information for each layer if the bitstream is a multiview bitstream, and 2) auxiliary IDs for each layer if there may be auxiliary information (e.g., depth or alpha) carried by one or more layers in the bitstream.
[0250] As an example, a bitstream (e.g., bitstreamInScope) may be a sequence of AUs consisting of one or more AUs that currently contain an SDI SEI message in decoding order, up to and including any subsequent AUs that contain an SDI SEI message, and which may be followed by zero or more subsequent AUs, including all subsequent AUs that do not contain an SEI message.
[0251] As an example, sdi_max_layers_minus1 plus 1 may indicate the maximum number of layers in the bitstream. If sdi_multiview_info_flag is 1, this may indicate that the bitstream may be a multiview bitstream and the sdi_view_id_val[] syntax element is present in the SDI SEI message. If sdi_multiview_flag is 0, this may indicate that the bitstream is not a multiview bitstream and the sdi_view_id_val[] syntax element is not present in the SDI SEI message.
[0252] As another example, an SDI SEI message may carry scalable dimension information for more layers than the actual number of layers present in the associated CVS or bitstream. That is, the value of sdi_max_layers_minus1 may not be limited to the number of layers in the CVS or bitstream. That is, it may be less than, equal to, or greater than the number of layers in the CVS or bitstream. Or, the value of sdi_max_layers_minus1+1 may not always be less than the actual number of layers in the CVS or bitstream.
[0253] If sdi_auxiliary_info_flag is 1, this may indicate that there may be auxiliary information carried by one or more layers in the bitstream and the sdi_aux_id[] syntax element is present in the SDI SEI message. If sdi_auxiliary_info_flag is 0, this may indicate that there is no auxiliary information carried by one or more layers in the bitstream and the sdi_aux_id[] syntax element is not present in the SDI SEI message.
[0254] sdi_view_id_len may represent the length in bits of the sdi_view_id_val[i] syntax element.
[0255] sdi_layer_id[i] may indicate information on the layer ID of the i-th layer in the bitstream (bitstreamInScope). The value of i may have a value from 0 to sdi_max_layers_minus1-1. In this case, the value of sdi_layer_id[i] may be sorted in descending order so as to be smaller than the value of sid_layer_id[i-1]. However, the value is not limited to this, and may be sorted in ascending order or other order.
[0256] As an example, sdi_view_id_val[i] may indicate layer identifier information of the i-th layer in the bitstream. The length of the sdi_view_id_val[i] syntax element may be sdi_view_id_len bits. If this element is not present, the value of sdi_view_id_val[i] may be considered to be 0.
[0257] As an example, if sdi_aux_id[i] is 0, this may indicate that the i-th layer in the bitstream does not contain an auxiliary picture. If sdi_aux_id[i] is greater than 0, this may indicate the type of auxiliary picture for the i-th layer in the bitstream, as specified in Table 4. Table 4 is as follows.
[0258] Mapping Sdi_aux_id[i] to Auxiliary Picture Type
[0259] [Table 4]
[0260] As shown in the table, the interpretation of the auxiliary picture type associated with sdi_aux_id in the range of 128 to 159 may not be specified. Thus, the auxiliary picture type may be specified via other means than the sdi_aux_id value.
[0261] As an example, for bitstream conformance, sdi_aux_id[i] must be in the range of 0 to 2, or it may be restricted to the range of 128 to 159. As an example, even if the value of sdi_aux_id[i] is restricted to the range of 0 to 2 or 128 to 159, the decoder must allow the value of sdi_aux_id[i] to have values in the range of 0 to 255.
[0262] As another example, a layer id may be signaled for each view id and / or auxiliary id. In other words, layer identifier information may be signaled in a repeated statement for a layer in the SEI message. In this case, the layer identifier information for a layer may be sorted or signaled in ascending order or in descending order with respect to the value of the layer identifier information.
[0263] According to the above embodiment, it is possible to improve coding efficiency by removing unnecessary dependency on the VPS related to the layer identifier.
[0264] Example 2: Image decoding method according to the present disclosure
[0265] Hereinafter, with reference to FIG. 16, an image decoding method according to an embodiment of the present disclosure will be described.
[0266] The image decoding method of Fig. 16 may be performed by an image decoding apparatus. The image decoding apparatus may first receive, as a bitstream, a supplemental enhancement information (SEI) message including information on one or more layers included in the bitstream (S1601). Thereafter, the image decoding apparatus may obtain information on the one or more layers based on the received SEI message (S1602), and restore an image in the bitstream based on the obtained information on the one or more layers (S1603).
[0267] In this case, the SEI message may be an SDI SEI message according to the present disclosure. As an example, the received SEI message may include layer identifier information for the one or more layers included in the bitstream. As an example, the layer identifier information may include the above-mentioned sdi_layer_id. Also, the layer identifier information may be obtained as much as the maximum number of layers in the bitstream. Information on the maximum number of layers in the bitstream may be obtained from the SEI message. That is, information on the maximum number of layers may be included in the SEI message. As an example, the SEI message may include information on layers other than one or more layers included in the bitstream. That is, it may include information on a number of layers greater than the actual number of layers included in the bitstream. That is, the maximum number of layers in the bitstream may not be limited to the number of one or more layers in the bitstream. As another example, the maximum number of layers in the bitstream may be limited to not have a value smaller than the number of one or more layers in the bitstream, or may be smaller than the number of one or more layers in the bitstream. Meanwhile, the layer identifier information may be signaled for each of view identifier information or auxiliary identifier information. In addition, the layer identifier information may be included in the SEI message in ascending order of layer identifier values so that the layer identifier of the i-th layer has a greater value than the layer identifier of the i-1-th layer, or conversely, in descending order of layer identifier values so that the layer identifier of the i-th layer has a smaller value than the layer identifier of the i-1-th layer.
[0268] Meanwhile, since FIG. 16 corresponds to one embodiment of the present disclosure, some steps may be added, modified, or deleted, and the order of steps may be changed.
[0269] Example 3: Image encoding method according to the present disclosure
[0270] Hereinafter, with reference to FIG. 17, an image encoding method according to an embodiment of the present disclosure will be described.
[0271] The image coding method of FIG. 17 may be performed by an image coding apparatus. The image coding apparatus may first code an image based on tisle information for one or more layers (S1701). That is, the image coding apparatus may code an image in a bitstream based on information for one or more layers in the bitstream. Then, the image coding apparatus may code an SEI message including information for one or more layers (S1702). That is, the image coding apparatus may code an SEI (supplemental enhancement information) message including information for one or more layers included in the bitstream into the bitstream.
[0272] In this case, the SEI message may be an SDI SEI message according to the present disclosure. As an example, the SEI message may include layer identifier information for one or more layers included in the bitstream. As an example, the layer identifier information may include the above-mentioned sdi_layer_id. Also, the layer identifier information may be included in the SEI message as many as the maximum number of layers in the bitstream. Information on the maximum number of layers in the bitstream may also be included in the SEI message. That is, information on the maximum number of layers may be included in the SEI message. As an example, the SEI message may further include information on layers other than the one or more layers included in the bitstream. As another example, the maximum number of layers in the bitstream may be limited to not have a value smaller than the number of one or more layers in the bitstream, or may be smaller than the number of one or more layers in the bitstream. Meanwhile, the layer identifier information may be included in the bitstream for each of view identifier information or auxiliary identifier information. Also, the layer identifier information may be included in the SEI message and encoded in ascending order of layer identifier values such that the layer identifier of the i-th layer has a value larger than the layer identifier of the i-1-th layer. Alternatively, they may be included in the SEI message in descending order.
[0273] Meanwhile, since FIG. 17 corresponds to one embodiment of the present disclosure, some steps may be added, modified, or deleted, and the order of steps may be changed.
[0274] Example 4: Image encoding / decoding device according to the present disclosure
[0275] As an embodiment, the image encoding / decoding device 1801 of Fig. 18 may include a memory 1802 for storing data and a processor 1803 for controlling the memory, and may perform the above-mentioned image encoding / decoding based on the processor. Also, the device 1801 of Fig. 18 is a simplified representation of the devices of Figs. 2 and 3, and may perform the functions described above.
[0276] 18 is an image decoding device, the processor 1803 may receive a supplemental enhancement information (SEI) message including information on one or more layers included in the bitstream as a bitstream, obtain information on one or more layers based on the received SEI message, and restore an image in the bitstream based on the obtained information on one or more layers. Also, the received SEI message may include layer identifier information for the one or more layers included in the bitstream.
[0277] As another example, when the device 1801 of FIG. 18 is an image encoding device, the processor 1803 may encode an image in the bitstream based on information for one or more layers in the bitstream, and may encode an SEI (supplemental enhancement information) message into the bitstream that includes information for one or more layers included in the bitstream. Meanwhile, the SEI message may include layer identifier information for the one or more layers included in the bitstream.
[0278] Various embodiments according to the present disclosure may be used alone or in combination with other embodiments.
[0279] Although the exemplary method of the present disclosure is expressed as a series of operations for clarity of explanation, this is not intended to limit the order in which the steps are performed, and each step may be performed simultaneously or in a different order, if necessary. To realize the method according to the present disclosure, the steps illustrated may include other steps, some steps may be omitted and the remaining steps may be omitted, or some steps may be omitted and additional steps may be included.
[0280] In the present disclosure, an image encoding device or an image decoding device that performs a predetermined operation (step) can perform the operation (step) to check the execution conditions or circumstances of the operation (step). For example, if it is described that a predetermined operation is performed when a predetermined condition is satisfied, the image encoding device or the image decoding device can perform the predetermined operation after performing an operation to check whether the predetermined condition is satisfied or not.
[0281] The various embodiments of the present disclosure are not intended to enumerate all possible combinations, but are intended to describe representative aspects of the present disclosure, and the matters described in the various embodiments may be applied independently or in combination of two or more.
[0282] Additionally, various embodiments of the present disclosure may be implemented using hardware, firmware, software, or a combination thereof, etc. In the case of a hardware implementation, the implementation may be implemented using one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), general processors, controllers, microcontrollers, microprocessors, etc.
[0283] In addition, the image decoding device and the image encoding device to which the embodiments of the present disclosure are applied may be included in a multimedia broadcast transmitting / receiving device, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video conversation device, a real-time communication device such as a video communication, a mobile streaming device, a storage medium, a camcorder, a custom video (VoD) service providing device, an over-the-top video (OTT) device, an internet streaming service providing device, a three-dimensional (3D) video device, an image telephone video device, and a medical video device, and may be used to process a video signal or a data signal. For example, the over-the-top video (OTT) device may include a game console, a Blu-ray player, an internet-connected TV, a home theater system, a smartphone, a tablet PC, a digital video recorder (DVR), and the like.
[0284] FIG. 19 is a diagram illustrating a content streaming system to which an embodiment of the present disclosure can be applied.
[0285] As shown in FIG. 19, a content streaming system to which an embodiment of the present disclosure is applied may broadly include an encoding server, a streaming server, a Web server, a media storage, a user device, and a multimedia input device.
[0286] The encoding server compresses content input from a multimedia input device such as a smartphone, camera, camcorder, etc. into digital data to generate a bitstream and transmits the bitstream to the streaming server. As another example, if a multimedia input device such as a smartphone, camera, video camera, etc. directly generates a bitstream, the server can be omitted.
[0287] The bitstream may be generated by an image encoding method and / or image encoding device to which an embodiment of the present disclosure is applied, and the streaming server may temporarily store the bitstream in the process of transmitting or receiving the bitstream.
[0288] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server can act as a medium to inform the user of available services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, and the streaming server can transmit the multimedia data to the user. At this time, the content streaming system can include a separate control server, and in this case, the control server can control commands / responses between devices in the content streaming system.
[0289] The streaming server may receive the content from a media storage and / or an encoding server. For example, when receiving the content from the encoding server, the content may be received in real time. In this case, the streaming server may store the bitstream for a certain period of time in order to provide a smooth streaming service.
[0290] Examples of the user devices include mobile phones, smart phones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices such as smartwatches, smart glass, head mounted displays (HMDs), digital TVs, desktop computers, and digital signage.
[0291] Each server in the content streaming system can be operated as a distributed server, in which case data received from each server can be processed in a distributed manner.
[0292] The scope of the present disclosure includes software or machine-executable commands (e.g., operating systems, applications, firmware, programs, etc.) that cause operations according to the methods of the various embodiments to be performed on a device or computer, and non-transitory computer-readable medium on which such software or commands can be stored and executed on a device or computer. [Industrial Applicability]
[0293] The embodiments of the present disclosure can be used to encode / decode images.
Claims
1. An image decoding method performed by an image decoding device, comprising: receiving a scalability dimension information (SDI) supplemental enhancement information (SEI) message containing information regarding one or more layers included in the bitstream; obtaining the information about the one or more layers based on the received SDI SEI message; reconstructing an image in the bitstream based on the obtained information about the one or more layers; the received SDI SEI message includes layer identifier information of the one or more layers included in the bitstream; the layer identifier information is obtained for a maximum number of layers in the bitstream; 11. The method of claim 10, wherein information regarding the maximum number of layers in the bitstream is obtained from the SDI SEI message.
2. The image decoding method of claim 1 , wherein the SDI SEI message includes information about layers other than the one or more layers included in the bitstream.
3. The image decoding method of claim 1 , wherein the maximum number of layers in the bitstream is not limited by the number of one or more layers in the bitstream.
4. The image decoding method of claim 1 , wherein the maximum number of layers in the bitstream is constrained to have no value less than the number of one or more layers in the bitstream.
5. The image decoding method of claim 1 , wherein the layer identifier information is signaled for each view identifier or auxiliary identifier.
6. 2. The image decoding method of claim 1, wherein the layer identifier information is included in the SDI SEI message in ascending order of layer identifier values, such that a layer identifier for an i-th layer has a higher value than a layer identifier for an (i-1)-th layer.
7. 2. The image decoding method of claim 1, wherein the layer identifier information is included in the SDI SEI message in descending order of layer identifier values, such that a layer identifier of an i-th layer has a smaller value than a layer identifier of an (i-1)-th layer.
8. An image coding method performed by an image coding device, comprising: encoding an image in a bitstream based on information about one or more layers in the bitstream; encoding a scalability dimension information (SDI) supplemental enhancement information (SEI) message into the bitstream, the SDI message including information about the one or more layers included in the bitstream; the SDISEI message includes layer identifier information of the one or more layers included in the bitstream; The layer identifier information is included in the SDI SEI message for the maximum number of layers in the bitstream; 11. The image coding method, wherein information regarding a maximum number of layers in the bitstream is included in the SDI SEI message.
9. A method for transmitting a bitstream, comprising: encoding an image in a bitstream based on information about one or more layers in the bitstream; encoding a scalability dimension information (SDI) supplemental enhancement information (SEI) message into the bitstream, the SEI message including information about the one or more layers included in the bitstream; transmitting the bitstream; The SDI SEI message includes layer identifier information of the one or more layers included in the bitstream, The layer identifier information is included in the SDI SEI message for the maximum number of layers in the bitstream; A method according to claim 1, wherein information regarding a maximum number of layers in the bitstream is included in the SDI SEI message.
Citation Information
Patent Citations
Encoding, storing and signaling scalability information
JP2008536420A
Encoding concept enabling efficient multi-view / layer encoding
JP2016519513A