Image encoding / decoding method and apparatus based on sub-layer level information, and recording medium storing bitstream
By using an image encoding/decoding method based on sub-level information, which encodes/decodes sub-level information flags in descending order of time identifier values, the problem of transmission and storage costs for high-resolution and high-quality images is solved, and the encoding/decoding efficiency is improved.
Patent Information
- Application Number
- CN202180041691.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-10
- Filing Date
- 2021-06-08
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2041-06-08
AI Technical Summary
Existing technologies, when transmitting and storing high-resolution and high-quality images, lead to increased transmission and storage costs due to the increased amount of information, necessitating efficient image compression technologies.
An image encoding/decoding method based on sub-level information is adopted. By encoding/decoding sub-level information flags in descending order of time identifier values, signal notification and inference processes are carried out in the same loop, and a storable bit stream is generated.
It improves the efficiency of image encoding/decoding, reduces transmission and storage costs, and achieves efficient image information processing.
Smart Images

Figure CN115702566B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to image encoding / decoding methods and apparatus, as well as recording media for storing bitstreams. More specifically, it relates to an image encoding / decoding method and apparatus based on sub-layer level information, and a recording medium for storing bitstreams generated by the image encoding method / apparatus of this disclosure. Background Technology
[0002] Recently, there has been an increasing demand across various fields for high-resolution and high-quality images, such as high-definition (HD) and ultra-high-definition (UHD) images. With the increase in image data resolution and quality, the amount of information or bits transmitted increases relative to existing image data. This increase in the amount of information or bits transmitted leads to increased transmission and storage costs.
[0003] Therefore, efficient image compression techniques are needed to effectively transmit, store, and reproduce information about high-resolution and high-quality images. Summary of the Invention
[0004] Technical issues
[0005] The purpose of this disclosure is to provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.
[0006] Another object of this disclosure is to provide an image encoding / decoding method and apparatus based on sub-layer level information flags encoded / decoded in descending order according to time identifier values.
[0007] Another object of this disclosure is to provide an image encoding / decoding method and apparatus based on sub-layer level information encoded / decoded in descending order according to time identifier values.
[0008] Another object of this disclosure is to provide an image encoding / decoding method and apparatus that performs signal notification / parsing and inference processes for sub-level information in the same loop.
[0009] Another object of this disclosure is to provide a non-transitory recording medium for storing bit streams generated by an image encoding method or device according to this disclosure.
[0010] Another object of this disclosure is to provide a non-transitory recording medium for storing a bitstream received, decoded and used to reconstruct an image by an image decoding device according to this disclosure.
[0011] Another object of this disclosure is to provide a method for transmitting a bitstream generated by an image encoding method or device according to this disclosure.
[0012] The technical problems solved by this disclosure are not limited to those described above. Other technical problems not described herein will become clear to those skilled in the art through the following description.
[0013] Technical solution
[0014] An image decoding method according to one aspect of this disclosure includes: obtaining from a bitstream a first flag indicating the presence of sublayer-level information for each of one or more sublayers in a specified current layer; and obtaining the sublayer-level information from the bitstream based on the first flag. The first flag may be obtained in descending order of the time identifier values of one or more sublayers.
[0015] An image decoding apparatus according to another aspect of this disclosure includes a memory and at least one processor. The at least one processor can obtain from a bitstream a first flag indicating the presence or absence of sublayer-level information for each of one or more sublayers in the current layer, and obtain the sublayer-level information from the bitstream based on the first flag. The first flag can be obtained in descending order of the time identifier values of one or more sublayers.
[0016] An image coding method according to another aspect of this disclosure includes: encoding a first flag indicating the presence or absence of sub-layer level information for each of one or more sub-layers in the current layer; and encoding the sub-layer level information based on the first flag. The first flag may be encoded in descending order of the time identifier values of one or more sub-layers.
[0017] Additionally, according to another aspect of this disclosure, a computer-readable recording medium can store a bitstream generated by the image encoding device or image encoding method of this disclosure.
[0018] In another aspect of the transmission method according to this disclosure, a bit stream generated by the image encoding method or image encoding device of this disclosure may be transmitted.
[0019] The features described above in this brief overview are merely exemplary aspects of the following detailed description of this disclosure and do not limit the scope of this disclosure.
[0020] Beneficial effects
[0021] According to this disclosure, an image encoding / decoding method and apparatus with improved encoding / decoding efficiency can be provided.
[0022] According to this disclosure, an image encoding / decoding method and apparatus can be provided based on sub-layer level information flags encoded / decoded in descending order according to time identifier values.
[0023] According to this disclosure, an image encoding / decoding method and apparatus can be provided based on sub-layer level information encoded / decoded in descending order according to time identifier values.
[0024] According to this disclosure, an image encoding / decoding method and apparatus can be provided that performs sub-level information information in the same loop, including signal notification / parsing and inference processes.
[0025] Furthermore, according to this disclosure, a computer-readable recording medium may be provided for storing a bitstream generated by an image encoding method or apparatus according to this disclosure.
[0026] Furthermore, according to this disclosure, a computer-readable recording medium may be provided that stores a bitstream received, decoded, and used by an image decoding device according to this disclosure for reconstructing an image.
[0027] Furthermore, according to this disclosure, a method for transmitting a bit stream generated by an image encoding method or device according to this disclosure can be provided.
[0028] Those skilled in the art will understand that the effects achievable through this disclosure are not limited to those specifically described above, and that other advantages of this disclosure will become clearer from the detailed description. Attached Figure Description
[0029] Figure 1 This is a diagram that schematically illustrates a video coding system to which embodiments of the present disclosure are applicable.
[0030] Figure 2 This is a diagram that schematically illustrates an image encoding device to which embodiments of the present disclosure are applicable.
[0031] Figure 3 This is a diagram that schematically illustrates an image decoding apparatus to which embodiments of the present disclosure are applicable.
[0032] Figure 4 This is a diagram illustrating an example of the layered structure for encoding images / videos.
[0033] Figure 5 This is a schematic block diagram of a multilayer coding apparatus to which embodiments of the present disclosure are applicable, and wherein the step of encoding multilayer video / image signals is performed.
[0034] Figure 6 This is a schematic block diagram of a decoding device to which embodiments of the present disclosure are applicable, and wherein the step of decoding multi-layer video / image signals is performed.
[0035] Figure 7 This is a diagram illustrating a method for encoding an image based on a multi-layer structure using an image encoding device according to an embodiment.
[0036] Figure 8 This is a diagram illustrating a method for decoding an image based on a multi-layer structure using an image decoding device according to an embodiment.
[0037] Figure 9 This is a diagram illustrating an example of VPS syntax that includes PTL information.
[0038] Figure 10 This is a diagram illustrating an example of SPS syntax that includes PTL information.
[0039] Figure 11 This is a diagram illustrating an example of the profile_tier_level(profileTierPresentFlag,maxNumSubLayersMinus1) syntax that includes PTL information.
[0040] Figure 12 This is a diagram illustrating the syntax of profile_tier_level(profileTierPresentFlag,maxNumSubLayersMinus1) according to an embodiment of this disclosure.
[0041] Figure 13 This is a flowchart illustrating an image encoding method according to an embodiment of the present disclosure.
[0042] Figure 14 This is a flowchart illustrating an image decoding method according to an embodiment of the present disclosure.
[0043] Figure 15 A diagram illustrating the content streaming system to which the embodiments of this disclosure are applicable. Detailed Implementation
[0044] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings to facilitate implementation by those skilled in the art. However, this disclosure can be implemented in various different forms and is not limited to the embodiments described herein.
[0045] In describing this disclosure, detailed descriptions of relevant known functions or constructions will be omitted if they unnecessarily obscure the scope of this disclosure. In the accompanying drawings, portions irrelevant to the description of this disclosure are omitted, and similar reference numerals are assigned to similar portions.
[0046] In this disclosure, when a component is "connected," "linked," or "coupled" to another component, it may include not only direct connections but also indirect connections where intermediate components exist. Furthermore, when a component "comprises" or "has" other components, unless otherwise stated, it means that other components may be included, not excluded.
[0047] In this disclosure, the terms first, second, etc., are used only for the purpose of distinguishing one component from other components and do not limit the order or importance of the components, unless otherwise stated. Accordingly, within the scope of this disclosure, a first component in one embodiment may be referred to as a second component in another embodiment, and similarly, a second component in one embodiment may be referred to as a first component in another embodiment.
[0048] In this disclosure, the components are distinguished from each other to clearly describe each feature, but this does not mean that the components must be separate. That is, multiple components may be integrated into a single hardware or software unit, or a single component may be distributed and implemented across multiple hardware or software units. Therefore, unless otherwise specified, implementations of these integrated or distributed components are included within the scope of this disclosure.
[0049] In this disclosure, the components described in the various embodiments are not necessarily essential components, and some components may be optional. Therefore, embodiments consisting of a subset of the components described in the embodiments are also included within the scope of this disclosure. Furthermore, embodiments that include other components besides those described in the various embodiments are also included within the scope of this disclosure.
[0050] This disclosure relates to the encoding and decoding of images. Unless redefined in this disclosure, the terms used herein may have the general meaning commonly used in the art to which this disclosure pertains.
[0051] In this disclosure, a "picture" generally refers to a unit representing an image within a specific time period, while a slice / tile is a coding unit that constitutes part of a picture. A picture can be composed of one or more slices / tiles. Furthermore, a slice / tile may include one or more coding tree units (CTUs).
[0052] In this disclosure, "pixel" or "pixel" can refer to the smallest unit that constitutes a frame (or image). Furthermore, "sample" can be used as a term corresponding to a pixel. A sample can generally represent a pixel or a pixel value, or it can represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component.
[0053] In this disclosure, "unit" can refer to a basic unit of image processing. A unit may include at least one of a specific region of a picture and information associated with that region. In some cases, the term "unit" may be used interchangeably with terms such as "sample array," "block," or "region." Generally, an M×N block may include a set (or array) of samples (or transform coefficients) with M columns and N rows.
[0054] In this disclosure, "current block" can mean one of "current coding block," "current coding unit," "coding target block," "decoding target block," or "processing target block." When performing prediction, "current block" can mean "current prediction block" or "prediction target block." When performing transform (inverse transform) / quantization (dequantization), "current block" can mean "current transform block" or "transform target block." When performing filtering, "current block" can mean "filter target block."
[0055] Furthermore, in this disclosure, unless explicitly stated as a chroma block, "current block" may mean a block that includes both luma component blocks and chroma component blocks, or "the luma block of the current block." The luma component block of the current block can be represented by an explicit description including terms such as "luma block" or "current luma block." Similarly, "the chroma component block of the current block" can be represented by an explicit description including terms such as "chroma block" or "current chroma block."
[0056] In this disclosure, the terms “ / ” or “,” can be interpreted as indicating “and / or”. For example, “A / B” and “A, B” can mean “A and / or B”. Furthermore, “A / B / C” and “A / B / C” can mean “at least one of A, B and / or C”.
[0057] In this disclosure, the term "or" should be interpreted to indicate "and / or". For example, the expression "A or B" can include 1) only "A", 2) only "B", or 3) both "A and B". In other words, in this disclosure, "or" should be interpreted to indicate "additionally or alternatively".
[0058] Overview of Video Encoding Systems
[0059] Figure 1 This is a diagram illustrating a video coding system to which embodiments of the present disclosure are applicable.
[0060] The video encoding system according to the embodiment may include an encoding device 10 and a decoding device 20. The encoding device 10 may deliver encoded video and / or image information or data to the decoding device 20 in the form of a file or stream via a digital storage medium or network.
[0061] The encoding device 10 according to an embodiment may include a video source generator 11, an encoding unit 12, and a transmitter 13. The decoding device 20 according to an embodiment may include a receiver 21, a decoding unit 22, and a renderer 23. The encoding unit 12 may be referred to as a video / image encoding unit, and the decoding unit 22 may be referred to as a video / image decoding unit. The transmitter 13 may be included in the encoding unit 12. The receiver 21 may be included in the decoding unit 22. The renderer 23 may include a display, and the display may be configured as a separate device or an external component.
[0062] The video source generator 11 can acquire video / images through a process of capturing, compositing, or generating video / images. The video source generator 11 may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, a video / image archive including previously captured video / images, etc. The video / image generation device may include, for example, a computer, tablet computer, and smartphone, and can generate video / images (electronically). For example, virtual video / images can be generated by a computer, etc. In this case, the video / image capture process can be replaced by a process of generating related data.
[0063] The encoding unit 12 can encode the input video / image. For compression and encoding efficiency, the encoding unit 12 can perform a series of processes, such as prediction, transformation, and quantization. The encoding unit 12 can output encoded data (encoded video / image information) in the form of a bitstream.
[0064] Transmitter 13 can transmit encoded video / image information or data, output in bitstream form, to receiver 21 of decoding device 20 in the form of a file or stream via digital storage medium or network. Digital storage medium can include various storage media, such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. Transmitter 13 can include elements for generating media files according to a predetermined file format and may include elements for transmission via broadcast / communication networks. Receiver 21 can extract / receive bitstreams from storage medium or network and transmit the bitstreams to decoding unit 22.
[0065] The decoding unit 22 can decode video / images by performing a series of processes corresponding to the operations of the encoding unit 12, such as dequantization, inverse transform, and prediction.
[0066] Renderer 23 can render decoded video / images. The rendered video / images can be displayed on a monitor.
[0067] Overview of Image Encoding Devices
[0068] Figure 2 This is a diagram schematically illustrating an image encoding apparatus to which embodiments of the present disclosure may be applied.
[0069] like Figure 2As shown, the image encoding device 100 may include an image segmenter 110, a subtractor 115, a transformer 120, a quantizer 130, a dequantizer 140, an inverse transformer 150, an adder 155, a filter 160, a memory 170, an inter-frame prediction unit 180, an intra-frame prediction unit 185, and an entropy encoder 190. The inter-frame prediction unit 180 and the intra-frame prediction unit 185 may be collectively referred to as "predictors". The transformer 120, quantizer 130, dequantizer 140, and inverse transformer 150 may be included in a residual processor. The residual processor may also include a subtractor 115.
[0070] In some implementations, all or at least some of the components configuring the image encoding device 100 may be configured by a single hardware component (e.g., an encoder or a processor). Furthermore, the memory 170 may include a decoded screen buffer (DPB) and may be configured by a digital storage medium.
[0071] Image segmenter 110 can segment an input image (or picture or frame) input to image encoding device 100 into one or more processing units. For example, a processing unit may be called an encoding unit (CU). Encoding units can be obtained by recursively segmenting encoding tree units (CTUs) or maximum encoding units (LCUs) according to a quadtree / binary tree / tritree (QT / BT / TT) structure. For example, an encoding unit can be segmented into multiple encoding units of greater depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. For the segmentation of encoding units, a quadtree structure can be applied first, followed by a binary tree structure and / or a ternary tree structure. The encoding process according to this disclosure can be performed based on the final encoding unit that is no longer segmented. The maximum encoding unit can be used as the final encoding unit, or a deeper encoding unit obtained by segmenting the maximum encoding unit can be used as the final encoding unit. Here, the encoding process may include prediction, transformation, and reconstruction processes, which will be described later. As another example, the processing unit of the encoding process may be a prediction unit (PU) or a transformation unit (TU). Prediction units and transform units can be partitioned or segmented from the final coding unit. Prediction units can be sample prediction units, and transform units can be units used to derive transform coefficients and / or units used to derive residual signals from transform coefficients.
[0072] The prediction unit (inter-frame prediction unit 180 or intra-frame prediction unit 185) can perform prediction on the block to be processed (the current block) and generate a prediction block that includes prediction samples of the current block. The prediction unit can determine whether to apply intra-frame prediction or inter-frame prediction based on the current block or CU. The prediction unit can generate various information related to the prediction of the current block and transmit the generated information to the entropy encoder 190. The information about the prediction can be encoded in the entropy encoder 190 and output as a bitstream.
[0073] Intra-prediction unit 185 can predict the current block by referencing samples in the current frame. Depending on the intra-prediction mode and / or intra-prediction technique, the reference samples may be located among the neighbors of the current block or may be placed separately. Intra-prediction modes may include multiple non-directional modes and multiple directional modes. Non-directional modes may include, for example, DC mode and planar mode. Depending on the level of detail in the prediction direction, directional modes may include, for example, 33 or 65 directional prediction modes. However, this is merely an example, and more or fewer directional prediction modes may be used depending on the settings. Intra-prediction unit 185 can determine the prediction mode to be applied to the current block by using prediction modes applied to neighboring blocks.
[0074] The inter-frame prediction unit 180 can deduce the prediction block of the current block based on a reference block (reference sample array) specified by motion vectors on a reference frame. In this case, to reduce the amount of motion information transmitted in the inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation between motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference frame indices. Motion information may also include inter-frame prediction direction (L0 prediction, L1 prediction, dual prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current frame and temporally neighboring blocks existing in the reference frame. The reference frame including the reference block and the reference frame including the temporally neighboring block may be the same or different. The temporally neighboring block may be referred to as a juxtaposed reference block, a juxtaposed CU (colCU), etc. The reference frame including the temporally neighboring block may be referred to as a juxtaposed frame (colPic). For example, the inter-frame prediction unit 180 can configure a motion information candidate list based on neighboring blocks and generate information indicating which candidate to use to deduce the motion vector and / or reference frame index of the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in skip mode and merge mode, the inter-frame prediction unit 180 can use motion information from neighboring blocks as motion information for the current block. In skip mode, unlike merge mode, residual signals may not be transmitted. In motion vector prediction (MVP) mode, motion vectors from neighboring blocks can be used as motion vector predictors, and the motion vector of the current block can be signaled by encoding motion vector differences and indicators of the motion vector predictors. The motion vector difference can refer to the difference between the motion vector of the current block and the motion vector predictor.
[0075] The prediction unit can generate a prediction signal based on various prediction methods and techniques described below. For example, the prediction unit can apply not only intra-frame prediction or inter-frame prediction, but also both intra-frame prediction and inter-frame prediction simultaneously to predict the current block. A prediction method that simultaneously applies both intra-frame prediction and inter-frame prediction to predict the current block can be called Combined Inter-Frame and Intra-Frame Prediction (CIIP). Furthermore, the prediction unit can perform Intra-Frame Block Copy (IBC) to predict the current block. Intra-Frame Block Copy can be used for content image / video coding in games, such as Screen Content Coding (SCC). IBC is a method of predicting the current frame using a previously reconstructed reference block in the current frame at a predetermined distance from the current block. When IBC is applied, the position of the reference block in the current frame can be encoded as a vector (block vector) corresponding to the predetermined distance. IBC essentially performs prediction in the current frame, but can be performed similarly to inter-frame prediction because the reference block is derived within the current frame. That is, IBC can use at least one inter-frame prediction technique described in this disclosure.
[0076] The prediction signal generated by the prediction unit can be used to generate a reconstructed signal or a residual signal. Subtractor 115 can generate a residual signal (residual block or residual sample array) by subtracting the prediction signal (prediction block or prediction sample array) output from the prediction unit from the input image signal (original block or original sample array). The generated residual signal can be transmitted to converter 120.
[0077] Transformer 120 can generate transform coefficients by applying transform techniques to the residual signal. For example, the transform techniques may include at least one of Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Karhunen-Loève Transform (KLT), Graph-Based Transform (GBT), or Conditional Nonlinear Transform (CNT). Here, GBT refers to a transform obtained from a graph when the relationship information between pixels is represented graphically. CNT refers to a transform obtained based on a prediction signal generated using all previously reconstructed pixels. Furthermore, the transform processing can be applied to square pixel blocks of the same size or to blocks of variable size instead of square.
[0078] Quantizer 130 can quantize the transform coefficients and transmit them to entropy encoder 190. Entropy encoder 190 can encode the quantized signal (information about the quantized transform coefficients) and output a bitstream. The information about the quantized transform coefficients can be referred to as residual information. Quantizer 130 can rearrange the block-type quantized transform coefficients into a one-dimensional vector form based on the coefficient scan order, and generate information about the quantized transform coefficients based on the one-dimensional vector form of the quantized transform coefficients.
[0079] The entropy encoder 190 can perform various encoding methods, such as exponential Columbus coding, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy encoder 190 can encode, either together or separately, the information required for video / image reconstruction other than the quantization transform coefficients (e.g., values of syntax elements). The encoded information (e.g., encoded video / image information) can be transmitted or stored in bitstream form at the Network Abstraction Layer (NAL) level. The video / image information may also include information about various parameter sets, such as Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). Furthermore, the video / image information may also include general constraint information. The signaled information, transmitted information, and / or syntax elements described in this disclosure can be encoded and included in the bitstream through the above encoding process.
[0080] The bitstream can be transmitted over a network or stored in a digital storage medium. The network may include broadcast networks and / or communication networks, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. A transmitter (not shown) for transmitting the signal output from the entropy encoder 190 and / or a storage unit (not shown) for storing the signal may be included as internal / external components of the image encoding device 100. Alternatively, a transmitter may be provided as a component of the entropy encoder 190.
[0081] The quantization transform coefficients output from quantizer 130 can be used to generate residual signals. For example, the residual signals (residual blocks or residual samples) can be reconstructed by applying dequantization and inverse transform to the quantization transform coefficients through dequantizer 140 and inverse transformer 150.
[0082] Adder 155 adds the reconstructed residual signal to the prediction signal output from inter-frame prediction unit 180 or intra-frame prediction unit 185 to generate a reconstructed signal (reconstructed frame, reconstructed block, reconstructed sample array). If the block to be processed has no residual, such as in the case of applying skip mode, the prediction block can be used as a reconstructed block. Adder 155 can be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current frame, and can be used for inter-frame prediction of the next frame by filtering as described below.
[0083] Filter 160 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 160 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image and store the modified reconstructed image in memory 170, specifically in the DPB of memory 170. Various filtering methods can include, for example, deblocking filtering, sample adaptive offsetting, adaptive loop filtering, bilateral filtering, etc. Filter 160 can generate various filtering-related information and transmit the generated information to entropy encoder 190, as described later in the description of each filtering method. The filtering-related information can be encoded by entropy encoder 190 and output as a bitstream.
[0084] The modified reconstructed frame transmitted to memory 170 can be used as a reference frame in inter-frame prediction unit 180. When inter-frame prediction is applied by image encoding device 100, prediction mismatch between image encoding device 100 and image decoding device can be avoided and coding efficiency can be improved.
[0085] The DPB of memory 170 can store modified reconstructed frames for use as reference frames in inter-frame prediction unit 180. Memory 170 can store motion information of blocks from which motion information in the current frame is derived (or encoded) and / or motion information of already reconstructed blocks in the frame. The stored motion information can be transmitted to inter-frame prediction unit 180 and used as motion information for spatially or temporally neighboring blocks. Memory 170 can store reconstructed samples of reconstructed blocks in the current frame and can transmit the reconstructed samples to intra-frame prediction unit 185.
[0086] Overview of image decoding devices
[0087] Figure 3 This is a diagram schematically illustrating an image decoding device to which embodiments of the present disclosure may be applied.
[0088] like Figure 3 As shown, the image decoding device 200 may include an entropy decoder 210, a dequantizer 220, an inverse transformer 230, an adder 235, a filter 240, a memory 250, an inter-frame prediction unit 260, and an intra-frame prediction unit 265. The inter-frame prediction unit 260 and the intra-frame prediction unit 265 may be collectively referred to as "predictors". The dequantizer 220 and the inverse transformer 230 may be included in a residual processor.
[0089] According to an implementation, all or at least some of the components of the image decoding device 200 can be configured by hardware components (e.g., a decoder or a processor). Furthermore, the memory 250 may include a decoded screen buffer (DPB) or may be configured by a digital storage medium.
[0090] The image decoding device 200, having received a bitstream including video / image information, can perform operations related to... Figure 2 The image is reconstructed by processing corresponding to the processing performed by the image encoding device 100. For example, the image decoding device 200 can perform decoding using a processing unit applied in the image encoding device. Therefore, the decoding processing unit can be, for example, an encoding unit. The encoding unit can be obtained by segmenting a coding tree unit or a maximum coding unit. The reconstructed image signal decoded and output by the image decoding device 200 can be reproduced by a reproduction device (not shown).
[0091] Image decoding device 200 can receive data in bitstream form from... Figure 2The signal output by the image encoding device. The received signal can be decoded by the entropy decoder 210. For example, the entropy decoder 210 can parse the bitstream to derive the information (e.g., video / image information) required for image reconstruction (or picture reconstruction). The video / image information may also include information about various parameter sets, such as adaptive parameter sets (APS), picture parameter sets (PPS), sequence parameter sets (SPS), or video parameter sets (VPS). In addition, the video / image information may also include general constraint information. The image decoding device can also decode the picture based on the information about the parameter sets and / or general constraint information. The information and / or syntax elements notified / received by signals described in this disclosure can be decoded and obtained from the bitstream through the decoding process. For example, the entropy decoder 210 decodes the information in the bitstream based on encoding methods such as exponential Golomb coding, CAVLC, or CABAC, and outputs the values of the syntax elements required for image reconstruction and the quantized values of the transform coefficients of the residuals. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element in the bitstream, determine the context model using information about the target syntax element, decoding information of neighboring blocks and the target block, or information about symbols / bins decoded in the previous stage, perform arithmetic decoding on the bins based on the determined context model by predicting the occurrence probability of the bins, and generate symbols corresponding to the value of each syntax element. In this case, the CABAC entropy decoding method can update the context model after determining the context model by using the information of the decoded symbols / bins for the context model of the next symbol / bin. The prediction-related information in the information decoded by the entropy decoder 210 can be provided to the prediction units (inter-frame prediction unit 260 and intra-frame prediction unit 265), and the residual value of entropy decoding performed in the entropy decoder 210, i.e., the quantization transform coefficients and related parameter information, can be input to the dequantizer 220. In addition, the filtering information in the information decoded by the entropy decoder 210 can be provided to the filter 240. Furthermore, the receiver (not shown) for receiving signals output from the image encoding device may be further configured as an internal / external element of the image decoding device 200, or the receiver may be a component of the entropy decoder 210.
[0092] Furthermore, the image decoding apparatus according to this disclosure can be referred to as a video / image / screen decoding apparatus. The image decoding apparatus can be divided into an information decoder (video / image / screen information decoder) and a sample decoder (video / image / screen sample decoder). The information decoder may include an entropy decoder 210. The sample decoder may include at least one of a dequantizer 220, an inverse transformer 230, an adder 235, a filter 240, a memory 250, an inter-frame prediction unit 260, or an intra-frame prediction unit 265.
[0093] Dequantizer 220 can dequantize the quantized transform coefficients and output transform coefficients. Dequantizer 220 can rearrange the quantized transform coefficients in the form of two-dimensional blocks. In this case, the rearrangement can be performed based on the coefficient scan order performed in the image encoding device. Dequantizer 220 can obtain transform coefficients by performing dequantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information).
[0094] The inverse transformer 230 can perform inverse transformation on the transformation coefficients to obtain the residual signal (residual block, residual sample array).
[0095] The prediction unit can perform prediction on the current block and generate a prediction block that includes prediction samples of the current block. The prediction unit can determine whether to apply intra-frame prediction or inter-frame prediction to the current block based on the prediction information output from the entropy decoder 210, and can determine a specific intra-frame / inter-frame prediction mode (prediction technique).
[0096] Similar to that described in the prediction unit of the image coding device 100, the prediction unit can generate a prediction signal based on various prediction methods (techniques) described later.
[0097] Intra-prediction unit 265 can predict the current block by referring to samples in the current frame. The description of intra-prediction unit 185 also applies to intra-prediction unit 265.
[0098] The inter-frame prediction unit 260 can deduce the prediction block of the current block based on a reference block (reference sample array) specified by a motion vector on a reference frame. In this case, to reduce the amount of motion information transmitted in the inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference frame indices. Motion information may also include inter-frame prediction direction (L0 prediction, L1 prediction, dual prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current frame and temporally neighboring blocks existing in the reference frame. For example, the inter-frame prediction unit 260 can configure a motion information candidate list based on neighboring blocks and deduce the motion vector and / or reference frame index of the current block based on the received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and the information about the prediction may include information indicating the inter-frame prediction mode of the current block.
[0099] Adder 235 generates a reconstruction signal (reconstructed frame, reconstruction block, reconstruction sample array) by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the prediction unit (including inter-frame prediction unit 260 and / or intra-frame prediction unit 265). If the block to be processed has no residual (e.g., in the case of applying skip mode), the prediction block can be used as a reconstruction block. The description of adder 155 also applies to adder 235. Adder 235 may be referred to as a reconstructor or reconstruction block generator. The generated reconstruction signal can be used for intra-frame prediction of the next block to be processed in the current frame, and can be used for inter-frame prediction of the next frame by filtering as described below.
[0100] Filter 240 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 240 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image and store the modified reconstructed image in memory 250, specifically in the DPB of memory 250. Various filtering methods may include, for example, deblocking filtering, adaptive sample shifting, adaptive loop filtering, bilateral filtering, etc.
[0101] The (modified) reconstructed frame stored in the DPB of memory 250 can be used as a reference frame in inter-frame prediction unit 260. Memory 250 can store motion information of blocks from which motion information in the current frame is derived (or decoded) and / or motion information of already reconstructed blocks in the frame. The stored motion information can be transmitted to inter-frame prediction unit 260 to be used as motion information for spatially or temporally neighboring blocks. Memory 250 can store reconstructed samples of reconstructed blocks in the current frame and transmit the reconstructed samples to intra-frame prediction unit 265.
[0102] In this disclosure, the embodiments described in the filter 160, inter-frame prediction unit 180 and intra-frame prediction unit 185 of the image encoding device 100 can be equally or correspondingly applied to the filter 240, inter-frame prediction unit 260 and intra-frame prediction unit 265 of the image decoding device 200.
[0103] Example of encoding layer structure
[0104] The encoded video / images according to this disclosure can be processed, for example, according to the coding layers and structures described below.
[0105] Figure 4 This is a diagram illustrating an example of the layered structure for encoding images / videos.
[0106] Encoded images / videos are classified into a Video Coding Layer (VCL) for image / video decoding and processing, a lower-layer system for sending and storing encoded information, and a Network Abstraction Layer (NAL) that exists between the VCL and the lower-layer system and is responsible for network adaptation functions.
[0107] In VCL, VCL data that includes compressed image data (slice data) can be generated, or additional enhancement information (SEI) messages required for image decoding processing or parameter sets that include information such as picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS) can be generated.
[0108] In NAL, header information (NAL unit header) can be added to the raw byte sequence payload (RBSP) generated in VCL to generate NAL units. In this case, RBSP refers to the slice data, parameter set, and SEI message generated in VCL. The NAL unit header may include NAL unit type information specified according to the RBSP data included in the corresponding NAL unit.
[0109] like Figure 4 As shown, NAL units can be classified into VCL NAL units and non-VCL NAL units based on the type of RBSP generated in the VCL. A VCL NAL unit can refer to a NAL unit that includes information about the image (slice data), while a non-VCL NAL unit can refer to a NAL unit that includes information required for decoding the image (parameter set or SEI message).
[0110] VCL NAL units and non-VCL NAL units can be appended with header information and sent over the network according to the data standard of the underlying system. For example, NAL units can be modified to have a data format with a predetermined standard (e.g., H.266 / VVC file format, RTP (Real-Time Transport Protocol), or TS (Transport Stream)) and sent over various networks.
[0111] As described above, within a NAL unit, the NAL unit type can be specified based on the RBSP data structure included in the corresponding NAL unit, and information about the NAL unit type can be stored in the NAL unit header and signaled. For example, this can be broadly categorized into VCL NAL unit types and non-VCL NAL unit types based on whether the NAL unit includes information about the image (slice data). VCL NAL unit types can be categorized based on the nature and type of the image included in the VCL NAL unit, while non-VCL NAL unit types can be categorized based on the type of parameter set.
[0112] Below are examples of NAL cell types specified based on the type of parameter set / information included in non-VCL NAL cell types.
[0113] - DCI (Decoding Capability Information) NAL Unit: Includes NAL unit types for DCI.
[0114] -VPS (Video Parameter Set) NAL Unit: Includes the NAL unit type of VPS.
[0115] -SPS (Sequence Parameter Set) NAL Unit: Includes NAL unit types related to SPS.
[0116] -PPS (Picture Parameter Set) NAL Unit: Includes the NAL unit type of PPS.
[0117] -APS (Adaptive Parameter Set) NAL Units: Includes NAL unit types related to APS.
[0118] -PH (Picture Header) NAL Unit: Includes NAL unit types for the picture header.
[0119] The aforementioned NAL unit type can have syntax information for the NAL unit type, and this syntax information can be stored in the NAL unit header and signaled. For example, this syntax information can be nal_unit_type, and the NAL unit type can be specified using the nal_unit_type value.
[0120] Furthermore, a frame can include multiple slices, and a slice can include a slice header and slice data. In this case, a frame header can be further added to the multiple slices (slice headers and slice data sets) within a frame. The frame header (frame header syntax) can include information / parameters common to the frame. The slice header (slice header syntax) can include information / parameters common to the slice. APS (APS syntax) or PPS (PPS syntax) can include information / parameters common to one or more slices or frames. SPS (SPS syntax) can include information / parameters common to one or more sequences. VPS (VPS syntax) can be information / parameters common to multiple layers. DCI (DCI syntax) can include information / parameters related to decoding capabilities.
[0121] In this disclosure, the high-level syntax (HLS) may include at least one of APS syntax, PPS syntax, SPS syntax, VPS syntax, DCI syntax, picture header syntax, or slice header syntax. Additionally, in this disclosure, the low-level syntax (LLS) may include, for example, slice data syntax, CTU syntax, coding unit syntax, transform unit syntax, etc.
[0122] In this disclosure, the image / video information encoded in the encoding device and signaled to the decoding device in the form of a bitstream may include not only intra-frame segmentation related information, intra / inter-frame prediction information, residual information, and intra-loop filtering information, but also information about slice headers, frame headers, APS, PPS, SPS, VPS, and / or DCI. Additionally, the image / video information may also include general constraint information and / or information about NAL unit headers.
[0123] Multi-layer coding
[0124] Image / video coding according to this disclosure may include multi-layer-based image / video coding. Multi-layer-based image / video coding may include scalable coding. In multi-layer-based coding or scalable coding, input signals can be processed for individual layers. Depending on the layer, the input signal (input image / video) may have different values in at least one aspect: resolution, frame rate, bit depth, color format aspect ratio, or view. In this case, performing inter-layer prediction by using the differences between layers (e.g., based on scalability) can reduce redundant information transmission / processing and increase compression efficiency.
[0125] Figure 5 This is a schematic block diagram of a multilayer encoding device 500 to which embodiments of the present disclosure are applicable for encoding multilayer video / image signals.
[0126] Figure 5 The multi-layer coding device 500 may include Figure 2 Encoding devices. With Figure 2 In comparison, Figure 5 The image segmenter 110 and adder 155 are not shown in the multi-layer coding device 500, but the multi-layer coding device 500 may include the image segmenter 110 and adder 155. In this case, the image segmenter 110 and adder 155 may be included on a layer-by-layer basis. Hereinafter, Figure 5 The description will focus on multi-layer-based prediction. For example, in addition to the following description, the multi-layer coding device 500 may include the above-mentioned references. Figure 2 The technical concept of the described encoding device.
[0127] For ease of description, Figure 5 The diagram illustrates a multi-layer structure consisting of two layers. However, embodiments of this disclosure are not limited to two layers, and multi-layer structures applying embodiments of this disclosure may include two or more layers.
[0128] Reference Figure 5The encoding device 500 includes an encoder 500-1 for layer 1 and an encoder 500-0 for layer 0. Layer 0 can be a base layer, a reference layer, or a lower layer, and layer 1 can be an enhancement layer, a current layer, or a higher layer.
[0129] The encoder 500-1 of layer 1 may include a predictor 520-1, a residual processor 530-1, a filter 560-1, a memory 570-1, an entropy encoder 540-1, and a multiplexer (MUX) 570. In an embodiment, the MUX may be included as an external component.
[0130] The encoder 500-0 of layer 0 may include a predictor 520-0, a residual processor 530-0, a filter 560-0, a memory 570-0, and an entropy encoder 540-0.
[0131] Predictors 520-0 and 520-1 can perform predictions on the input image based on various prediction schemes as described above. For example, predictors 520-0 and 520-1 can perform inter-frame prediction and intra-frame prediction. Predictors 520-0 and 520-1 can perform predictions in a predetermined processing unit. The prediction unit can be a coding unit (CU) or a transform unit (TU). Prediction blocks (including prediction samples) can be generated based on the prediction results, and based on this, the residual processor can derive residual blocks (including residual samples).
[0132] Inter-frame prediction can generate prediction blocks based on information about at least one of the previous and / or next frames. Intra-frame prediction can generate prediction blocks based on neighboring samples in the current frame.
[0133] As inter-frame prediction modes or methods, the various prediction modes or methods described above can be used. In inter-frame prediction, a reference frame can be selected for the current block to be predicted, and a reference block corresponding to the current block can be selected from the reference frame. Predictors 520-0 and 520-1 can generate prediction blocks based on the reference blocks.
[0134] Additionally, predictor 520-1 can use information about layer 0 to perform predictions for layer 1. In this disclosure, for ease of description, the method of using information about another layer to predict information about the current layer is referred to as inter-layer prediction.
[0135] Information about the current layer that is predicted using information about another layer (e.g., predicted via inter-layer prediction) can be at least one of texture, motion information, unit information, or predetermined parameters (e.g., filtering parameters, etc.).
[0136] Additionally, the information about another layer used for prediction of the current layer (e.g., for inter-layer prediction) can be at least one of texture, motion information, unit information, or predetermined parameters (e.g., filtering parameters, etc.).
[0137] Inter-layer prediction: The current block can be a block in the current frame of the current layer (e.g., layer 1) or the block to be encoded. The reference block is a block in the frame (reference frame) of the same access unit (AU) as the frame (current frame) to which the prediction of the current block belongs, on the layer (reference layer, e.g., layer 0) to which the prediction of the current block belongs, and can be a block corresponding to the current block.
[0138] As an example of inter-layer prediction, there exists inter-layer motion prediction that uses motion information from a reference layer to predict the motion information of the current layer. Based on inter-layer motion prediction, the motion information of the current block can be predicted using the motion information of the reference block. That is, when deriving motion information according to the inter-frame prediction mode described below, motion information candidates can be derived based on the motion information of the inter-layer reference block rather than temporally neighboring blocks.
[0139] When applying interlayer motion prediction, predictor 520-1 can scale and use motion information from the reference block of the reference layer (i.e., the interlayer reference block).
[0140] As another example of inter-layer prediction, inter-layer texture prediction can use the texture of a reconstructed reference block as the predicted value for the current block. In this case, predictor 520-1 can scale the texture of the reference block by up-scaling. Inter-layer texture prediction can be referred to as inter-layer (reconstructed) sample prediction or simply inter-layer prediction.
[0141] In inter-layer parameter prediction, which is another example of inter-layer prediction, the derivation parameters of the reference layer can be reused in the current layer, or the parameters of the current layer can be derived based on the parameters used in the reference layer.
[0142] In interlayer residual prediction, which is another example of interlayer prediction, residual information from another layer can be used to predict the residual information of the current layer, and based on this, prediction of the current block can be performed.
[0143] In inter-layer difference prediction, which is another example of inter-layer prediction, the prediction of the current block can be performed using the difference between the images obtained by upsampling or downsampling the reconstructed image of the current layer and the reconstructed image of the reference layer.
[0144] In inter-layer syntax prediction, another example of inter-layer prediction, the syntax information of a reference layer can be used to predict or generate the texture of the current block. In this case, the syntax information of the reference layer can include information about the intra-frame prediction mode and motion information.
[0145] The various prediction methods described above can be used when predicting specific blocks.
[0146] Here, as examples of interlayer prediction, although interlayer texture prediction, interlayer motion prediction, interlayer unit information prediction, interlayer parameter prediction, interlayer residual prediction, interlayer difference prediction, interlayer syntax prediction, etc. are described, the interlayer prediction applicable in this disclosure is not limited to these.
[0147] For example, inter-layer prediction can be applied as an extension of inter-frame prediction for the current layer. That is, inter-frame prediction can be performed for the current block by including the reference frame derived from the reference layer in the reference frames that the inter-frame prediction of the current block can refer to.
[0148] In this case, inter-layer reference frames can be included in the reference frame list for the current block. Predictor 520-1 can use inter-layer reference frames to perform inter-frame prediction for the current block.
[0149] Here, the inter-layer reference frame can be a reference frame constructed by sampling the reconstructed frame of a reference layer to correspond with the current layer. Therefore, when the reconstructed frame of the reference layer corresponds to the frame of the current layer, the reconstructed frame of the reference layer can be used as the inter-layer reference frame without sampling. For example, when the width and height of the samples in the reconstructed frames of the reference layer and the current layer are the same, and the offset between the top left, top right, bottom left, and bottom right edges of the reference layer and the top left, top right, bottom left, and bottom right edges of the current layer is 0, the reconstructed frame of the reference layer can be used as the inter-layer reference frame of the current layer without resampling.
[0150] Furthermore, the reconstructed frame of the reference layer of the interlayer reference frame can be a frame belonging to the same AU as the current frame to be encoded.
[0151] When performing inter-frame prediction for the current block by including inter-layer reference frames in the reference frame list, the positions of the inter-layer reference frames in the reference frame list can differ between reference frame lists L0 and L1. For example, in reference frame list L0, the inter-layer reference frame can be located after a short reference frame preceding the current frame, while in reference frame list L1, the inter-layer reference frame can be located at the end of the reference frame list.
[0152] Here, reference frame list L0 is a list of reference frames used for inter-frame prediction of P-slices or a list of reference frames used as the first reference frame list in inter-frame prediction of B-slices. Reference frame list L1 can be a second reference frame list used for inter-frame prediction of B-slices.
[0153] Therefore, the reference frame list L0 can be composed, in sequence, of short-term reference frames preceding the current frame, inter-layer reference frames, short-term reference frames following the current frame, and long-term reference frames. The reference frame list L1 can be composed, in sequence, of short-term reference frames following the current frame, short-term reference frames preceding the current frame, long-term reference frames, and inter-layer reference frames.
[0154] In this context, a prediction (P) slice is a slice in which intra-frame prediction is performed or inter-frame prediction is performed using at most one motion vector per prediction block and a reference frame index. A double prediction (B) slice is a slice in which intra-frame prediction is performed or prediction is performed using at most two motion vectors per prediction block and a reference frame index. In this respect, an intra-frame (I) slice is a slice in which intra-frame prediction is applied only.
[0155] Additionally, when performing inter-frame prediction for the current block based on a list of reference frames that includes inter-layer reference frames, the list of reference frames may include multiple inter-layer reference frames derived from multiple layers.
[0156] When multiple inter-layer reference frames are included, they can be arranged alternately in reference frame lists L0 and L1. For example, suppose two inter-layer reference frames, such as inter-layer reference frames ILRPi and ILRPiJ, are included in the reference frame list for inter-frame prediction of the current block. In this case, in reference frame list L0, ILRPi can be located after a short reference frame preceding the current frame, and ILRPi can be located at the end of the list. Similarly, in reference frame list L1, ILRPi can be located at the end of the list, and ILRPi can be located after a short reference frame following the current frame.
[0157] In this case, the reference frame list L0 can be composed, in sequence, of short-term reference frames preceding the current frame, interlayer reference frames (ILRPi), short-term reference frames following the current frame, long-term reference frames, and interlayer reference frames (ILRPj). The reference frame list L1 can be composed, in sequence, of short-term reference frames following the current frame, interlayer reference frames (ILRPj), short-term reference frames preceding the current frame, long-term reference frames, and interlayer reference frames (ILRPi).
[0158] Additionally, one of the two inter-layer reference frames can be an inter-layer reference frame derived from a scalable layer used for resolution, and the other can be an inter-layer reference frame derived from a layer used to provide another view. In this case, for example, if ILRPi is an inter-layer reference frame derived from a layer used to provide a different resolution and ILRPj is an inter-layer reference frame derived from a layer used to provide a different view, then in the case of scalability-only scalable video coding without view scalability support, the reference frame list L0 can be composed in sequence of short-term reference frames before the current frame, inter-layer reference frame ILRPi, short-term reference frames after the current frame, and long-term reference frames, and the reference frame list L1 can be composed in sequence of short-term reference frames after the current frame, short-term reference frames before the current frame, long-term reference frames, and inter-layer reference frame ILRPi.
[0159] Furthermore, in inter-layer prediction, information about the inter-layer reference frame can be obtained using only sample values, only motion information (motion vectors), or both. When the reference frame index indicates an inter-layer reference frame, based on information received from the encoding device, the predictor 520-1 can use only sample values of the inter-layer reference frame, only motion information (motion vectors) of the inter-layer reference frame, or both.
[0160] When using only sample values from the inter-frame reference frame, predictor 520-1 can derive samples of the block specified by the motion vector from the inter-frame reference frame as the prediction samples for the current block. Without considering scalable video coding of the view, the motion vectors in inter-frame prediction (inter-frame prediction) using the inter-frame reference frame can be set to a fixed value (e.g., 0).
[0161] When using only the motion information from the inter-layer reference frame, predictor 520-1 can use the motion vector specified by the inter-layer reference frame as a motion vector predictor for deriving the motion vector of the current block. Alternatively, predictor 520-1 can use the motion vector specified by the inter-layer reference frame as the motion vector of the current block.
[0162] When using both sample values and motion information from the inter-layer reference frame, the predictor 520-1 can use samples from the region in the inter-layer reference frame corresponding to the current block and the motion information (motion vector) specified in the inter-layer reference frame to predict the current block.
[0163] When applying interlayer prediction, the encoding device can send a reference index of the interlayer reference frame in the reference frame list to the decoding device, and can also send information to the decoding device to specify which information (sample information, motion information, or both) to be used from the interlayer reference frame, that is, information to specify the dependency type of the interlayer prediction dependency between the two layers.
[0164] Figure 6 This is a schematic block diagram of a decoding apparatus to which embodiments of the present disclosure are applicable for performing decoding of multi-layer video / image signals. Figure 6 The decoding device may include Figure 3 Decoding devices. Figure 6 The realigner shown can be omitted or included in the dequantizer. This diagram focuses on multi-level prediction. Alternatively, it may include... Figure 3 Description of the decoding device.
[0165] exist Figure 6 In the examples provided, for ease of description, a multi-layer structure consisting of two layers will be described. However, it should be noted that the embodiments of this disclosure are not limited thereto, and the multi-layer structure to which the embodiments of this disclosure are applied may include two or more layers.
[0166] Reference Figure 6 The decoding device 600 may include a layer 1 decoder 600-1 and a layer 1 decoder 600-0. The layer 1 decoder 600-1 may include an entropy decoder 610-1, a residual processor 620-1, a predictor 630-1, an adder 640-1, a filter 650-1, and a memory 660-1. The layer 0 decoder 600-2 may include an entropy decoder 610-0, a residual processor 620-0, a predictor 630-0, an adder 640-0, a filter 650-0, and a memory 660-0.
[0167] When a bitstream containing image information is received from an encoding device, the demultiplexer (DEMUX) 605 can demultiplex the information from each layer and send the information to the decoding devices of each layer.
[0168] Entropy decoders 610-1 and 610-0 can perform decoding corresponding to the encoding method used in the encoding device. For example, when the encoding device uses CABAC, entropy decoders 610-1 and 610-0 can perform entropy decoding using CABAC.
[0169] When the prediction mode of the current block is intra-prediction mode, predictors 630-1 and 630-0 can perform intra-prediction of the current block based on neighboring reconstructed samples in the current frame.
[0170] When the prediction mode for the current block is inter-frame prediction mode, predictors 630-1 and 630-0 can perform inter-frame prediction for the current block based on information included in at least one frame before or after the current frame. Some or all of the motion information required for inter-frame prediction can be derived by examining information received from the encoding device.
[0171] When the skip mode is applied as the inter-frame prediction mode, no residual is sent from the encoding device, and the prediction block can be a reconstruction block.
[0172] Furthermore, the predictor 630-1 of layer 1 can perform inter-frame prediction or intra-frame prediction using only information about layer 1 and perform inter-layer prediction using information about another layer (layer 0).
[0173] As information about the current layer that is predicted using information about another layer (e.g., predicted via inter-layer prediction), it may include at least one of texture, motion information, unit information, and predetermined parameters (e.g., filtering parameters, etc.).
[0174] As information about another layer used for prediction of the current layer (e.g., for inter-layer prediction), at least one of texture, motion information, unit information, and predetermined parameters (e.g., filtering parameters, etc.) may exist.
[0175] In inter-layer prediction, the current block can be a block in the current frame of the current layer (e.g., layer 1), and can be the block to be decoded. The reference block can be a block in the frame (reference frame) of the same access unit (AU) as the frame (current frame) to which the prediction of the current block belongs, on the layer (reference layer, e.g., layer 0) to which the prediction of the current block belongs, and can be the block corresponding to the current block.
[0176] The multilayer decoding device 600 can perform interlayer prediction as described in the multilayer coding device 500. For example, the multilayer decoding device 600 can perform interlayer texture prediction, interlayer motion prediction, interlayer unit information prediction, interlayer parameter prediction, interlayer residual prediction, interlayer difference prediction, interlayer syntax prediction, etc., as described in the multilayer coding device 500, and the interlayer prediction applicable in this disclosure is not limited thereto.
[0177] When a reference frame index is received from the encoding device or when a reference frame index derived from a neighboring block indicates an interlayer reference frame in the reference frame list, the predictor 630-1 can perform interlayer prediction using the interlayer reference frame. For example, when the reference frame index indicates an interlayer reference frame, the predictor 630-1 can deduce the sample values of the region specified by the motion vector in the interlayer reference frame as the prediction block for the current block.
[0178] In this case, inter-layer reference frames can be included in the reference frame list for the current block. Predictor 630-1 can use inter-layer reference frames to perform inter-frame prediction for the current block.
[0179] As described above in the multi-layer encoding device 500, in the operation of the multi-layer decoding device 600, the inter-layer reference frame can be a reference frame constructed by sampling the reconstructed frame of the reference layer to correspond with the current layer. The processing for the case where the reconstructed frame of the reference layer corresponds to the frame of the current layer can be performed in the same manner as the encoding processing.
[0180] Furthermore, as described above in the multilayer encoding device 500, in the operation of the multilayer decoding device 600, the reconstructed frame of the reference layer from which the interlayer reference frame is derived can be a frame belonging to the same AU as the current frame to be encoded.
[0181] Furthermore, as described above in the multilayer coding device 500, in the operation of the multilayer decoding device 600, when inter-frame prediction of the current block is performed by including interlayer reference frames in the reference frame list, the positions of the interlayer reference frames in the reference frame list L0 and L1 may be different.
[0182] Furthermore, as described above in the multilayer coding device 500, in the operation of the multilayer decoding device 600, when inter-frame prediction of the current block is performed based on a reference frame list including interlayer reference frames, the reference frame list may include multiple interlayer reference frames derived from multiple layers, and the arrangement of the interlayer reference frames may be performed to correspond to the encoding process described.
[0183] Furthermore, as described above in the multilayer encoding device 500, in the operation of the multilayer decoding device 600, the information about the interlayer reference image can be obtained by using only sample values, only motion information (motion vectors), or both sample values and motion information.
[0184] The multilayer decoding device 600 can receive reference indices from the multilayer encoding device 500 indicating interlayer reference frames in the reference frame list and perform interlayer prediction based on these indices. Additionally, the multilayer decoding device 600 can receive information from the multilayer encoding device 500 specifying which information (sample information, motion information, or both) is used from the interlayer reference frames; that is, information specifying the dependency type of the interlayer prediction dependency between two layers.
[0185] Reference Figure 7 and Figure 8This document describes an image encoding method and a decoding method performed by a multilayer image encoding device and a multilayer image decoding device according to embodiments. Hereinafter, for ease of description, the multilayer image encoding device may be referred to as an image encoding device. Additionally, the multilayer image decoding device may be referred to as an image decoding device.
[0186] Figure 7 This diagram illustrates a method for encoding an image based on a multi-layer structure using an image encoding device according to an embodiment. The image encoding device according to the embodiment can encode a first layer of the image (S710). Next, the image encoding device can encode a second layer of the image based on the first layer (S720). Next, the image encoding device can output a bitstream (for multiple layers) (S730).
[0187] Figure 8 This diagram illustrates a method for decoding an image based on a multi-layer structure using an image decoding device according to an embodiment. The image decoding device according to the embodiment can obtain video / image information from a bitstream (S810). Next, the image decoding device can decode the first layer of the image based on the video / image information (S820). Next, the image decoding device can decode the second layer of the image based on the video / image information and the first layer (S830).
[0188] In one embodiment, the video / image information may include the High-Level Syntax (HLS) described below. In one embodiment, the HLS may include the SPS and / or PPS disclosed in this disclosure. For example, the video / image information may include the information and / or syntax elements described in this disclosure. As described in this disclosure, the second layer of the image may be encoded based on the motion information / reconstruction samples / parameters of the first layer's image. In one embodiment, the layer may be a layer lower than the second layer. In one embodiment, when the second layer is the current layer, the first layer may be referred to as the reference layer.
[0189] Video Parameter Set Signaling
[0190] Information about multiple layers can be signaled in the Video Parameter Set (VPS). For example, for a multi-layer bitstream, information about dependencies between layers and / or the available set of decodeable layers can be signaled in the VPS. Here, the available set of decodeable layers may be referred to as the Output Layer Set (OLS). Additionally, the PTL (Profile, Layer, and Level) information, DPB (Decoded Picture Buffer) information, and / or HRD (Hypothetical Reference Decoder) information of the OLS can be signaled in the VPS.
[0191] Figure 9 Here is a diagram illustrating an example of VPS syntax that includes PTL information.
[0192] Reference Figure 9 A VPS can include the syntax element `vps_video_parameter_set_id`, which specifies the identifier of the VPS. The value of `vps_video_parameter_set_id` can be constrained to be greater than 0.
[0193] Additionally, a VPS can include vps_max_layers_minus1, vps_max_sublayer_minus1, and vps_all_layers_same_num_sublayer_flag as syntax elements relating to the number of layers or sublayers.
[0194] The value obtained by incrementing the syntax element vps_max_layers_minus1 by 1 can specify the maximum allowed number of layers in each encoded video sequence (CVS) of the reference VPS.
[0195] The value obtained by adding 1 to the syntax element vps_max_sublayer_minus1 specifies the maximum number of time sublayers that can exist in each CVS of the reference VPS. In one implementation, the value of vps_max_sublayer_minus1 can be constrained to the range of 0 to 6.
[0196] The syntax element `vps_all_layers_same_num_sublayer_flag` specifies whether the number of time sublayers is the same for all layers in each CVS of the reference VPS. For example, a first value (e.g., 1) for `vps_all_layers_same_num_sublayer_flag` specifies that the number of time sublayers is the same for all layers in each CVS of the reference VPS. In contrast, a second value (e.g., 0) for `vps_all_layers_same_num_sublayer_flag` specifies that the number of time sublayers can be the same or different for all layers in each CVS of the reference VPS. When `vps_all_layers_same_num_sublayer_flag` is not present, its value can be inferred as the first value (e.g., 1).
[0197] Additionally, a VPS can include vps_num_ptls_minus1, vps_pt_present_flag[i], vps_ptl_max_temporal_id[i], and vps_ols_ptl_idx[i] as pairs of syntax elements (profile, layer, and level) related to the PTL.
[0198] The value obtained by incrementing the syntax element `vps_num_ptls_minus1` by 1 specifies the number of `profile_tier_level()` syntax structures in the VPS. In one example, the value of `vps_num_ptls_minus1` can be constrained to be less than the value of the variable `TotalNumOls`. Here, `TotalNumOls` can specify the total number of OLS specified by the VPS.
[0199] The syntax element `vps_pt_present_flag[i]` specifies whether profile, layer, and general constraint information exists in the `i`th `profile_tier_level()` syntax structure of the VPS. For example, a first value (e.g., 1) of `vps_pt_present_flag[i]` specifies that profile, layer, and general constraint information exists in the `i`th `profile_tier_level()` syntax structure of the VPS. In contrast, a second value (e.g., 0) of `vps_pt_present_flag[i]` specifies that profile, layer, and general constraint information does not exist in the `i`th `profile_tier_level()` syntax structure of the VPS. When `vps_pt_present_flag[i]` has a second value (e.g., 0), the profile, layer, and general constraint information for the `i`th `profile_tier_level()` syntax structure of the VPS can be inferred to be the same as the profile, layer, and general constraint information for the `(i-1)`th `profile_tier_level()` syntax structure of the VPS. When vps_pt_present_flag[i] does not exist, the value of vps_pt_present_flag[i] can be inferred to be the first value (e.g., 1).
[0200] The syntax element `vps_ptl_max_temporal_id[i]` specifies the TemporalId of the highest sublayer representation in the `iprofile_tier_level()` syntax structure within the VPS where level information exists. In one example, the value of `vps_ptl_max_temporal_id[i]` can be constrained to the range of 0 to `vps_max_sublayer_minus1`. When `vps_ptl_max_temporal_id[i]` does not exist, its value can be inferred to be the same as the value of `vps_max_sublayer_minus1`.
[0201] The syntax element `vps_ols_ptl_idx[i]` specifies the index of the `profile_tier_level()` syntax structure applied to the `i`-th OLS in the list of `profile_tier_level()` syntax structures in the VPS. When `vps_ols_ptl_idx[i]` exists, its value can be constrained to the range of 0 to `vps_num_ptls_minus1`. When `vps_ols_ptl_idx[i]` does not exist, its value can be inferred using the following methods.
[0202] - When vps_num_ptls_minus1 equals 0, the value of vps_ols_ptl_idx[i] can be inferred to be equal to 0.
[0203] - In other cases (e.g., when vps_num_ptls_minus1 is greater than 0 and vps_num_ptls_minus1+1 equals TotalNumOlss), the value of vps_ols_ptl_idx[i] can be inferred to be equal to i.
[0204] When NumLayersInOls[i] equals 1, the profile_tier_level() syntax structure applied to the i-th OLS can also exist in the SPS referenced by the layer in the i-th OLS. When NumLayersInOls[i] equals 1, bitstream consistency can require that the profile_tier_level() signaled in the VPS and SPS should be the same for the i-th OLS.
[0205] Each profile_tier_level() syntax structure in a VPS can be constrained to be referenced by at least one value of vps_ols_ptl_idx[i]. Here, the value of i can be constrained to the range of 0 to TotalNumOlss-1.
[0206] In addition, PTL information can be communicated via signals in the SPS (Sequence Parameter Set).
[0207] Figure 10 This is a diagram illustrating an example of SPS syntax that includes PTL information.
[0208] Reference Figure 10 The SPS can include the syntax element `sps_seq_parameter_set_id` which specifies the identifier of the SPS, and the syntax element `sps_video_parameter_set_id` which specifies the identifier of the VPS referenced by the SPS. When the value of `sps_video_parameter_set_id` is greater than 0, `sps_video_parameter_set_id` can specify the value of `vps_video_parameter_set_id` of the VPS referenced by the SPS.
[0209] The presence of a VPS for a single-layer bitstream is optional. When the VPS does not exist, the value of sps_video_parameter_set_id can be inferred to be equal to 0. When the value of sps_video_parameter_set_id is equal to 0, the following can be applied.
[0210] - The SPS does not reference the VPS, and the VPS is not referenced when decoding each CVS that references the SPS.
[0211] The value of -vps_max_layers_minus1 is inferred to be equal to 0.
[0212] The value of -vps_max_sublayer_minus1 is inferred to be equal to 6.
[0213] - CVS (Encoded Video Sequence) is constrained to include only one layer (i.e., all VCL NAL units in CVS are constrained to have the same nuh_layer_id value).
[0214] The value of -GeneralLayerIdx[nuh_layer_id] is inferred to be equal to 0.
[0215] The value of -vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is inferred to be equal to 1.
[0216] Additionally, SPS can include the syntax element sps_max_sublayer_minus1, which relates to the number of sublayers.
[0217] The value obtained by incrementing the syntax element sps_max_sublayer_minus1 by 1 can specify the maximum number of temporal sublayers that can exist in each CLVS (coded layer video sequence) of the reference SPS.
[0218] When the value of sps_video_parameter_set_id is greater than 0, the value of sps_max_sublayer_minus1 can be constrained to the range of 0 to vps_max_sublayer_minus1.
[0219] In other cases (e.g., when the value of sps_video_parameter_set_id is equal to 0), the following can be applied.
[0220] The value of -sps_max_sublayer_minus1 can be constrained to a range of 0 to 6.
[0221] The value of -vps_max_sublayer_minus1 is inferred to be equal to sps_max_sublayer_minus1.
[0222] The value of -NumSubLayersInLayerInOLS can be inferred to be equal to sps_max_sublayer_minus1+1.
[0223] The value of -vps_ols_ptl_idx is inferred to be equal to 0, and the value of vps_ptl_max_tid (i.e., vps_ptl_max_tid) is inferred to be equal to sps_max_sublayer_minus1.
[0224] Additionally, the SPS can include a syntax element `sps_ptl_dpb_hrd_params_present_flag` that specifies whether a predefined syntax structure, including the `profile_tier_level()` syntax structure, exists in the SPS. For example, a first value (e.g., 1) of `sps_ptl_dpb_hrd_params_present_flag` can specify that the `profile_tier_level()` and `dpb_parameters()` syntax structures exist in the SPS, and that the `general_timing_hrd_parameters()` and `ols_timing_hrd_parameters()` syntax structures also exist in the SPS. In contrast, a second value (e.g., 0) of `sps_ptl_dpb_hrd_params_present_flag` can specify that all `profile_tier_level()`, `dpb_parameters()`, `general_timing_hrd_parameters()`, and `ols_timing_hrd_parameters()` syntax structures do not exist in the SPS.
[0225] In the example, when the OLS includes only one layer with a nuh_layer_id that is greater than 0 and has the same nuh_layer_id as the SPS, or when the value of sps_video_parameter_set_id is equal to 0, sps_ptl_dpb_hrd_params_present_flag can be constrained to have a first value (e.g., 1).
[0226] Furthermore, when `sps_ptl_dpb_hrd_params_present_flag` has a first value (e.g., 1), the `profile_tier_level(1,sps_max_sublayer_minus1)` syntax can be invoked in SPS. In this case, the first input value of the `profile_tier_level(1,sps_max_sublayer_minus1)` syntax can specify whether PTL information exists (e.g., `profileTierPresentFlag`). Additionally, the second input value of the `profile_tier_level(1,sps_max_sublayer_minus1)` syntax can specify the maximum number of time sublayers that can exist in each CVS (e.g., `maxNumSubLayersMinus1`).
[0227] Figure 11 This is a diagram illustrating an example of the profile_tier_level(profileTierPresentFlag, maxNumSubLayersMinus1) syntax that includes PTL information.
[0228] Reference Figure 11 The `profile_tier_level(profileTierPresentFlag, maxNumSubLayersMinus1)` syntax can include syntax elements `general_profile_idc`, `general_tier_flag`, and `general_level_idc` containing general information about the PTL (profile, tier, and level). Syntax elements can be signaled / resolved only when PTL information is present (i.e., `profileTierPresentFlag == 1`).
[0229] The syntax element `general_profile_idc` can specify profile information that OlsInScope adheres to, as specified in Appendix A of the VVC standard. The bitstream can be constrained to exclude values of `general_profile_idc` other than those specified in Appendix A above. Other values for `general_profile_idc` can be reserved for future use.
[0230] The syntax element general_tier_flag can specify layer context information used for parsing general_level_idc as specified in Appendix A of the VVC standard.
[0231] The syntax element `general_level_idc` can specify level information that OlsInScope adheres to, as specified in Appendix A of the VVC standard. The bitstream can be constrained to exclude values of `general_level_idc` other than those specified in Appendix A above. Other values for `general_level_idc` can be reserved for future use.
[0232] In one example, a larger value for `general_level_idc` can specify a higher level. The maximum level signaled by a signal in the DCI (Decoding Capability Information) NAL unit for OlsInScope can be higher than the level signaled by a signal in the SPS for CLVS in OlsInScope, but it can be no lower than it.
[0233] In one example, when OlsInScope conforms to multiple profiles, general_profile_idc can be constrained to specify a profile that provides preferred decoding results or preferred bitstream identification as determined by the encoder.
[0234] In one example, when the CVS of OlsInScope conforms to different profiles, since multiple profile_tier_level() syntax structures can be included in the DCI NAL unit, there can be at least one set of PTLs specified by a decoder capable of decoding the CVS for each CVS of OlsInScope.
[0235] Additionally, the `profile_tier_level(profileTierPresentFlag, maxNumSubLayersMinus1)` syntax can include syntax elements `ptl_num_sub_profiles` and `general_sub_profile_idc[i]` regarding subprofiles. Syntax elements can be signaled / resolved only if PTL information exists (i.e., `profileTierPresentFlag == 1`).
[0236] The syntax element ptl_num_sub_profiles can specify the number of general_sub_profile_idc[i] present in the profile_tier_level(profileTierPresentFlag,maxNumSubLayersMinus1) syntax.
[0237] The syntax element general_sub_profile_idx[i] can specify an indicator for the i-th interoperability metadata.
[0238] Additionally, the profile_tier_level(profileTierPresentFlag,maxNumSubLayersMinus1) syntax can include syntax elements ptl_sublayer_level_present_flag[i] and sublayer_level_idc[i] related to the sublayer level.
[0239] The syntax element `ptl_sublayer_level_present_flag[i]` can specify whether sublayer level information exists in the `profile_tier_level()` syntax structure representing a sublayer with a `TemporalId` equal to `i`. For example, a first value (e.g., 1) of `ptl_sublayer_level_present_flag[i]` can specify that the sublayer level information of the `i`-th time sublayer exists in the `profile_tier_level()` syntax structure. In contrast, a second value (e.g., 0) of `ptl_sublayer_level_present_flag[i]` can specify that the sublayer level information of the `i`-th time sublayer does not exist in the `profile_tier_level()` syntax structure. `ptl_sublayer_level_present_flag[i]` can be signaled / resolved in ascending order of the value of `i`, where `i` ranges from 0 to `maxNumSubLayersMinus1-1`.
[0240] The syntax element `sublayer_level_idc[i]` can specify the sublayer level index of the i-th time sublayer. `sublayer_level_idc[i]` can be signaled / resolved in ascending order of `i`, where `i` ranges from 0 to `maxNumSubLayersMinus1-1`. Additionally, `sublayer_level_idc[i]` can be signaled / resolved only if `ptl_sublayer_level_present_flag[i]` has a first value (e.g., 1).
[0241] When sublayer_level_idc[i] does not exist, the value of sublayer_level_idc[i] can be inferred as follows.
[0242] -sublayer_level_idc[maxNumSubLayersMinus1] can be inferred to be the same value as general_level_idc in the same profile_tier_level() syntax structure.
[0243] - For i from maxNumSubLayersMinus1-1 to 0 (i.e., in descending order of i values), sublayer_level_idc[i] is inferred to be the same value as sublayer_level_idc[i+1].
[0244] As described above, the presence of sublayer_level_idc[i] can be determined based on the value of ptl_sublayer_level_present_flag[i]. Specifically, when ptl_sublayer_level_present_flag[i] has a first value (e.g., 1), sublayer_level_idc[i] can be signaled / resolved in ascending order of TemporalId values (i.e., i). In contrast, when ptl_sublayer_level_present_flag[i] has a second value (e.g., 0), sublayer_level_idc[i] can be signaled / resolved without signaling.
[0245] Furthermore, when `sublayer_level_idc[i]` is not present, its value can be inferred to be the same as the sublayer level index of the higher sublayer (i.e., `sublayer_level_idc[i+1]`). In other words, the signaling / resolution of `sublayer_level_idc[i]` is performed in ascending order of the value of `i`, while the inference process (or inference rule) for `sublayer_level_idc[i]` can be performed in descending order of the value of `i`. Therefore, when `sublayer_level_idc[i+1]` is not signaled / resolved, the value of `sublayer_level_idc[i]` (which does not exist in the `profile_tier_level()` syntax structure) does not need to be inferred. Additionally, the inference process does not need to be performed in the same loop as the signaling / resolution process for `sublayer_level_idc[i]`.
[0246] To address these issues, according to embodiments of this disclosure, sub-level information can be signaled / parsed in descending order from higher to lower time sub-layers.
[0247] Figure 12 This is a diagram illustrating the syntax of profile_tier_level(profileTierPresentFlag,maxNumSubLayersMinus1) according to an embodiment of this disclosure.
[0248] Reference Figure 12The `profile_tier_level(profileTierPresentFlag, maxNumSubLayersMinus1)` syntax can include syntax elements `ptl_sublayer_level_present_flag[i]` and `sublayer_level_idc[i]` regarding sublayer level information. In one implementation, sublayer level information can be signaled / resolved from higher to lower time sublayers in descending order.
[0249] A time sublayer (or sublayer) can refer to a time-scalable layer of a time-scalable bitstream consisting of VCL NAL units with a specific value of TemporalId and associated non-VCL NAL units.
[0250] For inter-frame prediction, a (decoded) frame with a TemporalId value less than or equal to the current frame's TemporalId can be used as a reference frame.
[0251] In one example, TemporalId can be derived by subtracting 1 from the value of nuh_temporal_id_plus1, which is notified by the NAL unit signaling via the NAL unit (i.e., TemporalId = nuh_temporal_id_plus1 - 1).
[0252] Additionally, when nal_unit_type (in the NAL unit header) is in the range IDR_W_RADL to RSV_IRAP_12, TemporalId can be constrained to 0. In contrast, when nal_unit_type equals STSA_NUT and vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] has a first value (e.g., 1), TemporalId can be constrained to non-zero. Here, vps_independent_layer_flag[i] can specify whether the layer with index i can use inter-layer prediction. For example, a vps_independent_layer_flag[i] with a first value (e.g., 1) can specify that the layer with index i can not use inter-layer prediction. In contrast, a second value (e.g., 0) of `vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]]` can specify that the layer with index `i` can use inter-layer prediction and the syntax element `vps_direct_ref_layer_flag[i][j]` (where 0 ≤ j ≤ i-1) exists in the VPS syntax. `vps_independent_layer_flag[i]` can be included in the reference... Figure 9 In the described VPS syntax, and when vps_independent_layer_flag[i] does not exist in the VPS syntax, the value of vps_independent_layer_flag[i] can be inferred to be the first value (e.g., 1).
[0253] Additionally, the TemporalId value can be constrained to be the same for all VCL NAL units of an Access Unit (AU). The TemporalId value of an encoded screen, PU (screen unit), or AU can be equal to the TemporalId value of the VCL NAL units of the encoded screen, PU, or AU. The TemporalId value of a sublayer representation can be equal to the maximum value of the TemporalId of all VCL NAL units in the sublayer representation.
[0254] Additionally, the TemporalId value of non-VCL NAL units can be constrained as follows.
[0255] - When nal_unit_type is equal to DCI_NUT, VPS_NUT or SPS_NUT, the value of TemporalId can be constrained to 0, and the value of TemporalId of AU including NAL units is also constrained to 0.
[0256] - In contrast, when nal_unit_type is equal to PH_NUT, the value of TemporalId is constrained to be equal to the value of TemporalId of the PU including the NAL unit.
[0257] - In contrast, when nal_unit_type is equal to EOS_NUT or EOB_NUT, the value of TemporalId is constrained to 0.
[0258] - In contrast, when nal_unit_type is equal to AUD_NUT, FD_NUT, PREFIX_SEI_NUT or SUFFIX_SEI_NUT, the value of TemporalId is constrained to be equal to the value of TemporalId of the AU including the NAL unit.
[0259] - In contrast, when nal_unit_type is equal to PPS_NUT, PREFIX_APS_NUT, or SUFFIX_APS_NUT, the value of TemporalId is constrained to be greater than or equal to the value of TemporalId of the PU including the NAL unit.
[0260] In one example, when the NAL unit is a non-VCL NAL unit, the TemporalId value can be equal to the minimum TemporalId value of all AUs that apply the non-VCL NAL unit. When nal_unit_type is equal to PPS_NUT, PREFIX_APS_NUT, or SUFFIX_APS_NUT, the TemporalId value can be greater than or equal to the TemporalId value of the AUs that include the NAL unit. This is because all PPS and APS can be included at the beginning of the bitstream (e.g., this information is transmitted out of band and the receiver places it at the beginning of the bitstream). Here, the first encoded frame can have a TemporalId value equal to 0.
[0261] The syntax element ptl_sublayer_level_present_flag[i] can specify whether sublayer level information exists in the profile_tier_level() syntax structure of the sublayer representation (i.e., the i-th time sublayer) with a TemporalId equal to i (e.g., "1": present and "0": not present).
[0262] In one implementation, ptl_sublayer_level_present_flag[i] can be signaled / resolved in descending order of the TemporalId value (i.e., i). More specifically, ptl_sublayer_level_present_flag[i] can be signaled / resolved in the order from the second highest temporal sublayer (i.e., i = maxNumSubLayersMinus1-1) to the first temporal sublayer (i.e., i = 0).
[0263] The syntax element sublayer_level_idc[i] can specify the sublayer level index of the i-th time sublayer.
[0264] The semantics of the syntax element sublayer_level_idc[i] are equal to general_level_idc except for the inference process (or inference rule) when the value does not exist, but it applies to sublayer representations with a TemporalId equal to i.
[0265] In one implementation, `sublayer_level_idc[i]` can be signaled / resolved in descending order of the values of `TemporalId` (i.e., `i`). More specifically, `sublayer_level_idc[i]` can be signaled in the order of the second highest temporal sublayer (i.e., `i = maxNumSubLayersMinus1-1`) to the first temporal sublayer (i.e., `i = 0`). Additionally, `sublayer_level_idc[i]` can be signaled / resolved only if `ptl_sublayer_level_present_flag[i]` has a first value (e.g., 1). The value of `sublayer_level_idc[maxNumSubLayersMinus1]` can be inferred to be equal to the value of `general_level_idc` within the same `profile_tier_level()` syntax structure.
[0266] When sublayer_level_idc[i] does not exist, the value of sublayer_level_idc[i] (where 0≤i≤maxNumSubLayersMinus1-1) can be inferred to be the same as the value of sublayer_level_idc[i+1].
[0267] Furthermore, during performance evaluation simulations, to compare layer capabilities, a layer with a second value (e.g., 0) of `general_tier_flag` (i.e., the main layer) can be considered a lower layer than a layer with a first value (e.g., 1) of `general_tier_flag` (i.e., a higher layer). Additionally, to compare layer capabilities, when the value of `general_level_idc` or `sublayer_level_idc[i]` at a specific level is less than the value at another level, that specific level of a predetermined layer can be considered a lower level than that of the previous layer.
[0268] As described above, according to the embodiments of this disclosure, the signal notification / resolution of ptl_sublayer_level_present_flag[i] can be performed in descending order of the value of TemporalId (i.e., i). Similarly, the signal notification / resolution and inference process (or inference rule) of sublayer_level_idc[i] can also be performed in descending order of the value of i. That is, sublayer_level_idc[i+1] can be signal-notified / resolution before sublayer_level_idc[i]. Therefore, the inference process of the value of sublayer_level_idc[i] can be performed in the same loop as the signal notification / resolution process of sublayer_level_idc[i].
[0269] In the following text, reference will be made to Figure 13 and Figure 14 A detailed description of an image encoding / decoding method according to embodiments of the present disclosure is provided.
[0270] Figure 13 This is a flowchart of an image encoding method according to one embodiment of the present disclosure.
[0271] Figure 13 Image encoding methods can be derived from Figure 2 or Figure 5 The image encoding device performs these steps. For example, steps S1310 and S1320 can be performed by entropy encoders 190, 540-0, or 540-1.
[0272] Reference Figure 13 The image encoding device can encode a first flag specifying the presence or absence of sub-layer level information for each of one or more sub-layers in the current layer (S1310). The first flag may refer to the above-mentioned... Figure 12The description refers to `ptl_sublayer_level_present_flag[i]` in the syntax `profile_tier_level(profileTierPresentFlag, maxNumSubLayersMinus1)`. Additionally, sublayer level information can refer to the above-mentioned... Figure 12 The sublayer_level_idc[i] in the described profile_tier_level(profileTierPresentFlag, maxNumSubLayersMinus1) syntax.
[0273] In one implementation, the first flag (e.g., ptl_sublayer_level_present_flag[i]) can be encoded in descending order of the time identifier values of one or more sublayers. The time identifier value can be derived by subtracting 1 from the value of nuh_temporal_id_plus1 signaled via the NAL unit header (i.e., TemporalId = nuh_temporal_id_plus1 - 1).
[0274] In one implementation, the first flag can be encoded based on the maximum number of sublayers in the current layer (e.g., maxNumSubLayersMinus1). For example, as referenced above... Figure 12 The ptl_sublayer_level_present_flag[i] can be encoded / signed from the second highest time sublayer (i.e., i = maxNumSubLayersMinus1-1) to the first time sublayer (i.e., i = 0).
[0275] The image encoding device can encode sub-level information based on the first flag (S1320). For example, as described above... Figure 12 The sublayer_level_idc[i] can be encoded / signaled only if ptl_sublayer_level_present_flag[i] has a first value (e.g., 1) indicating the presence of sublayer level information.
[0276] In one implementation, sublayer-level information (e.g., sublayer_level_idc[i]) can be encoded in descending order of time identifier values.
[0277] In one implementation, sub-layer level information can be encoded based on the maximum number of sub-layers in the current layer. For example, as described above... Figure 12The sublayer_level_idc[i] can be encoded / signed in the order from the second highest time sublayer (i.e., i = maxNumSubLayersMinus1-1) to the first time sublayer (i.e., i = 0).
[0278] Figure 14 This is a flowchart of an image decoding method according to one embodiment of the present disclosure.
[0279] Figure 14 Image decoding methods can be derived from Figure 3 or Figure 6 The image decoding device performs the steps. For example, steps S1410 to S1420 can be performed by entropy decoder 210, 610-0, or 610-1.
[0280] Reference Figure 14 The image decoding device can obtain (or decode) a first flag (S1410) from the bitstream that specifies whether sub-layer level information exists for each of one or more sub-layers in the current layer.
[0281] In one implementation, the first flag (e.g., ptl_sublayer_level_present_flag[i]) can be obtained in descending order of the time identifier values of one or more sublayers. The time identifier value can be derived by subtracting 1 from the value of nuh_temporal_id_plus1 signaled via the NAL unit header (i.e., TemporalId = nuh_temporal_id_plus1 - 1).
[0282] In one implementation, the first flag can be obtained based on the maximum number of sublayers in the current layer (e.g., maxNumSubLayersMinus1). For example, as referenced above. Figure 12 The ptl_sublayer_level_present_flag[i] can be decoded / signed from the second highest time sublayer (i.e., i = maxNumSubLayersMinus1-1) to the first time sublayer (i.e., i = 0).
[0283] The image decoding device can obtain (or decode) sub-layer level information from the bitstream based on the first flag (S1420). For example, as described above. Figure 12 As stated above, sublayer_level_idc[i] can only be decoded / parsed if ptl_sublayer_level_present_flag[i] has a first value (e.g., 1) indicating the presence of sublayer level information.
[0284] In one implementation, sublayer-level information (e.g., sublayer_level_idc[i]) can be obtained in descending order of time identifier values.
[0285] In one implementation, sub-layer level information can be obtained based on the maximum number of sub-layers in the current layer. For example, as described above... Figure 12 The sublayer_level_idc[i] can be decoded / parsed in the order from the second highest time sublayer (i.e., i = maxNumSubLayersMinus1-1) to the first time sublayer (i.e., i = 0).
[0286] In one implementation, the step of obtaining the first sublayer level information can be skipped based on a first flag (e.g., ptl_sublayer_level_present_flag[i]) indicating that the first sublayer level information (e.g., sublayer_level_idc[i]) in the current layer does not exist. In this case, the first sublayer level information can be set (or inferred) to the same value as the second sublayer level information (e.g., sublayer_level_idc[i+1]) of the second sublayer, which has a second time identifier value that is 1 greater than the first time identifier value of the first sublayer.
[0287] In one implementation, the third sub-layer level information of the third sub-layer with the largest time identifier value among the one or more sub-layers can be set (or inferred) to the same value as the general level index information (e.g., general_level_idc) preset for one or more OLS. Here, the general level index information can specify the level that one or more OLS conforms to, as specified by the Video Parameter Set (VPS). In one example, a larger value for general_level_idc can specify a higher level.
[0288] As described above, based on the reference above... Figure 13 and Figure 14 The image encoding / decoding method according to embodiments of this disclosure describes a first flag (e.g., ptl_sublayer_level_present_flag[i]) indicating the presence of sublayer-level information for each of one or more sublayers in the current layer, obtained in descending order of the temporal identifier values of one or more sublayers. Similarly, sublayer-level information (e.g., sublayer_level_idc[i]) can also be obtained in descending order of the temporal identifier values. Therefore, the inference process for sublayer-level information can be performed in the same loop as the signal notification / resolution process for sublayer-level information.
[0289] Although the exemplary methods of this disclosure described above are represented as a series of operations for clarity of description, they are not intended to limit the order in which the steps are performed, and these steps may be performed simultaneously or in different orders if necessary. To implement the methods according to this disclosure, the described steps may further include other steps, including steps in addition to some steps, or may include additional steps in addition to some steps.
[0290] In this disclosure, the image encoding device or image decoding device that performs a predetermined operation (step) can perform an operation (step) that confirms the execution conditions or circumstances of the corresponding operation (step). For example, if it is described that a predetermined operation is performed when predetermined conditions are met, the image encoding device or image decoding device can perform the predetermined operation after determining whether the predetermined conditions are met.
[0291] The various embodiments of this disclosure are not a list of all possible combinations and are intended to describe representative aspects of this disclosure; the matters described in the various embodiments may be applied independently or in combination of two or more.
[0292] Various embodiments of this disclosure can be implemented in hardware, firmware, software, or a combination thereof. When this disclosure is implemented in hardware, it can be implemented using application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, etc.
[0293] Furthermore, the image decoding and image encoding devices applying the embodiments of this disclosure can be included in multimedia broadcasting transmission and receiving devices, mobile communication terminals, home theater video devices, digital cinema video devices, surveillance cameras, video chat devices, real-time communication devices such as video communication, mobile streaming devices, storage media, cameras, video-on-demand (VoD) service providers, OTT (over-the-top) video devices, internet streaming service providers, three-dimensional (3D) video devices, video telephony devices, medical video devices, etc., and can be used to process video signals or data signals. For example, OTT video devices can include game consoles, Blu-ray players, internet access televisions, home theater systems, smartphones, tablet PCs, digital video recorders (DVRs), etc.
[0294] Figure 15 This is a diagram illustrating a content flow system to which embodiments of the present disclosure can be applied.
[0295] like Figure 15As shown, the content streaming system applying the embodiments of this disclosure may mainly include an encoding server, a streaming server, a network server, a media storage device, a user device, and a multimedia input device.
[0296] The encoding server compresses content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream and then sends the bitstream to the streaming server. As another example, when multimedia input devices such as smartphones, cameras, and camcorders directly generate bitstreams, the encoding server can be omitted.
[0297] The bitstream can be generated by an image encoding method or image encoding device applying the embodiments of this disclosure, and the stream server can temporarily store the bitstream during the sending or receiving of the bitstream.
[0298] A streaming server sends multimedia data to a user's device based on a request from a web server, and the web server acts as a medium for informing the user of the service. When a user requests a service from the web server, the web server can deliver it to the streaming server, and the streaming server can send the multimedia data to the user. In this scenario, the content streaming system may include a separate control server. In this case, the control server is used to control the commands / responses between devices in the content streaming system.
[0299] A streaming server can receive content from media storage devices and / or encoding servers. For example, when receiving content from an encoding server, the content can be received in real time. In this case, to provide a smooth streaming service, the streaming server can store the bitstream for a predetermined period of time.
[0300] Examples of user devices may include mobile phones, smartphones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, board PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head-mounted displays), digital televisions, desktop computers, digital signage, etc.
[0301] In a content streaming system, each server can operate as a distributed server, in which case the data received from each server can be distributed.
[0302] The scope of this disclosure includes software or machine-executable commands (e.g., operating systems, applications, firmware, programs, etc.) for enabling the operation of methods according to various embodiments to be executed on a device or computer, and non-transitory computer-readable media having such software or commands stored thereon and executable on a device or computer.
[0303] Industrial applicability
[0304] The embodiments disclosed herein can be used to encode or decode images.
Claims
1. An image decoding method performed by an image decoding device, the image decoding method comprising the following steps: Obtain a first flag from the bitstream specifying whether sub-layer level information exists for each of one or more sub-layers in the current layer; as well as The sub-layer level information is obtained from the bitstream based on the value of the first flag. The first flag is obtained in descending order of the time identifier values of the one or more sub-layers.
2. The image decoding method according to claim 1, wherein, The sub-level information is obtained in descending order of the time identifier values.
3. The image decoding method according to claim 1, wherein, The first flag is obtained based on the maximum number of sub-layers in the current layer, and Furthermore, the sub-layer level information is obtained based on the maximum number of sub-layers in the current layer.
4. The image decoding method according to claim 1, wherein, Based on the first flag, the first sub-layer level information of the first sub-layer in the current layer does not exist. Skip the step of obtaining the first sub-layer level information, and set the first sub-layer level information to the same value as the second sub-layer level information of the second sub-layer that has a second time identifier value that is 1 greater than the first time identifier value of the first sub-layer.
5. The image decoding method according to claim 1, wherein, The third sub-layer level information of the third sub-layer with the largest time identifier value among the one or more sub-layers is set to the same value as the general level index information preset for the one or more output layer sets.
6. An image encoding method performed by an image encoding device, the image encoding method comprising the following steps: Encode a first flag specifying whether sub-layer level information exists for each of one or more sub-layers in the current layer; as well as The sub-level information is encoded based on the value of the first flag. The first flag is encoded in descending order of the time identifier values of the one or more sub-layers.
7. The image encoding method according to claim 6, wherein, The sub-level information is encoded in descending order of the time identifier values.
8. The image encoding method according to claim 6, wherein, The first flag is encoded based on the maximum number of sub-layers in the current layer, and Furthermore, the sub-layer level information is encoded based on the maximum number of sub-layers in the current layer.
9. A method for transmitting a bit stream, the method comprising the following steps: Encode a first flag specifying the presence or absence of sub-layer level information for each of one or more sub-layers in the current layer into the bitstream; The sub-layer level information is encoded into a bitstream based on the value of the first flag; as well as Send the bit stream, The first flag is encoded in descending order of the time identifier values of the one or more sub-layers.
Citation Information
Patent Citations
Method and apparatus for encoding multi layer video and method and apparatus for decoding multilayer video
US20160065983A1