Image encoding / decoding method and apparatus based on output layer set determining whether to refer to parameter set, and method of transmitting bitstream

CN115668927BActive Publication Date: 2026-09-04LG ELECTRONICS INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202180036136.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-04-06
Filing Date
2021-04-01
Publication Date
2026-09-04
Estimated Expiration
2041-04-01

AI Technical Summary

Technical Problem

传输信息量或比特量的增加导致传输成本和存储成本的增加

Benefits of technology

[0020] According to this disclosure, an image encoding/decoding method and apparatus with improved encoding/decoding efficiency can be provided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115668927B_ABST
    Figure CN115668927B_ABST
Patent Text Reader

Abstract

Image encoding / decoding methods and apparatuses are provided. An image decoding method performed by an image decoding apparatus according to the present disclosure can include the steps of determining reference availability of a parameter set used for decoding encoded image data, and decoding the encoded image data based on the reference availability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to image encoding / decoding methods and apparatus, and more specifically, to image encoding and decoding methods and apparatus for determining whether to reference a parameter set, and to a method for transmitting a bitstream generated by the image encoding method / apparatus of this disclosure. Background Technology

[0002] Recently, the demand for high-resolution and high-quality images, such as high-definition (HD) and ultra-high-definition (UHD) images, has been increasing across various fields. With the improvement in the resolution and quality of image data, the amount of information or bits transmitted increases relatively compared to existing image data. This increase in the amount of information or bits transmitted leads to increased transmission and storage costs.

[0003] Therefore, efficient image compression techniques are needed to effectively send, store, and reproduce information about high-resolution and high-quality images. Summary of the Invention

[0004] Technical issues

[0005] The purpose of this disclosure is to provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.

[0006] Another object of this disclosure is to provide an image encoding / decoding method and apparatus that improves encoding / decoding efficiency by determining whether to reference a parameter set.

[0007] Another object of this disclosure is to provide a method for transmitting a bitstream generated by an image encoding method or device according to this disclosure.

[0008] Another object of this disclosure is to provide a recording medium for storing bitstreams generated by an image encoding method or device according to this disclosure.

[0009] Another object of this disclosure is to provide a recording medium that stores a bitstream received, decoded, and used to reconstruct an image by an image decoding apparatus according to this disclosure. For example, a bitstream used to enable the image decoding apparatus according to this disclosure to perform an image decoding method according to this disclosure may be stored in the recording medium.

[0010] The technical problems solved by this disclosure are not limited to those described above, and other technical problems not described herein will become apparent to those skilled in the art based on the following description.

[0011] Technical solution

[0012] An image decoding method performed by an image decoding device according to one aspect of this disclosure may include the following steps: obtaining a parameter set and image encoded data from a bitstream; determining a reference availability of the parameter set for decoding the image encoded data; and decoding the image encoded data based on the reference availability. The reference availability may be determined based on whether an output layer set that includes a layer corresponding to the image encoded data includes a predetermined layer.

[0013] Additionally, an image decoding apparatus according to one aspect of this disclosure may include a memory and at least one processor. The at least one processor may be configured to obtain a parameter set and image encoded data from a bitstream, determine a reference availability of the parameter set for decoding the image encoded data, and decode the image encoded data based on the reference availability. The reference availability may be determined based on whether an output layer set that includes a layer corresponding to the image encoded data includes a predetermined layer.

[0014] Additionally, an image encoding method performed by an image encoding device according to one aspect of this disclosure may include the following steps: encoding an image to generate image-encoded data of a portion of the image and a parameter set for the image-encoded data; and generating a bitstream including the image-encoded data and the parameter set. The parameter set may be generated based on a reference availability of the parameter set used for decoding the image-encoded data. The reference availability may be determined based on whether an output layer set that includes a layer corresponding to the image-encoded data includes a predetermined layer.

[0015] Furthermore, according to another aspect of the transmission method of this disclosure, a bit stream generated by an image encoding device or method according to this disclosure can be transmitted.

[0016] Furthermore, according to another aspect of this disclosure, a computer-readable recording medium can store a bitstream generated by an image encoding method or apparatus according to this disclosure.

[0017] Furthermore, according to another aspect of this disclosure, a computer-readable recording medium may store a bitstream for enabling a decoding device to perform an image decoding method according to this disclosure.

[0018] The features briefly outlined above are merely exemplary aspects of the detailed description of this disclosure below and do not limit the scope of this disclosure.

[0019] Beneficial effects

[0020] According to this disclosure, an image encoding / decoding method and apparatus with improved encoding / decoding efficiency can be provided.

[0021] Furthermore, according to this disclosure, an image encoding / decoding method and apparatus can be provided to improve encoding / decoding efficiency by determining whether to reference a parameter set.

[0022] Furthermore, according to this disclosure, a method for transmitting a bitstream generated by an image encoding method or device according to this disclosure can be provided.

[0023] Furthermore, according to this disclosure, a recording medium may be provided for storing a bitstream generated by an image encoding method or apparatus according to this disclosure.

[0024] Furthermore, according to this disclosure, a recording medium can be provided for storing a bitstream that is received, decoded, and used to reconstruct an image by an image decoding device according to this disclosure.

[0025] Those skilled in the art will understand that the effects achievable through this disclosure are not limited to those specifically described above, and that other advantages of this disclosure will become clearer from the specific embodiments. Attached Figure Description

[0026] Figure 1 This is a view schematically illustrating a video encoding system to which embodiments of this disclosure are applicable.

[0027] Figure 2 This is a schematic view illustrating an image encoding device to which embodiments of the present disclosure are applicable.

[0028] Figure 3 This is a schematic view illustrating an image decoding device to which embodiments of the present disclosure are applicable.

[0029] Figure 4 and Figure 5 This is a view illustrating an example of the screen decoding and encoding process according to an implementation method.

[0030] Figure 6 This is a view showing the layer structure of an image for encoding according to an embodiment.

[0031] Figures 7 to 10 This is an example of a view based on multi-layered encoding and decoding.

[0032] Figures 11 to 12 This is a view illustrating a VPS NAL unit according to an implementation method.

[0033] Figure 13 This is a view illustrating an SPS NAL unit according to an embodiment.

[0034] Figure 14 This is a view illustrating a PPS NAL unit according to an embodiment.

[0035] Figure 15 This is a view illustrating an APS NAL unit according to an embodiment.

[0036] Figure 16 This is an example Figure 12 A view of another implementation of the pseudocode.

[0037] Figures 17 to 18 This is a view illustrating a decoding method according to an embodiment.

[0038] Figure 19 This is a view illustrating an encoding method according to an implementation method.

[0039] Figure 20 This is a view illustrating the content streaming system to which embodiments of this disclosure are applicable. Detailed Implementation

[0040] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings to facilitate implementation by those skilled in the art. However, this disclosure may be implemented in various different forms and is not limited to the embodiments described herein.

[0041] In describing this disclosure, detailed descriptions of relevant known functions or constructions will be omitted if they unnecessarily obscure the scope of this disclosure. In the accompanying drawings, portions irrelevant to the description of this disclosure are omitted, and similar reference numerals are assigned to similar portions.

[0042] In this disclosure, when a component is "connected," "linked," or "coupled" to another component, it may include not only a direct connection but also an indirect connection with intermediate components. Furthermore, when a component "comprises" or "has" other components, unless otherwise stated, it means that other components may be included, not excluded.

[0043] In this disclosure, the terms first, second, etc., may be used only for the purpose of distinguishing one component from other components and do not limit the order or importance of the components, unless otherwise stated. Accordingly, within the scope of this disclosure, a first component in one embodiment may be referred to as a second component in another embodiment, and similarly, a second component in one embodiment may be referred to as a first component in another embodiment.

[0044] In this disclosure, the distinguishing components are intended to clearly describe each feature and do not imply that the components must be separate. That is, multiple components may be integrated and implemented in a single hardware or software unit, or a single component may be distributed and implemented in multiple hardware or software units. Therefore, unless otherwise stated, these implementations, where components are integrated or distributed, are included within the scope of this disclosure.

[0045] In this disclosure, the components described in the various embodiments are not necessarily essential components, and some components may be optional. Therefore, embodiments consisting of a subset of the components described in the embodiments are also included within the scope of this disclosure. In addition, embodiments that include other components besides those described in the various embodiments are also included within the scope of this disclosure.

[0046] This disclosure relates to the encoding and decoding of images. Unless redefined in this disclosure, the terms used herein may have the general meaning commonly used in the art to which this disclosure pertains.

[0047] The methods / implementations disclosed in this disclosure are applicable to methods disclosed in the Universal Video Coding (VVC) standard. Additionally, the methods / implementations disclosed in this disclosure are applicable to methods disclosed in the Basic Video Coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the second-generation Audio Video Coding (AVS2) standard, or next-generation video / image coding standards (e.g., H.267 or H.268).

[0048] This disclosure provides various implementations of video / image encoding, and implementations not described herein can be performed in combination.

[0049] In this disclosure, "video" can mean a set of images over time. A "picture" generally refers to a unit representing an image at a specific time, and a tile is a coding unit that constitutes part of a picture in the encoding process. A tile may include one or more coding tree units (CTUs). A CTU may be divided into one or more CUs.

[0050] A frame can consist of one or more slices / tiles. A tile is a rectangular area within a specific tile row and a specific tile column in a frame, and can consist of multiple CTUs. A tile column can be defined as a rectangular area of ​​a CTU, and can have a height equal to the height of the frame and a width specified by a syntax element signaled from a bitstream portion such as a frame parameter set. A tile row can be defined as a rectangular area of ​​a CTU, and can have a width equal to the width of the frame and a height specified by a syntax element signaled from a bitstream portion such as a frame parameter set.

[0051] Tiled scanning is a specific ordering of CTUs that divide a frame. Here, CTUs are ordered consecutively within a tile using CTU raster scans, while the tiles in a frame are ordered consecutively using the raster scans of the frame's tiles. A slice consists of an integer number of consecutive complete CTU rows or an integer number of complete tiles within a tile of the frame. Slices can be proprietaryly included in a single NAL unit.

[0052] A frame can be divided into two or more sub-frames. A sub-frame can be a rectangular area of ​​one or more slices of the frame.

[0053] A frame can include one or more tile groups. A tile group can include one or more tiles. A tile can represent a rectangular area of ​​CTU rows within a tile in the frame. A tile can include one or more tiles. A tile can represent a rectangular area of ​​CTU rows within a tile. A tile can be divided into multiple tiles, and each tile can include one or more CTU rows belonging to the tile. Tiles that are not divided into multiple tiles can also be considered as tiles.

[0054] A “pixel” or “pixel” can refer to the smallest unit that makes up a picture (or image). Additionally, “sample” can be used as the term corresponding to a pixel. A sample can typically represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component.

[0055] In this disclosure, "unit" can refer to a basic unit of image processing. A unit may include a specific region of an image and at least one of the information associated with that region. A unit may include a luminance block and two chrominance blocks (e.g., Cb and Cr). In some cases, "unit" may be used interchangeably with terms such as "sample array," "block," or "region." In general, an M×N block may include a set (or array) of samples (or transform coefficients) with M columns and N rows.

[0056] In this disclosure, "current block" can mean one of "current coding block," "current coding unit," "coding target block," "decoding target block," or "processing target block." When performing prediction, "current block" can mean "current prediction block" or "prediction target block." When performing transform (inverse transform) / quantization (dequantization), "current block" can mean "current transform block" or "transform target block." When performing filtering, "current block" can mean "filter target block."

[0057] Additionally, in this disclosure, unless explicitly stated as a chroma block, "current block" may mean "the luminance block of the current block". "The chroma block of the current block" can be expressed by including an explicit description of a chroma block such as "chroma block" or "current chroma block".

[0058] In this disclosure, the forward slash " / " or "," should be interpreted as indicating "and / or". For example, the expressions "A / B" and "A, B" can mean "A and / or B". Furthermore, "A / B / C" and "A / B / C" can mean "at least one of A, B and / or C".

[0059] In this disclosure, the term "or" should be interpreted as indicating "and / or". For example, expressing "A or B" can include 1) only "A", 2) only "B", and / or 3) both "A and B". In other words, in this disclosure, the term "or" should be interpreted as indicating "additionally or alternatively".

[0060] In this disclosure, "at least one of A and B" may mean "only A", "only B" or "both A and B". Furthermore, in this disclosure, "at least one of A or B" or "at least one of A and / or B" may be interpreted as the same as "at least one of A and B".

[0061] Furthermore, in this disclosure, "at least one of A, B, and C" may mean "only A," "only B," "only C," or "any combination of A, B, and C." Additionally, in this disclosure, "at least one of A, B, or C" or "at least one of A, B, and / or C" may be interpreted as the same as "at least one of A, B, and C."

[0062] Furthermore, the parentheses used in this disclosure may mean "for example". Specifically, when describing "prediction (intra-frame prediction)", "intra-frame prediction" may be cited as an example of "prediction". In other words, the "prediction" in this disclosure is not limited to "intra-frame prediction", and "intra-frame prediction" may be cited as an example of "prediction". In addition, even when describing "prediction (i.e., intra-frame prediction)", "intra-frame prediction" may be cited as an example of "prediction".

[0063] In this disclosure, a technical feature described individually in a single figure may be implemented individually or simultaneously.

[0064] Overview of Video Encoding Systems

[0065] Figure 1 This is a view showing a video encoding system according to this disclosure.

[0066] The video encoding system according to the embodiment may include a source device 10 and a receiving device 20. The source device 10 may deliver encoded video and / or image information or data to the receiving device 20 in the form of a file or stream via a digital storage medium or network.

[0067] The source device 10 according to an embodiment may include a video source generator 11, an encoding device 12, and a transmitter 13. The receiving device 20 according to an embodiment may include a receiver 21, a decoding device 22, and a renderer 23. The encoding device 12 may be referred to as a video / image encoding device, and the decoding device 22 may be referred to as a video / image decoding device. The transmitter 13 may be included in the encoding device 12. The receiver 21 may be included in the decoding device 22. The renderer 23 may include a display, and the display may be configured as a separate device or an external component.

[0068] The video source generator 11 can acquire video / images through a process of capturing, compositing, or generating video / images. The video source generator 11 may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, a video / image archive including previously captured video / images, etc. The video / image generation device may include, for example, a computer, tablet computer, and smartphone, and can generate video / images (electronically). For example, virtual video / images can be generated by a computer, etc. In this case, the video / image capture process can be replaced by a process of generating related data.

[0069] The encoding device 12 can encode the input video / image. For compression and encoding efficiency, the encoding device 12 can perform a series of processes such as prediction, transformation, and quantization. The encoding device 12 can output the encoded data (encoded video / image information) in the form of a bitstream.

[0070] Transmitter 13 can transmit encoded video / image information or data, output in bitstream form, to receiver 21 of receiving device 20 in the form of a file or stream via digital storage medium or network. Digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. Transmitter 13 can include elements for generating media files according to a predetermined file format and may include elements for transmission via broadcast / communication networks. Receiver 21 can extract / receive bitstreams from storage medium or network and transmit the bitstreams to decoding device 22.

[0071] The decoding device 22 can decode video / images by performing a series of processes (such as dequantization, inverse transform, and prediction) corresponding to the operation of the encoding device 12.

[0072] Renderer 23 can render decoded video / images. The rendered video / images can be displayed on a monitor.

[0073] Overview of Image Encoding Devices

[0074] Figure 2This is a schematic view illustrating an image encoding device to which embodiments of the present disclosure are applicable.

[0075] like Figure 2 As shown, the image source device 100 may include an image segmenter 110, a subtractor 115, a transformer 120, a quantizer 130, a dequantizer 140, an inverse transformer 150, an adder 155, a filter 160, a memory 170, an inter-frame predictor 180, an intra-frame predictor 185, and an entropy encoder 190. The inter-frame predictor 180 and the intra-frame predictor 185 may be collectively referred to as "predictors". The transformer 120, quantizer 130, dequantizer 140, and inverse transformer 150 may be included in a residual processor. The residual processor may also include a subtractor 115.

[0076] In some embodiments, all or at least some of the multiple components configuring the image source device 100 may be configured by a single hardware component (e.g., an encoder or a processor). Additionally, the memory 170 may include a decoded screen buffer (DPB) and may be configured by a digital storage medium.

[0077] Image segmenter 110 can segment an input image (or picture or frame) input to image source device 100 into one or more processing units. For example, a processing unit may be called a coding unit (CU). A coding unit can be obtained by recursively segmenting a coding tree unit (CTU) or a maximum coding unit (LCU) according to a quadtree / binary tree / tritree (QT / BT / TT) structure. For example, a coding unit can be segmented into multiple coding units of greater depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. For the segmentation of coding units, a quadtree structure can be applied first, and a binary tree structure and / or a ternary tree structure can be applied later. The encoding process according to this disclosure can be performed based on the final coding unit that is no longer segmented. The maximum coding unit can be used as the final coding unit, or a deeper coding unit obtained by segmenting the maximum coding unit can be used as the final coding unit. Here, the encoding process may include the prediction, transformation, and reconstruction processes described later. As another example, the processing unit of the encoding process may be a prediction unit (PU) or a transformation unit (TU). Prediction units and transform units can be partitioned or segmented from the final coding unit. Prediction units can be sample prediction units, and transform units can be units used to derive transform coefficients and / or units used to derive residual signals from transform coefficients.

[0078] The predictor (inter-frame predictor 180 or intra-frame predictor 185) can perform prediction on the block to be processed (the current block) and generate a prediction block that includes prediction samples of the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction based on the current block or CU. The predictor can generate various information related to the prediction of the current block and send the generated information to the entropy encoder 190. The information about the prediction can be encoded in the entropy encoder 190 and output as a bitstream.

[0079] Intra-predictor 185 can predict the current block by referencing samples in the current frame. Depending on the intra-prediction mode and / or intra-prediction technique, the reference samples may be located among the neighbors of the current block or may be placed separately. Intra-prediction modes may include multiple non-directional modes and multiple directional modes. Non-directional modes may include, for example, DC mode and planar mode. Depending on the level of detail in the prediction direction, directional modes may include, for example, 33 or 65 directional prediction modes. However, this is merely an example, and more or fewer directional prediction modes may be used depending on the settings. Intra-predictor 185 can determine the prediction mode to be applied to the current block by using prediction modes applied to neighboring blocks.

[0080] Inter-frame predictor 180 can derive the predicted block of the current block based on a reference block (reference sample array) specified by motion vectors on a reference frame. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation between motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference frame indices. Motion information may also include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current frame and temporally neighboring blocks existing in the reference frame. The reference frame including the reference block and the reference frame including the temporally neighboring block may be the same or different. The temporally neighboring block may be referred to as a juxtaposed reference block, juxtaposed CU (colCU), etc. The reference frame including the temporally neighboring block may be referred to as a juxtaposed frame (colPic). For example, inter-frame predictor 180 can configure a motion information candidate list based on neighboring blocks and generate information indicating which candidate to use to derive the motion vector and / or reference frame index of the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in skip mode and merge mode, the inter-frame predictor 180 can use motion information from neighboring blocks as motion information for the current block. In skip mode, unlike merge mode, residual signals may not be sent. In motion vector prediction (MVP) mode, motion vectors from neighboring blocks can be used as motion vector predictors, and the motion vector of the current block can be signaled by encoding motion vector differences and indicators of the motion vector predictors. The motion vector difference can refer to the difference between the motion vector of the current block and the motion vector predictor.

[0081] The predictor can generate a prediction signal based on various prediction methods and techniques described below. For example, the predictor can apply not only intra-frame prediction or inter-frame prediction, but also both intra-frame prediction and inter-frame prediction simultaneously to predict the current block. A prediction method that simultaneously applies both intra-frame prediction and inter-frame prediction to predict the current block can be called Combined Intra-Frame and Inter-Frame Prediction (CIIP). Alternatively, the predictor can perform Intra-Frame Block Copy (IBC) to predict the current block. Intra-Frame Block Copy can be used for content image / video coding in games, for example, Screen Content Coding (SCC). IBC is a method of predicting the current frame using a previously reconstructed reference block in the current frame at a predetermined distance from the current block. When IBC is applied, the position of the reference block in the current frame can be encoded as a vector (block vector) corresponding to the predetermined distance. IBC essentially performs prediction in the current frame, but can be performed similarly to inter-frame prediction because the reference block is derived within the current frame. That is, IBC can use at least one of the inter-frame prediction techniques described in this disclosure.

[0082] The predicted signal generated by the predictor can be used to generate a reconstructed signal or a residual signal. Subtractor 115 generates a residual signal (residual block or residual sample array) by subtracting the predicted signal (predicted block or predicted sample array) output from the predictor from the input image signal (original block or original sample array). The generated residual signal can be sent to transformer 120.

[0083] Transformer 120 can generate transform coefficients by applying transform techniques to the residual signal. For example, the transform techniques may include at least one of Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Karhunen-Loève Transform (KLT), Graph-Based Transform (GBT), or Conditional Nonlinear Transform (CNT). Here, GBT refers to a transform obtained from a graph when the relationship information between pixels is represented by a graph. CNT refers to a transform obtained based on a prediction signal generated using all previously reconstructed pixels. Furthermore, the transform processing can be applied to square pixel blocks of the same size or to blocks of variable size instead of square.

[0084] Quantizer 130 can quantize the transform coefficients and send them to entropy encoder 190. Entropy encoder 190 can encode the quantized signal (information about the quantized transform coefficients) and output a bitstream. The information about the quantized transform coefficients can be referred to as residual information. Quantizer 130 can rearrange the block-type quantized transform coefficients into a one-dimensional vector form based on the coefficient scan order, and generate information about the quantized transform coefficients based on the one-dimensional vector form of the quantized transform coefficients.

[0085] The entropy encoder 190 can perform various encoding methods, such as exponential Columbus coding, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy encoder 190 can encode, together or separately, information required for video / image reconstruction other than quantization transform coefficients (e.g., values ​​of syntax elements). The encoded information (e.g., encoded video / image information) can be transmitted or stored in bitstream form at the Network Abstraction Layer (NAL) level. The video / image information may also include information about various parameter sets, such as Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). Additionally, the video / image information may include general constraint information. The signaled information, transmitted information, and / or syntax elements described in this disclosure can be encoded and included in the bitstream through the above encoding process.

[0086] The bitstream can be transmitted over a network or stored in a digital storage medium. The network may include broadcast networks and / or communication networks, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) for transmitting the signal output from the entropy encoder 190 and / or a storage unit (not shown) for storing the signal may be included as internal / external components of the image source device 100. Alternatively, a transmitter may be provided as a component of the entropy encoder 190.

[0087] The quantization transform coefficients output from quantizer 130 can be used to generate residual signals. For example, the residual signals (residual blocks or residual samples) can be reconstructed by applying dequantization and inverse transform to the quantization transform coefficients through dequantizer 140 and inverse transformer 150.

[0088] Adder 155 adds the reconstructed residual signal to the prediction signal output from inter-frame predictor 180 or intra-frame predictor 185 to generate a reconstructed signal (reconstructed frame, reconstructed block, reconstructed sample array). If the block to be processed has no residual, such as in the case of applying a skip mode, the prediction block can be used as a reconstructed block. Adder 155 can be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current frame, and can be used for inter-frame prediction of the next frame by filtering as described below.

[0089] Filter 160 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 160 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image and store the modified reconstructed image in memory 170, specifically in the DPB of memory 170. Various filtering methods can include, for example, deblocking filtering, sample adaptive shifting, adaptive loop filtering, bilateral filtering, etc. Filter 160 can generate various filtering-related information and send the generated information to entropy encoder 190, as described later in the description of each filtering method. The filtering-related information can be encoded by entropy encoder 190 and output as a bitstream.

[0090] The modified reconstructed frame sent to memory 170 can be used as a reference frame in inter-frame predictor 180. When inter-frame prediction is applied through image source device 100, prediction mismatch between image source device 100 and image decoding device can be avoided and coding efficiency can be improved.

[0091] The DPB of memory 170 can store modified reconstructed frames for use as reference frames in inter-frame predictor 180. Memory 170 can store motion information of blocks used to derive (or encode) motion information in the current frame and / or motion information of already reconstructed blocks in the frame. The stored motion information can be sent to inter-frame predictor 180 and used as motion information for spatially or temporally neighboring blocks. Memory 170 can store reconstructed samples of reconstructed blocks in the current frame and can transmit the reconstructed samples to intra-frame predictor 185.

[0092] Overview of image decoding devices

[0093] Figure 3 This is a schematic view illustrating an image decoding device to which embodiments of the present disclosure are applicable.

[0094] like Figure 3 As shown, the image receiving device 200 may include an entropy decoder 210, a dequantizer 220, an inverse transformer 230, an adder 235, a filter 240, a memory 250, an inter-frame predictor 260, and an intra-frame predictor 265. The inter-frame predictor 260 and the intra-frame predictor 265 may be collectively referred to as "predictors". The dequantizer 220 and the inverse transformer 230 may be included in a residual processor.

[0095] According to an embodiment, all or at least some of the multiple components configuring the image receiving device 200 may be configured by hardware components (e.g., a decoder or a processor). Additionally, the memory 250 may include a decoded screen buffer (DPB) or may be configured by a digital storage medium.

[0096] The image receiving device 200, which has already received a bitstream including video / image information, can perform operations related to... Figure 2 The image is reconstructed by processing corresponding to the processing performed by the image source device 100. For example, the image receiving device 200 can use a processing unit applied in an image encoding device to perform decoding. Therefore, the decoding processing unit can be, for example, an encoding unit. The encoding unit can be obtained by segmenting a coding tree unit or a maximum coding unit. The reconstructed image signal decoded and output by the image receiving device 200 can be reproduced by a reproduction device (not shown).

[0097] Image receiving device 200 can receive images in bitstream form from Figure 2The signal output by the image encoding device. The received signal can be decoded by the entropy decoder 210. For example, the entropy decoder 210 can parse the bitstream to derive the information (e.g., video / image information) required for image reconstruction (or picture reconstruction). The video / image information may also include information about various parameter sets, such as adaptive parameter sets (APS), picture parameter sets (PPS), sequence parameter sets (SPS), or video parameter sets (VPS). In addition, the video / image information may also include general constraint information. The image decoding device can also decode the picture based on the information about the parameter sets and / or general constraint information. The information and / or syntax elements notified / received by signals described in this disclosure can be decoded and obtained from the bitstream through the decoding process. For example, the entropy decoder 210 decodes the information in the bitstream based on encoding methods such as exponential Golomb coding, CAVLC, or CABAC, and outputs the values ​​of the syntax elements required for image reconstruction and the quantized values ​​of the transform coefficients of the residuals. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element in the bitstream, determine a context model using information about the target syntax element, decoding information of neighboring blocks and the target block, or information about symbols / bins decoded in the previous stage, perform arithmetic decoding on the bins based on the determined context model by predicting the occurrence probability of the bins, and generate symbols corresponding to the value of each syntax element. In this case, the CABAC entropy decoding method can update the context model after determining the context model by using the information of the decoded symbols / bins for the context model of the next symbol / bin. The prediction-related information in the information decoded by the entropy decoder 210 can be provided to the predictors (inter-frame predictor 260 and intra-frame predictor 265), and the residual values ​​of the entropy decoding performed in the entropy decoder 210 (that is, the quantization transform coefficients and related parameter information) can be input to the dequantizer 220. In addition, the filtering information in the information decoded by the entropy decoder 210 can be provided to the filter 240. Furthermore, the receiver (not shown) for receiving signals output from the image encoding device may be further configured as an internal / external element of the image receiving device 200, or the receiver may be a component of the entropy decoder 210.

[0098] Furthermore, the image decoding apparatus according to this disclosure can be referred to as a video / image / screen decoding apparatus. The image decoding apparatus can be divided into an information decoder (video / image / screen information decoder) and a sample decoder (video / image / screen sample decoder). The information decoder may include an entropy decoder 210. The sample decoder may include at least one of a dequantizer 220, an inverse transformer 230, an adder 235, a filter 240, a memory 250, an inter-frame predictor 260, or an intra-frame predictor 265.

[0099] Dequantizer 220 can dequantize the quantized transform coefficients and output transform coefficients. Dequantizer 220 can rearrange the quantized transform coefficients in the form of two-dimensional blocks. In this case, the rearrangement can be performed based on the coefficient scan order performed in the image encoding device. Dequantizer 220 can obtain transform coefficients by performing dequantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information).

[0100] The inverse transformer 230 can perform inverse transformation on the transformation coefficients to obtain the residual signal (residual block, residual sample array).

[0101] The predictor can perform prediction on the current block and generate a prediction block that includes prediction samples of the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction to the current block based on the prediction information output from the entropy decoder 210, and can determine a specific intra-frame / inter-frame prediction mode (prediction technique).

[0102] Similar to the predictor described in the image source device 100, the predictor can generate a prediction signal based on various prediction methods (techniques) that will be described later.

[0103] Intra-predictor 265 can predict the current block by referring to samples in the current frame. The description of intra-predictor 185 also applies to intra-predictor 265.

[0104] Inter-frame predictor 260 can deduce the predicted block of the current block based on a reference block (reference sample array) specified by a motion vector on a reference frame. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation between motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference frame indices. Motion information may also include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current frame and temporally neighboring blocks existing in the reference frame. For example, inter-frame predictor 260 can configure a motion information candidate list based on neighboring blocks and deduce the motion vector and / or reference frame index of the current block based on the received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and the information about the prediction may include information indicating the inter-frame prediction mode of the current block.

[0105] Adder 235 generates a reconstruction signal (reconstructed frame, reconstruction block, reconstruction sample array) by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor (including inter-frame predictor 260 and / or intra-frame predictor 265). If the block to be processed has no residual, such as when a skip mode is applied, the prediction block can be used as a reconstruction block. The description of adder 155 also applies to adder 235. Adder 235 may be referred to as a reconstructor or reconstruction block generator. The generated reconstruction signal can be used for intra-frame prediction of the next block to be processed in the current frame, and can be used for inter-frame prediction of the next frame by filtering as described below.

[0106] Filter 240 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 240 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image and store the modified reconstructed image in memory 250, specifically in the DPB of memory 250. Various filtering methods may include, for example, deblocking filtering, adaptive sample shifting, adaptive loop filtering, bilateral filtering, etc.

[0107] The (modified) reconstructed frame stored in the DPB of memory 250 can be used as a reference frame in inter-frame predictor 260. Memory 250 can store motion information of blocks used to derive (or decode) motion information in the current frame and / or motion information of already reconstructed blocks in the frame. The stored motion information can be sent to inter-frame predictor 260 to be used as motion information for spatially or temporally neighboring blocks. Memory 250 can store reconstructed samples of reconstructed blocks in the current frame and transmit the reconstructed samples to intra-frame predictor 265.

[0108] In this disclosure, the embodiments described in the filter 160, inter-frame predictor 180 and intra-frame predictor 185 of the image source device 100 can be applied equally or correspondingly to the filter 240, inter-frame predictor 260 and intra-frame predictor 265 of the image receiving device 200.

[0109] General image / video encoding process

[0110] In image / video coding, the frames that make up an image / video can be encoded / decoded according to the decoding order. The frame order corresponding to the output order of the decoded frames can be set to be different from the decoding order, and based on this, not only forward prediction but also backward prediction can be performed during inter-frame prediction.

[0111] Figure 4 An example of an illustrative screen decoding process to which embodiments of this disclosure apply is shown. Figure 4In this process, S410 can be executed in the entropy decoder 210 of the decoding device, S420 can be executed in the predictor including the intra-frame predictor 265 and the inter-frame predictor 260, S430 can be executed in the residual processor including the dequantizer 220 and the inverse transformer 230, S440 can be executed in the adder 235, and S450 can be executed in the filter 240. S410 can include the information decoding process described in this disclosure, S420 can include the inter-frame / intra-frame prediction process described in this disclosure, S430 can include the residual processing process described in this disclosure, S440 can include the block / frame reconstruction process described in this disclosure, and S450 can include the in-loop filtering process described in this disclosure.

[0112] refer to Figure 4 The image decoding process can schematically include a process of obtaining image / video information from the bitstream (through decoding) (S410), an image reconstruction process (S420 to S440), and an in-loop filtering process for the reconstructed image (S450). The image reconstruction process can be performed based on prediction samples and residual samples obtained through inter-frame / intra-frame prediction (S420) and residual processing (S430) (dequantization and inverse transform of quantization transform coefficients) as described in this disclosure. For the reconstructed image generated by the image reconstruction process, a modified reconstructed image can be generated through the in-loop filtering process. The modified reconstructed image can be used as the decoded image output, stored in the decoded image buffer or memory 250 of the decoding device, and used as a reference image in the inter-frame prediction process when decoding the image later. In some cases, the in-loop filtering process can be omitted. In this case, the reconstructed image can be used as the decoded image output, stored in the decoded image buffer or memory 250 of the decoding device, and used as a reference image in the inter-frame prediction process when decoding the image later. The in-loop filtering process (S450) may include a deblocking filtering process, a sample adaptive offset (SAO) process, an adaptive loop filter (ALF) process, and / or a bilateral filter process, some or all of which may be omitted as described above. Furthermore, one or more of the deblocking filtering process, the sample adaptive offset (SAO) process, the adaptive loop filter (ALF) process, and / or the bilateral filter process may be applied sequentially, or all of them may be applied sequentially. For example, the SAO process may be performed after the deblocking filtering process has been applied to the reconstructed frame. Alternatively, for example, the ALF process may be performed after the deblocking filtering process has been applied to the reconstructed frame. This can be performed similarly even in an encoding device.

[0113] Figure 5 An example of an illustrative screen encoding process to which embodiments of this disclosure apply is shown. Figure 5 In the above reference, S510 can be found... Figure 2The coding apparatus described herein includes a predictor comprising an intra-frame predictor 185 or an inter-frame predictor 180. S520 may be executed in a residual processor comprising a transformer 120 and / or a quantizer 130, and S530 may be executed in an entropy encoder 190. S510 may include the inter-frame / intra-frame prediction process described herein, S520 may include the residual processing process described herein, and S530 may include the information encoding process described herein.

[0114] refer to Figure 5 The image encoding process can schematically include not only the process of encoding information used for image reconstruction (e.g., prediction information, residual information, segmentation information, etc.) and outputting that information as a bitstream, but also the process of generating a reconstructed image of the current image and (optionally) applying in-loop filtering to the reconstructed image, as shown in the following example. Figure 2 As described, the encoding device can derive (modified) residual samples from the quantized transform coefficients via dequantizer 140 and inverse transformer 150, and generate a reconstructed frame based on the predicted sample as the output of S510 and the (modified) residual samples. The reconstructed frame generated in this way can be equal to the reconstructed frame generated in the decoding device. The modified reconstructed frame can be generated for the reconstructed frame through an in-loop filtering process, can be stored in the decoded frame buffer or memory 170, and can be used as a reference frame in the inter-frame prediction process when encoding frames later, similar to the case in the decoding device. As mentioned above, in some cases, some or all of the in-loop filtering process can be omitted. When the in-loop filtering process is performed, the (in-loop) filtering-related information (parameters) can be encoded in the entropy encoder 190 and output as a bitstream. The decoding device can perform the in-loop filtering process based on the filtering-related information using the same method as the encoding device.

[0115] This in-loop filtering process reduces noise generated during image / video encoding (such as block artifacts and ringing artifacts) and improves subjective / objective visual quality. Furthermore, by performing the in-loop filtering process in both the encoding and decoding devices, the encoding and decoding devices can derive the same prediction results, increasing the reliability of image encoding and reducing the amount of data transmitted for image encoding.

[0116] As described above, the image reconstruction process can be performed not only in the decoding device but also in the encoding device. Reconstructed blocks can be generated based on intra-frame prediction / inter-frame prediction on a block-by-block basis, and a reconstructed image including these blocks can be generated. When the current image / slice / patch group is an I-frame / slice / patch group, blocks included in the current image / slice / patch group can be reconstructed based solely on intra-frame prediction. Furthermore, when the current image / slice / patch group is a P-frame / slice / patch group or a B-frame / slice / patch group, blocks included in the current image / slice / patch group can be reconstructed based on either intra-frame prediction or inter-frame prediction. In this case, inter-frame prediction can be applied to some blocks in the current image / slice / patch group, and intra-frame prediction can be applied to the remaining blocks. The color components of the image can include luminance and chrominance components; unless explicitly defined in this disclosure, the methods and implementations of this disclosure can be applied to both luminance and chrominance components.

[0117] Examples of coding layers and structures

[0118] Video / images encoded according to this disclosure can be processed, for example, according to the encoding layers and structures described below.

[0119] Figure 6 This is a view showing the layer structure of an encoded image. Encoded images can be classified into the Video Coding Layer (VCL) for image decoding and its own processing, the lower system for transmitting and storing encoded information, and the Network Abstraction Layer (NAL) that exists between the VCL and the lower system and is responsible for network adaptation functions.

[0120] In VCL, VCL data that includes compressed image data (slice data) can be generated, or additional enhancement information (SEI) messages that include information such as picture parameter set (PPS), sequence parameter set (SPS) or video parameter set (VPS) can be generated for the decoding process of the image.

[0121] In NAL, header information (NAL cell header) can be added to the raw byte sequence payload (RBSP) generated in VCL to create a NAL cell. In this case, RBSP refers to the slice data, parameter set, and SEI message generated in VCL. The NAL cell header can include NAL cell type information specified based on the RBSP data included in the corresponding NAL cell.

[0122] As shown in the figure, NAL units can be classified into VCL NAL units and non-VCL NAL units based on the RBSP generated in the VCL. VCL NAL units can refer to NAL units that include information about the image (slice data), while non-VCL NAL units can refer to NAL units that include information needed to decode the image (parameter set or SEI message).

[0123] VCL NAL units and non-VCL NAL units can be appended with header information and transmitted over a network according to the data standard of the lower system. For example, NAL units can be modified to a predetermined standard data format such as H.266 / VVC file format, RTP (Real-Time Transport Protocol), or TS (Transport Streaming) and transmitted over various networks.

[0124] As described above, in a NAL unit, the NAL unit type can be specified according to the RBSP data structure included in the corresponding NAL unit, and information about the NAL unit type can be stored in the NAL unit header and notified by a signal.

[0125] For example, based on whether the NAL unit includes information about the image (slice data), it can be roughly classified into VCLNAL unit type and non-VCL NAL unit type. VCL NAL unit type can be classified according to the characteristics and type of the image included in the VCL NAL unit, while non-VCL NAL unit type can be classified according to the type of parameter set.

[0126] Below are examples of NAL cell types specified based on the type of parameter set / information included in non-VCL NAL cell types.

[0127] - DCI (Decoding Capability Information) NAL Unit: Includes the type of NAL unit for DCI.

[0128] - VPS (Video Parameter Set) NAL Unit: Includes the type of NAL unit for the VPS.

[0129] - SPS (Sequence Parameter Set) NAL Unit: The type of NAL unit that includes SPS.

[0130] - PPS (Picture Parameter Set) NAL Unit: Includes the types of NAL units for PPS.

[0131] - APS (Adaptive Parameter Set) NAL Unit: The type of NAL unit including APS.

[0132] - PH (Picture Header) NAL Unit: Types of NAL units including PH.

[0133] The aforementioned NAL unit type can have syntax information specific to the NAL unit type, and this syntax information can be stored in the NAL unit header and signaled. For example, the syntax information can be nal_unit_type, and the NAL unit type can be specified as the nal_unit_type value.

[0134] Furthermore, as mentioned above, a frame can include multiple slices, and a slice can include a slice header and slice data. In this case, a frame header can be further added to multiple slices within a frame (slice header and slice data set). The frame header (frame header syntax) can include information / parameters that are typically applicable to the frame.

[0135] A slice header (slice header syntax) may include information / parameters typically applicable to a slice. APS (APS syntax) or PPS (PPS syntax) may include information / parameters typically applicable to one or more slices or frames. SPS (SPS syntax) may include information / parameters typically applicable to one or more sequences. VPS (VPS syntax) may include information / parameters typically applicable to multiple layers. DCI (DCI syntax) may include information / parameters typically applicable to the entire video. DCI may include information / parameters related to decoding capabilities. In this disclosure, the High-Level Syntax (HLS) may include at least one of APS syntax, PPS syntax, SPS syntax, VPS syntax, DCI syntax, frame header syntax, or slice header syntax. Furthermore, in this disclosure, the Low-Level Syntax (LLS) may include, for example, slice data syntax, CTU syntax, coding unit syntax, transform unit syntax, etc.

[0136] In this disclosure, the image / video information encoded in the encoding device and signaled to the decoding device in the form of a bitstream can include not only intra-frame segmentation related information, intra / inter-frame prediction information, residual information, and intra-loop filtering information, but also information about slice headers, frame headers, APS, PPS, SPS, VPS, and / or DCI. Additionally, the image / video information may also include general constraint information and / or information about NAL unit headers.

[0137] Multi-layer coding

[0138] Image / video coding according to this disclosure may include multi-layer-based image / video coding. Multi-layer-based image / video coding may include scalable coding. In multi-layer-based coding or scalable coding, the input signal may be processed for each layer. Depending on the layers, the input signal (input image / video) may have different values ​​in at least one of resolution, frame rate, bit depth, color format, aspect ratio, or view. In this case, redundant information transmission / processing and compression efficiency can be reduced and increased by performing inter-layer prediction using the differences between layers (e.g., based on scalability).

[0139] Figure 7This is a schematic block diagram of a multilayer encoding device 700 to which embodiments of the present disclosure are applied and which performs encoding of multilayer video / image signals.

[0140] Figure 7 The multi-layer coding device 700 may include Figure 2 Encoding devices. With Figure 2 In comparison, Figure 7 The image segmenter 110 and adder 155 are not shown in the multi-layer coding device 700, but the multi-layer coding device 700 may include the image segmenter 110 and adder 155. In an embodiment, the image segmenter 110 and adder 155 may be included in the cells of a layer. In the following, multi-layer-based prediction will focus on... Figure 7 In the description. For example, in addition to the following description, the multilayer coding device 700 may include the above references. Figure 2 The technical concept of the described encoding device.

[0141] For ease of description, Figure 7 The diagram illustrates a multi-layer structure consisting of two layers. However, embodiments of this disclosure are not limited to two layers, and the multi-layer structure to which embodiments of this disclosure are applied may include two or more layers.

[0142] refer to Figure 7 The encoding device 700 includes an encoder 700-1 for layer 1 and an encoder 700-0 for layer 0. Layer 0 can be a base layer, a reference layer, or a lower layer, and layer 1 can be an enhancement layer, a current layer, or a higher layer.

[0143] The encoder 700-1 of layer 1 may include a predictor 720-1, a residual processor 730-1, a filter 760-1, a memory 770-1, an entropy encoder 740-1, and a multiplexer (MUX) 770. In some embodiments, the MUX may be included as an external component.

[0144] The encoder 700-0 of layer 0 may include a predictor 720-0, a residual processor 730-0, a filter 760-0, a memory 770-0, and an entropy encoder 740-0.

[0145] Predictors 720-0 and 720-1 can perform predictions on the input image based on various prediction schemes as described above. For example, predictors 720-0 and 720-1 can perform inter-frame prediction and intra-frame prediction. Predictors 720-0 and 720-1 can perform predictions within a predetermined processing unit. The prediction unit can be a coding unit (CU) or a transform unit (TU). Prediction blocks (including prediction samples) can be generated based on the prediction results, and based on this, a residual processor can derive residual blocks (including residual samples).

[0146] Inter-frame prediction can generate a prediction block by performing prediction based on information about at least one of the previous and / or next frames. Intra-frame prediction can generate a prediction block by performing prediction based on neighboring samples in the current frame.

[0147] As inter-frame prediction modes or methods, the various prediction modes or methods described above can be used. In inter-frame prediction, a reference frame can be selected for the current block to be predicted, and a reference block corresponding to the current block can be selected from the reference frame. Predictors 720-0 and 720-1 can generate prediction blocks based on the reference blocks.

[0148] Additionally, predictor 720-1 can use information about layer 0 to perform predictions for layer 1. In this disclosure, for ease of description, the method of using information about another layer to predict information about the current layer is referred to as inter-layer prediction.

[0149] Information about the current layer predicted using information about another layer (e.g., predicted via inter-layer prediction) can be at least one of texture, motion information, cell information, or predetermined parameters (e.g., filter parameters, etc.).

[0150] Additionally, the information about another layer used for prediction of the current layer (e.g., for inter-layer prediction) can be at least one of texture, motion information, cell information, or predetermined parameters (e.g., filtering parameters, etc.).

[0151] For inter-layer prediction, the current block can be a block in the current frame of the current layer (e.g., layer 1) and can be a block to be encoded. The reference block is a block in the frame (reference frame) of the same access unit (AU) as the frame (current frame) to which the current block belongs, on the layer (reference layer, e.g., layer 0) to which the prediction of the current block is referenced, and can be a block corresponding to the current block.

[0152] As an example of inter-layer prediction, there exists inter-layer motion prediction that uses motion information from a reference layer to predict the motion information of the current layer. Based on inter-layer motion prediction, the motion information of the current block can be predicted using the motion information of the reference block. That is, in deriving motion information according to the inter-frame prediction mode described below, motion information candidates can be derived based on the motion information of the inter-layer reference block rather than temporally neighboring blocks.

[0153] When applying interlayer motion prediction, predictor 720-1 can scale and use motion information from the reference block of the reference layer (that is, the interlayer reference block).

[0154] As another example of inter-layer prediction, inter-layer texture prediction can use the texture of a reconstructed reference block as the predicted value for the current block. In this case, predictor 720-1 can scale the texture of the reference block by up-scaling. Inter-layer texture prediction can be called inter-layer (reconstructed) sample prediction or simply inter-layer prediction.

[0155] In inter-layer parameter prediction, which is another example of inter-layer prediction, the derivation parameters of the reference layer can be reused in the current layer, or the parameters for the current layer can be derived based on the parameters used in the reference layer.

[0156] In interlayer residual prediction, which is another example of interlayer prediction, residual information from another layer can be used to predict the residual information of the current layer, and based on this, prediction of the current block can be performed.

[0157] In inter-layer difference prediction, which is another example of inter-layer prediction, the prediction of the current block can be performed using the difference between the images obtained by upsampling or downsampling the reconstructed image of the current layer and the reconstructed image of the reference layer.

[0158] In inter-layer syntax prediction, another example of inter-layer prediction, the syntax information of a reference layer can be used to predict or generate the texture of the current block. In this case, the syntax information of the referenced layer can include information about intra-frame prediction modes and motion information.

[0159] When predicting a specific block, various prediction methods using the inter-layer prediction described above can be utilized.

[0160] Here, as an example of interlayer prediction, although interlayer texture prediction, interlayer motion prediction, interlayer cell information prediction, interlayer parameter prediction, interlayer residual prediction, interlayer difference prediction, interlayer syntax prediction, etc. are described, the interlayer prediction applicable to this disclosure is not limited to these.

[0161] For example, inter-layer prediction can be applied as an extension of inter-frame prediction for the current layer. That is, inter-frame prediction can be performed on the current block by including a reference frame derived from the reference layer in the reference frame that can be referenced for inter-frame prediction of the current block.

[0162] In this case, inter-layer reference frames can be included in the reference frame list of the current block. Predictor 720-1 can use inter-layer reference frames to perform inter-frame prediction for the current block.

[0163] Here, the interlayer reference frame can be a reference frame constructed by sampling the reconstructed frame of a reference layer to correspond to the current layer. Therefore, when the reconstructed frame of the reference layer corresponds to the frame of the current layer, the reconstructed frame of the reference layer can be used as the interlayer reference frame without sampling. For example, when the width and height of the sample are the same in the reconstructed frames of the reference layer and the current layer, and the offsets between the top left, top right, bottom left, and bottom right edges of the reference layer and the top left, top right, bottom left, and bottom right edges of the current layer are 0, the reconstructed frame of the reference layer can be used as the interlayer reference frame of the current layer without being resampled.

[0164] Furthermore, the reconstructed frame of the reference layer of the interlayer reference frame can be a frame belonging to the same AU as the current frame to be encoded.

[0165] When inter-frame prediction of the current block is performed by including inter-layer reference frames in the reference frame list, the position of the inter-layer reference frames in the reference frame list can differ between reference frame lists L0 and L1. For example, in reference frame list L0, the inter-layer reference frame can be located after a short reference frame preceding the current frame, and in reference frame list L1, the inter-layer reference frame can be located at the end of the reference frame list.

[0166] Here, reference frame list L0 is a reference frame list used for inter-frame prediction of P-slices or a reference frame list used as the first reference frame list in inter-frame prediction of B-slices. Reference frame list L1 can be a second reference frame list used for inter-frame prediction of B-slices.

[0167] Therefore, the reference frame list L0 can be composed of short-term reference frames preceding the current frame, inter-layer reference frames, short-term reference frames following the current frame, and long-term reference frames in this order. The reference frame list L1 can be composed of short-term reference frames following the current frame, short-term reference frames preceding the current frame, long-term reference frames, and inter-layer reference frames in this order.

[0168] In this context, a predictive (P) slice is a slice for which inter-frame prediction or intra-frame prediction is performed using at most one motion vector and reference frame index per predictive block. A double-predictive (B) slice is a slice for which prediction or intra-frame prediction is performed using at most two motion vectors and reference frame indices per predictive block. At this point, an intra-frame (I) slice is a slice for which only intra-frame prediction is applied.

[0169] Additionally, when performing inter-frame prediction for the current block based on a list of reference frames that includes inter-layer reference frames, the list of reference frames may include multiple inter-layer reference frames derived from multiple layers.

[0170] When multiple inter-frame reference frames are included, they can be arranged alternately in reference frame lists L0 and L1. For example, suppose two inter-frame reference frames (such as inter-frame reference frame ILRPi and inter-frame reference frame ILRPj) are included in the reference frame list for inter-frame prediction of the current block. In this case, in reference frame list L0, ILRPi can be located after a short reference frame preceding the current frame, and ILRPj can be located at the end of the list. Similarly, in reference frame list L1, ILRPi can be located at the end of the list, and ILRPj can be located after a short reference frame following the current frame.

[0171] In this case, the reference frame list L0 can be composed of short-term reference frames preceding the current frame, interlayer reference frames (ILRPi), short-term reference frames following the current frame, long-term reference frames, and interlayer reference frames (ILRPj) in this order. The reference frame list L1 can be composed of short-term reference frames following the current frame, interlayer reference frames (ILRPj), short-term reference frames preceding the current frame, long-term reference frames, and interlayer reference frames (ILRPi) in this order.

[0172] Furthermore, one of the two interlayer reference frames can be an interlayer reference frame derived from a scalable layer used for resolution, and the other can be an interlayer reference frame derived from a layer used to provide another view. In this case, for example, if ILRPi is an interlayer reference frame derived from a layer used to provide a different resolution and ILRPj is an interlayer reference frame derived from a layer used to provide a different view, then in the case of scalable video coding that only supports scalability excluding the view, the reference frame list L0 can be composed in this order of short-term reference frames before the current frame, interlayer reference frame ILRPi, short-term reference frames after the current frame, and long-term reference frames, and the reference frame list L1 can be composed in this order of short-term reference frames after the current frame, short-term reference frames before the current frame, long-term reference frames, and interlayer reference frame ILRPi.

[0173] Furthermore, in inter-layer prediction, information about the inter-layer reference frame can be obtained using only sample values, only motion information (motion vectors), or both. When the reference frame index indicates an inter-layer reference frame, based on information received from the encoding device, the predictor 720-1 can use only sample values ​​of the inter-layer reference frame, only motion information (motion vectors) of the inter-layer reference frame, or both.

[0174] When using only sample values ​​from the inter-frame reference frame, predictor 720-1 can derive the predicted sample for the current block from the sample of the block specified by the motion vector from the inter-frame reference frame. Without considering scalable video coding of the view, the motion vector in the inter-frame prediction (inter-frame prediction) using the inter-frame reference frame can be set to a fixed value (e.g., 0).

[0175] When using only motion information from inter-layer reference frames, predictor 720-1 can use the motion vector specified by the inter-layer reference frames as a motion vector predictor for deriving the motion vector of the current block. Alternatively, predictor 720-1 can use the motion vector specified by the inter-layer reference frames as the motion vector of the current block.

[0176] When using both sample values ​​from the inter-layer reference image and motion information, the predictor 720-1 can use samples from the region corresponding to the current block in the inter-layer reference image and the motion information (motion vector) specified in the inter-layer reference image for the prediction of the current block.

[0177] When applying interlayer prediction, the encoding device can send the reference index of the interlayer reference frame in the reference frame list to the decoding device, and can also send information to the decoding device to specify which information (sample information, motion information, or a combination of sample information and motion information) to be used from the interlayer reference frame (that is, information to specify the dependency type of the interlayer prediction dependency between two layers).

[0178] Figure 8 This is a schematic block diagram of a decoding device to which embodiments of the present disclosure are applied and which performs decoding of multi-layer video / image signals. Figure 8 The decoding device may include Figure 3 Decoding devices. Figure 8 The realigner shown can be omitted or included in the dequantizer. The description of this figure will focus on multi-layer-based prediction. Additionally, it may include... Figure 3 Description of the decoding device.

[0179] exist Figure 8 In the examples provided, for ease of description, a multi-layer structure consisting of two layers will be described. However, it should be noted that the embodiments of this disclosure are not limited thereto, and the multi-layer structures to which the embodiments of this disclosure apply may include two or more layers.

[0180] refer to Figure 8The decoding device 800 may include a layer 1 decoder 800-1 and a layer 0 decoder 800-0. The layer 1 decoder 800-1 may include an entropy decoder 810-1, a residual processor 820-1, a predictor 830-1, an adder 840-1, a filter 850-1, and a memory 860-1. The layer 0 decoder 800-0 may include an entropy decoder 810-0, a residual processor 820-0, a predictor 830-0, an adder 840-0, a filter 850-0, and a memory 860-0.

[0181] When a bitstream containing image information is received from an encoding device, the demultiplexer 805 can demultiplex the information of each layer and send the information to a decoding device for each layer.

[0182] Entropy decoders 810-1 and 810-0 can perform decoding corresponding to the encoding method used in the encoding device. For example, when CABAC is used in the encoding device, entropy decoders 810-1 and 810-0 can perform entropy decoding using CABAC.

[0183] When the prediction mode of the current block is intra-prediction mode, predictors 830-1 and 830-0 can perform intra-prediction on the current block based on the neighboring reconstructed samples in the current frame.

[0184] When the prediction mode for the current block is inter-frame prediction mode, predictors 830-1 and 830-0 can perform inter-frame prediction for the current block based on information from at least one of the frames included before or after the current frame. Some or all of the motion information necessary for inter-frame prediction can be derived by examining information received from the encoding device.

[0185] When skip mode is applied as inter-frame prediction mode, no residual is sent from the coding device, and the prediction block can be a reconstructed block.

[0186] Furthermore, the predictor 830-1 of layer 1 can perform inter-frame prediction or intra-frame prediction using only information about layer 1, and perform inter-layer prediction using information about another layer (layer 0).

[0187] Information about the current layer that is predicted using information about another layer (e.g., predicted via inter-layer prediction) can exist in at least one of texture, motion information, cell information, and predetermined parameters (e.g., filter parameters, etc.).

[0188] Information about another layer that is used for prediction of the current layer (e.g., for inter-layer prediction) may exist in at least one of texture, motion information, cell information, and predetermined parameters (e.g., filter parameters, etc.).

[0189] In inter-layer prediction, the current block can be a block in the current frame of the current layer (e.g., layer 1) and can be a block to be decoded. The reference block can be a block in the frame (reference frame) belonging to the same access unit (AU) as the frame (current frame) to which the current block belongs, on the layer (reference layer, e.g., layer 0) to which the prediction of the current block is referenced, and can be a block corresponding to the current block.

[0190] The multilayer decoding device 800 can perform interlayer prediction, as described in the multilayer coding device 700. For example, the multilayer decoding device 800 can perform interlayer texture prediction, interlayer motion prediction, interlayer cell information prediction, interlayer parameter prediction, interlayer residual prediction, interlayer difference prediction, interlayer syntax prediction, etc., as described in the multilayer coding device 700, and the interlayer prediction applicable in this disclosure is not limited thereto.

[0191] When a reference frame index received from the encoding device or a reference frame index derived from a neighboring block indicates an inter-layer reference frame in the reference frame list, the predictor 830-1 can perform inter-layer prediction using the inter-layer reference frame. For example, when the reference frame index indicates an inter-layer reference frame, the predictor 830-1 can derive the sample values ​​of the region specified by the motion vector in the inter-layer reference frame as the prediction block for the current block.

[0192] In this case, inter-layer reference frames can be included in the reference frame list of the current block. Predictor 830-1 can use inter-layer reference frames to perform inter-frame prediction for the current block.

[0193] As described above in the multi-layer encoding device 700, in the operation of the multi-layer decoding device 800, the inter-layer reference frame can be a reference frame constructed by sampling the reconstructed frame of the reference layer to correspond to the current layer. Processing for the case where the reconstructed frame of the reference layer corresponds to the frame of the current layer can be performed in the same manner as in the encoding process.

[0194] Furthermore, as described above in the multi-layer encoding device 700, in the operation of the multi-layer decoding device 800, the reconstructed frame of the reference layer from which the inter-layer reference frame is derived can be a frame belonging to the same AU as the current frame to be encoded.

[0195] Furthermore, as described above in the multilayer coding device 700, in the operation of the multilayer decoding device 800, when inter-frame prediction of the current block is performed by including interlayer reference frames in the reference frame list, the position of the interlayer reference frames in the reference frame list can be different between reference frame lists L0 and L1.

[0196] Furthermore, as described above in the multilayer coding device 700, in the operation of the multilayer decoding device 800, when inter-frame prediction of the current block is performed based on a reference frame list including interlayer reference frames, the reference frame list may include multiple interlayer reference frames derived from multiple layers, and the arrangement of the interlayer reference frames may be performed to correspond to the arrangement of the interlayer reference frames described in the coding process.

[0197] Furthermore, as described above in the multilayer encoding device 700, in the operation of the multilayer decoding device 800, information about the interlayer reference frame can be obtained using only sample values, only motion information (motion vectors), or both sample values ​​and motion information.

[0198] The multilayer decoding device 800 can receive reference indices from the multilayer encoding device 700 that indicate interlayer reference frames in the reference frame list, and perform interlayer prediction based on the reference indices. Additionally, the multilayer decoding device 800 can receive information from the multilayer encoding device 700 specifying which information (sample information, motion information, or both) to use from the interlayer reference frames (i.e., information specifying the dependency type of the interlayer prediction dependency between two layers).

[0199] Reference Figure 9 and Figure 10 This document describes image encoding and decoding methods performed by a multilayer image encoding device and a multilayer image decoding device, respectively, according to embodiments. In the following text, for ease of description, the multilayer image encoding device is referred to as an image encoding device. Furthermore, the multilayer image decoding device is referred to as an image decoding device.

[0200] Figure 9 This is a view illustrating a method for encoding an image based on multiple layers by an image encoding device according to an embodiment. The encoding device according to the embodiment can encode a first layer of the image (S910). Next, the encoding device can encode a second layer of the image based on the first layer (S920). Next, the encoding device can output a bitstream (for multiple layers) (S930).

[0201] Figure 10 This is a view illustrating a method for decoding an image based on multiple layers by an image decoding device according to an embodiment. The decoding device according to the embodiment can obtain video / image information from a bitstream (S1010). Next, the decoding device can decode the first layer of the image based on the video / image information (S1020). Next, the decoding device can decode the second layer of the image based on the video / image information and the first layer (S1030).

[0202] In implementations, video / image information may include High-Level Syntax (HLS) as described below. In implementations, HLS may include SPS and / or PPS as disclosed in this disclosure. For example, video / image information may include the information and / or syntax elements described in this disclosure. As described in this disclosure, the second layer of the image may be encoded based on motion information / reconstruction samples / parameters of the first layer's image. In implementations, the first layer may be lower than the second layer. In implementations, when the second layer is the current layer, the first layer may be referred to as the reference layer.

[0203] High-Level Syntax (HLS) Signaling and Semantics

[0204] As described above, the HLS can be encoded and / or signaled for use in video and / or image encoding. As stated above, the video / image information disclosed herein can be included in the HLS. Furthermore, image / video encoding methods can be performed based on this image / video information.

[0205] Video Parameter Set Signaling

[0206] A Video Parameter Set (VPS) is a set of parameters used to carry layer information. Layer information may include, for example, information about the Output Layer Set (OLS), information about the profile layer hierarchy, information about the relationship between the OLS and the hypothetical reference decoder, and information about the relationship between the OLS and the Decoding Picture Buffer (DPB).

[0207] The VPS Raw Byte Sequence Payload (RBSP) will be available for the decoding process before it is referenced, including in at least one access unit with a TemporalId equal to 0 or provided by an external device. All VPS NAL units in the Encoded Video Sequence (CVS) with a specific value of vps_video_parameter_set_id will have the same content.

[0208] Figure 11 This is a view illustrating a portion of the syntax of a VPS according to an implementation method. The Video Parameter Set (VPS) is a set of parameters used to send layer information. In the following text, reference will be made to... Figure 11 The description is a syntax element that can be signaled via VPS.

[0209] `vps_video_parameter_set_id` provides an identifier for the VPS. Other syntax elements can be found in the documentation for VPSs using `vps_video_parameter_set_id`. The value of `vps_video_parameter_set_id` will be greater than 0.

[0210] `vps_max_layers_minus1` can specify the maximum number of layers allowed per CVS of the reference VPS. For example, `vps_max_layers_minus1` plus 1 specifies the maximum number of layers allowed per CVS of the reference VPS.

[0211] Increasing `vps_max_sublayer_minus1` by 1 specifies the maximum number of time-based sublayers that can exist in each CVS of the reference VPS.

[0212] A `vps_all_layers_same_num_sublayer_flag` value of 1 specifies that the number of time sublayers is the same for all layers in each CVS of the reference VPS. A `vps_all_layers_same_num_sublayer_flag` value of 0 specifies that the number of time sublayers can be the same or different for each CVS of the reference VPS. When a value for `vps_all_layers_same_num_sublayer_flag` is not provided in the bitstream, its value can be inferred to be 1.

[0213] A `vps_all_independent_layers_flag` value of 1 specifies that all layers in CVS are encoded independently without using inter-layer prediction. A `vps_all_independent_layers_flag` value of 0 specifies that one or more layers in CVS can be encoded using inter-layer prediction.

[0214] `vps_layer_id[i]` can specify the `nuh_layer_id` value of the `i`-th layer. For any two non-negative integer values ​​`m` and `n`, the value of `vps_layer_id[m]` will be less than `vps_layer_id[n]` when `m` is less than `n`. Here, `nuh_layer_id` is a syntax element signaled in the NAL cell header and can specify the identifier of the NAL cell.

[0215] A `vps_independent_layer_flag[i]` equal to 1 specifies that the layer at index `i` does not use inter-layer prediction. A `vps_independent_layer_flag[i]` equal to 0 specifies that the layer at index `i` can use inter-layer prediction, and the syntax element `vps_direct_ref_layer_flag[i][j]` can be obtained from the VPS. Here, `j` can be in the range of 0 to `i-1` (inclusive). When the value of `vps_independent_layer_flag[i]` is not present in the bitstream, the value of `vps_independent_layer_flag[i]` can be deduced to be equal to 1.

[0216] A `vps_direct_ref_layer_flag[i][j]` equal to 0 specifies that the layer with index `j` is not a direct reference layer to the layer with index `i`. A `vps_direct_ref_layer_flag[i][j]` equal to 1 specifies that the layer with index `j` is a direct reference layer to the layer with index `i`. When `vps_direct_ref_layer_flag[i][j]` is not obtained from the bitstream for `i` and `j` in the range 0 to `vps_max_layers_minus1` (inclusive), its value can be deduced to be 0. When `vps_independent_layer_flag[i]` equals 0, there should exist at least one `j` value in the range 0 to `i-1` (inclusive) such that the value of `vps_direct_ref_layer_flag[i][j]` is equal to 1.

[0217] In the implementation method, use Figure 12 The pseudocode derives the variables NumDirectRefLayers[i], DirectRefLayerIdx[i][d], NumRefLayers[i], RefLayerIdx[i][r], and LayerUsedAsRefLayerFlag[j].

[0218] The variable GeneralLayerIdx[i] can be derived to specify the layer index of the layer whose nuh_layer_id is equal to vps_layer_id[i], as shown in the following formula.

[0219] [Formula 1]

[0220] for (i=0; i<=vps_max_layers_minus1; i++)

[0221] GeneralLayerIdx[vps_layer_id[i]]=i

[0222] Sequence Parameter Set Signaling

[0223] Figure 13 This is a view illustrating a portion of the syntax of an SPS according to an implementation. A Sequence Parameter Set (SPS) is a set of parameters used to transmit information for encoding video sequences. All SPS RBSPs are available for the decoding process before being referenced. This can be included in at least one AU with a TemporalId equal to 0 or provided by an external device.

[0224] In the following text, reference will be made to Figure 13 The description is a syntax element that can be signaled via SPS.

[0225] The syntax element `sps_seq_parameter_set_id` provides an identifier for an SPS referenced from another syntax element. Regardless of the value of `nuh_layer_id`, SPS NAL units can share the same value space for `sps_seq_parameter_set_id`.

[0226] Suppose that the nuh_layer_id value of a particular SPS NAL cell is spsLayerId, and the nuh_layer_id value of a particular VCL NAL cell is vclLayerId. In this case, unless the layer whose spsLayerId is equal to or less than vclLayerId and whose nuh_layer_id is equal to spsLayerId is included in at least one set of output layers that includes the layer whose nuh_layer_id is equal to vclLayerId, the particular VCL NAL cell will not reference the particular SPS NAL cell.

[0227] When the value of sps_video_parameter_set_id is greater than 0, the syntax element sps_video_parameter_set_id can specify the value of vps_video_parameter_set_id of the VPS referenced by SPS.

[0228] The following applies when the value of sps_video_parameter_set_id is equal to 0.

[0229] - SPS does not reference VPS.

[0230] - During the decoding process of each CLVS of the reference SPS, no VPS is referenced.

[0231] The value of vps_max_layers_minus1 is deduced to be equal to 0.

[0232] - CVS consists of only one layer. For example, all VCL NAL units in CVS will have the same nuh_layer_id value.

[0233] The value of GeneralLayerIdx[nuh_layer_id] can be deduced to be equal to 0.

[0234] The value of - vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] can be deduced to be equal to 1.

[0235] When the value of vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 1, the SPS referenced by the CLVS with nuhLayerId as the nuh_layer_id value should have nuh_layer_id equal to nuhLayerId.

[0236] The value of sps_video_parameter_set_id can be set equally in all SPS references in CVS.

[0237] A `res_change_in_clvs_allowed_flag` value of 1 specifies that the screen resolution in a CLVS referencing the SPS can be changed. A `res_change_in_clvs_allowed_flag` value of 0 specifies that the screen resolution in any CLVS referencing the SPS can remain unchanged.

[0238] `pic_width_max_in_luma_samples` can specify the maximum width of each decoded frame in the reference SPS within the luminance sample unit. `pic_width_max_in_luma_samples` can be non-zero and can be an integer multiple of `max(8, MinCbSizeY)`.

[0239] When the value of sps_video_parameter_set_id is greater than 0, for bitstream consistency, for all OLS including at least one layer of the reference SPS, when the OLS index i of the OLS is i, the value of pic_width_max_in_luma_samples should be equal to or less than ols_dpb_pic_width[i].

[0240] `pic_height_max_in_luma_samples` can specify the maximum height of each decoded frame in the reference SPS within the luminance sample unit. `pic_height_max_in_luma_samples` can be non-zero and can be an integer multiple of `max(8, MinCbSizeY)`.

[0241] When the value of sps_video_parameter_set_id is greater than 0, for bitstream consistency, for all OLS including at least one layer of the reference SPS, when the OLS index i of the OLS is i, the value of pic_height_max_in_luma_samples should be equal to or less than ols_dpb_pic_width[i].

[0242] Screen parameter set signaling

[0243] Figure 14 This is a view illustrating a portion of the syntax of a PPS according to an implementation. A Picture Parameter Set (PPS) is a set of parameters used to transmit information for encoding a picture. The PPS RBSP will be available for the decoding process before it is referenced. This can be included in at least one AU where TemporalId is equal to or less than 0 or provided by an external device.

[0244] In the following text, reference will be made to Figure 14 This describes the syntax element that can be signaled via PPS. The syntax element `pps_pic_parameter_set_id` can specify the PPS that will be referenced by another syntax element. The value of `pps_pic_parameter_set_id` can be in the range of 0 to 63 (inclusive). Regardless of the value of `nuh_layer_id`, PPS NAL units can share the same value space of `pps_pic_parameter_set_id`.

[0245] Suppose that the nuh_layer_id value of a specific PPS NAL cell is ppsLayerId, and the nuh_layer_id value of a specific VCL NAL cell is vclLayerId. In this case, unless the layer whose ppsLayerId is equal to or less than vclLayerId and whose nuh_layer_id is equal to ppsLayerId is included in at least one set of output layers that includes the layer whose nuh_layer_id is equal to vclLayerId, the specific VCL NAL cell will not reference the specific PPS NAL cell.

[0246] `pps_seq_parameter_set_id` specifies the value of `sps_seq_parameter_set_id` for the SPS. The value of `pps_seq_parameter_set_id` can be in the range of 0 to 15 (inclusive). The value of `pps_seq_parameter_set_id` can be the same across all PPSs referenced in the encoding screen within CLVS.

[0247] `pic_width_in_luma_samples` specifies the width of each decoded frame in the reference PPS within the luma sample unit. `pic_width_in_luma_samples` can be non-zero and can be an integer multiple of `Max(8, MinCbSizeY)`, and should be equal to or less than `pic_width_max_in_luma_samples`. Here, `Max(8, MinCbSizeY)` indicates the larger of 8 and `MinCbSizeY`, and `MinCbSizeY` indicates the size of the smallest available coded block in the luma sample unit.

[0248] When the value of res_change_in_clvs_allowed_flag is equal to 0, the value of pic_width_in_luma_samples will be equal to pic_width_max_in_luma_samples.

[0249] When the value of sps_ref_wraparound_enabled_flag is equal to 1, the value of (CtbSizeY / Min_CbSizeY+1) should be equal to or less than (pic_width_in_luma_samples / MinCbSizeY-1).

[0250] `pic_height_in_luma_samples` can specify the height of each decoded frame of the reference PPS in the luminance sample unit. `pic_height_in_luma_samples` can be non-zero and can be an integer multiple of Max(8, MinCbSizeY), and will be equal to or less than `pic_height_max_in_luma_samples`.

[0251] When the value of res_change_in_clvs_allowed_flag is equal to 0, the value of pic_height_in_luma_samples will be equal to pic_height_max_in_luma_samples.

[0252] The variables PicWidthInCtbsY (indicating the width of the image of the luminance sample unit in the coding tree block unit of the luminance sample unit), PicHeightInCtbsY (indicating the height of the image of the luminance sample unit in the coding tree block unit of the luminance sample unit), PicSizeInCtbsY (indicating the size of the image of the luminance sample unit in the coding tree block unit of the luminance sample unit), PicWidthInMinCbsY (indicating the width of the image of the luminance sample unit in the minimum coding tree block unit of the luminance sample unit), PicHeightInMinCbsY (indicating the height of the image of the luminance sample unit in the minimum coding tree block unit of the luminance sample unit), PicSizeInMinCbsY (indicating the size of the image of the luminance sample unit in the minimum coding tree block unit of the luminance sample unit), PicSizeInSamplesY (indicating the size of the image of the luminance sample unit in the luminance sample unit), PicWidthInSamplesC (indicating the width of the image of the chrominance sample unit), and PicHeightInSamplesC (indicating the height of the image of the chrominance sample unit) can be derived as shown in the following formula.

[0253] [Equation 2]

[0254] PicWidthInCtbsY=Ceil (pic_width_in_luma_samples÷CtbSizeY)

[0255] PicHeightInCtbsY=Ceil (pic_height_in_luma_samples÷CtbSizeY)

[0256] PicSizeInCtbsY=PicWidthInCtbsY PicHeightInCtbsY

[0257] PicWidthInMinCbsY=pic_width_in_luma_samples / MinCbSizeY

[0258] PicHeightInMinCbsY=pic_height_in_luma_samples / MinCbSizeY

[0259] PicSizeInMinCbsY=PicWidthInMinCbsY PicHeightInMinCbsY

[0260] PicSizeInSamplesY=pic_width_in_luma_samples pic_height_in_luma_samples

[0261] PicWidthInSamplesC=pic_width_in_luma_samples / SubWidthC

[0262] PicHeightInSamplesC=pic_height_in_luma_samples / SubHeightC

[0263] In the above formula, Ceil() is a floor function. CtbSizeY represents the size of the coding tree block in the luma sample unit, which can be obtained from the bitstream. MinCbSizeY represents the size of the smallest coding block in the luma sample unit, which can also be obtained from the bitstream. SubWidthC represents the difference in the ratio between the luma block and the chroma block. The ratio of the luma block to the ratio of the chroma block can be represented by an integer multiple and can be obtained from the bitstream.

[0264] Adaptive Parameter Set Signaling

[0265] Figure 15 This is a view illustrating a portion of the syntax of the Adaptive Parameter Set (APS) according to an implementation. The Adaptive Parameter Set (APS) is a set of parameters used to transmit information for encoding the screen. In the following text, reference will be made to... Figure 15 The description is a syntax element that can be signaled via APS.

[0266] Each APS RBSP can be decoded before being referenced and can be included in the TemporalId of the coded slice NAL unit whose TemporalId is equal to or less than that of the referenced unit, or in at least one AU provided by an external device.

[0267] The adaptation_parameter_set_id can provide an identifier for the APS that will be referenced from another syntax element.

[0268] When the value of aps_params_type is equal to ALF_APS or SCALING_APS, the value of adaptation_parameter_set_id can be in the range of 0 to 7 (inclusive).

[0269] When the value of aps_params_type is equal to LMCS_APS, the value of adaptation_parameter_set_id can be in the range of 0 to 3 (inclusive).

[0270] Suppose that the nuh_layer_id value of a particular APS NAL cell is apsLayerId, and the nuh_layer_id value of a particular VCL NAL cell is vclLayerId. In this case, unless the layer whose apsLayerId is equal to or less than vclLayerId and whose nuh_layer_id is equal to apsLayerId is included in at least one set of output layers that includes the layer whose nuh_layer_id is equal to vclLayerId, the particular VCL NAL cell will not reference the particular APS NAL cell.

[0271] aps_params_type can specify the type of APS parameters delivered by APS, as shown in the table below.

[0272] [Table 1]

[0273]

[0274] Parameter reference issues The following problems may occur according to the above implementation method.

[0275] Question 1. Even when layer B is not a direct or indirect reference layer to layer A, a frame belonging to layer A, which is a predetermined layer, can still reference parameters belonging to layer B, which is another layer. In this case, even when there is no inter-layer reference, the occurrence of parameter reference may lead to the problem of obtaining parameter reference results with unexpected values.

[0276] Question 2. For single-layer bitstreams, a VPS may not be required. However, even in this case, reservation information about the VPS can be referenced using vps_layer_id[0]. Such a value needs to be derived so that the decoding device can avoid using an incorrect value even if a VPS is not provided.

[0277] Question 3. When the nuh_layer_id of the referenced parameter set is less than the nuh_layer_id of the referenced frame, parameter sets such as SPS and PPS can be shared across layers in the implementation. However, when the following constraints are imposed, reference frame resampling is not allowed, and parameter set sharing may not be properly performed for multi-layer bitstreams with different frame resolutions at different layers.

[0278] - When the value of res_change_in_clvs_allowed_flag is equal to 0, the value of pic_width_in_luma_samples will be equal to pic_width_max_in_luma_samples.

[0279] - When the value of res_change_in_clvs_allowed_flag is equal to 0, the value of pic_height_in_luma_samples will be equal to pic_height_max_in_luma_samples.

[0280] Improvement methods

[0281] The following embodiments provide improved methods for addressing the aforementioned problems. These embodiments can be implemented individually, or at least some of them can be combined and implemented.

[0282] For problems 1 and 2, the above implementation method can be improved as follows.

[0283] Improved method 1. When a frame reference with the same nuh_layer_id as layerA has the same nuh_layer_id as layerB, and layerA and layerB are not the same, layerB will be a direct or indirect reference layer of layerA.

[0284] Improved Method 2. When a frame reference with the same nuh_layer_id as layerA has the same nuh_layer_id as layerB, and layerA and layerB are not the same, the current OLS (e.g., the OLS currently being decoded) will include both layerA and layerB.

[0285] Improved Method 3. When a frame reference with the same nuh_layer_id as layerA has the same nuh_layer_id as layerB, and layerA and layerB are not the same, each OLS that includes layerA will also include layerB.

[0286] Improved Method 4. When no VPS is provided (e.g., the value of sps_video_parameter_set_id is 0), the value of vps_layer_id[0] will be derived to be equal to nuh_layer_id of the NAL unit including the SPS.

[0287] Improvement Method 5. Alternatively, when no VPS is provided (e.g., the value of sps_video_parameter_set_id is 0), the value of vps_layer_id[0] will be derived to be equal to 0.

[0288] To improve problems 1 and 3, the following improvement methods are applicable. These methods can be used individually or in combination.

[0289] Improvement Method 6. When Reference Frame Resampling (RPR) is not enabled, all frames in CLVS will have the same size. For example, when RPR is not enabled, the frame sizes signaled by the signal (e.g., pic_width_in_luma_samples and pic_height_in_luma_samples) will have the same value in all PPS referencing the same SPS.

[0290] Improvement Method 7. When all of the following conditions are true, the maximum frame size notified by signaling in the reference SPS (e.g., pic_width_max_in_luma_samples and pic_height_max_in_luma_samples) is the same, and the frame size notified by signaling may be limited.

[0291] - RPR is unavailable.

[0292] - The nuh_layer_id of the referenced SPS and PPS are therefore the same.

[0293] - The layer is an independently encoded layer; for example, the value of vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 1.

[0294] Improved method 8. When a frame reference with the same nuh_layer_id as layerA has the same nuh_layer_id as layerB, and layerA and layerB are not the same, layerB will be a direct or indirect reference layer of layerA.

[0295] Implementation Method 1

[0296] As an implementation method for improving method 1, the following method is applicable.

[0297] It is possible Figure 16are: variable NumDirectRefLayers[i] indicating the number of direct reference layers of a layer with an i-th index, variable DirectRefLayerIdx[i][d] indicating the direct reference layer of the layer with the i-th index represented by an index d from 0 to NumDirectRefLayers[i], variable NumRefLayers[i] indicating the number of direct reference layers and indirect reference layers of the layer with the i-th index, variable RefLayerIdx[i][r] indicating the direct reference layers and indirect reference layers of the layer with the i-th index represented by an index r from 1 to NumRefLayers[i], variable DependencyFlag[i][j] indicating whether a layer with index i directly or indirectly references a layer with index j, and variable LayerUsedAsRefLayerFlag[j] indicating whether the layer with index j is referenced from another layer.

[0298] In addition, SPS constraint 1 related to the above syntax element sps_seq_parameter_set_id may be replaced and applied together with SPS constraint 2.

[0299] <SPS constraint 1>

[0300] Assume that a value of nuh_layer_id of a specific SPS NAL unit is spsLayerId, and a value of nuh_layer_id of a specific VCL NAL unit is vclLayerId. In this case, the specific VCL NAL unit shall not reference the specific SPS NAL unit unless spsLayerId is equal to or less than vclLayerId and the layer with nuh_layer_id equal to spsLayerId is included in at least one output layer set that includes the layer with nuh_layer_id equal to vclLayerId.

[0301] <SPS constraint 2>

[0302] Assume that a value of nuh_layer_id of a predetermined SPS NAL unit is spsLayerId, and a value of nuh_layer_id of a predetermined VCL NAL unit is vclLayerId. When all the following conditions are not satisfied, the predetermined VCL NAL unit shall not reference the predetermined SPS NAL unit.

[0303] - spsLayerId is equal to or less than vclLayerId.

[0304] - A layer with nuh_layer_id equal to spsLayerId is included in at least one OLS that includes a layer with nuh_layer_id equal to vclLayerId.

[0305] - The value of DependencyFlag[vclLayerId][spsLayerId] is equal to 1.

[0306] In addition, the PPS 1 constraint related to the above syntax element pps_pic_parameter_set_id may be replaced and applied together with the PPS 2 constraint.

[0307] <PPS constraint 1>

[0308] Assume that the value of nuh_layer_id of a specific PPS NAL unit is ppsLayerId, and the value of nuh_layer_id of a specific VCL NAL unit is vclLayerId. In this case, the specific VCL NAL unit shall not reference the specific PPS NAL unit unless ppsLayerId is less than or equal to vclLayerId and a layer with nuh_layer_id equal to ppsLayerId is included in at least one output layer set that includes a layer with nuh_layer_id equal to vclLayerId.

[0309] <PPS constraint 2>

[0310] Assume that the value of nuh_layer_id of a specific PPS NAL unit is ppsLayerId, and the value of nuh_layer_id of a specific VCL NAL unit is vclLayerId. In this case, the specific VCL NAL unit shall not reference the specific PPS NAL unit unless all of the following conditions are satisfied.

[0311] - ppsLayerId is less than or equal to vclLayerId.

[0312] - A layer with nuh_layer_id equal to ppsLayerId is included in at least one OLS that includes a layer with nuh_layer_id equal to vclLayerId.

[0313] - The value of DependencyFlag[vclLayerId][ppsLayerId] is equal to 1.

[0314] As described in this embodiment, by replacing and imposing constraints on SPS and PPS, when a picture having the same nuh_layer_id as layerA references a parameter set having the same nuh_layer_id as layerB and layerA and layerB are different, layerB will be a direct or indirect reference layer of layerA. Therefore, the aforementioned problem 1 and problem 2 can be solved.

[0315] Implementation Method 2

[0316] As another embodiment for implementing improved method 1, the following method is applicable.

[0317] First, as described in embodiment 1, as Figure 16 shown, the variables NumDirectRefLayers[i], DirectRefLayerIdx[i][d], NumRefLayers[i], RefLayerIdx[i][r], DependencyFlag[i][j] and LayerUsedAsRefLayerFlag[j] can be derived.

[0318] In addition, SPS constraint 1 related to the above-mentioned syntax element sps_seq_parameter_set_id can be replaced and applied together with SPS constraint 3.

[0319] <SPS constraint 3>

[0320] Assume that the value of nuh_layer_id of a predetermined SPS NAL unit is spsLayerId, and the value of nuh_layer_id of a predetermined VCL NAL unit is vclLayerId. When not all of the following conditions are satisfied, the predetermined VCL NAL unit shall not reference the predetermined SPS NAL unit.

[0321] spsLayerId is equal to vclLayerId.

[0322] The value of DependencyFlag[vclLayerId][spsLayerId] is equal to 1.

[0323] In addition, PPS constraint 1 related to the above-mentioned syntax element pps_pic_parameter_set_id can be replaced and applied together with PPS constraint 3.

[0324] <PPS constraint 3>

[0325] Assume that the value of nuh_layer_id of a specific PPS NAL unit is ppsLayerId, and the value of nuh_layer_id of a specific VCL NAL unit is vclLayerId. In this case, the specific VCL NAL unit shall not reference the specific PPS NAL unit unless all the following conditions are met.

[0326] - ppsLayerId is equal to ppsLayerId.

[0327] - the value of DependencyFlag[vclLayerId][ppsLayerId] is equal to 1.

[0328] In addition, APS constraint 1 related to the above syntax element adaptation_parameter_set_id may be replaced and applied together with APS constraint 2.

[0329] <APS Constraint 1>

[0330] Assume that the value of nuh_layer_id of a specific APS NAL unit is apsLayerId, and the value of nuh_layer_id of a specific VCL NAL unit is vclLayerId. In this case, the specific VCL NAL unit shall not reference the specific APS NAL unit unless apsLayerId is less than or equal to vclLayerId and the layer with nuh_layer_id equal to apsLayerId is included in at least one output layer set including the layer with nuh_layer_id equal to vclLayerId.

[0331] <APS Constraint 2>

[0332] Assume that the value of nuh_layer_id of a specific APS NAL unit is apsLayerId, and the value of nuh_layer_id of a specific VCL NAL unit is vclLayerId. In this case, the specific VCL NAL unit shall not reference the specific APS NAL unit unless all the following conditions are met.

[0333] - apsLayerId is equal to vclLayerId.

[0334] - the value of DependencyFlag[vclLayerId][apsLayerId] is equal to 1.

[0335] As described in this embodiment, by replacing and imposing constraints on SPS, PPS and APS, when a picture having the same nuh_layer_id as layerA references a parameter set having the same nuh_layer_id as layerB and layerA and layerB are different, layerB will be a direct or indirect reference layer of layerA. Therefore, the aforementioned problem 1 and problem 2 can be solved.

[0336] Implementation Method 3

[0337] As an embodiment for implementing improvement 2 and improvement 3, the following method is applicable. For example, constraint SPS1 related to the aforementioned syntax element sps_seq_parameter_set_id can be replaced and applied together with SPS constraint 4 or SPS constraint 5.

[0338] <SPS constraint 4>

[0339] Assume that the value of nuh_layer_id of a predetermined SPS NAL unit is spsLayerId, and the value of nuh_layer_id of a predetermined VCL NAL unit is vclLayerId. When all the following conditions are not satisfied, the predetermined VCL NAL unit shall not reference the predetermined SPS NAL unit.

[0340] - spsLayerId is less than or equal to vclLayerId.

[0341] - the currently decoded OLS includes a layer with nuh_layer_id equal to spsLayerId and a layer with nuh_layer_id equal to vclLayerId.

[0342] <SPS constraint 5>

[0343] Assume that the value of nuh_layer_id of a predetermined SPS NAL unit is spsLayerId, and the value of nuh_layer_id of a predetermined VCL NAL unit is vclLayerId. When all the following conditions are not satisfied, the predetermined VCL NAL unit shall not reference the predetermined SPS NAL unit.

[0344] - spsLayerId is less than or equal to vclLayerId.

[0345] - for each OLS specified by the VPS that includes a layer with nuh_layer_id equal to vclLayerId, the OLS includes a layer with nuh_layer_id equal to spsLayerId.

[0346] In addition, the constraint PPS1 related to the aforementioned syntax element pps_pic_parameter_set_id may be replaced and applied in conjunction with PPS constraint 4 or PPS constraint 5.

[0347] <PPS constraint 4>

[0348] Assuming that the value of nuh_layer_id of a specific PPS NAL unit is ppsLayerId, and the value of nuh_layer_id of a specific VCL NAL unit is vclLayerId. In this case, the specific VCL NAL unit shall not reference the specific PPS NAL unit unless all of the following conditions are satisfied.

[0349] - ppsLayerId is less than or equal to vclLayerId.

[0350] - The currently decoded OLS includes a layer with nuh_layer_id equal to ppsLayerId and a layer with nuh_layer_id equal to vclLayerId.

[0351] <PPS constraint 5>

[0352] Assuming that the value of nuh_layer_id of a specific PPS NAL unit is ppsLayerId, and the value of nuh_layer_id of a specific VCL NAL unit is vclLayerId. In this case, the specific VCL NAL unit shall not reference the specific PPS NAL unit unless all of the following conditions are satisfied.

[0353] - ppsLayerId is less than or equal to vclLayerId.

[0354] - For each OLS specified by the VPS that includes the layer with nuh_layer_id equal to vclLayerId, the OLS includes the layer with nuh_layer_id equal to ppsLayerId.

[0355] In addition, the APS constraint 1 related to the aforementioned syntax element adaptation_parameter_set_id may be replaced and applied in conjunction with APS constraint 3 or APS constraint 4.

[0356] <APS constraint 3>

[0357] It is assumed that the value of nuh_layer_id of a specific APS NAL unit is apsLayerId, and the value of nuh_layer_id of a specific VCL NAL unit is vclLayerId. In this case, the specific VCL NAL unit shall not reference the specific APS NAL unit unless all the following conditions are satisfied.

[0358] - apsLayerId is less than or equal to vclLayerId.

[0359] - The currently decoded OLS includes the layer with nuh_layer_id equal to apsLayerId and the layer with nuh_layer_id equal to vclLayerId.

[0360] <APS Constraint 4>

[0361] It is assumed that the value of nuh_layer_id of a specific APS NAL unit is apsLayerId, and the value of nuh_layer_id of a specific VCL NAL unit is vclLayerId. In this case, the specific VCL NAL unit shall not reference the specific APS NAL unit unless all the following conditions are satisfied.

[0362] - apsLayerId is less than or equal to vclLayerId.

[0363] - For each OLS specified by the VPS that includes the layer with nuh_layer_id equal to vclLayerId, the OLS includes the layer with nuh_layer_id equal to apsLayerId.

[0364] In an implementation mode, constraint 4 related to SPS, constraint 4 related to PPS and constraint 3 related to APS may be applied together. In another implementation mode, constraint 5 related to SPS, constraint 5 related to PPS and constraint 4 related to APS may be applied together.

[0365] As described in this embodiment, by replacing and imposing constraints regarding SPS, PPS, and APS, when a frame reference with the same nuh_layer_id as layer A has the same nuh_layer_id as layer B and layer A and layer B are not the same, the current OLS (e.g., the currently decoded OLS) will include layer A and layer B; or when a frame reference with the same nuh_layer_id as layer A has the same nuh_layer_id as layer B and layer A and layer B are not the same, each OLS including layer A will also include layer B. Therefore, problems 1 and 2 described above can be solved.

[0366] Implementation Method 4

[0367] As an implementation method for improving method 4, the following methods are applicable. For example, when the value of sps_video_parameter_set_id is 0, the following applies.

[0368] - SPS does not reference VPS.

[0369] - During the decoding process of each CLVS of the reference SPS, no VPS is referenced.

[0370] The value of - vps_max_layers_minus1 is deduced to be equal to 0.

[0371] - CVS consists of only one layer. For example, all VCL NAL units in CVS will have the same nuh_layer_id value.

[0372] The value of GeneralLayerIdx[nuh_layer_id] can be deduced to be equal to 0.

[0373] The value of - vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] can be deduced to be equal to 1.

[0374] The value of - vps_layer_id[0] can be deduced to be equal to nuh_layer_id.

[0375] As described in this embodiment, by adding constraints regarding the SPS, when no VPS is provided, the value of vps_layer_id[0] will be derived to be equal to nuh_layer_id of the NAL unit including the SPS. Therefore, problems 1 and 2 mentioned above can be solved.

[0376] Implementation Method 5

[0377] The following methods are applicable as implementation methods for improving methods 6, 7 and 8.

[0378] First, as described in Implementation 1, such as Figure 16 As shown, the variables NumDirectRefLayers[i], DirectRefLayerIdx[i][d], NumRefLayers[i], RefLayerIdx[i][r], DependencyFlag[i][j] and LayerUsedAsRefLayerFlag[j] can be derived.

[0379] Additionally, SPS constraint 1, which is associated with the above syntax element sps_seq_parameter_set_id, can be replaced and applied together with SPS constraint 2 as described in Implementation 1.

[0380] Additionally, the PPS constraint 1 associated with the above syntax element pps_pic_parameter_set_id can be replaced and applied together with the PPS constraint 2 as described in Implementation 1.

[0381] Alternatively, unlike the above description, the value of pic_width_in_luma_samples can be derived as follows.

[0382] For example, when the value of res_change_in_clvs_allowed_flag is equal to 0, the value of pic_width_in_luma_samples will have the same value for all PPS of the encoded picture reference in CLVS.

[0383] The value of pic_width_in_luma_samples will be equal to the value of pic_width_max_in_luma_samples when all of the following conditions are true.

[0384] The value of res_change_in_clvs_allowed_flag is equal to 0.

[0385] ppsLayerId is equal to the nuh_layer_id value of the referenced SPS.

[0386] The value of vps_independent_layer_flag[GeneralLayerIdx[ppsLayerId]] is equal to 1.

[0387] Alternatively, unlike the above description, the value of pic_height_in_luma_samples can be derived as follows.

[0388] For example, when the value of res_change_in_clvs_allowed_flag is equal to 0, the value of pic_height_in_luma_samples will have the same value for all PPS of the encoded picture reference in CLVS.

[0389] When the value of res_change_in_clvs_allowed_flag is equal to 0 and the value of vps_independent_layer_flag[GeneralLayerIdx[ppsLayerId]] is equal to 1, the value of pic_height_in_luma_samples will be equal to the value of pic_height_max_in_luma_samples.

[0390] As described in this embodiment, improved methods 6 to 8 are applicable by adding constraints regarding SPS and PPS. Therefore, problems 1 and 3 described above can be solved.

[0391] Encoding and Decoding Methods

[0392] In the following, an image encoding method and an image decoding method performed by an image encoding device and an image decoding device according to an embodiment will be described.

[0393] Figure 17 This is a flowchart illustrating a method for decoding an image using an image decoding device by determining the availability of a reference parameter set according to an embodiment. The image decoding device according to the embodiment may include a memory and a processor, and the decoding device may perform decoding via the operation of the processor according to the embodiment described below.

[0394] First, the decoding device can obtain the parameter set and image encoded data from the bitstream (S1710). Next, the decoding device can determine the reference availability of the parameter set used to decode the image encoded data (S1720). Next, the decoding device can decode the image encoded data based on the reference availability (S1730).

[0395] Here, the decoding device can determine reference availability based on whether the output layer set, which includes the layer corresponding to the image-coded data, includes a predetermined layer. Furthermore, the layer corresponding to the image-coded data can be a layer whose layer identifier is equal to the layer identifier of the Video Coding Layer (VCL) Network Abstraction Layer (NAL) corresponding to the image-coded data. The output layer set can be determined based on the Video Parameter Set (VPS). For example, the VPS can signal information about the output layer set as individual syntax elements.

[0396] A predetermined layer can be a layer corresponding to a parameter set. For example, a predetermined layer is a layer whose layer identifier is equal to the layer identifier of the NAL unit corresponding to the parameter set, and the parameter set can be at least one of the Sequence Parameter Set (SPS), Picture Parameter Set (PPS), or Adaptive Parameter Set (APS).

[0397] Layer identifiers can be represented by the nuh_layer_id mentioned above. For example, the layer identifier for image-coded data can be the nuh_layer_id of the VCL NAL unit corresponding to the image-coded data. The layer identifier for SPS can be the nuh_layer_id of the SPS NAL unit. The layer identifier for PPS can be the nuh_layer_id of the PPS NAL unit. The layer identifier for APS can be the nuh_layer_id of the APS NAL unit.

[0398] The decoding device can determine reference availability based on the following criteria. This will refer to... Figure 18 The decoding device can determine whether to reference the parameter set based on whether a first condition is met (S1810). Here, the first condition may be whether the value of the layer identifier corresponding to the parameter set is not greater than the value of the layer identifier corresponding to the image encoded data. The decoding device can determine that the parameter set cannot be used for reference to decode the image encoded data when the value of the layer identifier corresponding to the parameter set is greater than the value of the layer identifier corresponding to the image encoded data (S1840).

[0399] Furthermore, when the value of the layer identifier corresponding to the parameter set is not greater than the value of the layer identifier corresponding to the image encoded data, the decoding device can determine whether the second condition (S1820) is satisfied. Here, the second condition may be whether the set of all output layers including layers whose layer identifiers are equal to the layer identifiers corresponding to the image encoded data includes layers whose layer identifiers are equal to the layer identifiers corresponding to the parameter set.

[0400] For example, when the set of all output layers, including layers whose layer identifiers are equal to the layer identifiers corresponding to the image encoded data, includes layers whose layer identifiers are equal to the layer identifiers corresponding to the parameter set, the decoding device can determine whether the parameter set can be used as a reference for decoding the image encoded data (S1830). If not, the decoding device can determine that the parameter set cannot be used as a reference for decoding the image encoded data (S1840).

[0401] For example, if the value of the layer identifier corresponding to the parameter set is greater than the value of the layer identifier corresponding to the image-coded data, it can be determined that the parameter set cannot be used as a reference for decoding the image-coded data. Similarly, if at least one of the sets of all output layers, including layers whose layer identifiers are equal to those corresponding to the image-coded data, does not include layers whose layer identifiers are equal to those corresponding to the parameter set, it can be determined that the parameter set cannot be used as a reference for decoding the image-coded data.

[0402] Furthermore, based on the fact that the value of the second layer identifier corresponding to the parameter set is no greater than the value of the first layer identifier corresponding to the image encoded data, and that the set of all output layers including layers whose layer identifiers are equal to the first layer identifiers corresponding to the image encoded data include layers whose layer identifiers are equal to the second layer identifiers, it can be determined that the parameter set can be used as a reference for decoding the image encoded data.

[0403] In an implementation, the parameter set may include a sequence parameter set (SPS), a picture parameter set (PPS), and an adaptive parameter set (APS), and the reference availability of the parameter set may be determined individually for SPS, PPS, and APS.

[0404] For example, based on the fact that the value of the second layer identifier corresponding to the SPS is no greater than the value of the first layer identifier corresponding to the image encoded data, and that the set of all output layers including layers whose layer identifiers are equal to the first layer identifiers includes layers whose layer identifiers are equal to the second layer identifiers, it can be determined that the SPS can be used as a reference for decoding the image encoded data.

[0405] Additionally (or alternatively), based on the fact that the value of the third layer identifier corresponding to the PPS is no greater than the value of the first layer identifier corresponding to the image encoded data, and that the set of all output layers including layers whose layer identifiers are equal to the first layer identifiers includes layers whose layer identifiers are equal to the third layer identifiers, it can be determined that the PPS can be used as a reference for decoding the image encoded data.

[0406] Additionally (or alternatively), based on the fact that the value of the fourth layer identifier corresponding to the APS is no greater than the value of the first layer identifier corresponding to the image-coded data, and that the set of all output layers including layers whose layer identifiers are equal to the first layer identifiers includes layers whose layer identifiers are equal to the fourth layer identifiers, it can be determined that the APS can be used as a reference for decoding the image-coded data.

[0407] For example, in one implementation, the image decoding method may include the following steps: obtaining Sequence Parameter Set (SPS) Network Abstraction Layer (NAL) units from the bitstream; obtaining Video Parameter Set (VPS) Network Abstraction Layer (NAL) units from the bitstream; obtaining Picture Parameter Set (PPS) Network Abstraction Layer (NAL) units from the bitstream; obtaining Adaptive Parameter Set (APS) Network Abstraction Layer (NAL) units from the bitstream; obtaining Video Coding Layer (VCL) Network Abstraction Layer (NAL) units from the bitstream; and determining, based on the VCL NAL units, whether the VCL NAL units do not reference at least one of the SPS NAL units, PPS NAL units, and APS NAL units. Here, determining whether the VCL NAL units do not reference can be performed based on the layer identifier of the VCL NAL units and whether the set of all output layers in the output layer set identified by the VPS NAL units, including the layer corresponding to the layer identifier of the VCL NAL units, also includes a predetermined layer.

[0408] For example, whether a VCL NAL cell does not reference an SPS NAL cell can be determined based on whether the value of the SPS NAL cell's layer identifier is less than or equal to the value of the VCL NAL cell's layer identifier, and whether the set of all output layers in the output layer set identified by the VPS NAL cell, which includes the layer corresponding to the VCL NAL cell's layer identifier, also includes the layer corresponding to the SPS NAL cell's layer identifier.

[0409] Furthermore, whether a VCL NAL cell does not reference a PPS NAL cell can be determined based on whether the value of the layer identifier of the PPS NAL cell is less than or equal to the value of the layer identifier of the VCL NAL cell, and whether the set of all output layers in the output layer set identified by the VPS NAL cell, which includes the layer corresponding to the layer identifier of the VCL NAL cell, also includes the layer corresponding to the layer identifier of the PPS NAL cell.

[0410] Furthermore, whether a VCL NAL cell does not reference an APS NAL cell can be determined based on whether the value of the layer identifier of the APS NAL cell is less than or equal to the value of the layer identifier of the VCL NAL cell, and whether the set of all output layers in the output layer set identified by the VPS NAL cell, which includes the layer corresponding to the layer identifier of the VCL NAL cell, also includes the layer corresponding to the layer identifier of the APS NAL cell.

[0411] Figure 19 This is a flowchart illustrating a method for encoding an image by an image encoding device according to an embodiment of determining the availability of a parameter set. The image encoding device according to the embodiment includes a memory and a processor, and the encoding device can perform encoding according to a method corresponding to the decoding method described below through the operation of the processor.

[0412] For example, the encoding device can encode an image to generate image-encoded data for a portion of the image and a parameter set for the encoded data (S1910). Additionally, the encoding device can generate a bitstream including the image-encoded data and the parameter set (S1920). Here, the parameter set can be generated based on a reference availability of the parameter set used for decoding the image-encoded data. Furthermore, the reference availability can be determined based on whether the output layer set, which includes the layer corresponding to the image-encoded data, includes a predetermined layer.

[0413] More specifically, corresponding to the above decoding method, the video parameter set (VPS) includes information about the set of output layers. The predetermined layer is the layer corresponding to the parameter set. The layer corresponding to the image coding data is the layer whose layer identifier is equal to the layer identifier of the video coding layer (VCL) network abstraction layer (NAL) unit corresponding to the image coding data. The predetermined layer is the layer whose layer identifier is equal to the layer identifier of the NAL unit corresponding to the parameter set. The parameter set can be at least one of the sequence parameter set (SPS), picture parameter set (PPS), or adaptive parameter set (APS).

[0414] For example, in one implementation, the image encoding method may include the following steps: generating Video Coding Layer (VCL) Network Abstraction Layer (NAL) units, Sequence Parameter Set (SPS) Network Abstraction Layer (NAL) units, Video Parameter Set (VPS) Network Abstraction Layer (NAL) units, Picture Parameter Set (PPS) Network Abstraction Layer (NAL) units, and Adaptive Parameter Set (APS) Network Abstraction Layer (NAL) units by encoding an image; and generating a bitstream including VCL NAL units, SPS NAL units, VPS NAL units, PPS NAL units, and APS NAL units.

[0415] Here, the layer identifiers for VPS NAL units and VCL NAL units can be generated based on whether the VCL NAL unit references at least one of the SPS NAL unit, PPS NAL unit, and APS NAL unit.

[0416] Furthermore, VPS NAL units can be generated such that, based on whether the VCL NAL unit references at least one of SPS NAL units, PPSNAL units, and APS NAL units, the set of all output layers identified by the VPS NAL unit, including the layer corresponding to the layer identifier of the VCL NAL unit, also includes a predetermined layer. For example, the predetermined layer can be a layer whose layer identifier is equal to the layer identifier corresponding to at least one of SPS, PPS, and APS.

[0417] Application and Implementation Methods

[0418] Although the exemplary methods of this disclosure described above are represented as a series of operations for clarity of description, they are not intended to limit the order in which the steps are performed, and these steps may be performed simultaneously or in different orders if necessary. To implement the methods according to this disclosure, the described steps may further include other steps, including steps in addition to some steps, or may include additional steps in addition to some steps.

[0419] In this disclosure, an image encoding device or an image decoding device that performs a predetermined operation (step) can perform an operation (step) that confirms the execution conditions or circumstances of the corresponding operation (step). For example, if it is described that a predetermined operation is performed when predetermined conditions are met, the image encoding device or the image decoding device can perform the predetermined operation after determining whether the predetermined conditions are met.

[0420] The various embodiments of this disclosure are not a list of all possible combinations and are intended to describe representative aspects of this disclosure; the matters described in the various embodiments may be applied independently or in combination of two or more.

[0421] Various embodiments of this disclosure can be implemented in hardware, firmware, software, or a combination thereof. When this disclosure is implemented in hardware, it can be implemented using application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, etc.

[0422] Furthermore, the image decoding and image encoding devices applying the embodiments of this disclosure can be included in multimedia broadcasting transmitting and receiving devices, mobile communication terminals, home theater video devices, digital cinema video devices, surveillance cameras, video chat devices, real-time communication devices such as video communication, mobile streaming devices, storage media, cameras, video-on-demand (VoD) service providers, OTT (over-the-top) video devices, internet streaming service providers, three-dimensional (3D) video devices, video telephony devices, medical video devices, etc., and can be used to process video signals or data signals. For example, OTT video devices can include game consoles, Blu-ray players, internet access televisions, home theater systems, smartphones, tablet PCs, digital video recorders (DVRs), etc.

[0423] Figure 20 This is a view illustrating a content streaming system to which embodiments of the present disclosure can be applied.

[0424] like Figure 20 As shown, the content streaming system using the embodiments of this disclosure may mainly include an encoding server, a streaming server, a network server, a media storage device, a user device, and a multimedia input device.

[0425] An encoding server compresses content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream, which is then sent to a streaming server. As another example, when multimedia input devices such as smartphones, cameras, and camcorders directly generate bitstreams, the encoding server can be omitted.

[0426] The bitstream can be generated by an image encoding method or image encoding device applying the embodiments of this disclosure, and the stream server can temporarily store the bitstream during the sending or receiving of the bitstream.

[0427] A streaming server sends multimedia data to a user's device based on a request from a web server, and the web server acts as a medium for informing the user of the service. When a user requests a service from the web server, the web server can deliver it to the streaming server, and the streaming server can send the multimedia data to the user. In this scenario, the content streaming system may include a separate control server. In this case, the control server is used to control the commands / responses between devices in the content streaming system.

[0428] A streaming server can receive content from media storage devices and / or encoding servers. For example, when receiving content from an encoding server, the content can be received in real time. In this case, to provide a smooth streaming service, the streaming server can store the bitstream for a predetermined period of time.

[0429] Examples of user devices may include mobile phones, smartphones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, tablet computers, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head-mounted displays), digital televisions, desktop computers, digital signage, etc.

[0430] In a content streaming system, each server can operate as a distributed server, in which case the data received from each server can be distributed.

[0431] The scope of this disclosure includes software or machine-executable commands (e.g., operating systems, applications, firmware, programs, etc.) for enabling the operation of methods according to various embodiments to be executed on a device or computer, and non-transitory computer-readable media having such software or commands stored thereon and executable on a device or computer.

[0432] Industrial applicability

[0433] The embodiments disclosed herein can be used to encode or decode images.

Claims

1. An image decoding method performed by an image decoding device, the image decoding method comprising the following steps: Obtain parameter sets and image encoded data from the bitstream; Determine the reference availability of the parameter set used to decode the image encoded data; as well as The image encoded data is decoded based on the reference availability. The reference availability is determined based on whether each output layer set specified by the video parameter set VPS includes the layer in the output layer set corresponding to the image encoded data and also includes the layer corresponding to the parameter set, or based on information obtained from the bitstream. Wherein, the layer corresponding to the image encoded data is a layer whose layer identifier is equal to the layer identifier of the Video Coding Layer (VCL) Network Abstraction Layer (NAL) unit corresponding to the image encoded data. Wherein, the layer corresponding to the parameter set is a layer whose layer identifier is equal to the layer identifier of the NAL unit corresponding to the parameter set, and Wherein, based on the information, the parameter set is specified not to reference the VPS, and the layer identifier of the first layer is inferred to be the same as the layer identifier of the VCL NAL unit.

2. The image decoding method according to claim 1, wherein, Based on the fact that the value of the layer identifier corresponding to the parameter set is greater than the value of the layer identifier corresponding to the image encoded data, it is determined that the parameter set cannot be used as a reference for decoding the image encoded data.

3. The image decoding method according to claim 1, wherein, Based on the fact that at least one of the sets of all output layers, including layers whose layer identifiers are equal to the layer identifiers corresponding to the image encoded data, does not include layers whose layer identifiers are equal to the layer identifiers corresponding to the parameter set, it is determined that the parameter set cannot be used as a reference for decoding the image encoded data.

4. The image decoding method according to claim 1, wherein, Based on the fact that the value of the second layer identifier corresponding to the parameter set is not greater than the value of the first layer identifier corresponding to the image encoded data, and including all output layer sets including layers whose layer identifiers are equal to the first layer identifiers, including layers whose layer identifiers are equal to the second layer identifiers, the parameter set is determined to be usable for reference in decoding the image encoded data.

5. The image decoding method according to claim 1, in, The parameter set is at least one of the sequence parameter set SPS, the picture parameter set PPS, and the adaptive parameter set APS.

6. An image encoding method performed by an image encoding device, the image encoding method comprising the following steps: The image is encoded to generate image-encoded data for a portion of the image and a parameter set for the image-encoded data; as well as Generate a bitstream including the image encoded data and the parameter set. The parameter set is generated based on the reference availability of the parameter set used for decoding the image encoded data. The reference availability is determined based on whether each output layer set included in the video parameter set (VPS) includes a layer in the output layer set corresponding to the image encoded data and also includes a layer corresponding to the parameter set, or based on whether the parameter set references the VPS. Wherein, the layer corresponding to the image encoded data is a layer whose layer identifier is equal to the layer identifier of the Video Coding Layer (VCL) Network Abstraction Layer (NAL) unit corresponding to the image encoded data. Wherein, the layer corresponding to the parameter set is a layer whose layer identifier is equal to the layer identifier of the NAL unit corresponding to the parameter set, and Since the parameter set does not reference the VPS, the layer identifier of the first layer is inferred to be the same as the layer identifier of the VCL NAL unit.

7. The image encoding method according to claim 6, in, The parameter set is at least one of the sequence parameter set SPS, the picture parameter set PPS, or the adaptive parameter set APS.

8. A method for transmitting a bit stream, the method comprising the following steps: Image encoding is performed to generate the bitstream, the bitstream comprising image-encoded data for a portion of an image and a set of parameters for the image-encoded data; as well as Send the generated bit stream, The parameter set is generated based on the reference availability of the parameter set used for decoding the image encoded data. The reference availability is determined based on whether each output layer set included in the video parameter set (VPS) includes a layer in the output layer set corresponding to the image encoded data and also includes a layer corresponding to the parameter set, or based on whether the parameter set references the VPS. Wherein, the layer corresponding to the image encoded data is a layer whose layer identifier is equal to the layer identifier of the Video Coding Layer (VCL) Network Abstraction Layer (NAL) unit corresponding to the image encoded data. Wherein, the layer corresponding to the parameter set is a layer whose layer identifier is equal to the layer identifier of the NAL unit corresponding to the parameter set, and Since the parameter set does not reference the VPS, the layer identifier of the first layer is inferred to be the same as the layer identifier of the VCL NAL unit.

Citation Information

Patent Citations

  • Image decoding device, image decoding method, recoding medium, image coding device, and image coding method

    US20170019673A1