Image Encoding / Decoding Method and Apparatus Based on Dependent Random Access Point Picture, and Bitstream Transmission Method
The image encoding/decoding method and apparatus address the challenge of high-resolution image data efficiency by employing a GDR picture as a reference for decoding, enhancing encoding/decoding efficiency and reducing transmission/storage costs.
Patent Information
- Application Number
- JP2025503419
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-09-05
- Filing Date
- 2023-09-04
- Publication Date
- 2025-07-17
AI Technical Summary
The increasing demand for high-resolution, high-quality images has led to a need for more efficient image compression techniques to reduce transmission and storage costs, as conventional methods face challenges with the increased amount of information required.
An image encoding/decoding method and apparatus that utilize a dependent random access point picture, specifically a GDR picture with a zero recovery point, to enhance encoding/decoding efficiency and enable random access.
Improves encoding/decoding efficiency and enables efficient transmission and storage of high-resolution images by utilizing a GDR picture as a reference for decoding, reducing data requirements and costs.
Smart Images

Figure 2025523262000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an image encoding / decoding method and apparatus, and a recording medium for storing a bitstream. More particularly, the present disclosure relates to an image encoding / decoding method and apparatus based on a dependent random access point picture, and a method for transmitting a bitstream generated by the image encoding method / apparatus of the present disclosure.
Background Art
[0002] Recently, the demand for high-resolution, high-quality images, such as HD (High Definition) images and UHD (Ultra High Definition) images, has been increasing in various fields. As the image data becomes higher in resolution and quality, the amount of information or bits to be transmitted increases relatively compared to conventional image data. The increase in the amount of information or bits to be transmitted results in an increase in transmission costs and storage costs.
[0003] Accordingly, there is a need for a highly efficient image compression technique for effectively transmitting, storing, and reproducing information of high-resolution, high-quality images.
Summary of the Invention
Problems to be Solved by the Invention
[0004] An object of the present disclosure is to provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.
[0005] Another object of the present disclosure is to provide an image encoding / decoding method and apparatus based on a dependent random access point picture.
[0006] Another object of the present disclosure is to provide an image encoding / decoding method and apparatus for performing random access to a dependent random access point picture based on a GDR picture having a zero recovery point picture.
[0007] In addition, an object of the present disclosure is to provide a non-transitory computer-readable recording medium that stores a bitstream generated by an image encoding method or apparatus according to the present disclosure.
[0008] In addition, an object of the present disclosure is to provide a non-transitory computer-readable recording medium that stores a bitstream received by an image decoding apparatus according to the present disclosure, decoded, and used for restoring an image.
[0009] In addition, an object of the present disclosure is to provide a method for transmitting a bitstream generated by an image encoding method or apparatus according to the present disclosure.
[0010] The technical problems to be solved by the present disclosure are not limited to the above-described technical problems, and other technical problems not described above will be clearly understood by those of ordinary skill in the technical field to which the present disclosure pertains from the following description.
Means for Solving the Problems
[0011] An image decoding method according to an aspect of the present disclosure includes: obtaining a second picture decoded before the first picture, which is associated with the first picture, in response to a random access request for the first picture; and decoding the first picture with reference to the second picture, where the first picture is a dependent random access point (RAP) picture, and the second picture may be a gradual decoding refresh (GDR) picture having a zero recovery point picture.
[0012] An image decoding apparatus according to another aspect of the present disclosure includes a memory and at least one processor. The at least one processor responds to a random access request for a first picture, obtains a second picture associated with the first picture and decoded before the first picture, and decodes the first picture with reference to the second picture. The first picture is a dependent random access point (RAP) picture, and the second picture may be a gradual decoding refresh (GDR) picture having a zero recovery point picture.
[0013] An image encoding method according to another aspect of the present disclosure includes a step of deriving a second picture associated with a first picture that is a start point of random access, and a step of encoding information regarding the first picture and the second picture. The first picture is a dependent random access point (RAP) picture, and the second picture may be a gradual decoding refresh (GDR) picture having a zero recovery point picture.
[0014] A computer-readable recording medium according to another aspect of the present disclosure can store a bitstream generated by the image encoding method or the image encoding apparatus of the present disclosure.
[0015] A transmission method according to another aspect of the present disclosure can transmit a bitstream generated by the image encoding apparatus or the image encoding method of the present disclosure.
[0016] The features briefly summarized and described above with respect to the present disclosure are merely exemplary aspects of the detailed description of the present disclosure to be described later, and do not limit the scope of the present disclosure.
Advantages of the Invention
[0017] According to the present disclosure, it is possible to provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.
[0018] Also, according to the present disclosure, it is possible to provide an image encoding / decoding method and apparatus based on a dependent random access point picture.
[0019] Also, according to the present disclosure, it is possible to provide an image encoding / decoding method and apparatus for performing random access to a dependent random access point picture based on a GDR picture having a zero recovery point picture.
[0020] Also, the present disclosure can provide a non-transitory computer-readable recording medium for storing a bitstream generated by the image encoding method or apparatus according to the present disclosure.
[0021] Also, according to the present disclosure, it is possible to provide a non-transitory computer-readable recording medium for storing a bitstream received by the image decoding apparatus according to the present disclosure, decoded, and used for image restoration.
[0022] Also, according to the present disclosure, it is possible to provide a method for transmitting a bitstream generated by an image encoding method or apparatus.
[0023] The effects obtained by the present disclosure are not limited to the above-described effects, and other effects not described above will be clearly understood by those of ordinary skill in the technical field to which the present disclosure pertains from the following description.
Brief Description of the Drawings
[0024]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
DETAILED DESCRIPTION OF THE INVENTION
[0025] Hereinafter, with reference to the accompanying drawings, embodiments of the present disclosure will be described in detail so that those having ordinary knowledge in the technical field to which the present disclosure pertains can easily implement them. However, the present disclosure can be realized in various different forms and is not limited to the embodiments described herein.
[0026] When it is determined that a specific description of a known configuration or function may obscure the gist of the present disclosure in explaining an embodiment of the present disclosure, a detailed description thereof will be omitted. In the drawings, parts not related to the description of the present disclosure are omitted, and the same reference numerals are given to the same parts.
[0027] In the present disclosure, when a component is "connected", "coupled", or "joined" to another component, this can include not only a direct connection relationship but also an indirect connection relationship in which another component exists between them. Also, when a component "includes" or "has" another component, this means that, unless otherwise stated to the contrary, it does not exclude other components but can further include other components.
[0028] In the present disclosure, terms such as "first", "second", etc. are used only for the purpose of distinguishing one component from another and do not limit the order or importance, etc. between components unless otherwise specifically mentioned. Therefore, within the scope of the present disclosure, the first component of one embodiment may be referred to as the second component in another embodiment, and similarly, the second component of one embodiment may be referred to as the first component in another embodiment.
[0029] In the present disclosure, components that are distinct from each other are for the purpose of clearly explaining their respective features and do not necessarily mean that the components are separated. That is, a plurality of components may be integrated and configured as one hardware or software unit, or one component may be distributed and configured as a plurality of hardware or software units. Therefore, without further mention, such integrated or distributed embodiments are also included in the scope of the present disclosure.
[0030] In the present disclosure, the components described in various embodiments do not necessarily mean essential components, and some may be optional components. Therefore, embodiments constituted by a subset of the components described in one embodiment are also included in the scope of the present disclosure. Also, embodiments that further include other components in addition to the components described in various embodiments are included in the scope of the present disclosure.
[0031] The present disclosure relates to image encoding and decoding, and the terms used in the present disclosure can have the ordinary meaning in the technical field to which the present disclosure belongs unless newly defined in the present disclosure.
[0032] In the present disclosure, "video" can mean a collection of a series of images over time.
[0033] In the present disclosure, "picture" generally means a unit indicating any one image in a specific time period. A slice / tile is an encoding unit that constitutes a part of a picture, and one picture can be composed of one or more slices / tiles. Also, a slice / tile can include one or more CTUs (coding tree units).
[0034] In the present disclosure, "pixel" or "pel" can mean the smallest unit that constitutes one picture (or image). Also, the term "sample" can be used as a term corresponding to a pixel. A sample can generally indicate a pixel or the value of a pixel, and can also indicate only the pixel / pixel value of the luma component, or only the pixel / pixel value of the chroma component.
[0035] In the present disclosure, "unit" can indicate the basic unit of image processing. A unit can include at least one of a specific region of a picture and information related to the region. Depending on the case, a unit can be used interchangeably with terms such as "sample array", "block", or "area". In general, an M×N block can include a set (or array) of samples (or sample arrays) or transform coefficients consisting of M columns and N rows.
[0036] In the present disclosure, "current block" can mean any one of "current coding block", "current coding unit", "block to be encoded", "block to be decoded", or "block to be processed". When prediction is performed, "current block" can mean "current prediction block" or "block to be predicted". When transformation (inverse transformation) / quantization (inverse quantization) is performed, "current block" can mean "current transformation block" or "block to be transformed". When filtering is performed, "current block" can mean "block to be filtered".
[0037] Also, in the present disclosure, "current block" can mean a block that includes all luma component blocks and chroma component blocks or "luma block of the current block" unless explicitly stated as a chroma block. The luma component block of the current block can be expressed explicitly including an explicit description of the luma component block such as "luma block" or "current luma block". Also, the chroma component block of the current block can be expressed explicitly including an explicit description of the chroma component block such as "chroma block" or "current chroma block".
[0038] In the present disclosure, " / " and "," can be interpreted as "and / or". For example, "A / B" and "A, B" can be interpreted as "A and / or B". Also, "A / B / C" and "A, B, C" can mean "at least one of A, B, and / or C".
[0039] In the present disclosure, "or" can be interpreted as "and / or". For example, "A or B" can mean 1) only "A", 2) only "B", or 3) "A and B". Alternatively, in the present disclosure, "or" can mean "additionally or alternatively".
[0040] In the present disclosure, "at least one of A, B, and C" can mean "only A", "only B", "only C", or "any and all combinations of A, B, and C". Also, "at least one of A, B, or C" or "at least one of A, B, and / or C" can mean "at least one of A, B, and C".
[0041] The parentheses used in the present disclosure can mean "for example". For example, when displayed as "prediction (intra prediction)", "intra prediction" can be proposed as an example of "prediction". In other words, "prediction" in the present disclosure is not limited to "intra prediction", and "intra prediction" can be proposed as an example of "prediction". Also, when displayed as "prediction (that is, intra prediction)", "intra prediction" can be proposed as an example of "prediction".
[0042] Overview of the Video Coding System
[0043] FIG. 1 is a diagram schematically showing a video coding system to which an embodiment according to the present disclosure can be applied.
[0044] A video coding system according to an embodiment can include an encoding device 10 and a decoding device 20. The encoding device 10 can transmit encoded video and / or image information or data in a file or streaming format to the decoding device 20 via a digital storage medium or a network.
[0045] An encoding device 10 according to an embodiment may include a video source generation unit 11, an encoding unit 12, and a transmission unit 13. A decoding device 20 according to an embodiment may include a reception unit 21, a decoding unit 22, and a rendering unit 23. The encoding unit 12 may be referred to as a video / image encoding unit, and the decoding unit 22 may be referred to as a video / image decoding unit. The transmission unit 13 may be included in the encoding unit 12. The reception unit 21 may be included in the decoding unit 22. The rendering unit 23 may also include a display unit, and the display unit may be configured as a separate device or an external component.
[0046] The video source generation unit 11 can obtain a video / image through processes such as capture, synthesis, or generation of the video / image. The video source generation unit 11 may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, a video / image archive including previously captured video / images, etc. The video / image generation device may include, for example, a computer, a tablet, and a smartphone, etc., and can (electronically) generate a video / image. For example, a virtual video / image can be generated via a computer or the like, and in this case, the video / image capture process can be replaced by a process in which related data is generated.
[0047] The encoding unit 12 can encode the input video / image. The encoding unit 12 can perform a series of procedures such as prediction, transformation, quantization, etc. for compression and encoding efficiency. The encoding unit 12 can output the encoded data (encoded video / image information) in the form of a bitstream.
[0048] The transmission unit 13 can transmit the encoded video / image information or data output in bitstream format to the receiving unit 21 of the decoding device 20 via a digital storage medium or a network in file or streaming format. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray (registered trademark), HDD, SSD, etc. The transmission unit 13 can include elements for generating a media file via a predetermined file format and can include elements for transmission via a broadcast / communication network. The receiving unit 21 can extract / receive the bitstream from the storage medium or the network and transmit it to the decoding unit 22.
[0049] The decoding unit 22 can decode a video / image by performing a series of procedures such as inverse quantization, inverse transformation, and prediction corresponding to the operation of the encoding unit 12.
[0050] The rendering unit 23 can render the decoded video / image. The rendered video / image can be displayed via the display unit.
[0051] Overview of the Image Encoding Device
[0052] FIG. 2 is a diagram schematically showing an image encoding device to which an embodiment according to the present disclosure can be applied.
[0053] As shown in FIG. 2, the image encoding device 100 can include an image division unit 110, a subtraction unit 115, a conversion unit 120, a quantization unit 130, an inverse quantization unit 140, an inverse conversion unit 150, an addition unit 155, a filtering unit 160, a memory 170, an inter prediction unit 180, an intra prediction unit 185, and an entropy encoding unit 190. The inter prediction unit 180 and the intra prediction unit 185 can be collectively referred to as a "prediction unit". The conversion unit 120, the quantization unit 130, the inverse quantization unit 140, and the inverse conversion unit 150 can be included in a residual processing unit. The residual processing unit can further include the subtraction unit 115.
[0054] All or at least a part of the plurality of components constituting the image encoding apparatus 100 can be realized by one hardware component (for example, an encoder or a processor) according to an embodiment. Further, the memory 170 can include a DPB (decoded picture buffer) and can be realized by a digital storage medium.
[0055] The image segmentation unit 110 can divide an input image (or picture, frame) input to the image encoding device 100 into one or more processing units. As an example, the processing unit can be called a coding unit (CU). The coding unit can be obtained by recursively dividing a coding tree unit (CTU) or a largest coding unit (LCU) in a QT / BT / TT (Quad-tree / Binary-tree / Ternary-tree) structure. For example, one coding unit can be divided into a plurality of coding units with a deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. For the division of the coding unit, the quadtree structure can be applied first, and the binary tree structure and / or the ternary tree structure can be applied later. Based on the final coding unit that cannot be further divided, the coding procedure according to the present disclosure can be performed. The largest coding unit can be used as the final coding unit, and the coding units with a lower depth obtained by dividing the largest coding unit can also be used as the final coding unit. Here, the coding procedure can include procedures such as prediction, conversion, and / or restoration, which will be described later. As another example, the processing unit of the coding procedure can be a prediction unit (PU: Prediction Unit) or a transform unit (TU: Transform Unit). The prediction unit and the transform unit can be divided or partitioned from the final coding unit respectively. The prediction unit can be a unit of sample prediction, and the transform unit can be a unit of deriving transform coefficients and / or a unit of deriving a residual signal from the transform coefficients.
[0056] The prediction unit (inter prediction unit 180 or intra prediction unit 185) can perform a prediction on a processing target block (current block) and generate a predicted block including prediction samples for the current block. The prediction unit can determine whether intra prediction is applied in units of the current block or CU, or whether inter prediction is applied. The prediction unit can generate various information regarding the prediction of the current block and transmit it to the entropy encoding unit 190. The information regarding the prediction can be encoded by the entropy encoding unit 190 and output in the form of a bitstream.
[0057] The intra prediction unit 185 can predict the current block by referring to samples within the current picture. The samples to be referred to can be located in the neighborhood of the current block or at a distance from it according to the intra prediction mode and / or intra prediction technique. The intra prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes can include, for example, the DC mode and the planar mode. The directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes according to the degree of fineness of the prediction direction. However, this is only an example, and more or fewer directional prediction modes can be used based on the setting. The intra prediction unit 185 can also determine the prediction mode to be applied to the current block using the prediction mode applied to the neighboring blocks.
[0058] The inter prediction unit 180 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on the reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the surrounding blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the surrounding blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different from each other. The temporal neighboring block can be called by names such as a collocated reference block, a collocated CU (colCU), etc. The reference picture including the temporal neighboring block can be called a collocated picture (colPic). For example, the inter prediction unit 180 can construct a motion information candidate list based on the surrounding blocks, and generate information indicating which candidate is used to derive the motion vector and / or the reference picture index of the current block. Inter prediction can be performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the inter prediction unit 180 can use the motion information of the surrounding blocks as the motion information of the current block. In the case of the skip mode, different from the merge mode, the residual signal cannot be transmitted.In the case of the motion information prediction (MVP) mode, the motion vectors of neighboring blocks are used as motion vector predictors, and the motion vector difference and an indicator for the motion vector predictor are encoded to signal the motion vector of the current block. The motion vector difference can mean the difference between the motion vector of the current block and the motion vector predictor.
[0059] The prediction unit can generate a prediction signal based on various prediction methods and / or prediction techniques described later. For example, the prediction unit can apply intra prediction or inter prediction for predicting the current block, and can also apply intra prediction and inter prediction simultaneously. A prediction method that applies intra prediction and inter prediction simultaneously for predicting the current block can be called CIIP (combined inter and intra prediction). In addition, the prediction unit can also perform intra block copy (IBC) for predicting the current block. Intra block copy can be used for content image / video coding such as games, for example, like SCC (screen content coding). IBC is a method of predicting the current block using a restored reference block within the current picture at a position a predetermined distance away from the current block. When IBC is applied, the position of the reference block within the current picture can be encoded as a vector (block vector) corresponding to the predetermined distance. IBC basically performs prediction within the current picture, but can be performed in the same way as inter prediction in terms of deriving a reference block within the current picture. That is, IBC can use at least one of the inter prediction techniques described in the present disclosure.
[0060] The prediction signal generated by the prediction unit can be used to generate a restored signal or can be used to generate a residual signal. The subtraction unit 115 can subtract the prediction signal (predicted block, predicted sample array) output from the prediction unit from the input image signal (original block, original sample array) to generate a residual signal (residual signal, residual block, residual sample array). The generated residual signal can be transmitted to the conversion unit 120.
[0061] The conversion unit 120 can apply a conversion technique to the residual signal to generate transform coefficients. For example, the conversion technique can include at least one of DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), KLT (Karhunen-Loeve Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT means the conversion obtained from this graph when representing the relationship information between pixels as a graph. CNT means the conversion obtained based on generating a prediction signal using all previously reconstructed pixels. The conversion process can also be applied to pixel blocks having the same size of a square and can also be applied to blocks of a variable size that are not square.
[0062] The quantization unit 130 can quantize the transform coefficients and transmit them to the entropy encoding unit 190. The entropy encoding unit 190 can encode the quantized signal (information regarding the quantized transform coefficients) and output it in the form of a bit stream. The information regarding the quantized transform coefficients can be referred to as residual information. The quantization unit 130 can reorder the quantized transform coefficients in block form into a one-dimensional vector form based on the coefficient scan order, and can also generate the information regarding the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form.
[0063] The entropy encoding unit 190 can perform various encoding methods such as, for example, exponential Golomb, CAVLC (context-adaptive variable length coding), CABAC (context-adaptive binary arithmetic coding), etc. The entropy encoding unit 190 can also encode, together or separately, information necessary for video / image restoration (such as the values of syntax elements, etc.) in addition to the quantized transform coefficients. The encoded information (such as the encoded video / image information) can be transmitted or stored in the form of a bit stream in units of NAL (network abstraction layer) units. The video / image information can further include information regarding various parameter sets such as an adaptive parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Also, the video / image information can further include general constraint information. The signaling information, the transmitted information, and / or the syntax elements referred to in the present disclosure can be encoded through the above-described encoding procedure and included in the bit stream.
[0064] The bitstream can be transmitted via a network or stored in a digital storage medium. Here, the network can include a broadcast network and / or a communication network, etc., and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmission unit (not shown) for transmitting and / or a storage unit (not shown) for storing the signal output from the entropy encoding unit 190 can be provided as internal / external elements of the image encoding apparatus 100, or the transmission unit can also be provided as a component of the entropy encoding unit 190.
[0065] The quantized transform coefficients output from the quantization unit 130 can be used to generate a residual signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients via the inverse quantization unit 140 and the inverse transformation unit 150, a residual signal (residual block or residual sample) can be restored.
[0066] The addition unit 155 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the restored residual signal to the prediction signal output from the inter prediction unit 180 or the intra prediction unit 185. When there is no residual for the block to be processed as in the case where the skip mode is applied, the predicted block can be used as the reconstructed block. The addition unit 155 can be called a restoration unit or a reconstructed block generation unit. The generated reconstructed signal can be used for intra prediction of the next block to be processed within the current picture, and can also be used for inter prediction of the next picture after passing through filtering as described later.
[0067] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied in the picture encoding and / or restoration process.
[0068] The filtering unit 160 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 160 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be stored in the memory 170, specifically in the DPB of the memory 170. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, and the like. The filtering unit 160 can generate various information related to filtering as described later in the description of each filtering method and transmit it to the entropy encoding unit 190. The information related to filtering can be encoded by the entropy encoding unit 190 and output in the form of a bit stream.
[0069] The modified restored picture transmitted to the memory 170 can be used as a reference picture in the inter prediction unit 180. When inter prediction is applied through this, the image encoding device 100 can avoid prediction mismatches between the image encoding device 100 and the image decoding device, and can also improve the encoding efficiency.
[0070] The DPB in the memory 170 can store the modified restored picture for use as a reference picture in the inter prediction unit 180. The memory 170 can store the motion information of the blocks for which the motion information in the current picture has been derived (or encoded) and / or the motion information of the blocks in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 180 for utilization as the motion information of the spatial neighboring blocks or the motion information of the temporal neighboring blocks. The memory 170 can store the restored samples of the restored blocks in the current picture and transmit them to the intra prediction unit 185.
[0071] Overview of the Image Decoding Device
[0072] FIG. 3 is a diagram schematically showing an image decoding apparatus to which an embodiment according to the present disclosure can be applied.
[0073] As shown in FIG. 3, the image decoding apparatus 200 can be configured to include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an addition unit 235, a filtering unit 240, a memory 250, an inter prediction unit 260, and an intra prediction unit 265. The inter prediction unit 260 and the intra prediction unit 265 can be collectively referred to as a "prediction unit". The inverse quantization unit 220 and the inverse transform unit 230 can be included in a residual processing unit.
[0074] All or at least a part of a plurality of components constituting the image decoding apparatus 200 can be realized by one hardware component (for example, a decoder or a processor) according to an embodiment. Further, the memory 170 can include a DPB and can be realized by a digital storage medium.
[0075] The image decoding apparatus 200 that has received a bitstream including video / image information can execute a process corresponding to the process performed by the image encoding apparatus 100 of FIG. 2 to restore an image. For example, the image decoding apparatus 200 can perform decoding using the processing unit applied in the image encoding apparatus. Therefore, the decoding processing unit can be, for example, a coding unit. The coding unit can be obtained by dividing a coding tree unit or a maximum coding unit. Then, the restored image signal decoded and output via the image decoding apparatus 200 can be played back via a playback apparatus (not shown).
[0076] The image decoding device 200 can receive the signal output from the image encoding device of FIG. 2 in bitstream format. The received signal can be decoded via the entropy decoding unit 210. For example, the entropy decoding unit 210 can parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information can further include information regarding various parameter sets such as an Adaptive Parameter Set (APS), a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), or a Video Parameter Set (VPS). Also, the video / image information can further include general constraint information. The image decoding device can further use the information regarding the parameter set and / or the general constraint information to decode the image. The signaling information, the received information, and / or the syntax elements referred to in the present disclosure can be obtained from the bitstream by being decoded through the decoding procedure. For example, the entropy decoding unit 210 can decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the value of the syntax element necessary for image restoration, the quantized value of the conversion coefficient regarding the residual. More specifically, the CABAC entropy decoding method receives the bin corresponding to each syntax element from the bitstream, determines a context model using the syntax element information to be decoded, the information of the surrounding blocks and the decoded information of the block to be decoded, or the information of the symbol / bin decoded in the previous step, predicts the occurrence probability of the bin based on the determined context model, and performs arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element. At this time, the CABAC entropy decoding method can update the context model using the information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model.Among the information decoded by the entropy decoding unit 210, the information related to prediction is provided to the prediction units (inter prediction unit 260 and intra prediction unit 265), and the residual values entropy decoded by the entropy decoding unit 210, that is, the quantized transform coefficients and related parameter information, can be input to the inverse quantization unit 220. Also, among the information decoded by the entropy decoding unit 210, the information related to filtering can be provided to the filtering unit 240. On the other hand, a receiving unit (not shown) that receives a signal output from the image encoding device can be further provided as an internal / external element of the image decoding device 200, or the receiving unit can also be provided as a component of the entropy decoding unit 210.
[0077] On the other hand, the image decoding device according to the present disclosure can be called a video / image / picture decoding device. The image decoding device can also include an information decoder (video / image / picture information decoder) and / or a sample decoder (video / image / picture sample decoder). The information decoder can include the entropy decoding unit 210, and the sample decoder can include at least one of the inverse quantization unit 220, the inverse transform unit 230, the addition unit 235, the filtering unit 240, the memory 250, the inter prediction unit 260, and the intra prediction unit 265.
[0078] In the inverse quantization unit 220, the quantized transform coefficients can be inverse quantized to output transform coefficients. The inverse quantization unit 220 can reorder the quantized transform coefficients in a two-dimensional block format. In this case, the reordering can be performed based on the coefficient scan order performed in the image encoding device. The inverse quantization unit 220 can perform inverse quantization on the quantized transform coefficients using a quantization parameter (for example, quantization step size information) to obtain transform coefficients.
[0079] In the inverse conversion unit 230, the conversion coefficients can be inversely converted to obtain a residual signal (residual block, residual sample array).
[0080] The prediction unit can perform prediction on the current block and generate a predicted block including predicted samples for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied to the current block based on the information regarding the prediction output from the entropy decoding unit 210, and can determine a specific intra / inter prediction mode (prediction technique).
[0081] The ability of the prediction unit to generate a prediction signal based on various prediction methods (techniques) described later is the same as that described in the explanation of the prediction unit of the image encoding device 100.
[0082] The intra prediction unit 265 can predict the current block by referring to samples within the current picture. The explanation of the intra prediction unit 185 can be similarly applied to the intra prediction unit 265.
[0083] The inter prediction unit 260 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the neighboring blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit 260 can construct a motion information candidate list based on neighboring blocks, and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction can be performed based on various prediction modes (techniques), and the information regarding the prediction can include information indicating the mode (technique) of inter prediction for the current block.
[0084] The adder 235 can generate a restored signal (restored picture, restored block, restored sample array) by adding the obtained residual signal to a predicted signal (predicted block, predicted sample array) output from a prediction unit (including the inter prediction unit 260 and / or the intra prediction unit 265). When there is no residual for the processing target block as in the case where the skip mode is applied, the predicted block can be used as the restored block. The description of the adder 155 can be similarly applied to the adder 235. The adder 235 may also be referred to as a restoration unit or a restored block generation unit. The generated restored signal can be used for intra prediction of the next processing target block in the current picture, and can also be used for inter prediction of the next picture through filtering as described later.
[0085] The filtering unit 240 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 240 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be stored in the memory 250, specifically, in the DPB of the memory 250. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0086] The (modified) restored picture stored in the DPB of the memory 250 can be used as a reference picture by the inter prediction unit 260. The memory 250 can store the motion information of the block where the motion information in the current picture has been derived (or decoded) and / or the motion information of the block in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 260 for utilization as the motion information of the spatial neighboring blocks or the motion information of the temporal neighboring blocks. The memory 250 can store the restored samples of the restored blocks in the current picture and transmit them to the intra prediction unit 265.
[0087] In this specification, the embodiments described in the filtering unit 160, the inter prediction unit 180, and the intra prediction unit 185 of the image encoding apparatus 100 can also be applied to the filtering unit 240, the inter prediction unit 260, and the intra prediction unit 265 of the image decoding apparatus 200 in the same or corresponding manner.
[0088] On the other hand, the prediction units of the above-described image encoding / decoding apparatuses 100 and 200 can derive reference samples for the current block according to the intra prediction mode of the current block among the peripheral reference samples of the current block, and can generate prediction samples for the current block based on the reference samples.
[0089] For example, a predicted sample can be derived based on the average or interpolation of neighboring reference samples of the current block, and (ii) the predicted sample can also be derived based on reference samples existing in a specific (predicted) direction with respect to the predicted sample among the neighboring reference samples of the current block. In the case of (i), it may be called a non-directional mode or a non-angle mode, and in the case of (ii), it may be called a directional mode or an angular mode. Also, among the neighboring reference samples, the predicted sample can be generated through interpolation between the second neighboring sample and the first neighboring sample located in the direction opposite to the prediction direction of the intra prediction mode of the current block, based on the predicted sample of the current block. The above-mentioned case may be called linear interpolation intra prediction (LIP). Further, a temporary predicted sample of the current block is derived based on the filtered neighboring reference samples, and a weighted sum of at least one reference sample derived according to the intra prediction mode and the temporary predicted sample among the existing neighboring reference samples, i.e., the unfiltered neighboring reference samples, is used to derive the predicted sample of the current block. The above-mentioned case may be called PDPC (Position dependent intra prediction). Also, from among the neighboring multi-reference sample lines of the current block, the reference sample line with the highest prediction accuracy is selected, and a predicted sample is derived using the reference sample located in the prediction direction from this line. At this time, intra prediction coding can be performed by signaling the used reference sample line to the decoder. The above-mentioned case may be called multi-reference line intra prediction (MRL) or MRL-based intra prediction.Also, the current block is divided into vertical or horizontal sub-partitions, and intra prediction is performed based on the same intra prediction mode. However, the surrounding reference samples can be derived and used in units of the sub-partitions. That is, in this case, the intra prediction mode for the current block is applied identically to the sub-partitions, but by deriving and using the surrounding reference samples in units of the sub-partitions, the intra prediction performance can be enhanced as appropriate. Such a prediction method may be referred to as intra sub-partitions (ISP) or ISP-based intra prediction. Specific details will be described later. Also, when the prediction direction based on the prediction sample points between the surrounding reference samples, that is, when the prediction direction points to a fractional sample position, the value of the prediction sample can also be derived through interpolation of a plurality of reference samples located around the prediction direction (around the fractional sample position).
[0090] The intra prediction method described above may be called an intra prediction type separately from the intra prediction mode. The intra prediction type may be called by various terms such as an intra prediction technique or an additional intra prediction mode. For example, the intra prediction type (or an additional intra prediction mode, etc.) may include at least one of the above-described LIP, PDPC, MRL, and ISP. Information regarding the intra prediction type can be encoded by an encoding device, included in a bitstream, and signaled to a decoding device. The information regarding the intra prediction type can be realized in various forms such as flag information indicating whether each intra prediction type is applied, or index information indicating one of various intra prediction types.
[0091] Inter Prediction
[0092] The prediction units of the image encoding device 100 and the image decoding device 200 can derive prediction samples by performing inter prediction in block units. The inter prediction can be a prediction derived in a method that depends on data elements (e.g., sample values, motion information, etc.) of pictures other than the current picture. When inter prediction is applied to the current block, a predicted block (prediction sample array) for the current block can be derived based on a reference block (reference sample array) specified by a motion vector on the reference picture indicated by the reference picture index. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, based on the correlation of the motion information between the peripheral block and the current block, the motion information of the current block can be predicted in block, sub-block, or sample units. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction type (L0 prediction, L1 prediction, Bi prediction, etc.) information. When inter prediction is applied, the peripheral block can include a spatial neighboring block existing in the current picture and a temporal neighboring block existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different from each other. The temporal neighboring block can be called by names such as a collocated reference block, a collocated CU (colCU), etc., and the reference picture including the temporal neighboring block can also be called a collocated picture (colPic). For example, a motion information candidate list can be configured based on the peripheral blocks of the current block, and flag or index information indicating which candidate is selected (used) to derive the motion vector and / or reference picture index of the current block can be signaled.Inter prediction can be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the motion information of the current block can be the same as that of the selected neighboring block. In the case of skip mode, different from merge mode, the residual signal may not be transmitted. In the case of motion vector prediction (MVP) mode, the motion vector of the selected neighboring block can be used as a motion vector predictor, and the motion vector difference can be signaled. In this case, the motion vector of the current block can be derived using the sum of the motion vector predictor and the motion vector difference.
[0093] The motion information can include L0 motion information and / or L1 motion information according to the inter-prediction type (such as L0 prediction, L1 prediction, Bi prediction, etc.). The motion vector in the L0 direction can be called the L0 motion vector or MVL0, and the motion vector in the L1 direction can be called the L1 motion vector or MVL1. The prediction based on the L0 motion vector can be called L0 prediction, the prediction based on the L1 motion vector can be called L1 prediction, and the prediction based on both the L0 motion vector and the L1 motion vector can be called bi (Bi) prediction. Here, the L0 motion vector can represent the motion vector related to the reference picture list L0 (L0), and the L1 motion vector can represent the motion vector related to the reference picture list L1 (L1). The reference picture list L0 can include previous pictures as reference pictures in the output order before the current picture, and the reference picture list L1 can include subsequent pictures in the output order after the current picture. The previous picture can be called the forward (reference) picture, and the subsequent picture can be called the backward (reference) picture. The reference picture list L0 can further include subsequent pictures as reference pictures in the output order after the current picture. In this case, the previous picture can be indexed first within the reference picture list L0, and the subsequent picture can be indexed next. The reference picture list L1 can further include previous pictures as reference pictures in the output order before the current picture. In this case, the subsequent picture can be indexed first within the reference picture list 1, and the previous picture can be indexed next. Here, the output order can correspond to the POC (picture order count) order (order).
[0094] General of the Image / Video Coding Procedure
[0095] In image / video coding, pictures constituting an image / video can be encoded / decoded according to a series of decoding orders. The picture order corresponding to the output order of the decoded pictures can be set to be different from the decoding order, and based on this, not only forward prediction but also backward prediction can be performed during inter prediction.
[0096] FIG. 4 shows an example of a schematic picture decoding procedure to which the embodiment of the present disclosure can be applied. In FIG. 4, S400 can be performed by the entropy decoding unit 310 of the decoding apparatus described above with reference to FIG. 3, S410 can be performed by the prediction unit 330, S420 can be performed by the residual processing unit 320, S430 can be performed by the addition unit 340, and S440 can be performed by the filtering unit 350. S400 may include the information decoding procedure described in the present disclosure, S410 may include the inter / intra prediction procedure described in the present disclosure, S420 may include the residual processing procedure described in the present disclosure, S430 may include the block / picture restoration procedure described in the present disclosure, and S440 may include the in-loop filtering procedure described in the present disclosure.
[0097] Referring to FIG. 4, the picture decoding procedure can generally include an image / video information acquisition procedure (S400) (by decoding) from a bitstream, a picture restoration procedure (S410 to S430), and an in-loop filtering procedure (S440) for the restored picture. The picture restoration procedure can be performed based on the predicted samples and residual samples obtained through the inter / intra prediction (S410) and residual processing (S420, inverse quantization and inverse transformation for the quantized transform coefficients) processes described in the present disclosure. A modified restored picture can be generated through the in-loop filtering procedure for the restored picture generated through the picture restoration procedure, and the modified restored picture can be output as the decoded picture, and can also be stored in the decoded picture buffer or memory 360 of the decoding device and used as a reference picture in the inter prediction procedure when decoding subsequent pictures. Optionally, the in-loop filtering procedure can be omitted. In this case, the restored picture can be output as the encoded picture, and can also be stored in the decoded picture buffer or memory 360 of the decoding device and used as a reference picture in the inter prediction procedure when decoding subsequent pictures. The in-loop filtering procedure (S440) can include, as described above, a deblocking filtering procedure, a SAO (sample adaptive offset) procedure, an ALF (adaptive loop filter) procedure, and / or a bilateral filter procedure, etc., and some or all of them can be omitted. Also, one or some of the deblocking filtering procedure, SAO (sample adaptive offset) procedure, ALF (adaptive loop filter) procedure, and bilateral filter procedure can be sequentially applied, or all of them can be sequentially applied. For example, after the deblocking filtering procedure is applied to the restored picture, the SAO procedure can be performed.Alternatively, for example, after the deblocking filtering procedure is applied to the reconstructed picture, the ALF procedure can be performed. This can be similarly performed in the encoding apparatus.
[0098] FIG. 5 shows an example of a schematic picture encoding procedure to which the embodiments of the present disclosure can be applied. In FIG. 5, S500 can be performed by the prediction unit 220 of the encoding apparatus described above with reference to FIG. 2, S510 can be performed by the residual processing unit 230, and S520 can be performed by the entropy encoding unit 240. S500 can include the inter / intra prediction procedure described in the present disclosure, S510 can include the residual processing procedure described in the present disclosure, and S520 can include the information encoding procedure described in the present disclosure.
[0099] Referring to FIG. 5, the picture encoding procedure can generally include not only a procedure of encoding information for picture restoration (e.g., prediction information, residual information, partitioning information, etc.) and outputting it in the form of a bitstream, but also a procedure of generating a restored picture for the current picture, and a procedure (optional) of applying in-loop filtering to the restored picture. The encoding device can derive (modified) residual samples from the quantized transform coefficients via the inverse quantization unit 234 and the inverse transform unit 235, and can generate a restored picture based on the prediction samples that are the output of S500 and the (modified) residual samples. The restored picture generated in this way can be the same as the restored picture generated by the decoding device described above. A restored picture modified via the in-loop filtering procedure for the restored picture can be generated, which can be stored in the decoded picture buffer or memory 270, and can be used as a reference picture in the inter-prediction procedure when encoding subsequent pictures, similar to the case of the decoding device. As described above, in some cases, part or all of the in-loop filtering procedure can be omitted. When the in-loop filtering procedure is performed, (in-loop) filtering-related information (parameters) can be encoded by the entropy encoding unit 240 and output in the form of a bitstream, and the decoding device can perform the in-loop filtering procedure in the same way as the encoding device based on the filtering-related information.
[0100] Through such an in-loop filtering procedure, noises generated during image / video coding, such as blocking artifacts and ringing artifacts, can be reduced, and the subjective / objective visual quality can be improved. Also, by performing the in-loop filtering procedure in both the encoding device and the decoding device, the encoding device and the decoding device can derive the same prediction result, improve the reliability of picture coding, and reduce the amount of data to be transmitted for picture coding.
[0101] As described above, not only the decoding apparatus but also the encoding apparatus can perform a picture restoration procedure. Restoration blocks can be generated based on intra prediction / inter prediction in each block unit, and a restored picture including the restoration blocks can be generated. When the current picture / slice / tile group is an I picture / slice / tile group, the blocks included in the current picture / slice / tile group can be restored based only on intra prediction. On the other hand, when the current picture / slice / tile group is a P or B picture / slice / tile group, the blocks included in the current picture / slice / tile group can be restored based on intra prediction or inter prediction. In this case, inter prediction can be applied to some blocks within the current picture / slice / tile group, and intra prediction can also be applied to some of the remaining blocks. The color components of a picture can include a luma component and a chroma component, and unless explicitly limited in the present disclosure, the methods and examples proposed in the present disclosure can be applied to the luma component and the chroma component.
[0102] Example of the Coding Hierarchy and Structure
[0103] The coded video / image according to the present disclosure can be processed, for example, by the coding hierarchy and structure described later.
[0104] FIG. 6 is a diagram showing a hierarchical structure for a coded image.
[0105] The coded image is divided into a VCL (video coding layer) that performs decoding processing of the image and handles itself, a lower system that transmits and stores the coded information, and a NAL (network abstraction layer) that exists between the VCL and the lower system and is responsible for network adaptation functions.
[0106] In VCL, it is possible to generate VCL data including compressed image data (slice data), or to generate a parameter set including information such as a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), a Video Parameter Set (VPS), or a Supplemental Enhancement Information (SEI) message that is additionally required in the process of decoding an image.
[0107] In NAL, a NAL unit can be generated by adding header information (NAL unit header) to the RBSP (Raw Byte Sequence Payload) generated in VCL. At this time, the RBSP refers to slice data, parameter sets, SEI messages, etc. generated in VCL. The NAL unit header can include NAL unit type information specified by the RBSP data included in the NAL unit.
[0108] As shown in FIG. 6, the NAL unit can be divided into a VCL NAL unit and a Non-VCL NAL unit by the RBSP generated in VCL. The VCL NAL unit can mean a NAL unit including information (slice data) for an image, and the Non-VCL NAL unit can mean a NAL unit including information (parameter set or SEI message) required for decoding an image.
[0109] The above-mentioned VCL NAL units and non-VCL NAL units can be transmitted via a network with header information according to the data standard of the lower system. For example, the NAL unit can be transformed into a data form of a predetermined standard such as the H.266 / VVC file format, RTP (Real-time Transport Protocol), or TS (Transport Stream) and transmitted via various networks.
[0110] As described above, the NAL unit type can be specified according to the RBSP data structure included in the NAL unit, and information regarding such NAL unit type can be stored in the NAL unit header and signaled. For example, the NAL unit can be largely classified into a VCL NAL unit type and a non-VCL NAL unit type depending on whether it includes image information (slice data). The VCL NAL unit type can be further subdivided according to the nature / type of the picture included in the VCL NAL unit, and the non-VCL NAL unit type can be further subdivided according to the type of parameter set included in the non-VCL NAL unit.
[0111] An example of the VCL NAL unit type according to the picture type is as follows.
[0112] - "IDR_W_RADL", "IDR_N_LP": VCL NAL unit type for an IDR (Instantaneous Decoding Refresh) picture, which is a type of IRAP (Intra Random Access Point) picture
[0113] - "CRA_Nut": VCL NAL unit type for a CRA (Clean Random Access) picture, which is a type of IRAP picture
[0114] - "GDR_NUT": VCL NAL unit type for randomly accessible GDR (Gradual Decoding Refresh) pictures
[0115] - "STSA_NUT": VCL NAL unit type for randomly accessible STSA (Step-wise Temporal Sublayer Access) pictures
[0116] - "RADL_NUT": VCL NAL unit type for RADL pictures which are leading pictures
[0117] - "RASL_NUT": VCL NAL unit type for RASL pictures which are leading pictures
[0118] - "TRAIL_NUT": VCL NAL unit type for Trailing pictures. Trailing pictures are non-IRAP pictures, and for the trailing pictures associated with IRAP or GDR pictures, they follow the IRAP or GDR pictures in the decoding order. Trailing pictures that follow the associated IRAP pictures in the output order but precede them in the decoding order are not allowed.
[0119] An example of non-VCL NAL unit types according to the type of parameter set is as follows.
[0120] - "DCI_NUT": Non-VCL NAL unit type containing DCI (Decoding capability information)
[0121] - "VPS_NUT": Non-VCL NAL unit type containing VPS (Video Parameter Set)
[0122] - "SPS_NUT": A non-VCL unit type containing the SPS (Sequence Parameter Set).
[0123] - "PPS_NUT": A non-VCL NAL unit type containing the PPS (Picture Parameter Set).
[0124] - "PREFIX_APS_NUT", "SUFFIX_APS_NUT": Non-VCL NAL unit types containing the APS (Adaptation Parameter Set).
[0125] - "PH_NUT": A non-VCL NAL unit type containing the Picture Header.
[0126] The above NAL unit types can be identified by predetermined syntax information (e.g., nal_unit_type) contained within the NAL unit header.
[0127] High-Level Syntax
[0128] A coded picture can be composed of one or more slices. The parameters describing the coded picture can be signaled within the picture header, and the parameters describing the slice can be signaled within the slice header. The picture header is carried in a specific NAL unit type. Also, the slice header is present at the start of the NAL unit containing the slice payload, i.e., the slice data.
[0129] Table 1 is a diagram showing an example of a picture header.
[0130]
Table 1
[0131] Referring to Table 1, the picture header may include the syntax element ph_gdr_or_irap_pic_flag. The ph_gdr_or_irap_pic_flag can indicate whether the current picture is a GDR (Gradual Decoding Refresh) or IRAP (Intra Random Access Point) picture. For example, a ph_gdr_or_irap_pic_flag with a first value (e.g., 1) can indicate that the current picture is a GDR or IRAP picture. In contrast, a ph_gdr_or_irap_pic_flag with a second value (e.g., 0) can indicate that the current picture is not a GDR picture but is an IRAP picture.
[0132] Also, the picture header can include the syntax element ph_gdr_pic_flag. The ph_gdr_pic_flag can indicate whether the current picture is a GDR picture. For example, a ph_gdr_pic_flag with a first value (e.g., 1) can indicate that the current picture is a GDR picture. In contrast, a ph_gdr_pic_flag with a second value (e.g., 0) can indicate that the current picture is not a GDR picture. When the above-mentioned ph_gdr_or_irap_pic_flag has the first value (e.g., 1) and the ph_gdr_pic_flag has the second value (e.g., 0), the current picture is determined to be an IRAP picture. On the other hand, when the ph_gdr_pic_flag does not exist, the value of the ph_gdr_pic_flag can be inferred to be the second value (e.g., 0).
[0133] Also, the picture header can include the syntax element ph_recovery_poc_cnt. ph_recovery_poc_cnt can indicate the recovery point according to the output order of the decoded picture. If i) the current picture is a GDR picture and ii) ph_recovery_poc_cnt is 0, the current picture itself can be referred to as the recovery point picture. In contrast, if i) the current picture is a GDR picture and ii) there is a subsequent picture in the decoding order within the CLVS that follows the current picture and has a POC value such as the POC value of the recovery point picture, the subsequent picture can be referred to as the recovery point picture. Or, the picture with the first POC value in the output order that is greater than the POC value of the recovery point picture within the CLVS can also be referred to as the recovery point picture. The recovery point picture shall not precede the current GDR picture in the decoding order. A picture associated with the current GDR picture and having a POC value smaller than the POC value of the recovery point picture may be called the recovering picture of the GDR picture.
[0134] On the other hand, based on the pps_rpl_info_in_ph_flag having a first value (e.g., 1), ref_pic_lists() including reference picture list information can be called within the picture header. pps_rpl_info_in_ph_flag is signaled / obtained via a picture parameter set and can indicate the call position of ref_pic_lists(). For example, the pps_rpl_info_in_ph_flag with the first value (e.g., 1) can indicate that ref_pic_lists() is called within the picture header rather than within the slice header. In contrast, the pps_rpl_info_in_ph_flag with the second value (e.g., 0) can indicate that ref_pic_lists() is not called within the picture header but can be called within the slice header.
[0135] Random Access and RAP Picture
[0136] Random access means that the decoding process for a bitstream starts at a point other than the start point of the bitstream (i.e., the first picture in the decoding order). A picture that can be a random access start point is called a random access point (RAP) picture, and can include an IRAP (Intra Random Access Point) picture and a GDR (Gradual Decoding Refresh) picture.
[0137] (1) IRAP picture
[0138] An IRAP picture is a picture that is randomly accessible and has a VCL NAL unit type within the IDR_W_RADL~CRA_NUT range.
[0139] An IRAP picture can include an IDR (Instantaneous decoding refresh) picture and a CRA (Clean random access) picture. An IRAP picture does not necessarily need to use inter prediction based on reference pictures within the same layer during the decoding process. In the decoding order, the first picture in the bitstream can be an IRAP picture or a GDR (Gradual Decoding Refresh) picture. For a single-layer bitstream, if the parameter set that needs to be referred to is available, any picture preceding the IRAP picture in the decoding order does not need to be decoded / referred to, and all non-RASL (Random Access Skipped Leading) pictures following the IRAP picture in the decoding order within the IRAP picture and the CLVS (coded layer video sequence) can be correctly decoded.
[0140] (2) CRA picture
[0141] A CRA (Clean Random Access) picture is a type of IRAP picture, which refers to a picture having a VCL NAL unit type such as CRA_NUT.
[0142] For a CRA picture, it is not necessary to use inter prediction in the decoding process. A CRA picture may be the first picture in the bitstream in the decoding order, or may be a picture after the first one. A CRA picture may be associated with a RADL or RASL picture. When the NoOutputBeforeRecoveryFlag has a first value (e.g., 1) for a CRA picture, the RASL picture associated with the CRA picture cannot be decoded because it refers to a picture that does not exist in the bitstream, and as a result, it may not be output by the image decoding device. Here, the NoOutputBeforeRecoveryFlag can indicate whether a picture preceding the recovery point picture in the decoding order is output earlier than the recovery point picture. For example, the NoOutputBeforeRecoveryFlag with the first value (e.g., 1) can indicate that a picture preceding the recovery point picture in the decoding order cannot be output earlier than the recovery point picture. In this case, the CRA picture can be the first picture in the bitstream, or the first picture following the EOS (End Of Sequence) NAL unit in the decoding order. This can mean that a random access has occurred. In contrast, the NoOutputBeforeRecoveryFlag having a second value (e.g., 0) can indicate that a picture preceding the recovery point picture in the decoding order can be output earlier than the recovery point picture. In this case, the CRA picture may not be the first picture in the bitstream, or the first picture following the EOS NAL unit in the decoding order, which can mean that no random access has occurred.
[0143] (3) IDR Picture
[0144] An IDR (Instantaneous Decoding Refresh) picture is a type of IRAP picture and refers to a picture having a VCL NAL unit type such as IDR_W_RADL or IDR_N_LP.
[0145] An IDR picture does not necessarily need to utilize inter prediction in the decoding process. An IDR picture may be the first picture in the bitstream in the decoding order, or may be a picture after the first one. Each IDR picture may be the first picture of a CVS (coded video sequence) in the decoding order. When each VCL NAL unit for an IDR picture has a NAL unit type such as IDR_W_RADL, the IDR picture can have associated RADL pictures. In contrast, when each VCL NAL unit for an IDR picture has a NAL unit type such as IDR_N_LP, the IDR picture does not necessarily need to have associated leading pictures. On the other hand, an IDR picture does not necessarily need to have associated RASL pictures.
[0146] (4) GDR Picture
[0147] A GDR (Gradual Decoding Refresh) picture is a picture that is randomly accessible and refers to a picture having a VCL NAL unit type such as GDR_NUT.
[0148] The GDR function means that when the decoding process starts from a picture where all parts of the restored picture may not be correctly decoded, until the whole picture is correctly decoded, the correctly decoded parts of the restored picture in the pictures following the said picture gradually increase. A picture having such a GDR function is called a GDR picture, and among the pictures following the GDR picture, the first picture that can correctly decode the whole picture is called a recovery point picture.
[0149] A GDR picture can include a refreshed area and a dirty area. The refreshed area refers to an area having an exact matching of the decoded sample values when the decoding process starts from the GDR picture as compared with the case where the decoding process starts from a previous IRAP picture in the decoding order. The refreshed area can include coded blocks and can be decoded without referring to other pictures using intra prediction. In contrast, the dirty area refers to an area having no exact matching of the decoded sample values when the decoding process starts from the GDR picture as compared with the case where the decoding process starts from a previous IRAP picture in the decoding order. The dirty area can include coded blocks and can be decoded only by referring to other pictures using inter prediction. As a result, the dirty area can be decoded only when the GDR picture is not the start point of random access.
[0150] On the other hand, different from the above-mentioned IRAP pictures, a RAP picture can further include DRAP (Dependent Random Access Point) and EDRAP (extended DRAP) pictures that cannot be independently decoded.
[0151] DRAP and EDRAP pictures refer to pictures that cannot be correctly decoded without referring to other pictures but can be starting points for random access. A DRAP picture can be correctly decoded if it can refer to the nearest picture preceding the DRAP picture in the decoding order. In contrast, an EDRAP picture can be correctly decoded if it can refer to the nearest picture preceding the EDRAP picture in the decoding order and one or more other identified EDRAP pictures preceding the EDRAP picture in the decoding order. Since DRAP and EDRAP pictures can be encoded and represented using considerably fewer bits compared to other RAP pictures, especially IRAP pictures, they have the advantage of increasing the total number of RAPs while maintaining the overall bit volume.
[0152] Information about a DRAP picture can be signaled via a DRAP presentation SEI message defined in the VSEI (Versatile Supplemental Enhancement Information) spec. Also, information about an EDRAP picture can be signaled via an EDRAP presentation SEI message defined in the VSEI spec. Conversely, a picture associated with a DRAP presentation SEI message is called a DRAP picture, and a picture associated with an EDRAP presentation SEI message is called an EDRAP picture.
[0153] The DRAP or EDRAP picture indicated by the SEI message can be associated with an anchor picture that is actually an independent picture for random access. In the current design, the anchor picture only includes IRAP pictures called related IRAP (intra random access point) pictures. The basic concept of the current design is that before decoding the DRAP picture, the related IRAP picture is provided, and as long as it is decoded first, random access from the DRAP picture is possible. In this case, no other pictures (i.e., pictures between the related IRAP picture and the DRAP picture) are required for random access from the DRAP picture. Such a concept can be similarly applied to the EDRAP picture if it is added that the EDRAP picture can further reference zero or more previous EDRAP pictures.
[0154] As described above, according to the current design, only IRAP pictures are specified and used as anchor pictures for DRAP and EDRAP pictures. However, a GDR picture with a zero recovery point picture (e.g., a GDR picture having the same ph_recovery_poc_cnt as 0) can also be a random access point and can be correctly decoded without referring to other pictures. In this regard, the current design that uses only IRAP pictures as related pictures has a problem that it cannot use all other types of pictures that can actually perform as anchor pictures.
[0155] To solve such a problem, according to the present disclosure, the following aspects regarding DRAP and EDRAP pictures can be provided.
[0156] -(Configuration 1): For DRAP and EDRAP pictures, an anchor picture associated with the DRAP and / or EDRAP picture can be defined. Here, the anchor picture refers to a picture from which random access can be performed and can be correctly decoded without referring to other pictures.
[0157] -(Configuration 2): The list of pictures that can be used as or considered as anchor pictures for DRAP and / or EDRAP pictures can include the following.
[0158] a. IRAP picture
[0159] b. IRAP-like picture. Here, the IRAP-like picture means a picture that is not an existing IRAP picture but from which random access can be performed. As an example of such a picture, there is a GDR (Gradual Decoding Refresh) picture with a zero recovery point picture (that is, in the H.266 / VVC standard, a GDR picture having the same ph_recovery_poc_cnt as 0).
[0160] Examples of the present disclosure can be provided based on these configurations. By the examples, the above configurations may be applied individually or in combinations of two or more. Hereinafter, examples of the present disclosure will be described in detail.
[0161] 1. DRAP Picture
[0162] According to an example of the present disclosure, in order to assist random access from a DRAP picture, a new anchor picture associated with the DRAP picture can be defined. And the DRAP picture can be correctly decoded with reference to the associated anchor picture when random access occurs.
[0163] In one embodiment, the associated anchor picture of the DRAP picture can be a GDR picture having a zero recovery point picture. In other words, the IRAP picture associated with an existing DRAP picture can be newly defined as a GDR picture having a zero recovery point picture.
[0164] Information about the DRAP picture can be signaled via a DRAP presentation SEI message. The DRAP presentation SEI message can have, for example, the form of Table 2.
[0165] [Table 2]
[0166] The presence of the DRAP presentation SEI message indicates that certain constraints regarding picture order and picture referencing apply. These constraints allow the decoder to properly decode the DRAP picture and pictures that exist within the same layer and follow the DRAP picture in both decoding order and output order, excluding the anchor picture associated with the DRAP picture and without the need to decode any other pictures within the same layer.
[0167] The associated anchor picture for a DRAP picture must be one of the following pictures.
[0168] - The associated IRAP picture of the DRAP picture.
[0169] - The associated GDR picture of the DRAP picture having the same ph_recovery_poc_cnt as 0. This means a GDR picture that is itself a recovery point picture.
[0170] On the other hand, all of the constraints indicated by the presence of the DRAP presentation SEI message must apply and are as follows.
[0171] - The DRAP picture is a trailing picture.
[0172] - The DRAP picture has the same temporal sublayer identifier as 0.
[0173] - The DRAP picture does not contain any other pictures of the same layer in the active entry of its reference picture list, except for the associated anchor picture of the DRAP picture.
[0174] - All pictures that exist in the same layer and follow the DRAP picture in both the decoding order and the output order do not belong to the same layer in the active entry of their reference picture lists, except for the associated anchor picture of the DRAP picture, and do not contain any pictures that precede the DRAP picture in both the decoding order and the output order.
[0175] As described above, according to the embodiments of the present disclosure, when a random access occurs from a DRAP picture, a GDR picture having a zero recovery point picture can be used as the anchor picture associated with the DRAP picture. In this case, when a random access to the DRAP picture occurs, the DRAP picture can be correctly decoded by referring to the GDR picture having a zero recovery point picture as shown in FIG. 7.
[0176] 2. EDRAP Picture
[0177] According to the embodiments of the present disclosure, in order to assist random access from an EDRAP picture, a new anchor picture associated with the EDRAP picture can be defined. And the EDRAP picture can be correctly decoded by referring to the associated anchor picture and the previous EDRAP picture that precedes it in the decoding order when a random access occurs.
[0178] In one embodiment, the associated anchor picture of the EDRAP picture can be a GDR picture having a zero recovery point picture. In other words, the IRAP picture associated with the existing EDRAP picture can be newly defined as a GDR picture having a zero recovery point picture.
[0179] Information about the EDRAP picture can be signaled via the EDRAP presentation SEI message. The presence of the EDRAP presentation SEI message indicates that certain constraints regarding picture order and picture referencing apply. These constraints enable the decoder to properly decode the EDRAP picture and pictures that follow the EDRAP picture in both decoding order and output order within the same layer, without having to decode any other pictures within the same layer except for the list of referenceable pictures. Here, the list of referenceable pictures belongs to the same CLVS as the EDRAP picture and is composed of the associated anchor picture or EDRAP picture indicated by the edrap_ref_rap_id[i] syntax element.
[0180] The associated anchor picture for the EDRAP picture must be one of the following pictures.
[0181] - The associated IRAP picture of the EDRAP picture.
[0182] - The associated GDR picture of the EDRAP picture having the same ph_recovery_poc_cnt as 0. This means a GDR picture that is itself a recovery point picture.
[0183] On the other hand, all of the constraints indicated by the presence of the EDRAP presentation SEI message must apply and are as follows.
[0184] - The EDRAP picture is a trailing picture.
[0185] - The EDRAP picture has the same temporal sublayer identifier as 0.
[0186] - The EDRAP picture does not contain any pictures of the same layer other than the referenceable pictures in the active entries of its reference picture list.
[0187] - All pictures following the EDRAP picture in both the decoding order and the output order and existing in the same layer do not contain any pictures of the same layer other than the referenceable pictures in the active entries of their reference picture lists and do not contain any pictures preceding the said EDRAP picture in both the decoding order and the output order.
[0188] - None of the pictures in the list of referenceable pictures contain any pictures that belong to the same layer and are not in an earlier position in the active entry of its reference picture list than the pictures in the list of referenceable pictures. As a result, even if the first picture in the list of referenceable pictures is an EDRAP picture rather than an anchor picture, its active entry in the reference picture list does not contain any pictures of the same layer.
[0189] An example of the EDRAP presentation SEI message is as shown in Table 3.
[0190]
Table 3
[0191] Referring to Table 3, the EDRAP display SEI message extended_drap_indication() can include the syntax elements edrap_rap_id_minus1, edrap_leading_pictures_decodable_flag, edrap_reserved_zero_12bits, edrap_num_ref_rap_pics_minus1, and edrap_ref_rap_id[i].
[0192] One plus edmg_rap_id_minus1 can indicate the RAP picture identifier denoted as RapPicId of the EDRAP picture. Each anchor picture or EDRAP picture can be associated with a RapPicId value. The RapPicId value for an anchor picture can be inferred as 0. The RapPicId values for two EDRAP pictures associated with the same anchor picture should be different from each other.
[0193] The edrap_leading_pictures_decodable_flag can indicate whether certain constraints apply. Specifically, an edrap_leading_pictures_decodable_flag such as 1 can indicate that all of the following constraints apply.
[0194] - All pictures belonging to the same layer and following the EDRAP picture in decoding order shall follow in output order all pictures belonging to the same layer and preceding the said EDRAP picture in decoding order.
[0195] - All pictures belonging to the same layer, following the EDRAP picture in decoding order and preceding the EDRAP picture in output order shall not include, within the active entries of their reference picture lists, any pictures belonging to the same layer and preceding the said EDRAP picture in decoding order, except for referenceable pictures.
[0196] In contrast, the edrap_leading_pictures_decodable_flag equal to 0 can indicate that the above-mentioned constraints do not apply.
[0197] edrap_reserved_zero_12bits shall be equal to 0 within the bitstream according to the video codec standard, for example, the H.266 / VVC standard. Other values of edrap_reserved_zero_12bits are reserved for future use, and the decoder shall ignore the edrap_reserved_zero_12bits value.
[0198] edrap_num_ref_rap_pics_minus1 plus 1 can indicate the number of anchor or EDRAP pictures that belong to the same CLVS as the EDRAP picture and can also be included in the active entries of the reference picture list of the EDRAP picture.
[0199] edrap_ref_rap_id[i] can indicate the RapPicId of the i-th RAP picture that can also be included in the active entries of the reference picture list of the EDRAP picture. The i-th RAP picture shall be an anchor picture associated with the current EDRAP picture or an EDRAP picture associated with an IRAP picture identical to the current EDRAP picture.
[0200] As described above, when a random access occurs from an EDRAP picture, a GDR picture with a zero recovery point picture can be used as the anchor picture associated with the EDRAP picture. In this case, when a random access occurs to the EDRAP picture, the EDRAP picture can be correctly decoded by referring to the GDR picture with a zero recovery point picture and other EDRAP pictures that precede it in the decoding order, as shown in Figure 8.
[0201] According to the embodiments of the present disclosure, an anchor picture can be newly defined as a related picture of DRAP and EDRAP pictures. The anchor picture means a picture that can be correctly decoded within the same layer without the need to decode any other pictures that precede it in the decoding order, and can be a GDR picture having a zero recovery point picture. That is, a GDR picture that is itself a recovery point picture can be newly defined as an IRAP picture associated with a DRAP picture. As a result, a GDR picture that can exhibit the performance as an actual anchor picture can be used at random access, enabling more flexible and extended random access support.
[0202] Hereinafter, with reference to FIGS. 9 and 10, an image encoding / decoding method according to an embodiment of the present disclosure will be described in detail.
[0203] FIG. 9 is a flowchart showing an image encoding method according to an embodiment of the present invention. The image encoding method of FIG. 9 can be performed by the image encoding apparatus of FIG. 2.
[0204] Referring to FIG. 9, the image encoding apparatus can derive a second picture associated with a first picture that is a starting point of random access (S910). Further, the image encoding apparatus can encode information regarding the first picture and the second picture (S920). Here, the first picture is a dependent random access point (RAP) picture, and the second picture can be a GDR (Gradual Text Decoding Refresh) picture having a zero recovery point picture.
[0205] In one embodiment, the first picture can be a DRAP (Dependent Random Access Point) picture or an EDRAP (extended DRAP) picture.
[0206] In one embodiment, when the first picture is a DRAP (Dependent Random Access Point) picture, the active entries of the reference picture list for the first picture may not include any other pictures within the same layer, except for the GDR picture having the zero recovery point picture.
[0207] In one embodiment, when the first picture is an EDRAP (extended DRAP) picture, the random access point (RAP) picture identifier indicating the second picture can be encoded in a SEI (supplemental enhancement information) message.
[0208] In one embodiment, for the GDR picture having the zero recovery point picture, the RAP picture identifier value can be determined to be 0.
[0209] The bitstream generated by the above-described image encoding method can be stored in a non-transitory computer-readable recording medium and transmitted to an image decoding device.
[0210] FIG. 10 is a flowchart showing an image decoding method according to an embodiment of the present invention. The image decoding method of FIG. 10 can be performed by the image decoding device of FIG. 3.
[0211] Referring to FIG. 10, the image decoding apparatus can obtain a second picture associated with the first picture and decoded before the first picture in response to a random access request for the first picture (S1010). Then, the image decoding apparatus can decode the first picture with reference to the second picture (S1020). Here, the first picture is a dependent random access point (RAP) picture, and the second picture can be a GDR (Gradual Decoding Refresh) picture having a zero recovery point picture.
[0212] In one embodiment, the first picture can be a DRAP (Dependent Random Access Point) picture or an EDRAP (extended DARP) picture.
[0213] In one embodiment, when the first picture is a DRAP (Dependent Random Access Point) picture, the active entries of the reference picture list for the first picture may not include any other pictures within the same layer except for the GDR picture having the zero recovery point picture.
[0214] In one embodiment, when the first picture is an EDRAP (extended DRAP) picture, the second picture can be identified based on a random access point (RAP) picture identifier obtained from an SEI (supplemental enhancement information) message.
[0215] In one embodiment, the RAP picture identifier value for the GDR picture having the zero recovery point picture can be determined to be 0.
[0216] Exemplary methods of the present disclosure are presented as a series of operations for clarity of explanation, but this is not intended to limit the order in which the steps are performed. If necessary, each step can also be performed simultaneously or in a different order. To implement the method according to the present disclosure, it is also possible to include additional steps in the exemplary steps, or include the remaining steps excluding some steps, or include additional other steps excluding some steps.
[0217] In the present disclosure, an image encoding device or an image decoding device that performs a predetermined operation (step) can perform an operation (step) of checking the execution conditions and situations of the operation (step). For example, when it is described that a predetermined operation is performed when a predetermined condition is satisfied, the image encoding device or the image decoding device can perform the predetermined operation after performing an operation of checking whether the predetermined condition is satisfied.
[0218] The various embodiments of the present disclosure do not enumerate all possible combinations, but are for explaining representative aspects of the present disclosure. Matters described in the various embodiments may be applied independently or in combinations of two or more.
[0219] Also, the various embodiments of the present disclosure can be realized by hardware, firmware, software, or a combination thereof. In the case of realization by hardware, it can be realized by one or more ASICs (Application Specific Integrated Circuits), DSPs (Digital Signal Processors), DSPDs (Digital Signal Processing Devices), PLDs (Programmable Logic Devices), FPGAs (Field Programmable Gate Arrays), general processors, controllers, microcontrollers, microprocessors, etc.
[0220] In addition, the image decoding device and the image encoding device to which the embodiments of the present disclosure are applied can be included in a multimedia broadcast transceiver device, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video conversation device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camcorder, an on-demand video (VoD) service providing device, an over-the-top video (OTT) device, an Internet streaming service providing device, a three-dimensional (3D) video device, an image phone video device, and a medical video device, etc., and can be used to process video signals or data signals. For example, as an over-the-top video (OTT) device, it can include a game console, a Blu-ray player, an Internet-connected TV, a home theater system, a smartphone, a tablet PC, a digital video recorder (DVR), etc.
[0221] FIG. 11 is a diagram illustrating a content streaming system to which an embodiment of the present disclosure can be applied.
[0222] As shown in FIG. 11, the content streaming system to which the embodiments of the present disclosure are applied can generally include an encoding server, a streaming server, a Web server, a media storage, a user device, and a multimedia input device.
[0223] The encoding server serves to compress content input from a multimedia input device such as a smartphone, a camera, or a camcorder into digital data to generate a bitstream and transmit this to the streaming server. As another example, when a multimedia input device such as a smartphone, a camera, or a video camera directly generates a bitstream, the server can be omitted.
[0224] The bitstream can be generated by an image encoding method and / or an image encoding apparatus to which the embodiments of the present disclosure are applied, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.
[0225] The streaming server transmits multimedia data to a user device based on a user request via a Web server, and the Web server can serve as a medium for notifying the user of what services are available. When the user requests a desired service from the Web server, the Web server transmits this to the streaming server, and the streaming server can transmit multimedia data to the user. At this time, the content streaming system can include a separate control server. In this case, the control server can play a role in controlling commands / responses between each device within the content streaming system.
[0226] The streaming server can receive content from a media storage and / or an encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.
[0227] Examples of the user device may include mobile phones, smart phones, laptop computers, digital broadcast terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation devices, slate PCs, tablet PCs, ultrabooks, wearable devices, for example, smartwatches, smart glasses, head-mounted displays (HMDs), digital TVs, desktop computers, digital signage, and the like.
[0228] Each server in the content streaming system can be operated as a distributed server. In this case, the data received from each server can be distributedly processed.
[0229] The scope of the present disclosure includes software or machine-executable commands (for example, operating systems, applications, firmware, programs, etc.) that enable the operations of various example methods to be executed on a device or computer, and non-transitory computer-readable media on which such software or commands are stored and can be executed on a device or computer.
Industrial Applicability
[0230] Examples according to the present disclosure can be used for encoding / decoding images.
Claims
1. An image decoding method performed by an image decoding apparatus, the image decoding method comprising: responding to a random access request for a first picture by obtaining a second picture that is associated with the first picture and that was decoded before the first picture; decoding the first picture with reference to the second picture; wherein the first picture is a dependent random access point (RAP) picture; wherein the second picture is a GDR (Gradual Decoding Refresh) picture having a zero recovery point picture, the image decoding method.
2. The image decoding method according to claim 1, wherein the first picture is a DRAP (Dependent Random Access Point) picture or an EDRAP (extended DRAP) picture.
3. Based on the first picture being a DRAP (Dependent Random Access Point) picture, the active entries of the reference picture list for the first picture do not include any other pictures in the same layer except for the GDR picture having the zero recovery point picture, the image decoding method according to claim 1.
4. Based on the first picture being an EDRAP (extended DRAP) picture, the second picture is identified based on a random access point (RAP) picture identifier obtained from a SEI (supplemental enhancement information) message associated with the second picture, the image decoding method according to claim 1.
5. The image decoding method according to claim 4, wherein the RAP picture identifier value for the GDR picture having the zero recovery point picture is determined to be 0.
6. An image encoding method performed by an image encoding apparatus, the image encoding method comprising: deriving a second picture associated with a first picture that is a start point of random access; encoding information regarding the first picture and the second picture. The first picture is a dependent random access point picture, The second picture is a GDR (Gradual Decoding Refresh) picture having a zero recovery point picture, an image encoding method. **Claim 7** The image encoding method according to claim 6, wherein the first picture is a DRAP (Dependent Random Access Point) picture or an EDRAP (extended DRAP) picture. **Claim 8** Based on the first picture being a DRAP (Dependent Random Access Point) picture, The active entry of the reference picture list for the first picture does not include any other picture within the same layer except the GDR picture having the zero recovery point picture, the image encoding method according to claim 6. **Claim 9** Based on the first picture being an EDRAP (extended DRAP) picture, The random access point (Random Access Point, RAP) picture identifier indicating the second picture is encoded in a SEI (supplemental enhancement information) message associated with the second picture, the image encoding method according to claim 6. **Claim 10** For the GDR picture having the zero recovery point picture, the RAP picture identifier value is determined to be 0, the image encoding method according to claim 9. **Claim 11** A non-transitory computer-readable recording medium storing a bitstream generated by the image encoding method according to claim 6. **Claim 12** A method for transmitting a bitstream generated by an image encoding method, The image encoding method is, Deriving a second picture associated with a first picture that is a starting point of random access; and Encoding information regarding the first picture and the second picture, and the first picture is a dependent random access point picture, The transmission method is such that the second picture is a GDR (Gradual Decoding Refresh) picture having a zero recovery point picture.
Citation Information
Patent Citations
Method and System of NAL Unit Header Structure for Signaling New Elements
US20200177923A1
Syntax for dependent random access point indication in video bitstreams
US20220103781A1
Signaling-based image or video coding of information related to recovery point for gdr
WO2021201559A1
Cross random access point signaling in video coding
WO2022143614A1
Cross random access point signaling enhancements
WO2022148269A1