Image or video encoding based on information related to picture output

CN115668940BActive Publication Date: 2026-09-11LG ELECTRONICS INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202180035900.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-05-22
Filing Date
2021-05-17
Publication Date
2026-09-11
Estimated Expiration
2041-05-17

AI Technical Summary

Technical Problem

因此,如果使用诸如现有有线或无线宽带线路的介质来传输图像数据,或者使用现有存储介质来存储图像和视频数据,则传输成本和存储成本增加

Benefits of technology

[0024] This document can have various effects. For example, according to the embodiments described in this document, the overall image/video compression efficiency can be improved. Furthermore, according to the embodiments described in this document, in image decoding processing, the image output flag is derived after all slices in the image have been decoded, thereby improving the efficiency of image output-related operations. Additionally, according to the embodiments described in this document, since the image output flag can be effectively derived even for images containing slices of different NAL unit types, the efficiency of image output-related operations can be improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115668940B_ABST
    Figure CN115668940B_ABST
Patent Text Reader

Abstract

According to the disclosure herein, a slice included in a current picture is decoded, a picture output flag for the current picture is derived after the decoding of all slices in the current picture is completed, and output for the current picture can be determined based on the picture output flag. Based on a value of the picture output flag being 0, the current picture can be displayed as "no need to output", and based on a value of the picture output flag being 1, the current picture can be displayed as "need to output".
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This technology relates to video or image encoding, for example, encoding techniques based on information related to the output of the image. Background Technology

[0002] Recently, there has been an increasing demand for high-resolution and high-quality images and videos, such as 4K, 8K, or even higher Ultra High Definition (UHD) images and videos, across various fields. As image and video data becomes higher resolution and higher quality, the amount of information or bits transmitted increases relatively compared to existing image and video data. Therefore, if media such as existing wired or wireless broadband lines are used to transmit image data, or existing storage media are used to store image and video data, transmission and storage costs increase.

[0003] Furthermore, there is a growing interest in and demand for immersive media such as virtual reality (VR), artificial reality (AR) content, or holograms. Broadcasting of images and videos with image features that differ from those of real-world images, such as game graphics, is also increasing.

[0004] Therefore, efficient image and video compression technologies are needed to effectively compress and transmit, or store and play back, high-resolution and high-quality image and video information with these various characteristics.

[0005] In addition, a method is needed to improve the efficiency of image / video encoding, and for this purpose, an encoding method is needed that can efficiently perform derivation processing of information related to the image output. Summary of the Invention

[0006] Technical issues

[0007] This document provides a method and apparatus for improving the efficiency of video / image encoding.

[0008] This document also provides a method and device for effectively deriving relevant information about the screen output during screen decoding processing.

[0009] This paper also provides a method and apparatus for efficiently deriving image output information for images that include slices of different NAL unit types.

[0010] Technical solution

[0011] According to the implementation method described in this document, decoding is performed on the slices included in the current frame, and the frame output flag for the current frame can be derived after all slices in the current frame have been decoded. The output of the current frame can be determined based on the frame output flag. For example, based on a frame output flag value of 0, the current frame is marked as "no output required", while based on a frame output flag value of 1, the current frame is marked as "output required".

[0012] Based on the premise that all slices included in the current frame have been decoded, the frame output flag can be derived according to at least one of the following conditions.

[0013] The value of the image output flag can be deduced to be 0 based on the fact that the value of the syntax element associated with the Video Parameter Set (VPS) ID is greater than 0 and the current layer is not the first condition of the output layer.

[0014] Based on the second condition that the current frame is a Random Access Skip Preview (RASL) frame and the NoOutputBeforeRecoveryFlag of the associated Intra-Random Access Point (IRAP) frame is 1, the value of the frame output flag can be deduced to be 0.

[0015] The third condition for restoring a screen based on the current screen being either a Progressive Decoding Refresh (GDR) screen with NoOutputBeforeRecoveryFlag equal to 1 or a GDR screen with NoOutputBeforeRecoveryFlag equal to 1, can be deduced to be 0.

[0016] When at least one of the first, second, and third conditions is not met, the value of the screen output flag can be deduced as the value of the screen output-related syntax element to be signaled.

[0017] According to embodiments of this document, a video / image decoding method performed by a decoding device is provided. The video / image decoding method may include the methods disclosed in the embodiments of this document.

[0018] According to embodiments of this document, a decoding device is provided for performing video / image decoding. The decoding device may include the methods disclosed in embodiments of this document.

[0019] According to embodiments of this document, a video / image encoding method performed by an encoding device is provided. The video / image encoding method may include the methods disclosed in embodiments of this document.

[0020] According to embodiments of this document, an encoding apparatus for performing video / image encoding is provided. The encoding apparatus may include the methods disclosed in embodiments of this document.

[0021] According to embodiments of this document, a computer-readable digital storage medium is provided for storing encoded video / image information generated by a video / image encoding method disclosed in at least one embodiment of this document.

[0022] According to embodiments of this document, a computer-readable digital storage medium is provided for storing encoded information or encoded video / image information, which enables a decoding device to perform the video / image decoding method disclosed in at least one embodiment of this document.

[0023] Beneficial effects

[0024] This document can have various effects. For example, according to the embodiments described in this document, the overall image / video compression efficiency can be improved. Furthermore, according to the embodiments described in this document, in image decoding processing, the image output flag is derived after all slices in the image have been decoded, thereby improving the efficiency of image output-related operations. Additionally, according to the embodiments described in this document, since the image output flag can be effectively derived even for images containing slices of different NAL unit types, the efficiency of image output-related operations can be improved.

[0025] The effects achievable through the detailed examples in this document are not limited to those listed above. For instance, those skilled in the art can understand or derive various technical effects from this document. Therefore, the detailed effects of this document are not limited to those explicitly stated herein, but may include various effects that can be understood or derived from the technical features of this document. Attached Figure Description

[0026] Figure 1 Examples of video / image encoding apparatuses to which the implementation methods described in this document are applicable are briefly illustrated.

[0027] Figure 2 This is a schematic diagram illustrating the configuration of a video / image encoding device to which the implementation methods described in this document are applicable.

[0028] Figure 3 This is a schematic diagram illustrating the configuration of a video / image decoding device to which the implementation methods described in this document are applicable.

[0029] Figure 4 These are illustrative examples of video / image encoding processing to which one or more embodiments of this document may be applied.

[0030] Figure 5 This document provides illustrative examples of video / image decoding processing to which one or more embodiments of this document may be applied.

[0031] Figure 6Examples of entropy coding methods to which the implementation methods described in this document can be applied are illustrated schematically, and Figure 7 An entropy encoder in an encoding device is illustrated schematically.

[0032] Figure 8 Examples of entropy decoding methods to which the implementation methods described in this document are applicable are illustrated schematically, and Figure 9 An entropy decoder in a decoding device is illustrated schematically.

[0033] Figure 10 An example is shown to represent the hierarchical structure of an encoded image / video.

[0034] Figure 11 This is a diagram illustrating the temporal layer structure of NAL cells in a bitstream that supports temporal scalability.

[0035] Figure 12 It is a diagram used to describe screens that can be accessed randomly.

[0036] Figure 13 It is a diagram used to describe the IDR screen.

[0037] Figure 14 It is a diagram used to describe CRA screens.

[0038] Figure 15 Examples of video / image encoding methods to which one or more of the above embodiments of this document may be applied are illustrated.

[0039] Figure 16 Examples of video / image decoding methods to which one or more of the above embodiments of this document may be applied are illustrated.

[0040] Figure 17 Examples of video / image encoding methods to which one or more of the above embodiments of this document may be applied are illustrated.

[0041] Figure 18 Examples of video / image decoding methods to which one or more of the above embodiments of this document may be applied are illustrated.

[0042] Figure 19 and Figure 20 Examples of video / image encoding methods and related components according to the implementation of this document are illustrated schematically.

[0043] Figure 21 and Figure 22 Examples of video / image decoding methods and related components according to the embodiments of this document are illustrated schematically.

[0044] Figure 23Examples of content streaming systems to which the implementation methods disclosed in this document may be applied are shown. Detailed Implementation

[0045] This document can be modified in various ways and can have various implementations, and specific implementations will be illustrated and described in detail in the accompanying drawings. However, this is not intended to limit this document to a particular implementation. The terminology commonly used in this specification is used to describe particular implementations and not to limit the technical spirit of this document. Unless explicitly stated otherwise in the context, singular expressions include plural expressions. Terms such as “comprising” or “having” in this specification should be understood to indicate the presence of the features, quantities, steps, operations, elements, components, or combinations thereof described in the specification, without excluding the presence or possibility of adding one or more other features, quantities, steps, operations, elements, parts, or combinations thereof.

[0046] Furthermore, for ease of description in relation to different features and functions, the elements in the accompanying drawings described in this document are illustrated independently. This does not mean that each element is implemented as a separate piece of hardware or separate piece of software. For example, at least two elements may be combined to form a single element, or a single element may be divided into multiple elements. Unless it departs from the spirit of this document, implementations that combine and / or separate elements are also included within the scope of this document.

[0047] In this document, “A or B” can mean “A only,” “B only,” or “both A and B.” In other words, “A or B” in this document can be interpreted as “A and / or B.” For example, in this document, “A, B, or C” means “A only,” “B only,” “C only,” or “any combination of A, B, and C.”

[0048] The forward slash ( / ) or comma (,) used in this document can mean "and / or". For example, "A / B" can mean "A and / or B". Therefore, "A / B" can mean "A only", "B only", or "both A and B". For example, "A, B, C" can mean "A, B, or C".

[0049] In this document, "at least one of A and B" may mean "A only", "B only" or "both A and B". Furthermore, in this document, the expression "at least one of A or B" or "at least one of A and / or B" may be interpreted as the same as "at least one of A and B".

[0050] Furthermore, in this document, "at least one of A, B, and C" means "A only", "B only", "C only" or "any combination of A, B, and C". Additionally, "at least one of A, B, or C" or "at least one of A, B, and / or C" may mean "at least one of A, B, and C".

[0051] Furthermore, the parentheses used in this document may mean "for example". Specifically, when indicating "prediction (intra-prediction)", "intra-prediction" can be cited as an example of "prediction". In other words, "prediction" in this document is not limited to "intra-prediction", and "intra-prediction" can be cited as an example of "prediction". Moreover, even when indicating "prediction (i.e., intra-prediction)", "intra-prediction" can be cited as an example of "prediction".

[0052] This document relates to video / image coding. For example, the methods / implementations disclosed in this document can be applied to methods disclosed in the Universal Video Coding (VVC) standard. Additionally, the methods / implementations disclosed in this document can be applied to methods disclosed in the Basic Video Coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the second-generation Audio Video Coding (AVS2) standard, or next-generation video / image coding standards (e.g., H.267, H.268, etc.).

[0053] This document presents various implementations of video / image coding, and unless otherwise specified, the above implementations may also be combined with each other.

[0054] In this document, video can refer to a series of images over time. A picture typically refers to a unit representing an image at a specific time frame, while a slice / tile refers to a unit that constitutes part of a picture in terms of encoding. A slice / tile can include one or more Code Tree Units (CTUs). A picture can include one or more slices / tiles. A tile is a rectangular area of ​​a CTU within a specific tile column and row in a picture. A tile column is a rectangular area of ​​a CTU with a height equal to the height of the picture and a width that can be specified by a syntax element in the picture parameter set. A tile row is a rectangular area of ​​a CTU with a height specified by a syntax element in the picture parameter set and a width that can be equal to the width of the picture. A tile scan can represent a specific order of CTUs in a segmented image, and CTUs can be sequentially ordered in a raster scan of CTUs within a tile, and tiles in an image can be sequentially ordered in a raster scan of tiles within an image (a tile scan is a specific order of CTUs in a segmented image, where CTUs are sequentially ordered in a raster scan of CTUs within a tile, and tiles in an image are sequentially ordered in a raster scan of tiles within an image). A slice comprises an integer number of complete tiles or an integer number of consecutive complete CTU rows that can be exclusively contained within a tile of an image in a single NAL unit.

[0055] Furthermore, a frame can be divided into two or more sub-frames. A sub-frame can be a rectangular area of ​​one or more slices within the frame.

[0056] A pixel, or image unit, can refer to the smallest unit that makes up a picture (or image). Additionally, the term "sample" can be used as the counterpart to a pixel. A sample can typically represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component.

[0057] A unit can represent a basic unit of image processing. A unit may include a specific region of the image and at least one of the information associated with that region. A unit may include a luminance block and two chrominance (e.g., cb, cr) blocks. In some cases, the terms "unit" and "block" or "region" may be used interchangeably. Typically, an M×N block may include a set (or array) of samples (or transform coefficients) with M columns and N rows.

[0058] Furthermore, in this document, at least one of quantization / dequantization and / or transform / inverse transform may be omitted. When quantization / dequantization is omitted, the quantization transform coefficients may be referred to as transform coefficients. When transform / inverse transform is omitted, the transform coefficients may be referred to as coefficients or residual coefficients, or, for consistency, may still be referred to as transform coefficients.

[0059] In this document, quantization transform coefficients and transform coefficients can be referred to as transform coefficients and scaling transform coefficients, respectively. In this case, residual information can include information about the transform coefficients, and this information can be signaled using residual coding syntax. Transform coefficients can be derived based on residual information (or information about the transform coefficients), and scaling transform coefficients can be derived by performing an inverse transform (scaling) on ​​the transform coefficients. Residual samples can be derived based on the inverse transform (scaling) of the scaling transform coefficients. This can also be applied / expressed in other parts of this document.

[0060] The technical features described individually in one of the accompanying figures in this document can be implemented individually or simultaneously.

[0061] In the following, preferred embodiments of this document are described in more detail with reference to the accompanying drawings. In the following drawings, the same reference numerals are used for the same elements, and redundant descriptions of the same elements may be omitted.

[0062] Figure 1 Examples of video / image coding systems to which the implementation methods described in this document are applicable are illustrated.

[0063] Reference Figure 1 A video / image encoding system may include a source device and a receiving device. The source device may transmit encoded video / image information or data to the receiving device in the form of a file or stream via a digital storage medium or network.

[0064] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.

[0065] Video sources can acquire video / images through processes that capture, synthesize, or generate video / images. Video sources may include video / image capture devices and / or video / image generation devices. Video / image capture devices may include, for example, one or more cameras, video / image archives including previously captured video / images, etc. Video / image generation devices may include, for example, computers, tablets, and smartphones, and can generate video / images (electronically). For example, virtual video / images can be generated by computers, etc. In this case, video / image capture processing can be replaced by processing that generates related data.

[0066] Encoding devices can encode input video / images. They can perform a series of processes such as prediction, transformation, and quantization for compression and encoding efficiency. The encoded data (encoded video / image information) can be output as a bitstream.

[0067] A transmitter can send encoded video / image information or data, output in bitstream form, to a receiver via a digital storage medium or network, either as a file or a stream. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter can include elements for generating media files according to a predetermined file format and may include elements for transmission over a broadcast / communication network. The receiver can receive / extract the bitstream and send the received bitstream to a decoding device.

[0068] Decoding devices can decode video / images by performing a series of processes, such as dequantization, inverse transform, and prediction, that correspond to the operations of encoding devices.

[0069] The renderer can render decoded video / images. Rendered video / images can then be displayed on a monitor.

[0070] Figure 2 This is a diagram schematically illustrating the configuration of a video / image encoding device to which the embodiments described in this document may be applied. In the following text, the term "encoding device" may include image encoding devices and / or video encoding devices.

[0071] Reference Figure 2The encoding device 200 may include and be configured with an image segmenter 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may also include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstruction block generator. The image segmenter 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 described above may be configured by one or more hardware components (e.g., an encoder chipset or processor) according to an embodiment. Additionally, the memory 270 may include a decoded picture buffer (DPB) and may also be configured by a digital storage medium. The hardware components may also include the memory 270 as an internal / external component.

[0072] Image segmenter 210 can segment an input image (or picture, frame) input to encoding device 200 into one or more processing units. As an example, a processing unit may be referred to as a coding unit (CU). In this case, coding units can be recursively segmented from coding tree units (CTUs) or maximum coding units (LCUs) according to a quadtree-binary-tritree (QTBTTT) structure. For example, a coding unit can be segmented into multiple deeper coding units based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. In this case, for example, a quadtree structure can be applied first, followed by a binary tree structure and / or a ternary tree structure. Alternatively, a binary tree structure can be applied first. The encoding process according to this document can be performed based on the final coding units that are no longer segmented. In this case, based on encoding efficiency according to image characteristics, etc., the maximum coding unit can be directly used as the final coding unit, or, as needed, the coding unit can be recursively segmented into deeper coding units such that a coding unit with an optimal size can be used as the final coding unit. Here, the encoding process may include processes such as prediction, transformation, and reconstruction, as described later. As another example, the processing unit may also include a prediction unit (PU) or a transform unit (TU). In this case, each of the prediction unit and the transform unit can be separated or partitioned from the final encoding unit described above. The prediction unit may be a unit for predicting samples, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving residual signals from transform coefficients.

[0073] In some cases, a unit can be used interchangeably with terms such as a block or region. Typically, an M×N block can represent a sample or a set of transform coefficients consisting of M columns and N rows. A sample can typically represent a pixel or pixel value, and can also represent only the pixel / pixel value of the luminance component, and can also represent only the pixel / pixel value of the chrominance component. A sample can be used as a term corresponding to the pixels or picometers that configure a frame (or image).

[0074] Encoding device 200 generates a residual signal (residual block, residual sample array) by subtracting the prediction signal (prediction block, prediction sample array) output from inter-frame predictor 221 or intra-frame predictor 222 from the input image signal (original block, original sample array), and the generated residual signal is sent to converter 232. In this case, as shown, the unit within encoding device 200 for subtracting the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) can be called subtractor 231. The predictor can perform prediction on the block to be processed (hereinafter referred to as the current block) and generate a prediction block including the prediction samples of the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction on a unit of the current block or CU. As described later in the description of each prediction mode, the predictor can generate various information about the prediction, such as prediction mode information, to transmit the generated information to entropy encoder 240. The information about the prediction can be encoded by entropy encoder 240 and output as a bitstream.

[0075] Intra-predictor 222 can refer to samples within the current frame to predict the current block. Depending on the prediction mode, the reference sample can be located near or far from the current block. The prediction modes in intra-prediction can include multiple non-directional modes and multiple directional modes. Non-directional modes can include, for example, DC mode or planar mode. Depending on the fineness of the prediction direction, the directional modes can include, for example, 33 or 65 directional prediction modes. However, this is exemplary, and more or fewer directional prediction modes can be used depending on the settings. Intra-predictor 222 can also determine the prediction mode applied to the current block by using prediction modes applied to neighboring blocks.

[0076] Inter-frame predictor 221 can deduce the predicted block of the current block based on a reference block (reference sample array) specified by motion vectors on a reference frame. Here, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation between motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference frame indices. Motion information may also include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current frame and temporally neighboring blocks existing in the reference frame. The reference frame including the reference block and the reference frame including the temporally neighboring block may be the same or different. The temporally neighboring block may be called a co-located reference block, a co-located CU (colCU), etc., and the reference frame including the temporally neighboring block may be called a co-located frame (colPic). For example, inter-frame predictor 221 can configure a motion information candidate list based on neighboring blocks and generate information indicating which candidate to use to deduce the motion vector and / or reference frame index of the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in skip mode and merge mode, the inter-frame predictor 221 can use motion information from neighboring blocks as motion information for the current block. In skip mode, unlike merge mode, residual signals may not be transmitted. In motion vector prediction (MVP) mode, the motion vector of the current block can be indicated by using motion vectors from neighboring blocks as motion vector predictors and signaling the motion vector difference.

[0077] Predictor 220 can generate prediction signals based on various prediction methods described later. For example, the predictor can not only apply intra-frame prediction or inter-frame prediction to predict a block, but can also apply intra-frame prediction and inter-frame prediction simultaneously. This can be referred to as combined intra-frame and inter-frame prediction (CIIP). Furthermore, the predictor can perform prediction on blocks based on an intra-block copy (IBC) prediction mode or a palette mode. The IBC prediction mode or palette mode can be used for content image / video coding (e.g., screen content coding (SCC)) such as games. IBC essentially performs prediction in the current frame, but can be performed similarly to inter-frame prediction because it derives a reference block in the current frame. That is, IBC can use at least one of the inter-frame prediction techniques described in this document. The palette mode can be considered as an example of intra-frame coding or intra-frame prediction. When applying a palette mode, sample values ​​in the frame can be signaled based on information about the palette index and palette table.

[0078] The predicted signal generated by the predictor (including inter-frame predictor 221 and / or intra-frame predictor 222) can be used to generate a reconstructed signal or a residual signal. Transformer 232 can generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Karhunen-Loève Transform (KLT), Graph-Based Transform (GBT), or Conditional Nonlinear Transform (CNT). Here, GBT refers to a transform obtained from a graph when the relationship information between pixels is represented by a graph. CNT refers to a transform generated based on the predicted signal generated using all previously reconstructed pixels. Furthermore, the transform processing can be applied to square pixel blocks of the same size, or it can be applied to blocks of variable size instead of square.

[0079] Quantizer 233 can quantize the transform coefficients to send the quantized transform coefficients to entropy encoder 240, which can encode the quantized signal (information about the quantized transform coefficients) into a bitstream and output the encoded quantized signal. The information about the quantized transform coefficients can be referred to as residual information. Quantizer 233 can rearrange the block-form quantized transform coefficients into a one-dimensional vector form based on the coefficient scan order, and can also generate information about the quantized transform coefficients based on the one-dimensional vector form. Entropy encoder 240 can perform various encoding methods such as Golomb coding, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). Entropy encoder 240 can also encode information necessary for the video / image other than the quantized transform coefficients (e.g., values ​​of syntax elements, etc.) together or separately. Encoded information (e.g., encoded video / image information) can be sent or stored in bitstream form at the Network Abstraction Layer (NAL) unit level. The video / image information may also include information about various parameter sets such as Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). Additionally, the video / image information may also include general constraint information. Information and / or syntax elements that are signaled / transmitted, as described later in this document, can be encoded through the aforementioned encoding process and thus included in the bitstream. The bitstream can be transmitted over a network or stored in a digital storage medium. Here, the network may include broadcast networks and / or communication networks, etc., and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. A transmitter (not illustrated) for transmitting the signal output from the entropy encoder 240 and / or a memory (not illustrated) for storing the signal can be configured as internal / external components of the encoding device 200, or the transmitter may also be included within the entropy encoder 240.

[0080] The quantization transform coefficients output from quantizer 233 can be used to generate a prediction signal. For example, dequantizer 234 and inverse transform 235 apply inverse quantization and inverse transform to the quantization transform coefficients, allowing the residual signal (residual block or residual sample) to be reconstructed. Adder 250 adds the reconstructed residual signal to the prediction signal output from inter-frame predictor 221 or intra-frame predictor 222, thereby generating a reconstructed signal (reconstructed frame, reconstructed block, reconstructed sample array). If no residual exists for the block to be processed, such as when a skip mode is applied, the prediction block can be used as the reconstructed block. Adder 250 can be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current frame, and, as described later, can also be used for inter-frame prediction of the next frame by filtering.

[0081] In addition, luminance mapping with chroma scaling (LMCS) can be applied in screen encoding and / or reconstruction processing.

[0082] Filter 260 can apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, filter 260 can apply various filtering methods to the reconstructed image to generate a modified reconstructed image, and store the modified reconstructed image in memory 270, specifically in the DPB of memory 270. Various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filtering, bilateral filtering, etc. Filter 260 can generate various filtering-related information to transmit the generated information to entropy encoder 240, as described later in the description of each filtering method. The filtering-related information can be encoded by entropy encoder 240 to output as a bitstream.

[0083] The modified reconstructed frame sent to memory 270 can be used as a reference frame in inter-frame predictor 221. If the inter-frame predictor applies inter-frame prediction, the encoding device can avoid prediction mismatch between encoding device 200 and decoding device, and also improve encoding efficiency.

[0084] The DPB of memory 270 can store modified reconstructed frames to be used as reference frames in inter-frame predictor 221. Memory 270 can store motion information of blocks in which motion information within the current frame is derived (or encoded) and / or motion information of blocks in previously reconstructed frames. The stored motion information can be transmitted to inter-frame predictor 221 to be used as motion information for spatially or temporally adjacent blocks. Memory 270 can store reconstructed samples of reconstructed blocks within the current frame and can transmit the reconstructed samples to intra-frame predictor 222.

[0085] Figure 3This diagram is an illustrative representation of the configuration of a video / image decoding device to which the embodiments described in this document may be applied. In the following text, the term "decoding device" may include image decoding devices and / or video decoding devices.

[0086] Reference Figure 3 The decoding device 300 may include and be configured with an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an intra-frame predictor 331 and an inter-frame predictor 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 322. The entropy decoder 310, residual processor 320, predictor 330, adder 340, and filter 350 described above may be configured by one or more hardware components (e.g., a decoder chipset or processor) according to an embodiment. Furthermore, the memory 360 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware components may also include the memory 360 as an internal / external component.

[0087] When the input includes a bitstream containing video / image information, the decoding device 300 can respond to... Figure 2 The encoding device shown reconstructs an image by processing video / image information. For example, the decoding device 300 can deduce units / blocks based on block segmentation information obtained from the bitstream. The decoding device 300 can perform decoding using processing units applied to the encoding device. Therefore, the processing unit used for decoding can be, for example, an encoding unit, and the encoding unit can be separated from the encoding tree unit or the maximum encoding unit according to a quadtree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units can be derived from the encoding unit. Furthermore, the reconstructed image signal decoded and output by the decoding device 300 can be reproduced by a reproduction device.

[0088] Decoding device 300 can receive data from... in the form of a bitstream. Figure 2The signal output by the encoding device shown can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can deduce the information (e.g., video / image information) required for image reconstruction (or picture reconstruction) by parsing the bitstream. The video / image information may also include information about various parameter sets such as Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), and Video Parameter Set (VPS). In addition, the video / image information may also include general constraint information. The decoding device can also decode the picture based on the information about the parameter sets and / or general constraint information. The information and / or syntax elements notified / received by signals, as described later in this document, can be decoded by the decoding process and obtained from the bitstream. For example, the entropy decoder 310 can decode the information within the bitstream based on encoding methods such as exponential Golomb coding, CAVLC, or CABAC, and output the values ​​of the syntax elements necessary for image reconstruction and the quantized values ​​of the residual correlation transform coefficients. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element from the bitstream, use the information of the syntax element to be decoded and decode information of neighboring blocks and the block to be decoded, or information of symbols / bins decoded in the previous stage, to determine a context model, and generate symbols corresponding to the values ​​of each syntax element by predicting the generation probability of bins according to the determined context model, thereby performing arithmetic decoding of bins. At this time, the CABAC entropy decoding method can determine the context model and then update the context model using the information of decoded symbols / bins for the context model of the next symbol / bin. The prediction information in the information decoded by the entropy decoder 310 can be provided to the predictors (inter-frame predictor 332 and intra-frame predictor 331), and the residual values ​​(i.e., quantization transform coefficients and related parameter information) of the entropy decoding performed by the entropy decoder 310 can be input to the residual processor 320. The residual processor 320 can derive the residual signals (residual blocks, residual samples, and residual sample arrays). In addition, the filtering information in the information decoded by the entropy decoder 310 can be provided to the filter 350. Furthermore, a receiver (not illustrated) for receiving signals output from the encoding device may be configured as an internal / external element of the decoding device 300, or the receiver may be a component of the entropy decoder 310. Additionally, the decoding device according to this document may be referred to as a video / image / picture decoding device, and the decoding device may also be classified as an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoder 310, and the sample decoder may include at least one of a dequantizer 321, an inverse transformer 322, an adder 340, a filter 350, a memory 360, an inter-frame predictor 332, and an intra-frame predictor 331.

[0089] Dequantizer 321 can dequantize the quantized transform coefficients to output transform coefficients. Dequantizer 321 can rearrange the quantized transform coefficients in the form of two-dimensional blocks. In this case, the rearrangement can be performed based on the coefficient scan order performed by the encoding device. Dequantizer 321 can perform dequantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain the transform coefficients.

[0090] The inverse transformer 322 performs an inverse transformation on the transformation coefficients to obtain the residual signal (residual block, residual sample array).

[0091] Predictor 330 can perform prediction on the current block and generate a prediction block that includes prediction samples of the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction to the current block based on the prediction information output from entropy decoder 310, and determine the specific intra-frame / inter-frame prediction mode.

[0092] The predictor can generate prediction signals based on various prediction methods described later. For example, the predictor can apply not only intra-frame prediction or inter-frame prediction to the prediction of a block, but also simultaneous intra-frame prediction and inter-frame prediction. This can be referred to as combined intra-frame and inter-frame prediction (CIIP). Furthermore, the predictor can perform prediction on blocks based on an intra-block copy (IBC) prediction mode or a palette mode. The IBC prediction mode or palette mode can be used for content image / video coding (e.g., screen content coding (SCC)) such as games. IBC essentially performs prediction within the current frame, but can be performed similarly to inter-frame prediction because it derives a reference block in the current frame. That is, IBC can use at least one of the inter-frame prediction techniques described in this document. The palette mode can be considered an example of intra-frame coding or intra-frame prediction. When a palette mode is applied, information about the palette table and palette index can be included in the video / image information and signaled.

[0093] The intra-predictor 331 can refer to samples within the current frame to predict the current block. Depending on the prediction mode, the reference samples can be located near the current block or far from it. The prediction modes in intra-prediction can include multiple non-directional modes and multiple directional modes. The intra-predictor 331 can also use prediction modes applied to neighboring blocks to determine the prediction mode applied to the current block.

[0094] Inter-frame predictor 332 can deduce the predicted block for the current block based on a reference block (reference sample array) specified by motion vectors on a reference frame. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation between motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference frame indices. Motion information may also include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing within the current frame and temporally neighboring blocks existing in the reference frame. For example, inter-frame predictor 332 can configure a motion information candidate list based on neighboring blocks and deduce the motion vector and / or reference frame index for the current block based on received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and the information about the prediction may include information indicating the mode of inter-frame prediction for the current block.

[0095] Adder 340 can add the acquired residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor (including inter-frame predictor 332 and / or intra-frame predictor 331) to generate a reconstruction signal (reconstructed frame, reconstruction block, reconstruction sample array). If no residual exists for the block to be processed when skip mode is applied, the prediction block can be used as the reconstruction block.

[0096] Adder 340 can be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current frame, and as described later, it can also be filtered out or used for inter-frame prediction of the next frame.

[0097] In addition, luminance mapping with chroma scaling (LMCS) can also be applied in the image decoding process.

[0098] Filter 350 can apply filtering to the reconstructed signal, thereby improving the subjective / objective image quality. For example, filter 350 can apply various filtering methods to the reconstructed image to generate a modified reconstructed image, and send the modified reconstructed image to memory 360, specifically to the DPB in memory 360. Various filtering methods may include, for example, deblocking filtering, adaptive sample shifting, adaptive loop filtering, bidirectional filtering, etc.

[0099] The (modified) reconstructed frame stored in the DPB of memory 360 can be used as a reference frame in inter-frame predictor 332. Memory 360 can store motion information of blocks in which motion information within the current frame is derived (decoded) and / or motion information of blocks within previously reconstructed frames. The stored motion information can be transmitted to inter-frame predictor 260 so that it can be used as motion information for spatially or temporally adjacent blocks. Memory 360 can store reconstructed samples of reconstructed blocks within the current frame and can transmit the stored reconstructed samples to intra-frame predictor 331.

[0100] In this document, the exemplary embodiments described in the encoding device 200 filter 260, inter-frame predictor 221 and intra-frame predictor 222 can be equivalently applied to or correspond to the decoding device 300 filter 350, inter-frame predictor 332 and intra-frame predictor 331, respectively.

[0101] As described above, prediction is performed during video encoding to improve compression efficiency. This allows the generation of prediction blocks that include prediction samples of the current block (i.e., the target block to be encoded), which is the block to be encoded. Here, the prediction block includes prediction samples in the spatial domain (or pixel domain). The prediction block is derived in the same manner in both the encoding and decoding devices, and the encoding device can signal information about the residual between the original block and the prediction block (residual information) instead of the original sample values ​​of the original block to the decoding device, thereby improving image encoding efficiency. The decoding device can derive a residual block including residual samples based on the residual information, add the residual block and the prediction block to generate a reconstruction block including reconstructed samples, and generate a reconstructed image including the reconstruction block.

[0102] Residual information can be generated through transformation and quantization processes. For example, the encoding device can derive a residual block between the original block and the prediction block, perform a transformation process on the residual samples (residual sample array) included in the residual block to derive transform coefficients, perform a quantization process on the transform coefficients to derive quantized transform coefficients, and (via a bitstream) signal the relevant residual information to the decoding device. Here, the residual information may include the value information, position information, transform technique, transform kernel, quantization parameters, etc., of the quantized transform coefficients. The decoding device can perform dequantization / inverse transform processes based on the residual information and derive residual samples (or residual blocks). The decoding device can generate a reconstructed frame based on the prediction block and the residual block. Furthermore, for reference in future inter-frame prediction of the frame, the encoding device can also perform dequantization / inverse transform on the quantized transform coefficients to derive residual blocks and generate a reconstructed frame based on these residual blocks.

[0103] Furthermore, as described above, when performing prediction on the current block, intra-frame prediction or inter-frame prediction can be applied. In an implementation, when inter-frame prediction is applied to the current block, the predictor of the encoding / decoding device (more specifically, the inter-frame predictor) can derive prediction samples by performing inter-frame prediction on a block-by-block basis. Inter-frame prediction can indicate a prediction derived in a method that depends on data elements (e.g., sample values, motion information, etc.) of one or more frames other than the current frame. When applying inter-frame prediction to the current block, a prediction block (prediction sample array) for the current block can be derived based on a reference block (reference sample array) specified by a motion vector on a reference frame indicated by a reference frame index. In this case, to reduce the amount of motion information transmitted in the inter-frame prediction mode, the motion information of the current block can be predicted on a block, sub-block, or sample-by-sample basis based on the correlation of motion information between neighboring blocks and the current block. Motion information can include motion vectors and reference frame indices. Motion information can also include inter-frame prediction type (L0 prediction, L1 prediction, Bi prediction, etc.) information. When applying inter-frame prediction, neighboring blocks can include spatially neighboring blocks existing in the current frame and temporally neighboring blocks existing in a reference frame. The reference frame including the reference block and the reference frame including the temporally neighboring block can be the same or different from each other. The temporally neighboring block can be referred to as a juxtaposed reference block, juxtaposed CU (colCU), etc., and the reference frame including the temporally neighboring block can be referred to as a juxtaposed frame (colPic). For example, a motion information candidate list can be configured based on the neighboring blocks of the current block, and a flag or index information indicating which candidate to select (use) for deriving the motion vector of the current block and / or the reference frame index can be signaled. Inter-frame prediction can be performed based on various prediction modes. For example, in skip mode and merge mode, the motion information of the current block can be the same as the motion information of the selected neighboring block. In skip mode, unlike merge mode, residual signals may not be sent. In motion information prediction (motion vector prediction (MVP)) mode, the motion vector of the selected neighboring block can be used as a motion vector predictor, and the motion vector difference can be signaled. In this case, the motion vector of the current block can be derived using the sum of the motion vector predictor and the motion vector difference.

[0104] Depending on the inter-frame prediction type (L0 prediction, L1 prediction, Bi prediction, etc.), motion information may also include L0 motion information and / or L1 motion information. The motion vector in the L0 direction can be referred to as the L0 motion vector or MVL0, and the motion vector in the L1 direction can be referred to as the L1 motion vector or MVL1. Prediction based on the L0 motion vector is called L0 prediction, prediction based on the L1 motion vector is called L1 prediction, and prediction based on both L0 and L1 motion vectors is called bidirectional prediction. Here, the L0 motion vector may indicate the motion vector associated with the reference frame list L0, and the L1 motion vector may indicate the motion vector associated with the reference frame list L1. The reference frame list L0 may include frames that precede the current frame in the output order, and the reference frame list L1 may include frames that follow the current frame in the output order, as reference frames. The previous frame may be referred to as the forward (reference) frame, and the subsequent frame may be referred to as the reverse (reference) frame. The reference frame list L0 may also include frames that follow the current frame in the output order, as reference frames. In this scenario, the previous frame can be indexed first in the reference frame list L0, and then the subsequent frames can be indexed. The reference frame list L1 can also include frames that precede the current frame in the output order as reference frames. In this case, the subsequent frames can be indexed first in the reference frame list L1, and then the previous frames can be indexed. Here, the output order can correspond to the Frame Order Count (POC) order.

[0105] Reference Figure 4 The video / image encoding process can schematically include the process of generating a reconstructed frame for the current frame and the process of applying in-loop filtering to the reconstructed frame (optionally), as well as encoding information used for reconstructing the frame (e.g., prediction information, residual information, or segmentation information) to output as shown in the reference. Figure 2 The process of encoding information in the form of a bitstream.

[0106] The encoding device can derive (modified) residual samples from the quantization transform coefficients using dequantizer 234 and inverse transformer 235, and generate a reconstructed frame based on the prediction samples output in S400 and the (modified) residual samples. The reconstructed frame thus generated can be identical to the aforementioned reconstructed frame generated by the decoding device. The modified reconstructed frame can be generated through an in-loop filtering process for the reconstructed frame and can be stored in the decoded frame buffer or memory 270, and, as in the case of the decoding device, used as a reference frame in the inter-frame prediction process when the frame is later encoded. As mentioned above, in some cases, some or all of the in-loop filtering process can be omitted. If an in-loop filtering process is performed, the (in-loop) filtering-related information (parameters) is encoded by entropy encoder 240 and output as a bitstream, and the decoding device can perform the in-loop filtering process based on the filtering-related information in the same way as the encoding device.

[0107] In-loop filtering can reduce noise (e.g., block artifacts and ringing artifacts) generated during image / video encoding and enhance subjective / objective visual quality. Furthermore, by performing in-loop filtering in both the encoding and decoding devices, the encoding and decoding devices can derive the same prediction results, improving the reliability of image encoding and reducing the amount of data to be transmitted for image encoding.

[0108] As described above, the image reconstruction process can be performed in both the encoding and decoding devices. Reconstructed blocks can be generated on a per-block basis based on intra-frame prediction / inter-frame prediction, and a reconstructed image including these blocks can be generated. If the current image / slice / patch group is an I-frame / slice / patch group, the blocks included in the current image / slice / patch group can be reconstructed based solely on intra-frame prediction. Furthermore, if the current image / slice / patch group is a P-frame / slice / patch group or a B-frame / slice / patch group, the blocks included in the current image / slice / patch group can be reconstructed based on either intra-frame prediction or inter-frame prediction. In this case, inter-frame prediction can be applied to some blocks in the current image / slice / patch group, and intra-frame prediction can be applied to other blocks. The color components of the image can include luma and chroma components, and the methods and exemplary implementations presented in this document can be applied to both luma and chroma components, unless explicitly limited herein.

[0109] Figure 5 Examples illustrating a video / image decoding process to which one or more embodiments of this document may be applied. Figure 5 In the above, the S500 can be... Figure 3S500 can be executed in the entropy decoder 310 of the decoding device; S510 can be executed in the predictor 330; S520 can be executed in the residual processor 320; S530 can be executed in the adder 340; and S540 can be executed in the filter 350. S500 can include the information decoding process described in this document; S510 can include the inter-frame / intra-frame prediction process described in this document; S520 can include the residual processing process described in this document; S530 can include the block / frame reconstruction process described in this document; and S540 can include the in-loop filtering process described in this document.

[0110] Reference Figure 5 , such as regarding Figure 3 As described herein, the image decoding process can schematically include (through decoding) an image / video information acquisition process S500 from the bitstream, image reconstruction processes S510 to S530, and an in-loop filtering process S540 for the reconstructed image. The image reconstruction process can be performed based on residual samples and prediction samples obtained through the inter-frame / intra-frame prediction S510 and residual processing S520 (dequantization and inverse transform for quantization transform coefficients) processes described in this document. By performing in-loop filtering on the reconstructed image generated through the image reconstruction process, a modified reconstructed image can be generated. This modified reconstructed image can be output as a decoded image and can also be stored in the decoded image buffer or memory 360 of the decoding device, and can be used as a reference image in the inter-frame prediction process for later image decoding.

[0111] Depending on the circumstances, the in-loop filtering process can be skipped, and in this case, the reconstructed frame can be output as the decoded frame and stored in the decoded frame buffer or memory 360 of the decoding device, and can be used as a reference frame in the inter-frame prediction process for later frame decoding. The in-loop filtering process S540 may include the deblocking filtering process, the sample adaptive offset (SAO) process, the adaptive loop filter (ALF) process, and / or the bidirectional filtering process as described above, and all or some of them may be skipped. Furthermore, one or a portion of the deblocking filtering process, the sample adaptive offset (SAO) process, the adaptive loop filter (ALF) process, and the bidirectional filtering process may be applied sequentially, or all of them may be applied sequentially. For example, the SAO process may be performed on the reconstructed frame after the deblocking filtering process has been applied. Alternatively, for example, the ALF process may be performed on the reconstructed frame after the deblocking filtering process has been applied. This can also be performed in the encoding device.

[0112] Furthermore, as mentioned above, the encoding device performs entropy encoding based on various encoding methods such as Exponential Golomb coding, Context Adaptive Variable Length Coding (CAVLC), and Context Adaptive Binary Arithmetic Coding (CABAC). Similarly, the decoding device can perform entropy decoding based on encoding methods such as Exponential Golomb coding, CAVLC, or CABAC. The entropy encoding / decoding process will be described below.

[0113] Figure 6 Examples of entropy coding methods to which the implementation methods described in this document can be applied are illustrated schematically. Figure 7 An entropy encoder in an encoding device is illustrated schematically. Figure 7 The entropy encoder in the encoding device can also be applied equivalently or correspondingly to... Figure 2 The aforementioned entropy encoder 240 of the encoding device 200.

[0114] Reference Figure 6 and Figure 7 The encoding device (entropy encoder) performs entropy coding processing on image / video information. Image / video information may include partitioning-related information, prediction-related information (e.g., inter-frame / intra-frame prediction differentiation information, intra-frame prediction mode information, inter-frame prediction mode information, etc.), residual information, in-loop filtering-related information, or may include various syntax elements related to them. Entropy coding can be performed on a syntax element-by-syntax basis. S600 and S610 can be... Figure 2 The above-mentioned entropy encoder 240 of the encoding device 200 is used to perform the encoding.

[0115] The encoding device can perform binary conversion on the target syntax element (S600). Here, binary conversion can be based on various binary conversion methods, such as a truncated Rice binary conversion process, a fixed-length binary conversion process, etc., and a binary conversion method for the target syntax element can be predefined. The binary conversion process can be performed by the binary converter 242 in the entropy encoder 240.

[0116] The encoding device can perform entropy encoding on the target syntax element (S610). The encoding device can perform conventional (context-based) or bypass-based encoding on the bin string of the target syntax element based on an entropy encoding scheme (e.g., context-adaptive arithmetic coding (CABAC) or context-adaptive variable-length coding (CAVLC)), and can incorporate its output into the bitstream. The entropy encoding process can be executed by the entropy encoding processor 243 in the entropy encoder 240. As described above, the bitstream can be transmitted to the decoding device via a (digital) storage medium or a network.

[0117] Figure 8 Examples of entropy decoding methods to which the implementation methods described in this document are applicable are illustrated schematically, and Figure 9 An entropy decoder in a decoding device is illustrated schematically. Figure 9 The entropy decoder in the decoding device can also be applied equivalently or correspondingly to Figure 3 The aforementioned entropy decoder 310 of the decoding device 300.

[0118] Reference Figure 8 and Figure 9 The decoding device (entropy decoder) can decode encoded image / video information. Image / video information may include segmentation-related information, prediction-related information (e.g., inter-frame / intra-frame prediction differentiation information, intra-frame prediction mode information, inter-frame prediction mode information, etc.), residual information, in-loop filtering-related information, or may include various syntax elements associated with them. Entropy coding can be performed on a syntax element-by-syntax basis. S800 and S810 can be... Figure 3 The aforementioned entropy decoder 310 of the decoding device 300 is used to perform the decoding.

[0119] The decoding device can perform binary conversion on the target syntax element (S800). Here, binary conversion can be based on various binary conversion methods such as a truncated Rice binary conversion process, a fixed-length binary conversion process, etc., and a predefined binary conversion method for the target syntax element can be used. The decoding device can deduce the enable bin string (bin string candidate) of the enable value of the target syntax element through the binary conversion process. The binary conversion process can be performed by the binary converter 312 in the entropy decoder 310.

[0120] The decoding device can perform entropy decoding (S810) on the target syntax element. While sequentially decoding and parsing each bin of the target syntax element from the input bits in the bitstream, the decoding device compares the derived bin string with the enabled bin string of the corresponding syntax element. When the derived bin string matches one of the enabled bin strings, the value corresponding to the bin string can be derived as the value of the syntax element. If not, the above process can be repeated after further parsing the next bit in the bitstream. Through these processes, even if no start or end bit is used for specific information (specific syntax element) in the bitstream, the decoding device can use variable-length bits to signal information. Therefore, relatively fewer bits can be allocated to low values, thereby improving overall encoding efficiency.

[0121] The decoding device can perform context-based or bypass-based decoding on the corresponding bins in the bin string from the bitstream based on entropy coding techniques such as CABAC and CAVLC. In this regard, the bitstream can include various information for image / video decoding as described above. As mentioned above, the bitstream can be transmitted to the decoding device via (digital) storage media or a network.

[0122] In this document, a table including syntax elements (syntax table) is used to indicate signaling of information from an encoding device to a decoding device. The order of syntax elements in the table including syntax elements used in this document indicates the parsing order of syntax elements from the bitstream. The encoding device can construct and encode the syntax table such that the decoding device can parse the syntax elements in the parsing order, and the decoding device can obtain the values ​​of syntax elements by parsing and decoding the syntax elements of the corresponding syntax table from the bitstream according to the parsing order.

[0123] Figure 10 An example of a layered structure for encoded images / videos is shown.

[0124] Reference Figure 10 The encoded image / video is divided into the VCL (Video Coding Layer) which handles the image / video decoding process and its own subsystem, the subsystem for sending and storing encoded information, and the Network Abstraction Layer (NAL) which exists between the VCL and the subsystem and is responsible for network adaptation functions.

[0125] VCL can generate VCL data that includes compressed image data (slice data), or parameter sets that include picture parameter sets (PPS), sequence parameter sets (SPS), video parameter sets (VPS), etc., or supplementary enhancement information (SEI) messages required for the image decoding process.

[0126] In NAL, NAL units can be generated by adding header information (NAL unit header) to the raw byte sequence payload (RBSP) generated in VCL. In this case, RBSP refers to slice data, parameter sets, SEI messages, etc., generated in VCL. The NAL unit header can include NAL unit type information specified based on the RBSP data included in the corresponding NAL unit.

[0127] Furthermore, based on the RBSP generated in the VCL, NAL units can be divided into VCL NAL units and non-VCL NAL units. VCL NAL units can refer to NAL units that include information about the image (slice data), while non-VCL NAL units can refer to NAL units that include information required for decoding the image (parameter set or SEI message).

[0128] VCL NAL units and non-VCL NAL units can be transmitted over a network by appending header information according to the subsystem's data standard. For example, NAL units can be converted into predetermined standard data formats (e.g., H.266 / VVC file format, Real-Time Transport Protocol (RTP), and Transport Stream (TS), etc.) and transmitted over various networks.

[0129] As described above, in a NAL unit, the NAL unit type can be specified according to the RBSP data structure included in the corresponding NAL unit, and information about the NAL unit type can be stored in the NAL unit header and signaled.

[0130] For example, based on whether a NAL unit includes information about the image (slice data), NAL units can be roughly classified into VCL NAL unit types and non-VCL NAL unit types. VCL NAL unit types can be classified according to the attributes and type of the image included in the VCL NAL unit, while non-VCL NAL unit types can be classified according to the type of parameter set.

[0131] The following is an example of a NAL cell type specified based on the type of the parameter set included in a non-VCL NAL cell type.

[0132] -DCI (Decoding Capability Information) NAL Unit: Includes the type of NAL unit for DCI.

[0133] -VPS (Video Parameter Set) NAL Unit: Includes the type of NAL unit for the VPS.

[0134] -SPS (Sequence Parameter Set) NAL Unit: The type of NAL unit that includes SPS.

[0135] -PPS (Picture Parameter Set) NAL Unit: Includes the types of NAL units for PPS.

[0136] -APS (Adaptive Parameter Set) NAL Units: Types of NAL units including APS.

[0137] -PH (Picture Header) NAL Unit: Includes the type of NAL unit for PH.

[0138] The aforementioned NAL unit type has syntax information specific to the NAL unit type, which can be stored in the NAL unit header and signaled. For example, the syntax information can be nal_unit_type, and the NAL unit type can be specified by the nal_unit_type value.

[0139] Furthermore, as mentioned above, a frame can include multiple slices, and a slice can include a slice header and slice data. In this case, a frame header can also be added to multiple slices in a frame (slice header and slice data set). The frame header (frame header syntax) can include information / parameters that can be applied together to the frame. The slice header (slice header syntax) can include information / parameters that can be applied together to the slice. APS (APS syntax) or PPS (PPS syntax) can include information / parameters that can be applied together to one or more slices or frames. SPS (SPS syntax) can include information / parameters that can be applied together to one or more sequences. VPS (VPS syntax) can include information / parameters that can be applied together to multiple layers. DCI (DCI syntax) can include information / parameters that can be applied together on the video. DCI can include information / parameters related to decoding capabilities. In this document, the High-Level Syntax (HLS) can include at least one of the following: APS syntax, PPS syntax, SPS syntax, VPS syntax, DCI syntax, frame header syntax, and slice header syntax. In addition, in this document, low-level syntax (LLS) may include, for example, slice data syntax, CTU syntax, coding unit syntax, transform unit syntax, etc.

[0140] Furthermore, in this document, tile groups can be used interchangeably or replaced with slices or screens. Additionally, in this document, tile group headers can be used interchangeably or replaced with slice headers or screen headers.

[0141] In this paper, the image / video information encoded from the encoding device in the form of a bitstream and signaled to the decoding device may include not only information related to intra-frame segmentation, intra / inter-frame prediction information, residual information, and in-loop filtering information, but also slice header information, frame header information, APS information, PPS information, SPS information, and / or DCI information. Additionally, the image / video information may also include general constraint information and / or NAL unit header information.

[0142] As described above, High-Level Syntax (HLS) can be encoded / signaled for use in video / image coding. In this document, video / image information can include HLS, and video / image coding methods can be performed based on this information. For example, a frame to be encoded can be constructed using one or more slices. Parameters describing the frame to be encoded can be signaled in the Frame Header (PH), and parameters describing the slices can be signaled in the Slice Header (SH). The PH can be sent with its own NAL unit type. The SH can be present at the beginning of a NAL unit that includes the slice's payload (i.e., slice data). Details of the syntax and semantics of the PH and SH can be as disclosed in the VVC standard. Each frame can be associated with a PH. Frames can be constructed using different types of slices: intra-frame coded slices (i.e., I-slices) and inter-frame coded slices (i.e., P-slices and B-slices). As a result, the PH can include the syntax elements necessary for intra-frame and inter-frame slices of a frame.

[0143] In addition, typically, a single NAL unit type can be set for a single frame. The NAL unit type can be indicated by a signal in the NAL unit header, specifically the `nal_unit_type` parameter, which includes the slice's NAL unit. `nal_unit_type` is syntax information used to specify the NAL unit type; that is, as shown in Table 1 or Table 2 below, it specifies the type of RBSP data structures included in the NAL unit.

[0144] Table 1 below shows examples of NAL unit type codes and NAL unit type classes.

[0145] [Table 1]

[0146]

[0147] Alternatively, as an example, NAL unit type codes and NAL unit type classes can be defined as shown in Table 2 below.

[0148] [Table 2]

[0149]

[0150]

[0151] As shown in Table 1 or Table 2, the name and value of the NAL unit type can be specified based on the RBSP data structure included in the NAL unit, and can be classified into VCLNAL unit type and non-VCL NAL unit type based on whether the NAL unit includes information about the image (slice data). VCL NAL unit types can be classified based on the nature and type of the image, while non-VCL NAL unit types can be classified based on the type of parameter set, etc. For example, the NAL unit type can be specified based on the nature and type of the image included in the VCL NAL unit as follows.

[0152] TRAIL: This indicates the type of NAL unit that includes the coded slice data of the end frame / subframe. For example, nal_unit_type can be defined as TRAIL_NUT, and the value of nal_unit_type can be specified as 0.

[0153] Here, the final screen refers to the screen that can be randomly accessed in both output and decoding order. The final screen can be a non-IRAP screen that follows the associated IRAP screen in output order, and it is not an STSA screen. For example, the final screen associated with an IRAP screen may follow the IRAP screen in decoding order. Screens that are not allowed to follow the associated IRAP screen in output order and precede the associated IRAP screen in decoding order are not permitted.

[0154] STSA (Step-by-Step Temporal Sublayer Access): This indicates the type of NAL unit that contains the coded slice data of the STSA frame / subframe. For example, nal_unit_type can be defined as STSA_NUT, and the value of nal_unit_type can be specified as 1.

[0155] Here, an STSA frame is a frame that can be switched between temporal sub-layers in a bitstream that supports temporal scalability, and it indicates the position from which a switch can be made from a lower sub-layer upwards to an upper sub-layer. STSA frames do not use the same TemporalId as STSA frames or frames in the same layer as STSA frames for inter-frame prediction reference. Frames following STSA frames in the same layer in decoding order, and frames with the same TemporalId as STSA frames, do not use the same TemporalId as STSA frames preceding STSA frames in the same layer in decoding order. STSA frames enable switching from the immediately following lower sub-layer upwards within an STSA frame to a sub-layer that includes the STSA frame. In this case, the encoded frame must not belong to the lowest sub-layer. That is, the STSA frame must always have a TemporalId greater than 0.

[0156] RADL (Random Access Decodeable Precursor (Frame)): This indicates the type of NAL unit that contains the encoded slice data of the RADL frame / subframe. For example, nal_unit_type can be defined as RADL_NUT, and the value of nal_unit_type can be specified as 2.

[0157] Here, all RADL frames are leading frames. RADL frames are not used as reference frames for decoding the final frame of the same associated IRAP frame. Specifically, a RADL frame with a nuh_layer_id equal to its layerId is the frame that follows the IRAP frame associated with it in output order, and is not used as a reference frame for decoding the frame with the nuh_layer_id equal to its layerId. When field_seq_flag (i.e., sps_field_seq_flag) is 0, all RADL frames (i.e., if any) precede all non-leading frames of the same associated IRAP frame in decoding order. Furthermore, a leading frame is a frame that precedes the associated IRAP frame in output order.

[0158] RASL (Random Access Skip Preview (Picture)): This indicates the type of NAL unit in the encoded slice data that includes the RASL picture / subpicture. For example, nal_unit_type can be defined as RASL_NUT, and the value of nal_unit_type can be specified as 3.

[0159] Here, all RASL frames are preceding frames of associated CRA frames. When an associated CRA frame has a NoOutputBeforeRecoveryFlag value of 1, the RASL frame cannot be output or correctly decoded because it may include references to frames not present in the bitstream. RASL frames are not used as reference frames for decoding non-RASL frames of the same layer. However, RADL sub-frames in RASL frames of the same layer can be used for inter-frame prediction of juxtaposed RADL sub-frames in RADL frames associated with the same CRA frame as the RASL frame. When field_seq_flag (i.e., sps_field_seq_flag) is 0, all RASL frames (i.e., if RASL frames exist) precede all non-preceding frames of the same associated CRA frame in decoding order.

[0160] A reserved nal_unit_type may exist for non-IRAP VCL NAL unit types. For example, nal_unit_type can be defined as RSV_VCL_4 to RSV_VCL_6, and the value of nal_unit_type can be specified as 4 to 6 respectively.

[0161] Here, Intra-Frame Random Access Point (IRAP) is information indicating the NAL unit of a frame that can be randomly accessed. An IRAP frame can be a CRA frame or an IDR frame. For example, as shown in Table 1 or Table 2 above, an IRAP frame refers to a frame with NAL unit types where nal_unit_type is defined as IDR_W_RADL, IDR_N_LP, and CRA_NUT, and the value of nal_unit_type can be specified as 7 to 9 respectively.

[0162] IRAP frames do not use any reference frames in the same layer for inter-frame prediction during decoding. In other words, IRAP frames do not reference any frames other than themselves for inter-frame prediction during decoding. The first frame in the bitstream in decoding order becomes an IRAP or GDR frame. For a single-layer bitstream, if the necessary parameter set is available when reference is needed, all subsequent non-RASL and IRAP frames in the coded layer video sequence (CLVS) in decoding order can be decoded accurately without performing decoding processing on frames that precede the IRAP frames in decoding order.

[0163] The value of `mixed_nalu_types_in_pic_flag` for an IRAP frame is 0. When the value of `mixed_nalu_types_in_pic_flag` for a frame is 0, one slice in the frame can have a NAL unit type (nal_unit_type) in the range from IDR_W_RADL to CRA_NUT (e.g., a value of 7 to 9 for NAL unit types in Table 1 or Table 2), and all other slices in the frame can have the same NAL unit type (nal_unit_type). In this case, the frame can be considered an IRAP frame.

[0164] Instantaneous Decoding Refresh (IDR): This indicates the type of NAL unit in the encoded slice data that includes the IDR frame / subframe. For example, the nal_unit_type for the IDR frame / subframe can be defined as IDR_W_RADL or IDR_N_LP, and the value of nal_unit_type can be specified as 7 or 8, respectively.

[0165] Here, an IDR frame may not use inter-frame prediction during decoding (i.e., it does not reference frames other than itself for inter-frame prediction), but it can be the first frame in the bitstream in decoding order, or it may appear later in the bitstream (i.e., not first, but later). Each IDR frame is the first frame of a coded video sequence (CVS) in decoding order. For example, when an IDR frame is associated with a decorable breezy frame, its NAL unit type can be represented as IDR_W_RADL, while when an IDR frame is not associated with a breezy frame, its NAL unit type can be represented as IDR_N_LP. That is, an IDR frame with NAL unit type IDR_W_RADL may not have an associated RASL frame in the bitstream, but it may have an associated RADL frame in the bitstream. An IDR frame with NAL unit type IDR_N_LP does not have an associated breezy frame in the bitstream.

[0166] Clean Random Access (CRA): This indicates the type of NAL unit that contains the encoded slice data of the CRA frame / subframe. For example, nal_unit_type can be defined as CRA_NUT, and the value of nal_unit_type can be specified as 9.

[0167] Here, a CRA frame may not use inter-frame prediction during decoding (i.e., it does not reference frames other than itself for inter-frame prediction), but it can be the first frame in the bitstream in decoding order, or it may appear later in the bitstream (i.e., not first, but later). A CRA frame can have associated RADL or RASL frames present in the bitstream. For a CRA frame where the NoOutputBeforeRecoveryFlag value is 1, the decoder may not output the associated RASL frame. This is because decoding is impossible in this case due to the inclusion of references to frames not present in the bitstream.

[0168] Gradual Decoding Refresh (GDR): This indicates the type of NAL unit in the encoded slice data that includes the GDR frame / subframe. For example, nal_unit_type can be defined as GDR_NUT, and the value of nal_unit_type can be specified as 10.

[0169] Here, the value of pps_mixed_nalu_types_in_pic_flag for a GDR frame can be 0. When the value of pps_mixed_nalu_types_in_pic_flag for a frame is 0 and one slice in the frame has a GDR_NUT NAL unit type, all other slices in the frame have the same NAL unit type (nal_unit_type), and in this case, the frame can become a GDR frame after the first slice is received.

[0170] Furthermore, for example, the NAL unit type can be specified based on the types of parameters included in non-VCL NAL units, and as shown in Table 1 or Table 2 above, it can include NAL unit types such as VPS_NUT indicating the type of NAL unit including video parameter sets, SPS_NUT indicating the type of NAL unit including sequence parameter sets, PPS_NUT indicating the type of NAL unit including picture parameter sets, and PH_NUT indicating the type of NAL unit including picture headers.

[0171] Furthermore, the bitstream supporting time scalability (or time-scalable bitstream) includes information about the time layer regarding time scaling. This time layer information can be identification information for the time layer specified according to the time scalability of the NAL unit. For example, the time layer identification information can use the temporal_id syntax information, and this temporal_id syntax information can be stored in the NAL unit header in the encoding device and signaled to the decoding device. In the following, in this specification, the time layer may be referred to as a sublayer, time sublayer, time-scalable layer, etc.

[0172] Figure 11 This is a diagram illustrating the temporal layer structure of NAL cells in a bitstream that supports temporal scalability.

[0173] When a bitstream supports temporal scalability, the NAL units included in the bitstream have temporal layer identification information (e.g., temporal_id). As an example, a temporal layer constructed from NAL units with a temporal_id value of 0 provides the lowest temporal scalability, while a temporal layer constructed from NAL units with a temporal_id value of 2 provides the highest temporal scalability.

[0174] exist Figure 11 In the diagram, boxes marked with an "I" represent screen I, and boxes marked with a "B" represent screen B. Additionally, arrows indicate whether a screen references another screen.

[0175] like Figure 11As shown, a NAL unit in a time layer with a temporal_id value of 0 is a reference frame that can be referenced by NAL units in time layers with temporal_id values ​​of 0, 1, or 2. A NAL unit in a time layer with a temporal_id value of 1 is a reference frame that can be referenced by NAL units in time layers with temporal_id values ​​of 1 or 2. A NAL unit in a time layer with a temporal_id value of 2 can be a reference frame that can be referenced by NAL units in the same time layer (i.e., the time layer with a temporal_id value of 2), or it can be a non-reference frame that is not referenced by other frames.

[0176] If, as Figure 11 As shown, NAL units in the time layer (i.e., the highest time layer) with a temporal_id value of 2 are non-reference frames. These NAL units are extracted (or removed) from the bitstream during decoding without affecting other frames.

[0177] Furthermore, among the NAL unit types mentioned above, the IDR and CRA types indicate information about NAL units that include frames capable of random access (or splicing) (i.e., random access point (RAP) frames used as random access points or internal random access point (IRAP) frames). In other words, an IRAP frame can be an IDR or CRA frame and can consist only of I-slices. In the bitstream, the first frame in the decoding order becomes the IRAP frame.

[0178] If IRAP frames (IDR, CRA frames) are included in the bitstream, there may be frames that appear before the IRAP frames in output order but after the IRAP frames in decoding order. These frames are called lead frames (LPs).

[0179] Figure 12 It is a diagram used to describe screens that can be accessed randomly.

[0180] A frame that can be randomly accessed (i.e., a RAP or IRAP frame used as a random access point) is the first frame in the bitstream during random access in decoding order and consists only of I slices.

[0181] Figure 12 The output (or display) order and decoding order of the screens are shown. As illustrated, the output and decoding orders of the screens can differ from each other. For convenience, the screens are described while being divided into predetermined groups.

[0182] The screens belonging to Group 1 (I) are screens that precede the IRAP screens in both output and decoding order. The screens belonging to Group 2 (II) are screens that precede the IRAP screens in output order but follow the IRAP screens in decoding order. The screens belonging to Group 3 (III) are screens that follow the IRAP screens in both output and decoding order.

[0183] The first group (I) of images can be decoded and output independently of IRAP images.

[0184] The frame that is output before the IRAP frame and belongs to the second group (II) is called the lead frame, and the lead frame may cause problems in the decoding process when the IRAP frame is used as a random access point.

[0185] The third group (III) of the screens following the IRAP screens in both output and decoding order is called the normal screen. The normal screen is not used as a reference screen for the lead screen.

[0186] The random access point in the bitstream becomes the IRAP screen, and random access begins when the first screen of the second group (II) is output.

[0187] Figure 13 It is a diagram used to describe the IDR screen.

[0188] An IDR (Independent Rendering) frame is a frame that becomes a random access point when the frame group has a closed structure. As mentioned above, since an IDR frame is an IRAP frame, it only consists of I-slices and can be the first frame in the bitstream in decoding order, or it can appear in the middle of the bitstream. When an IDR frame is decoded, all reference frames stored in the Decoded Frame Buffer (DPB) are marked as "not used for reference".

[0189] Figure 13 The bar indicates the screen, and the arrow indicates the reference relationship between the screen and another screen. The 'x' mark on the arrow indicates that the screen cannot reference the screen indicated by the arrow.

[0190] As shown in the figure, a frame with a POC of 32 is an IDR frame. A frame with a POC of 25 to 31, and output before an IDR frame, is a pilot frame 1310. A frame with a POC equal to or greater than 33 corresponds to a normal frame 1320.

[0191] The lead screen 1310, which precedes the IDR screen in the output order, can use a lead screen different from the IDR screen as a reference screen, but it is not necessary to use the past screen 1330, which precedes the lead screen 1310 in the output and decoding order, as a reference screen.

[0192] You can refer to the IDR screen, the lead screen, and other normal screens to decode the normal screen 1320, which is after the IDR screen in the output and decoding order.

[0193] Figure 14 It is a diagram used to describe CRA screens.

[0194] A CRA (Cross-Area Rendering) frame is a frame that becomes a random access point when the frame group has an open structure. As mentioned above, since a CRA frame is also an IRAP frame, it only consists of I-slices and can be the first frame in the bitstream in decoding order, or it can appear in the middle of the bitstream for normal playback.

[0195] Figure 14 The bar indicates the screen, and the arrow indicates the reference relationship between the screens, suggesting whether another screen can be used as a reference screen. The 'x' mark on the arrow indicates that the screen cannot reference one or more screens indicated by the arrow.

[0196] The preview screen 1410, which precedes the CRA screen in the output order, can use all CRA screens, other preview screens, and past screens 1430, which precede the preview screen 1410 in the output and decoding order, as reference screens.

[0197] Conversely, a normal screen, which is different from the CRA screen, can be referenced to decode the normal screen 1420, which follows the CRA screen in both output and decoding order. The normal screen 1420 does not need to use the preceding screen 1410 as a reference screen.

[0198] Furthermore, the VVC standard allows the frame being encoded (i.e., the current frame) to include slices of different NAL unit types. Whether the current frame includes slices of different NAL unit types can be indicated based on the syntax element `pps_mixed_nalu_types_in_pic_flag`. For example, when the current frame includes slices of different NAL unit types, the value of the syntax element `pps_mixed_nalu_types_in_pic_flag` can be represented as 1. In this case, the current frame must reference a PPS that includes a `pps_mixed_nalu_types_in_pic_flag` with a value of 1. The semantics of the flag (`pps_mixed_nalu_types_in_pic_flag`) are as follows:

[0199] When the value of the syntax element pps_mixed_nalu_types_in_pic_flag is 1, it can indicate that each frame of the reference PPS has one or more VCL NAL units, the VCL NAL units do not have the same NAL unit type (nal_unit_type), and the frame is not an IRAP frame.

[0200] When the value of the syntax element pps_mixed_nalu_types_in_pic_flag is 0, it can indicate that each frame of the reference PPS has one or more VCL NAL units, and that the VCL NAL units of each frame of the reference PPS have the same value of NAL unit type (nal_unit_type).

[0201] As mentioned above, slices in the image can have different NAL unit types. For example, when the value of pps_mixed_nalu_types_in_pic_flag is 1, the following method can be applied.

[0202] - The frame must include at least two sub-frames.

[0203] - The VCL NAL unit of the screen must have two or more different NAL unit type (nal_unit_type) values.

[0204] - There must be no VCLNAL unit in the frame that has the same NAL unit type (nal_unit_type) as GDR_NUT.

[0205] - When at least one sub-frame in a frame has a VCL NAL unit with a specific value of a NAL unit type (nal_unit_type) such as IDR_W_RADL, IDR_N_LP, or CRA_NUT, all VCL NAL units in other sub-frames within the frame must have the same NAL unit type (nal_unit_type) as TRAIL_NUT.

[0206] Furthermore, if the picture contains a mixed slice of RASL_NUT and RADL_NUT, the picture can be treated as a RASL picture and can be processed as described below to derive PictureOutputFlag.

[0207] Here, PictureOutputFlag can be flag information related to the output of the picture. For example, if the PictureOutputFlag of the currently decoded picture is 1, it is marked as "output required"; if the PictureOutputFlag of the currently decoded picture is 0, it is marked as "output not required". RASL (Random Access Skip Precursor) pictures include at least one VCL NAL unit with a NAL unit type (nal_unit_type) of RASL_NUT, and all other VCL NAL units refer to encoded pictures with a NAL unit type (nal_unit_type) of RASL_NUT or RADL_NUT.

[0208] For example, the PictureOutputFlag of the current screen can be derived as follows.

[0209] - If sps_video_parameter_set_id is greater than 0 and the current layer is not an output layer (i.e., nuh_layer_id is not equal to OutputLayerIdInOls[TargetOlsIdx][i] for a value of i ranging from 0 to NumOutputLayersInOls[TargetOlsIdx]-1), or if one of the following conditions is true, then PictureOutputFlag is set to 0.

[0210] If the current screen is a RASL screen and the associated IRAP screen's NoOutputBeforeRecoveryFlag is 1

[0211] If the current screen is a GDR screen where NoOutputBeforeRecoveryFlag is 1, or a recovery screen where NoOutputBeforeRecoveryFlag is 1.

[0212] - Otherwise, PictureOutputFlag can be set to equal ph_pic_output_flag.

[0213] For reference, in terms of implementation, the decoder can output frames that do not belong to the output layer. For example, if a frame of the output layer within an AU is unavailable (e.g., due to loss or layer downswitching) and only one output layer exists, the decoder can set PictureOutpFlag to 1 for the frame with the highest nuh_layer_id value and ph_pic_output_flag equal to 1 among all frames of the AU available in the decoder, and set PictureOutputFlag to 0 for all other frames of the AU available in the decoder.

[0214] In addition, the VVC standard supports Progressive Decode & Refresh (GDR) functionality. This feature allows decoding to begin from a point where all parts of the reconstructed frame have not been correctly decoded. However, the correctly decoded portions of the reconstructed frame are gradually added to the subsequence of frames until the entire frame is correctly decoded. A frame from which decoding can begin using GDR is called a GDR frame, and the first frame after a GDR frame where the entire frame is correctly decoded is called a recovery point frame (or recovery frame).

[0215] The aforementioned NoOutputBeforeRecoveryFlag can be information indicating whether a decoded screen of a GDR screen preceding a recovery point screen, in either POC or decoding order, can be output. For example, it can be determined that a (decoded) screen between a GDR screen and a recovery point screen will not be output, and in this case, the value of NoOutputBeforeRecoveryFlag can be set to 1. Alternatively, it can be determined that a (decoded) GDR screen will not be output, and in this case, the value of NoOutputBeforeRecoveryFlag can be set to 1. However, it can be determined that a (decoded) recovery point screen will be output, and in this case, the value of NoOutputBeforeRecoveryFlag can be set to 0. That is, in the case of a screen with a POC from the GDR screen to a POC preceding the recovery point POC, the value of NoOutputBeforeRecoveryFlag can be determined to be 1.

[0216] On the other hand, the current VVC standard has the following problems with the above-mentioned output of the image.

[0217] 1) Before decoding the current frame (but after parsing the slice header of the first slice in the current frame), processing related to frame output is invoked. However, it is unclear whether the derivation / determination of the frame output flags is invoked only for the first slice of the frame.

[0218] 2) In the current standard, the decoder derives the value of the picture output flag when it receives the first slice of a picture. When the decoder receives the first slice of a picture where the value of pps_mixed_nalu_types_in_pic_flag becomes 1, and the NAL unit type of the slice is RASL_NUT or RADL_NUT, the decoder cannot determine whether the picture is a RASL picture. Because the remaining slices of the picture can become TRAIL_NUT, the picture may become a non-leading picture.

[0219] In other words, when a picture includes a mixed NAL unit type, one of which is RASL_NUT, it is difficult to determine the value of a picture output flag (e.g., PictureOutputFlag) after only receiving the first slice of the picture. Therefore, the process of determining the value of the picture output flag (e.g., PictureOutputFlag) requires constraints to operate correctly. That is, this paper proposes a method for setting the value of a picture output flag (e.g., PictureOutputFlag) that specifies whether a picture should be output when the picture includes different NAL unit types (particularly when the leading picture (i.e., RASL_NUT and / or RADL_NUT includes a mixed NAL unit type)).

[0220] For example, the following implementation methods are proposed to solve the above problems, and the following implementation methods can be applied individually or in combination of one or more.

[0221] 1. The derivation of the value of the picture output flag (e.g., PictureOutputFlag) can be applied only once per frame. That is, the value of the picture output flag (e.g., PictureOutputFlag) can be derived only for the first slice of the frame.

[0222] a) Processing related to screen output can be invoked during the parsing of the slice header of the first slice in the screen.

[0223] 2. When the screen includes NAL unit types other than RADL_NUT mixed with RASL_NUT, or when RADL_NUT is mixed with NAL unit types other than RASL_NUT, the following conditions can be applied.

[0224] a) When interleaving coding is not used (i.e., when sps_field_seq_flag is 0), there must be at least one non-leading screen between the screen and the associated IRAP screen.

[0225] b) Otherwise, when interleaving coding is used (i.e., when sps_field_seq_flag is 1), there must be at least two non-leading frames between the frame and the associated IRAP frame.

[0226] 3. If interleaving is not used, if the decoder receives the first slice of the picture and determines that the value of pps_mixed_nalu_types_in_pic_flag is equal to 1, and if the NAL unit type of the slice is RASL_NUT or RADL_NUT at this time, and the decoder has not received a non-leading picture from the last IRAP picture, then the decoder can determine that the picture including the first slice is a RASL picture.

[0227] 4. If interleaving is used, if the decoder receives the first slice of the picture and determines that the value of pps_mixed_nalu_types_in_pic_flag is equal to 1, and if the NAL unit type of the slice is RASL_NUT or RADL_NUT at this time, and the decoder receives exactly one non-leading picture from the last IRAP picture, then the decoder can determine that the picture including the first slice is a RASL picture.

[0228] The embodiments described above in this document can be implemented in the forms disclosed in Table 3 below. Table 3 shows examples of implementations of the above embodiments related to the output of a screen in the VVC specification.

[0229] [Table 3]

[0230]

[0231]

[0232]

[0233]

[0234]

[0235] Figure 15 Examples of video / image coding methods to which one or more of the above embodiments of this document can be applied are illustrated. Figure 15 The method disclosed in the article can be derived from Figure 2 The encoding device 200 disclosed herein performs the operation. Furthermore, according to the embodiment, the following can be omitted. Figure 15 It can be one or more steps, and additional steps can be added.

[0236] Reference Figure 15The encoding device can determine the NAL unit type of the slice in the image (S1500) and generate NAL unit type related information (S1510).

[0237] Here, NAL unit type related information may include information / syntax elements related to the NAL unit types disclosed in Table 1 or Table 2 above. For example, NAL unit type related information may include the mixed_nalu_types_in_pic_flag syntax element of PPS and / or the nal_unit_type syntax element of the NAL unit header that includes information about the encoded slice.

[0238] When the value of mixed_nalu_types_in_pic_flag is 0, the slices in the frame associated with PPS use the same NAL unit type. That is, when the value of mixed_nalu_types_in_pic_flag is 0, the NAL unit type defined in the first NAL unit header of the first NAL unit that includes information about the first slice of the frame is the same as the NAL unit type defined in the second NAL unit header of the second NAL unit that includes information about the second slice in the same frame.

[0239] When the value of mixed_nalu_types_in_pic_flag is 1, according to the implementation described above in this document, slices in a frame associated with PPS can use other NAL unit types. For example, when the value of mixed_nalu_types_in_pic_flag is 1, slices in a frame (associated with PPS) can use different NAL unit types, but slices in the same sub-frame can use the same NAL unit type.

[0240] For example, a picture may include sub-picture A and sub-picture B. In this case, slices in sub-picture A use the same NAL unit type (NAL unit type A), and slices in sub-picture B use the same NAL unit type (NAL unit type B), but when the value of mixed_nalu_types_in_pic_flag is 1, NAL unit type A and NAL unit type B are different. The NAL unit type of a slice (i.e., the NAL unit that includes information about the slice) can be one of the indices 0 to 10 disclosed in Table 1 or Table 2 above.

[0241] The encoding device can generate a bitstream that includes at least one NAL unit containing information about the encoded slice (S1520).

[0242] Here, the bitstream may include NAL units containing information about the encoded slices. The bitstream may include PPS.

[0243] Figure 16 Examples of video / image decoding methods to which one or more of the above embodiments of this document may be applied are illustrated. Figure 16 The method disclosed in the article can be derived from Figure 3 The decoding device 300 disclosed herein performs the operation. Furthermore, according to the embodiment, the following can be omitted. Figure 16 It can be one or more steps, and additional steps can be added.

[0244] Reference Figure 16 The decoding device can receive a bit stream including at least one NAL unit containing information about the encoded slice (S1600), and can obtain information related to the NAL unit type (S1610).

[0245] As described above, NAL unit type related information may include information / syntax elements related to the NAL unit types disclosed in Table 1 or Table 2 above. For example, NAL unit type related information may include the mixed_nalu_types_in_pic_flag syntax element of PPS and / or the nal_unit_type syntax element of the NAL unit header that includes information about the encoded slice.

[0246] The decoding device can determine the NAL unit type of a slice in the image based on NAL unit type related information (S1620).

[0247] For example, as mentioned above, the NAL unit type of a slice in the image can be determined based on the value of `mixed_nalu_types_in_pic_flag`. At this point, the NAL unit type of the slice (i.e., the NAL unit containing information about the slice) can be determined as one of indices 0 to 10 disclosed in Table 1 or Table 2 above. Since already... Figure 15 The implementation details specific examples related to it, so these specific examples are omitted in this implementation.

[0248] The decoding device can decode / reconstruct samples / blocks / slices in the image based on the NAL unit type of the slice (S1630).

[0249] For example, samples / blocks within a slice can be decoded / reconstructed based on the slice's NAL unit type. When a first NAL unit type is set for a first slice in the current frame and a second NAL unit type (different from the first NAL unit type) is set for a second slice in the current frame, samples / blocks in the first slice or the first slice itself can be decoded / reconstructed based on the first NAL unit type, and samples / blocks in the second layer or the second layer itself can be decoded / reconstructed based on the second NAL unit type.

[0250] Furthermore, for example, the first slice can be in the first sub-screen, and the second slice can be in the second sub-screen. In this case, when the third slice is located in the first sub-screen, the NAL unit type of the third slice can be the same as the first NAL unit type. In this case, when the fourth slice is located in the second sub-screen, the NAL unit type of the fourth slice can be the same as the second NAL unit type. When the fifth slice is located in the third sub-screen, the NAL unit type of the fifth slice can be different from the first NAL unit type and the second NAL unit type.

[0251] Furthermore, as mentioned above, there is a problem with deriving screen output-related flags (e.g., PictureOutputFlag) in the current VVC standard. That is, the derivation of screen output-related flags (e.g., PictureOutputFlag) is invoked for all screens. However, there is a problem that it is unclear when the derivation process is invoked. According to the current VVC standard, it appears to be invoked at the start of screen decoding. However, screen output-related flags (e.g., PictureOutputFlag) themselves are used or required at the end of screen decoding (i.e., during the additional collision process).

[0252] Therefore, in order to provide a solution to the above problems, the following implementation methods are proposed in this document. These implementation methods can be applied individually or in combination with the implementation methods presented in Table 3 above.

[0253] As an implementation, the process of deriving the picture output flag (e.g., PictureOutputFlag) can be modified to be called at the end of decoding the current frame. That is, the process of deriving the picture output flag (e.g., PictureOutputFlag) can be called after all slices in the frame have been decoded.

[0254] The above-described implementations in this document can be implemented in the form shown in Table 4 below. Table 4 illustrates examples of implementations of the above-described implementations related to the output of a screen in the VVC specification.

[0255] [Table 4]

[0256]

[0257]

[0258]

[0259] Referring to Table 4, the decoding process for the current screen can be performed as follows.

[0260] First, NAL units can be decoded, and slices can be decoded using syntax elements included in the slice header and higher levels. During slice decoding, POC-related information (variables and functions) can be derived, and decoding for constructing a list of reference frames can be performed for each slice in the frame at the start of decoding. Additionally, decoding processing for reference frame marking can be performed, where reference frames can be marked as "not used for reference" or "used for long-term reference." Furthermore, decoding processing for generating unusable reference frames can be performed.

[0261] Next, decoding processing can be performed based on the syntax elements included in all syntaxes (i.e., the aforementioned inter-frame prediction or intra-frame prediction, residual processing, intra-loop filtering, etc.). At this point, the requirement for bitstream consistency is that the encoded slices of the picture include slice data for each CTU of the picture, such that dividing the picture into slices and dividing the slices into CTUs respectively form the segmentation of the picture.

[0262] Next, after all slices in the current frame have been decoded, the PictureOutputFlag can be derived. For example, if the currently decoded frame is marked as "for short-term reference", each ILRP entry in the reference frame list (RefPicList[0] or RefPicList[1]) can be marked as "for short-term reference". Furthermore, the PictureOutputFlag can be derived based on at least one of the following conditions.

[0263] The PictureOutputFlag is deduced to be 0 based on the first condition that the value of the syntax element associated with the Video Parameter Set (VPS) ID is greater than 0 and the current layer is not the output layer. Here, the syntax element associated with the VPS ID can be the `sps_video_parameter_set_id` syntax element, and can be signaled from the Sequence Parameter Set (SPS) mentioned above. For example, when the value of `sps_video_parameter_set_id` is greater than 0, the value of `sps_video_parameter_set_id` can represent the value of the VPS ID (e.g., `vps_video_parameter_set_id`) which serves as identification information for the Video Parameter Set (VPS). In the case where the bitstream is single-layered, the presence of a VPS can be optional.

[0264] Based on the second condition that the current frame is a Random Access Skip Preview (RASL) frame and the associated Intra-Random Access Point (IRAP) frame has a NoOutputBeforeRecoveryFlag of 1, the value of the PictureOutputFlag can be deduced to be 0. In this respect, NoOutputBeforeRecoveryFlag can be a flag indicating whether a decoded frame can be output, either in frame order count (POC) or in decoding order from the Progressive Decode Refresh (GDR) frame to the recovery point frame. For example, it can be determined that the (decoded) frame between the GDR frame and the recovery point frame will not be output, and in this case, the value of NoOutputBeforeRecoveryFlag can be set to 1. Alternatively, it can be determined that the (decoded) GDR frame will not be output, and in this case, the value of NoOutputBeforeRecoveryFlag can be set to 1. However, it can be determined that the (decoded) recovery point frame will be output, and in this case, the value of NoOutputBeforeRecoveryFlag can be set to 0. In other words, in the case of a frame with a POC from the GDR frame to the POC before the recovery point POC, the value of NoOutputBeforeRecoveryFlag can be determined to be 1.

[0265] Based on the third condition of a recovery screen where the current screen is a Progressive Decoding Refresh (GDR) screen with NoOutputBeforeRecoveryFlag equal to 1 or a GDR screen with NoOutputBeforeRecoveryFlag equal to 1, the value of the PictureOutputFlag can be deduced to be 0.

[0266] Otherwise, if at least one of the first, second, and third conditions is not met, the value of the PictureOutputFlag can be deduced as the value of the picture output-related syntax element to be signaled. In this case, as shown in Table 4 above, the picture output-related syntax element to be signaled can be ph_pic_output_flag. The ph_pic_output_flag syntax element can be information affecting the output and removal processing of the decoded picture, and for example, if the value of ph_pic_output_flag is 1, it can indicate that the decoded picture is output, while if the value of ph_pic_output_flag is 0, it can indicate that no decoded picture is output.

[0267] Figure 17 Examples of video / image encoding methods to which one or more of the above embodiments of this document may be applied are illustrated. Figure 17 The method disclosed in the article can be derived from Figure 2 The encoding device 200 disclosed herein performs the operation. Furthermore, according to the embodiment, the following can be omitted. Figure 17 It can be one or more steps, and additional steps can be added.

[0268] Reference Figure 17 The encoding device can decode / reconstruct the image (S1700). The encoding device can manage the DPB based on the DPB parameters (S1710). In other words, the encoding device can mark and / or remove images decoded from the DPB based on the DPB parameters.

[0269] Decoded frames can be used as a reference for inter-frame prediction of a sequence of frames. Each decoded frame can essentially be inserted (stored) in the DPB. The DPB can typically be updated before decoding the current frame. When the layer associated with the DPB is not an output layer (or the DPB parameters are not associated with an output layer) but a reference layer, the decoded frames in the DPB may not be output. When the layer associated with the DPB (or DPB parameters) is an output layer, the decoded frames in the DPB can be output based on the DPB and / or DPB parameters.

[0270] Managing the DPB can be referred to as updating the DPB. Managing the DPB may include decoding the screen from the DPB output. DPB parameter-related information may include information / syntax elements related to DPB management, and may include, for example, syntax elements included in the DPB parameter syntax disclosed in the VVC standard. DPB management may be performed based on the value of the screen output flag (e.g., PictureOutputFlag) derived in the process of deriving the screen output flag (e.g., PictureOutputFlag) in the above-described embodiments.

[0271] Other DPB parameters can be notified by signaling based on whether the current layer is an output layer or a reference layer, or based on whether the DPB (or DPB parameter) is mapped to an OLS.

[0272] The encoding device can encode video / image information (S1720). In this embodiment, the encoding device can encode video / image information including DPB parameter-related information. The DPB parameter-related information may include information / syntax elements related to DPB management as described above.

[0273] although Figure 17 As not shown, the encoding device can decode the current frame based on the DPB updated / managed after step S1710. Alternatively, the decoded current frame can be inserted into the DPB, and the DPB, including the decoded current frame, can be updated based on the DPB parameters before decoding the sequence of frames.

[0274] Figure 18 Examples of video / image decoding methods to which one or more of the above embodiments of this document may be applied are illustrated. Figure 18 The method disclosed in the article can be derived from Figure 3 The decoding device 300 disclosed herein performs the operation. Furthermore, according to the embodiment, the following can be omitted. Figure 18 It can be one or more steps, and additional steps can be added.

[0275] Reference Figure 18 The decoding device can obtain video / image information from the bitstream (S1800). In this embodiment, the decoding device can obtain video / image information including DPB parameter-related information from the bitstream. The DPB parameter-related information may include information / syntax elements related to DPB management as described above.

[0276] The decoding device can manage the DPB based on the DPB parameters (S1810). In other words, the decoding device can mark and / or remove frames decoded from the DPB based on the DPB parameters.

[0277] Decoded frames can be used as a reference for inter-frame prediction of a sequence of frames. Each decoded frame can essentially be inserted (stored) in the DPB. The DPB can typically be updated before decoding the current frame. When the layer associated with the DPB is not an output layer (or the DPB parameters are not associated with an output layer) but a reference layer, the decoded frames in the DPB may not be output. When the layer associated with the DPB (or DPB parameters) is an output layer, the decoded frames in the DPB can be output based on the DPB and / or DPB parameters.

[0278] Managing the DPB can be referred to as updating the DPB. Managing the DPB may include decoding the screen from the DPB output. DPB parameter-related information may include information / syntax elements related to DPB management, and may include, for example, syntax elements included in the DPB parameter syntax disclosed in the VVC standard. DPB management may be performed based on the value of the screen output flag (e.g., PictureOutputFlag) derived in the process of deriving the screen output flag (e.g., PictureOutputFlag) in the above-described embodiments.

[0279] Other DPB parameters can be signaled based on whether the current layer is an output layer or a reference layer, or based on whether the DPB (or DPB parameter) is used for OLS (mapped to OLS).

[0280] The decoding device can decode / output the current frame based on the DPB (S1820). In an implementation, the (previously) decoded frame in the DPB can be used as a reference to decode the blocks / pieces in the current frame based on inter-frame prediction.

[0281] The following figures are provided to illustrate specific examples of this document. Since the names or names of specific terms or devices described in the figures (e.g., names of grammars / grammatical elements, etc.) are presented as examples, the technical features of this document are not limited to the specific names used in the following figures.

[0282] Figure 19 and Figure 20 Examples of video / image encoding methods and associated components according to the implementation of this document are illustrated schematically.

[0283] Figure 19 The method disclosed in the article can be derived from Figure 2 or Figure 20 The publicly disclosed encoding device 200 is executed. Here, Figure 20 The publicly disclosed encoding device 200 is Figure 2 A simplified representation of the encoding device 200 disclosed herein. Specifically, Figure 19 Step S1900 can be performed by Figure 2 The image segmenter 210, predictor 220, residual processor 230, adder 340, etc. disclosed herein shall be used to perform the steps; steps S1910 to S1920 may be performed by Figure 2 The DPB is executed as disclosed in the document; and S1930 can be executed by... Figure 2 The entropy encoder 240 disclosed in the document is executed. Additionally, it can be executed... Figure 19 The methods disclosed herein include the embodiments described above. Therefore, in Figure 19 In this document, detailed descriptions of content that corresponds to repetitions of the above-described embodiments will be omitted or simplified.

[0284] Reference Figure 19 The encoding device can decode the slices included in the current frame (S1900).

[0285] In the implementation, the encoding device can decode (reconstruct) the slices in the current frame through the above-described frame segmentation process, intra-frame prediction or inter-frame prediction, residual processing, and in-loop filtering.

[0286] Decoded frames can be used as a reference for inter-frame prediction of a sequence of frames. For this purpose, each decoded frame can essentially be inserted (stored) in the DPB. As mentioned above, the DPB can typically be updated before decoding the current frame. When the layer associated with the DPB is not an output layer (or the DPB parameters are not associated with an output layer) but a reference layer, the decoded frames in the DPB may not be output. When the layer associated with the DPB (or DPB parameters) is an output layer, the decoded frames in the DPB can be output based on the DPB and / or DPB parameters. In this case, it can be determined whether to output the decoded frames stored in the DPB. Whether to output a frame can be determined by deriving the frame output flag.

[0287] The encoding device can deduce the image output flag for the current image based on the fact that all slices included in the current image have been decoded (S1910). The encoding device can determine the output of the current image based on the image output flag (S1920).

[0288] As described above, the picture output flag can have a value related to whether the current picture is output, and can be represented as, for example, a PictureOutputFlag value. For example, based on a picture output flag (e.g., PictureOutputFlag) value of 0, the current picture can be marked as "not required to output". Alternatively, based on a picture output flag (e.g., PictureOutputFlag) value of 1, the current picture can be marked as "required to output". That is, if the picture output flag (e.g., PictureOutputFlag) value is 1, the current picture can be output from the DPB, while if the picture output flag (e.g., PictureOutputFlag) value is 0, the current picture can be stored or removed instead of being output from the DPB.

[0289] In an implementation, the picture output flag (e.g., PictureOutputFlag) can be derived to a value of 0 or 1 based on at least one of the following conditions.

[0290] Based on the first condition that the value of the syntax element associated with the Video Parameter Set (VPS) ID is greater than 0 and the current layer is not the output layer, the value of the picture output flag (e.g., PictureOutputFlag) is deduced to be 0. Here, the syntax element associated with the VPS ID can be the sps_video_parameter_set_id syntax element, and can be signaled from the Sequence Parameter Set (SPS) mentioned above. For example, when the value of sps_video_parameter_set_id is greater than 0, the value of sps_video_parameter_set_id can represent the value of the VPS ID (e.g., vps_video_parameter_set_id) as identification information for the Video Parameter Set (VPS). In the case where the bitstream is single-layered, the presence of a VPS can be optional.

[0291] Based on the second condition that the current frame is a Random Access Skip Preview (RASL) frame and the associated Intra-Random Access Point (IRAP) frame has a NoOutputBeforeRecoveryFlag of 1, the value of the picture output flag (e.g., PictureOutputFlag) can be deduced to be 0. In this respect, NoOutputBeforeRecoveryFlag can be a flag indicating whether a decoded frame can be output, either in frame order count (POC) or in decoding order from the Progressive Decode Refresh (GDR) frame to the recovery point frame. For example, it can be determined that a (decoded) frame between a GDR frame and a recovery point frame will not be output, and in this case, the value of NoOutputBeforeRecoveryFlag can be set to 1. Alternatively, it can be determined that a (decoded) GDR frame will not be output, and in this case, the value of NoOutputBeforeRecoveryFlag can be set to 1. However, it can be determined that a (decoded) recovery point frame will be output, and in this case, the value of NoOutputBeforeRecoveryFlag can be set to 0. In other words, in the case of a frame with a POC from the GDR frame to the POC before the recovery point POC, the value of NoOutputBeforeRecoveryFlag can be determined to be 1.

[0292] Based on the third condition of a recovery screen where the current screen is a Progressive Decoding Refresh (GDR) screen with NoOutputBeforeRecoveryFlag equal to 1 or a GDR screen with NoOutputBeforeRecoveryFlag equal to 1, the value of the screen output flag (e.g., PictureOutputFlag) can be deduced to be 0.

[0293] Otherwise, when at least one of the first, second, and third conditions is not met, the value of the picture output flag (e.g., PictureOutputFlag) can be deduced as the value of the picture output-related syntax element to be signaled. In this case, as shown in Table 4 above, the picture output-related syntax element to be signaled can be ph_pic_output_flag. The ph_pic_output_flag syntax element can be information affecting the output and removal processing of the decoded picture, and for example, if the value of ph_pic_output_flag is 1, it can indicate that the decoded picture is output, while if the value of ph_pic_output_flag is 0, it can indicate that no decoded picture is output.

[0294] In the above embodiments, the process of deriving the picture output flag (e.g., PictureOutputFlag) is based on the conditional application of the canonical algorithm disclosed in Table 4 above. In another embodiment, the value of the picture output flag (e.g., PictureOutputFlag) can be derived conditionally based on the canonical algorithm disclosed in Table 3 above.

[0295] Furthermore, in one embodiment, regarding a single-layer bitstream, the encoding device can also derive a picture output flag (e.g., PictureOutputFlag) for the current frame based on the fact that all slices included in the current frame have been decoded. Additionally, in one embodiment, even when the bitstream supports multiple layers, the picture output flag (e.g., PictureOutputFlag) can be derived after all slices in the current frame have been decoded in the same manner as with a single layer. Thus, when the picture output flag (e.g., PictureOutputFlag) is derived after all slices in each frame have been decoded, regardless of whether the bitstream is single-layer or multi-layer, unnecessary processing of deriving output-related information (i.e., picture output flag) is not required, regardless of whether the bitstream is single-layer or multi-layer. Furthermore, unnecessary updating of output-related information (i.e., picture output flag) is not required, regardless of whether the frame is the last frame in the access unit (AU).

[0296] When the value of the output flag (e.g., PictureOutputFlag) derived as described above is 1, the current (decoded) screen can be output according to the output order or the POC order. Alternatively, when the value of the output flag (e.g., PictureOutputFlag) is 0, the current (decoded) screen is not output.

[0297] The encoding device can encode image information about the current frame (S1930).

[0298] In one implementation, the encoding device can generate various information / syntax elements derived during the decoding process of slices in the current frame as image / video information, and can encode such various information. For example, the image / video information may include a slice header for a slice in the current frame.

[0299] Image / video information, including the various types of information described above, can be encoded and output as a bitstream. The bitstream can be sent to a decoding device via a network or (digital) storage medium. Here, the network can include broadcast networks, communication networks, etc., and the digital storage medium can include various storage media such as Universal Serial Bus (USB), Secure Digital (SD), Optical Disc (CD), Digital Video Disc (DVD), Blu-ray, Hard Disk Drive (HDD), Solid State Drive (SSD), etc.

[0300] Figure 21 and Figure 22 Examples of video / image decoding methods and associated components according to the implementation of this document are illustrated schematically.

[0301] Figure 21 The method disclosed in the article can be derived from Figure 3 or Figure 22 The publicly disclosed decoding device 300 is executed. Here, Figure 22 The publicly disclosed decoding device 300 is Figure 3 A simplified representation of the decoding device 300 disclosed in the document. Specifically, Figure 21 Step S2100 can be performed by Figure 3 The entropy decoder 310, residual processor 320, predictor 330, adder 340, etc., disclosed in the document are used for execution, and steps S2110 to S2120 can be performed by... Figure 3 It can be executed using the publicly available DPB. Alternatively, it can be executed... Figure 21 The methods disclosed herein include the embodiments described above. Therefore, in Figure 21 In this document, detailed descriptions of content that corresponds to repetitions of the above-described embodiments will be omitted or simplified.

[0302] Reference Figure 21 The decoding device can decode the slices included in the current frame (S2100).

[0303] In the implementation, the decoding device can decode (reconstruct) the slices in the current frame through the above-mentioned frame segmentation process, intra-frame prediction or inter-frame prediction, residual processing, and in-loop filtering.

[0304] Decoded frames can be used as a reference for inter-frame prediction of a sequence of frames. For this purpose, each decoded frame can essentially be inserted (stored) in the DPB. As mentioned above, the DPB can typically be updated before decoding the current frame. When the layer associated with the DPB is not an output layer (or the DPB parameters are not associated with an output layer) but a reference layer, the decoded frames in the DPB may not be output. When the layer associated with the DPB (or DPB parameters) is an output layer, the decoded frames in the DPB can be output based on the DPB and / or DPB parameters. In this case, it can be determined whether to output the decoded frames stored in the DPB. Whether to output a frame can be determined by deriving the frame output flag.

[0305] The decoding device can deduce the image output flag for the current image based on the fact that all slices included in the current image have been decoded (S2110). The decoding device can determine the output for the current image based on the image output flag (S2120).

[0306] As described above, the picture output flag can have a value related to whether the current picture is output, and can be represented as, for example, a PictureOutputFlag value. For example, based on a picture output flag (e.g., PictureOutputFlag) value of 0, the current picture can be marked as "not required to output". Alternatively, based on a picture output flag (e.g., PictureOutputFlag) value of 1, the current picture can be marked as "required to output". That is, if the picture output flag (e.g., PictureOutputFlag) value is 1, the current picture can be output from the DPB, while if the picture output flag (e.g., PictureOutputFlag) value is 0, the current picture can be stored or removed instead of being output from the DPB.

[0307] In an implementation, the picture output flag (e.g., PictureOutputFlag) can be derived to a value of 0 or 1 based on at least one of the following conditions.

[0308] Based on the first condition that the value of the syntax element associated with the Video Parameter Set (VPS) ID is greater than 0 and the current layer is not the output layer, the value of the picture output flag (e.g., PictureOutputFlag) is deduced to be 0. Here, the syntax element associated with the VPS ID can be the sps_video_parameter_set_id syntax element, and can be signaled from the Sequence Parameter Set (SPS) mentioned above. For example, when the value of sps_video_parameter_set_id is greater than 0, the value of sps_video_parameter_set_id can represent the value of the VPS ID (e.g., vps_video_parameter_set_id) as identification information for the Video Parameter Set (VPS). In the case where the bitstream is single-layered, the presence of a VPS can be optional.

[0309] Based on the second condition that the current frame is a Random Access Skip Preview (RASL) frame and the associated Intra-Random Access Point (IRAP) frame has a NoOutputBeforeRecoveryFlag of 1, the value of the picture output flag (e.g., PictureOutputFlag) can be deduced to be 0. In this respect, NoOutputBeforeRecoveryFlag can be flag information indicating whether a decoded frame can be output, either in frame order count (POC) or in decoding order from the Progressive Decode Refresh (GDR) frame to the recovery point frame. For example, it can be determined that a (decoded) frame between a GDR frame and a recovery point frame will not be output, and in this case, the value of NoOutputBeforeRecoveryFlag can be set to 1. Alternatively, it can be determined that a (decoded) GDR frame will not be output, and in this case, the value of NoOutputBeforeRecoveryFlag can be set to 1. However, it can be determined that a (decoded) recovery point frame will be output, and in this case, the value of NoOutputBeforeRecoveryFlag can be set to 0. In other words, in the case of a frame with a POC from the GDR frame to the POC before the recovery point POC, the value of NoOutputBeforeRecoveryFlag can be determined to be 1.

[0310] Based on the third condition of a recovery screen where the current screen is a Progressive Decoding Refresh (GDR) screen with NoOutputBeforeRecoveryFlag equal to 1 or a GDR screen with NoOutputBeforeRecoveryFlag equal to 1, the value of the screen output flag (e.g., PictureOutputFlag) can be deduced to be 0.

[0311] Otherwise, when at least one of the first, second, and third conditions is not met, the value of the picture output flag (e.g., PictureOutputFlag) can be deduced as the value of the picture output-related syntax element to be signaled. In this case, as shown in Table 4 above, the picture output-related syntax element to be signaled can be ph_pic_output_flag. The ph_pic_output_flag syntax element can be information affecting the output and removal processing of the decoded picture, and for example, if the value of ph_pic_output_flag is 1, it can indicate that the decoded picture is output, while if the value of ph_pic_output_flag is 0, it can indicate that no decoded picture is output.

[0312] In the above embodiments, the process of deriving the picture output flag (e.g., PictureOutputFlag) is based on the conditional application of the canonical algorithm disclosed in Table 4 above. In another embodiment, the value of the picture output flag (e.g., PictureOutputFlag) can be derived conditionally based on the canonical algorithm disclosed in Table 3 above.

[0313] Furthermore, in one embodiment, regarding a single-layer bitstream, the encoding device can also derive a picture output flag (e.g., PictureOutputFlag) for the current frame based on the fact that all slices included in the current frame have been decoded. Additionally, in one embodiment, even when the bitstream supports multiple layers, the picture output flag (e.g., PictureOutputFlag) can be derived after decoding all slices in the current frame in the same manner as in a single-layer bitstream. Thus, when the picture output flag (e.g., PictureOutputFlag) is derived after all slices in each frame have been decoded, regardless of whether the bitstream is single-layer or multi-layer, unnecessary processing of deriving output-related information (i.e., picture output flag) is not required, regardless of whether the bitstream is single-layer or multi-layer. Furthermore, unnecessary updating of output-related information (i.e., picture output flag) is not required, regardless of whether the frame is the last frame in the access unit (AU).

[0314] When the value of the output flag (e.g., PictureOutputFlag) derived as described above is 1, the current (decoded) screen can be output according to the output order or the POC order. Alternatively, when the value of the output flag (e.g., PictureOutputFlag) is 0, the current (decoded) screen is not output.

[0315] Although the method has been described based on a flowchart listing the steps or blocks in the above embodiments, the steps in this document are not limited to a specific order, and specific steps may be performed in different steps or in a different order or simultaneously relative to the steps described above. Furthermore, those skilled in the art will understand that the steps in the flowchart are not exclusive, and one or more steps may be included or removed from the flowchart without affecting the scope of this document.

[0316] The methods described above according to this disclosure may be in the form of software, and the encoding and / or decoding devices according to this document may be included in an apparatus for performing image processing (e.g., TV, computer, smartphone, set-top box, display device, etc.).

[0317] When the embodiments described in this document are implemented in software, the methods described above can be implemented by modules (processes or functions) that perform the functions described above. Modules can be stored in memory and executed by a processor. Memory can be installed inside or outside the processor and can be connected to the processor via various known means. The processor may include application-specific integrated circuits (ASICs), other chipsets, logic circuits, and / or data processing devices. Memory may include read-only memory (ROM), random access memory (RAM), flash memory, memory cards, storage media, and / or other storage devices. In other words, the embodiments described in this document can be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units shown in the various figures can be implemented and executed on a computer, processor, microprocessor, controller, or chip. In this case, information about the implementation (e.g., information about instructions) or algorithms can be stored in a digital storage medium.

[0318] Furthermore, the decoding and encoding devices used in this document can include multimedia broadcast transceivers, mobile communication terminals, home theater video devices, digital cinema video devices, surveillance cameras, video chat devices, real-time communication devices such as video communication, mobile streaming devices, storage media, cameras, video-on-demand (VoD) service providers, over-the-top (OTT) video devices, internet streaming service providers, three-dimensional (3D) video devices, virtual reality (VR) devices, augmented reality (AR) devices, video telephony devices, vehicle terminals (e.g., vehicle (including autonomous vehicles) terminals, aircraft terminals, ship terminals, etc.), and medical video devices, and can be used to process video signals or data signals. For example, over-the-top (OTT) video devices can include game consoles, Blu-ray players, internet access TVs, home theater systems, smartphones, tablet PCs, digital video recorders (DVRs), etc.

[0319] Furthermore, the processing methods described in this document can be generated in the form of a computer-executable program and can be stored in a computer-readable recording medium. Multimedia data with data structures according to the embodiments of this document can also be stored in a computer-readable recording medium. Computer-readable recording media include various storage devices and distributed storage devices in which computer-readable data is stored. Computer-readable recording media can include, for example, Blu-ray discs (BD), Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disks, and optical data storage devices. Computer-readable recording media also include media implemented in carrier wave form (e.g., transmission via the Internet). Additionally, bitstreams generated by encoding methods can be stored in computer-readable recording media or transmitted via wired or wireless communication networks.

[0320] Furthermore, the embodiments described in this document can be implemented as a computer program product based on program code, and the program code can be executed on a computer according to the embodiments described in this document. The program code can be stored on a computer-readable medium.

[0321] Figure 23 Examples of content streaming systems to which the implementation methods described in this document are applicable.

[0322] Reference Figure 23 The content streaming system that applies the implementation methods described in this document can typically include an encoding server, a streaming server, a network server, a media storage device, a user device, and a multimedia input device.

[0323] An encoding server is used to compress content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream, which is then sent to a streaming server. As another example, in cases where multimedia input devices such as smartphones, cameras, and camcorders directly generate bitstreams, the encoding server can be omitted.

[0324] Bitstreams can be generated using the encoding methods or bitstream generation methods described in this document. Furthermore, the streaming server can temporarily store the bitstream during transmission or reception.

[0325] The streaming server sends multimedia data to the user's device via a web server based on the user's request. The web server acts as a tool to notify the user of available services. When a user requests a desired service, the web server forwards the request to the streaming server, which then delivers the multimedia data to the user. In this context, the content streaming system may include a separate control server, which in this case controls the commands / responses between the various devices within the content streaming system.

[0326] A streaming server can receive content from media storage and / or encoding servers. For example, when receiving content from an encoding server, the content can be received in real time. In this case, the streaming server can store the bitstream for a predetermined period of time to smoothly provide streaming services.

[0327] For example, user equipment may include mobile phones, smartphones, laptops, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, tablet PCs, tablet PCs, ultrabooks, wearable devices (e.g., watch-type terminals (smartwatches), glasses-type terminals (smart glasses), head-mounted displays (HMDs)), digital televisions, desktop computers, digital signage, etc.

[0328] Each server in a content streaming system can operate as a distributed server, and in this case, the data received by each server can be processed in a distributed manner.

[0329] The claims in this document can be combined in various ways. For example, technical features in the method claims can be combined to be implemented or performed in a device, and technical features in the device claims can be combined to be implemented or performed in a method. Furthermore, technical features in both the method claims and the device claims can be combined to be implemented or performed in a device.

Claims

1. An image decoding method performed by a decoding device, the image decoding method comprising the following steps: Decode the slices included in the current frame; After all the slices included in the current frame have been decoded, derive the frame output flag for the current frame; as well as The output of the current screen is determined based on the screen output flag. The image output flag has a value related to whether or not the current image is output. Wherein, based on the value of the screen output flag being equal to 0, the current screen is marked as "no output required," and Wherein, based on the value of the output flag being equal to 1, the current screen is marked as "required for output". Specifically, based on the first condition that the value of the syntax element associated with the video parameter set VPS ID is greater than 0 and the current layer is not the output layer, the value of the screen output flag is deduced to be 0.

2. The method according to claim 1, wherein, For a bitstream with a single layer, the image output flag for the current image is derived based on the fact that all the slices included in the current image have been decoded.

3. The method according to claim 1, wherein, Based on the second condition that the current frame is a random access skipping leading RASL frame and the associated intra-frame random access point (IRAP) frame has a NoOutputBeforeRecoveryFlag of 1, the value of the frame output flag is deduced to be 0, and The NoOutputBeforeRecoveryFlag is a flag indicating whether the decoded screen can be output in the order of POC counting or in the order of decoding from the GDR screen to the recovery point screen.

4. The method according to claim 1, wherein, Based on the third condition that the current screen is a GDR screen that is gradually decoded and refreshed with NoOutputBeforeRecoveryFlag equal to 1 or a GDR screen that is restored with NoOutputBeforeRecoveryFlag equal to 1, the value of the screen output flag is deduced to be 0.

5. The method according to claim 1, wherein, Since the first condition is not met, the value of the screen output flag is deduced to be the value of the screen output related syntax element to be notified by signal.

6. The method according to claim 3, wherein, Since the second condition is not met, the value of the screen output flag is deduced to be the value of the screen output related syntax element to be notified by signal.

7. The method according to claim 4, wherein, Since the third condition is not met, the value of the screen output flag is deduced to be the value of the screen output related syntax element to be notified by signal.

8. An image encoding method performed by an encoding device, the image encoding method comprising the following steps: Decode the slices included in the current frame; After all the slices included in the current frame have been decoded, derive the frame output flag for the current frame; The output of the current screen is determined based on the screen output flag; as well as The image information of the current screen is encoded. The image output flag has a value related to whether or not the current image is output. Specifically, if the value of the output flag is 0, the current screen is marked as "no output required". Wherein, based on the value of the output flag being equal to 1, the current screen is marked as "required for output," and Specifically, based on the first condition that the value of the syntax element associated with the video parameter set VPS ID is greater than 0 and the current layer is not the output layer, the value of the screen output flag is deduced to be 0.

9. The method according to claim 8, wherein, For a bitstream with a single layer, the output flag for the current frame is derived based on the fact that all the slices included in the current frame have been decoded.

10. The method according to claim 8, wherein, Based on the second condition that the current frame is a random access skipping leading RASL frame and the associated intra-frame random access point (IRAP) frame has a NoOutputBeforeRecoveryFlag of 1, the value of the frame output flag is deduced to be 0, and The NoOutputBeforeRecoveryFlag is a flag indicating whether the decoded screen can be output in the order of POC counting or in the order of decoding from the GDR screen to the recovery point screen.

11. The method according to claim 8, wherein, Based on the third condition that the current screen is a GDR screen that is gradually decoded and refreshed with NoOutputBeforeRecoveryFlag equal to 1 or a GDR screen that is restored with NoOutputBeforeRecoveryFlag equal to 1, the value of the screen output flag is deduced to be 0.

12. The method according to claim 8, wherein, Since the first condition is not met, the value of the screen output flag is deduced to be the value of the screen output related syntax element to be notified by signal.

13. The method according to claim 10, wherein, Since the second condition is not met, the value of the screen output flag is deduced to be the value of the screen output related syntax element to be notified by signal.

14. The method according to claim 11, wherein, Since the third condition is not met, the value of the screen output flag is deduced to be the value of the screen output related syntax element to be notified by signal.

15. A method for transmitting a bit stream, the method comprising the following steps: Perform the image encoding method according to claim 8 to generate the bitstream; as well as Send the bit stream.

Citation Information

Patent Citations

  • Signaling change in output layer sets

    CN105379285A