Image or video coding based on NAL unit related information

Through a method based on NAL unit type related information, the problem of low efficiency in high-resolution and high-quality image/video encoding is solved, efficient encoding and flexible signal notification of mixed NAL unit type pictures are achieved, and encoding efficiency and flexibility are improved.

CN115104315BActive Publication Date: 2025-09-05LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202080096753.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-12-23
Filing Date
2020-12-10
Publication Date
2025-09-05
Estimated Expiration
2040-12-10

AI Technical Summary

Technical Problem

Existing image/video coding technologies are inefficient in compressing high-resolution and high-quality images/videos, making them difficult to compress and send or store effectively, especially when transmission and storage costs increase, and are difficult to handle pictures with mixed NAL unit types.

Method used

By determining the NAL unit type of a picture slice based on NAL unit type related information and allowing signaling of reference picture list related information, video/image coding efficiency is improved and picture coding of mixed NAL unit types is supported.

Benefits of technology

It improves the overall image/video compression efficiency, flexibly handles pictures of mixed NAL unit types, effectively encodes reference picture list related information, and enhances the flexibility and efficiency of encoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115104315B_ABST
    Figure CN115104315B_ABST
Patent Text Reader

Abstract

According to the disclosure of this document, based on the NAL unit type related information about whether the picture has a mixed NAL unit type, the NAL unit type of the slice within the picture can be determined, and for the picture with a mixed NAL unit type, not only IRAP can be provided, but also a form mixed with other types of NAL units can be provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present technology relates to video or image coding, for example, to a coding technology based on information related to a Network Abstraction Layer (NAL) unit. Background Art

[0002] Recently, there has been an increasing demand for high-resolution, high-quality images / videos, such as 4K or 8K ultra-high-definition (UHD) images / videos, in various fields. As image / video resolution or quality becomes higher, relatively more information or bits are transmitted compared to conventional image / video data. Therefore, if image / video data is transmitted via a medium such as an existing wired / wireless broadband line or stored in a conventional storage medium, the cost of transmission and storage is likely to increase.

[0003] In addition, there is growing interest and demand for virtual reality (VR) and artificial reality (AR) content and immersive media such as holograms; and there is also growing broadcasting of images / videos that exhibit image / video characteristics that are different from actual images / videos (e.g., game images / videos).

[0004] Therefore, highly efficient image / video compression technology is required to effectively compress and transmit, store, or play high-resolution, high-quality images / videos exhibiting various characteristics as described above.

[0005] In addition, a method for improving image / video encoding efficiency is required, and to this end, a method for efficiently signaling and encoding information related to a Network Abstraction Layer (NAL) unit is necessary. Summary of the Invention

[0006] Technical issues

[0007] This document provides methods and devices for improving video / image coding efficiency.

[0008] This document also provides a method and apparatus for improving video / image coding efficiency based on NAL unit related information.

[0009] This document also provides methods and apparatus for improving video / image coding efficiency for pictures with mixed (hybrid) NAL unit types.

[0010] This document also provides methods and apparatus for allowing the signaling or presence of reference picture list related information for slices in a picture having a specific NAL unit type relative to a picture having a mixed NAL unit type.

[0011] Technical Solution

[0012] According to an embodiment of this document, the NAL unit type of a slice in a picture can be determined based on NAL unit type related information about whether the picture has a mixed NAL unit type. For example, based on the NAL unit type related information of a picture with a mixed NAL unit type, the first NAL unit of the first slice of the picture and the second NAL unit of the second slice of the picture can have different NAL unit types. Alternatively, based on the NAL unit type related information of a picture without a mixed NAL unit type, the first NAL unit of the first slice of the picture and the second NAL unit of the second slice of the picture can have the same NAL unit type.

[0013] According to an embodiment of this document, based on the case where a picture is allowed to have mixed NAL unit types, information related to signaling a reference picture list may exist for slices with a specific NAL unit type in the picture.

[0014] According to an embodiment of the present document, a video / image decoding method performed by a decoding device is provided. The video / image decoding method may include the method disclosed in the embodiment of the present document.

[0015] According to an embodiment of the present document, a decoding device for performing video / image decoding is provided. The decoding device can execute the method disclosed in the embodiment of the present document.

[0016] According to an embodiment of the present document, a video / image encoding method performed by an encoding device is provided. The video / image encoding method may include the method disclosed in the embodiment of the present document.

[0017] According to an embodiment of the present document, an encoding device for performing video / image encoding is provided. The encoding device can execute the method disclosed in the embodiment of the present document.

[0018] According to an embodiment of this document, a computer-readable digital storage medium is provided that stores encoded video / image information generated according to the video / image encoding method disclosed in at least one of the embodiments of this document.

[0019] According to an embodiment of this document, there is provided a computer-readable digital storage medium storing encoded information or encoded video / image information that causes a decoding device to perform the video / image decoding method disclosed in at least one of the embodiments of this document.

[0020] Technical Effects

[0021] This document can have various effects. For example, according to the embodiments of this document, the overall image / video compression efficiency can be improved. In addition, according to the embodiments of this document, the video / image coding efficiency can be improved based on the NAL unit related information. In addition, according to the embodiments of this document, the video / image coding efficiency of pictures with mixed NAL unit types can be improved. In addition, according to the embodiments of this document, for pictures with mixed NAL unit types, reference picture list related information can be efficiently signaled and encoded. In addition, according to the embodiments of this document, by allowing pictures to include leading picture NAL unit types (e.g., RASL_NUT, RADL_NUT) and other non-IRAP NAL unit types (e.g., TRAIL_NUT, STSA, NUT) in a mixed form, for pictures with mixed NAL unit types, a form mixed not only with IRAP but also with other types of NAL units can be provided, and accordingly, more flexible characteristics can be provided.

[0022] The effects that can be achieved through the detailed examples of this document are not limited to the effects listed above. For example, there may be various technical effects that can be understood or derived from this document by a person skilled in the relevant art. Therefore, the detailed effects of this document are not limited to the effects explicitly described in this document, but may include various effects that can be understood or derived from the technical features of this document. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 An example of a video / image coding system to which the embodiments of this document are applicable is schematically illustrated.

[0024] Figure 2 is a diagram schematically illustrating the configuration of a video / image encoding device to which an embodiment of this document is applied.

[0025] Figure 3 is a schematic diagram illustrating a configuration of a video / image decoding device to which the embodiments of this document can be applied.

[0026] Figure 4 An example of an illustrative video / image encoding process to which one or more embodiments of this document are applicable is shown.

[0027] Figure 5 An example of an exemplary video / image decoding process to which one or more embodiments of this document are applicable is shown.

[0028] Figure 6 An example of an entropy coding method to which the embodiments of this document are applicable is schematically illustrated, and Figure 7 An entropy encoder in an encoding device is schematically illustrated.

[0029] Figure 8 An example of an entropy decoding method to which the embodiments of this document are applicable is schematically illustrated, and Figure 9 An entropy decoder in a decoding device is schematically illustrated.

[0030] Figure 10 Schematically represents the hierarchical structure of coded images / videos.

[0031] Figure 11 is a diagram showing a temporal layer structure of a NAL unit in a bitstream supporting temporal scalability.

[0032] Figure 12 A diagram used to describe pictures that can be randomly accessed.

[0033] Figure 13 This is a diagram used to describe an IDR picture.

[0034] Figure 14 A diagram used to describe a CRA picture.

[0035] Figure 15 An example of a video / image encoding method to which the embodiments of this document are applied is schematically shown.

[0036] Figure 16 An example of a video / image decoding method to which the embodiments of this document are applied is schematically shown.

[0037] Figure 17 and Figure 18 An example of a video / image encoding method and related components according to an embodiment of this document is schematically illustrated.

[0038] Figure 19 and Figure 20 An example of a video / image decoding method and related components according to an embodiment of this document is schematically illustrated.

[0039] Figure 21 An example of a content streaming system to which the embodiments disclosed in this document are applicable is illustrated. DETAILED DESCRIPTION

[0040] The present disclosure can be modified in various forms, and its specific embodiments will be described and illustrated in the accompanying drawings. However, these embodiments are not intended to limit the present disclosure. The terms used in the following description are only used to describe specific embodiments and are not intended to limit the present disclosure. Singular expressions include plural expressions as long as they are clearly interpreted differently. Terms such as "including" and "having" are intended to indicate the presence of features, quantities, steps, operations, elements, parts, or combinations thereof used in the following description, so it should be understood that the possibility of the presence or addition of one or more different features, quantities, steps, operations, elements, parts, or combinations thereof is not excluded.

[0041] In addition, the various configurations of the drawings described in this document are independent illustrations for the purpose of illustrating the functions of different features, and do not imply that the various configurations are implemented by different hardware or different software. For example, two or more of the configurations may be combined to form a single configuration, and a single configuration may also be divided into multiple configurations. Without departing from the subject matter of this document, embodiments in which the configurations are combined and / or separated are included within the scope of the claims.

[0042] In this document, the term "A or B" may mean "only A," "only B," or "both A and B." In other words, in this document, the term "A or B" may be interpreted to mean "A and / or B." For example, in this document, the term "A, B, or C" may mean "only A," "only B," "only C," or "any combination of A, B, and C."

[0043] As used in this document, a slash ( / ) or a comma may mean "and / or." For example, "A / B" may mean "A and / or B." Thus, "A / B" may mean "only A," "only B," or "both A and B." For example, "A, B, C" may mean "A, B, or C."

[0044] In this document, "at least one of A and B" may mean "only A", "only B", or "both A and B". In addition, in this document, the expression "at least one of A or B" or "at least one of A and / or B" may be interpreted as the same as "at least one of A and B".

[0045] In addition, in this document, “at least one of A, B, and C” may mean “only A,” “only B,” “only C,” or “any combination of A, B, and C.” In addition, “at least one of A, B, or C” or “at least one of A, B, and / or C” may mean “at least one of A, B, and C.”

[0046] In addition, the parentheses used in this document may mean "for example." Specifically, when the expression "prediction (intra-frame prediction)" is used, it can indicate an example in which "intra-frame prediction" is proposed as "prediction." In other words, the term "prediction" in this document is not limited to "intra-frame prediction" and can indicate an example in which "intra-frame prediction" is proposed as "prediction." In addition, even when the expression "prediction (i.e., intra-frame prediction)" is used, it can indicate an example in which "intra-frame prediction" is proposed as "prediction."

[0047] This document relates to video / image coding. For example, the methods / implementations disclosed in this document may be applied to methods disclosed in the Versatile Video Coding (VVC) standard. Furthermore, the methods / implementations disclosed in this document may be applied to methods disclosed in the Essential Video Coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the second-generation Audio Video Coding standard (AVS2), or next-generation video / image coding standards (e.g., H.267 or H.268).

[0048] This document proposes various embodiments of video / image coding, and unless otherwise mentioned, the above embodiments may also be performed in combination with each other.

[0049] In this document, video may refer to a collection of images over time. A picture generally refers to a unit representing an image in a specific time period, and a slice / tile is a unit that constitutes a part of a picture when encoded. A slice / tile may include one or more coding tree units (CTUs). A picture may be composed of one or more slices / tiles. A tile is a rectangular area of ​​CTUs within a specific tile column and a specific tile row in a picture. A tile column is a rectangular area of ​​CTUs with a height equal to the height of the picture and a width specified by a syntax element in the picture parameter set. A tile row is a rectangular area of ​​CTUs with a width specified by a syntax element in the picture parameter set and a height equal to the picture height. Tile scan is a specific sequential ordering of CTUs of a partitioned picture: CTUs are ordered consecutively in a CTU raster scan within a tile, and tiles within a picture are ordered consecutively in a raster scan of the tiles of the picture. A slice includes an integer number of complete tiles of a picture, or an integer number of consecutive complete CTU rows within a tile, that can be exclusively contained in a single NAL unit.

[0050] In addition, a picture can be divided into two or more sub-pictures. A sub-picture can be a rectangular area of ​​one or more slices within a picture.

[0051] A pixel or picture element (pel) may refer to the smallest unit constituting a picture (or image). In addition, "sample" may be used as a term corresponding to a pixel. A sample may generally represent a pixel or a pixel value, and may represent only a pixel / pixel value of a luma component or only a pixel / pixel value of a chroma component. Alternatively, a sample may refer to a pixel value in the spatial domain, or may refer to a transform coefficient in the frequency domain when a pixel value is transformed into the frequency domain.

[0052] A unit may represent a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to the region. A unit may include a luma block and two chroma (e.g., CB, CR) blocks. In some cases, a unit may be used interchangeably with terms such as block or region. In general, an M×N block may include M columns and N rows of samples (or sample arrays) or a set (or array) of transform coefficients.

[0053] In addition, in this document, at least one of quantization / dequantization and / or transform / inverse transform may be omitted. When quantization / dequantization is omitted, the quantized transform coefficient may be referred to as a transform coefficient. When transform / inverse transform is omitted, the transform coefficient may be referred to as a coefficient or a residual coefficient, or may still be referred to as a transform coefficient for the sake of consistency.

[0054] In this document, quantized transform coefficients and transform coefficients may be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, residual information may include information about the transform coefficients, and information about the transform coefficients may be signaled via residual coding syntax. The transform coefficients may be derived based on the residual information (or information about the transform coefficients), and the scaled transform coefficients may be derived by inverse transforming (scaling) the transform coefficients. Residual samples may be derived based on the inverse transform (transform) of the scaled transform coefficients. This may also be applied / expressed in other parts of this document.

[0055] In this document, technical features described individually in one drawing may be implemented individually or simultaneously.

[0056] Hereinafter, preferred embodiments of the present document will be described in more detail with reference to the accompanying drawings. Hereinafter, in the accompanying drawings, the same reference numerals are used for the same elements, and redundant descriptions of the same elements may be omitted.

[0057] Figure 1 An example of a video / image encoding system to which the embodiments of this document can be applied is illustrated.

[0058] Reference Figure 1 The video / image coding system may include a source device and a receiving device. The source device may send the coded video / image information or data to the receiving device in the form of a file or stream transmission via a digital storage medium or a network.

[0059] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.

[0060] A video source may acquire video / images through a process of capturing, synthesizing, or generating video / images. A video source may include a video / image capture device and / or a video / image generation device. For example, a video / image capture device may include one or more cameras, a video / image archive including previously captured video / images, etc. A video / image generation device may include, for example, a computer, a tablet computer, and a smartphone, and may (electronically) generate video / images. For example, a virtual video / image may be generated by a computer, etc. In this case, the video / image capture process may be replaced by a process of generating relevant data.

[0061] An encoding device encodes input video / images. For compression and coding efficiency, the encoding device performs a series of processes such as prediction, transformation, and quantization. The encoded data (encoded video / image information) is output as a bitstream.

[0062] The transmitter can transmit the encoded image / image information or data, output as a bitstream, to a receiver of a receiving device in the form of a file or stream transmission via a digital storage medium or network. Digital storage media may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter may include components for generating a media file in a predetermined file format and may also include components for transmission via a broadcast / communication network. The receiver may receive / extract the bitstream and transmit the received bitstream to a decoding device.

[0063] The decoding device may decode a video / image by performing a series of processes such as inverse quantization, inverse transformation, and prediction corresponding to the operations of the encoding device.

[0064] The renderer may render the decoded video / image, and the rendered video / image may be displayed on a display.

[0065] Figure 2 Schematically illustrates a configuration of a video / image encoding device to which embodiments of the present document can be applied. Hereinafter, the so-called encoding device may include an image encoding device and / or a video encoding device.

[0066] Reference Figure 2, the encoding device 200 may include and be configured with an image segmenter 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, an inverse quantizer 234, and an inverse transformer 235. The residual processor 230 may also include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstructed block generator. According to an embodiment, the image segmenter 210, the predictor 220, the residual processor 230, the entropy encoder 240, the adder 250, and the filter 260 described above may be configured by one or more hardware components (e.g., an encoder chipset or processor). In addition, the memory 270 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware components may also include the memory 270 as an internal / external component.

[0067] The image splitter 210 may split the input image (or picture or frame) input to the encoding device 200 into one or more processing units. For example, a processing unit may be referred to as a coding unit (CU). In this case, the coding unit may be recursively split from a coding tree unit (CTU) or a maximum coding unit (LCU) based on a quadtree, binary tree, and / or ternary tree (QTBTTT) structure. For example, a coding unit may be split into multiple coding units of increasing depth based on a quadtree, binary tree, and / or ternary tree structure. In this case, for example, the quadtree structure may be applied first, followed by the binary tree and / or ternary tree structure. Alternatively, the binary tree structure may be applied first. The encoding process according to the present disclosure may be performed based on the final coding unit that is no longer split. In this case, the maximum coding unit may be used as the final coding unit based on image characteristics, coding efficiency, and the like. Alternatively, if necessary, the coding unit may be recursively split into coding units of increasing depth so that a coding unit of optimal size is used as the final coding unit. The encoding process may include processes such as prediction, transformation, and reconstruction (described later). As another example, a processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, each of the prediction unit and the transform unit may be split or partitioned from the final coding unit. A prediction unit may be a unit for sample prediction, and a transform unit may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.

[0068] In some cases, the term "unit" may be used interchangeably with terms such as "block" or "region." Typically, an M×N block may represent a set of samples or transform coefficients consisting of M columns and N rows. A sample may generally represent a pixel or pixel value, and may represent only the pixel / pixel value of the luma component or only the pixel / pixel value of the chroma component. A sample may be used as a term corresponding to a pixel or picture element that constitutes a picture (or image).

[0069] The encoding device 200 can generate a residual signal (residual block, residual sample array) by subtracting the prediction signal (prediction block, prediction sample array) output from the inter predictor 221 or the intra predictor 222 from the input image signal (original block, original sample array), and the generated residual signal is transmitted to the transformer 232. In this case, as illustrated, the unit for subtracting the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) within the encoding device 200 can be referred to as a subtractor 231. The predictor can perform prediction on a block to be processed (hereinafter referred to as a current block) and generate a prediction block including prediction samples of the current block. The predictor can determine whether to apply intra prediction or inter prediction in units of the current block or CU. The predictor can generate various prediction-related information, such as prediction mode information, and transmit the generated information to the entropy encoder 240, as described later when describing each prediction mode. The prediction-related information can be encoded by the entropy encoder 240 and output in the form of a bitstream.

[0070] The intra-frame predictor 222 can predict the current block with reference to samples in the current picture. Depending on the prediction mode, the referenced samples can be located near the current block or far away from the current block. In intra-frame prediction, the prediction mode may include multiple non-directional modes and multiple directional modes. For example, the non-directional mode may include a DC mode and a planar mode. For example, depending on the degree of refinement of the prediction direction, the directional mode may include 33 directional prediction modes or 65 directional prediction modes. However, this is only an example, and more or fewer directional prediction modes may be used depending on the settings. The intra-frame predictor 222 can use the prediction mode applied to the neighboring blocks to determine the prediction mode applied to the current block.

[0071] The inter-frame predictor 221 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector in a reference picture. To reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. It can also include information about the inter-frame prediction direction (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, neighboring blocks can include spatially neighboring blocks in the current picture and temporally neighboring blocks in a reference picture. The reference picture containing the reference block and the reference picture containing the temporally neighboring block can be the same or different. Temporally neighboring blocks can be referred to as collocated reference blocks, collocated CUs (colCUs), etc., and reference pictures containing temporally neighboring blocks can be referred to as collocated pictures (colPics). For example, the inter-frame predictor 221 can configure a motion information candidate list based on the neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index for the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the inter-frame predictor 221 can use the motion information of the neighboring block as the motion information of the current block. In skip mode, unlike merge mode, a residual signal may not be transmitted. The motion vector prediction (MVP) mode can indicate the motion vector of the current block by using the motion vector of the neighboring block as a motion vector predictor and signaling the motion vector difference.

[0072] The predictor 220 can generate a prediction signal based on various prediction methods to be described later. For example, the predictor can not only apply intra prediction or inter prediction to predict a block, but also apply intra prediction and inter prediction simultaneously. This can be referred to as combined inter and intra prediction (CIIP). In addition, the predictor can perform prediction on the block based on an intra-block copy (IBC) prediction mode or a palette mode. The IBC prediction mode or palette mode can be used for content image / video coding such as games such as screen content coding (SCC). IBC basically performs prediction in the current picture, but is similar to inter prediction in deriving a reference block in the current picture. That is, IBC can use at least one of the inter prediction techniques described in this document. The palette mode can be considered an example of intra coding or intra prediction. When the palette mode is applied, the sample values ​​in the picture can be signaled based on information about the palette index and the palette table.

[0073] The prediction signal generated by the predictor (including the inter-frame predictor 221 and / or the intra-frame predictor 222) can be used to generate a reconstructed signal or to generate a residual signal. The transformer 232 can generate a transform coefficient by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loève transform (KLT), a graph-based transform (GBT), or a conditional nonlinear transform (CNT). Here, when the relationship information between pixels is illustrated as a curve graph, GBT means a transform obtained from the curve graph. CNT means a transform obtained based on a prediction signal generated by using all previously reconstructed pixels. In addition, the transform process can also be applied to square pixel blocks of the same size, and can also be applied to variable-sized blocks that are not square.

[0074] The quantizer 233 can quantize the transform coefficients and transmit the quantized transform coefficients to the entropy encoder 240. The entropy encoder 240 can encode the quantized signal (information about the quantized transform coefficients) into an encoded quantized signal and output it as a bitstream. The information about the quantized transform coefficients can be referred to as residual information. The quantizer 233 can rearrange the quantized transform coefficients in block form into a one-dimensional vector form based on the coefficient scanning order, and also generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. The entropy encoder 240 can perform various encoding methods such as exponential Golomb coding, context-adaptive variable length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy encoder 240 can also encode information necessary for video / image reconstruction (e.g., syntax element values, etc.) in addition to the quantized transform coefficients, either together or separately. The encoded information (e.g., encoded video / image information) can be transmitted in the form of a bitstream or stored in units of network abstraction layer (NAL) units. The video / image information may also include information about various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may also include conventional constraint information. The signaled / sent information and / or syntax elements described later in this document may be encoded by the encoding process mentioned above and thus included in the bitstream. The bitstream may be sent over a network or stored in a digital storage medium. Here, the network may include a broadcast network and / or a communication network, etc., and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. A transmitter (not shown) for sending a signal output from the entropy encoder 240 and / or a memory (not shown) for storing the signal may be configured as an internal / external element of the encoding device 200, or the transmitter may also be included in the entropy encoder 240.

[0075] The quantized transform coefficients output from the quantizer 233 can be used to generate a prediction signal. For example, the inverse quantizer 234 and the inverse transformer 235 apply inverse quantization and inverse transformation to the quantized transform coefficients so that the residual signal (residual block or residual sample) can be reconstructed. The adder 250 adds the reconstructed residual signal to the prediction signal output from the inter-frame predictor 221 or the intra-frame predictor 222 so that a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) can be generated. For example, in the case of applying the skip mode, if there is no residual in the block to be processed, the prediction block can be used as a reconstructed block. The adder 250 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current picture, and as described later, it is also used for inter-frame prediction of the next picture by filtering.

[0076] Furthermore, luma mapping with chroma scaling (LMCS) may be applied in the picture encoding and / or reconstruction process.

[0077] The filter 260 can apply filtering to the reconstructed signal to improve the subjective / objective image quality. For example, the filter 260 can apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture, and store the modified reconstructed picture in the memory 270, specifically, in the DPB of the memory 270. Various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filter 260 can generate various information related to filtering to transmit the generated information to the entropy encoder 240, as described subsequently in the description of each filtering method. The information related to filtering can be encoded by the entropy encoder 240 and output in the form of a bitstream.

[0078] The modified reconstructed picture transmitted to the memory 270 may be used as a reference picture in the inter predictor 221. If inter prediction is applied through the inter predictor, prediction mismatch between the encoding apparatus 200 and the decoding apparatus may be avoided, and encoding efficiency may be improved.

[0079] The DPB of the memory 270 can store the modified reconstructed picture for use as a reference picture in the inter-frame predictor 221. The memory 270 can store the motion information of the block from which the motion information in the current picture is derived (or encoded) and / or the motion information of the block in the previously reconstructed picture. The stored motion information can be transmitted to the inter-frame predictor 221 to be used as the motion information of the spatially neighboring block or the motion information of the temporally neighboring block. The memory 270 can store the reconstructed samples of the reconstructed block in the current picture and can transmit the reconstructed samples to the intra-frame predictor 222.

[0080] Figure 3 Schematically illustrates the configuration of a video / image decoding device to which this document is applicable. Hereinafter, the so-called decoding device may include an image decoding device and / or a video decoding device.

[0081] Reference Figure 3, the decoding device 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-frame predictor 332 and an intra-frame predictor 331. The residual processor 320 may include an inverse quantizer 321 and an inverse transformer 322. According to an embodiment, the entropy decoder 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 described above may be configured by hardware components (e.g., a decoder chipset or processor). In addition, the memory 360 may include a decoded picture buffer (DPB) or may be configured by a digital storage medium. The hardware components may also include the memory 360 as an internal / external component.

[0082] When a bit stream including video / image information is input, the decoding apparatus 300 may generate a signal in response to the bit stream received in the video / image processing. Figure 2 , the image is reconstructed by processing the video / image information in the encoding device illustrated in . For example, the decoding device 300 may derive a unit / block based on block segmentation related information obtained from the bit stream. The decoding device 300 may perform decoding using a processing unit applied to the encoding device. Thus, for example, the processing unit of decoding may be a coding unit, and the coding unit may be split from a coding tree unit or a maximum coding unit according to a quadtree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units may be derived from the coding unit. In addition, the reconstructed image signal decoded and output by the decoding device 300 may be reproduced by a reproduction device.

[0083] The decoding device 300 may receive the data from the Figure 2The signal output by the encoding device illustrated in the example, and the received signal can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can derive the information necessary for image reconstruction (or picture reconstruction) (e.g., video / image information) by parsing the bitstream. The video / image information may also include information about various parameter sets such as the Adaptive Parameter Set (APS), the Picture Parameter Set (PPS), the Sequence Parameter Set (SPS), and the Video Parameter Set (VPS). In addition, the video / image information may also include conventional constraint information. The decoding device can further decode the picture based on the information about the parameter sets and / or the conventional constraint information. The signaled / received information and / or syntax elements described later in this document can be decoded and obtained from the bitstream through a decoding process. For example, the entropy decoder 310 can decode the information in the bitstream based on a coding method such as Exponential Golomb coding, CAVLC, or CABAC, and output the values ​​of the syntax elements necessary for image reconstruction and the quantized values ​​of the residual-related transform coefficients. More specifically, the CABAC entropy decoding method can receive the bin corresponding to each syntax element from the bitstream, use the syntax element information to be decoded and the decoded information of the neighboring blocks and the block to be decoded or the information of the symbol / bin decoded in the previous stage to determine the context model, and generate the symbol corresponding to the value of each syntax element by predicting the bin generation probability according to the determined context model to perform arithmetic decoding of the bin. At this time, the CABAC entropy decoding method can determine the context model and then update the context model using the decoded symbol / bin information of the context model for the next symbol / bin. The information about prediction among the information decoded by the entropy decoder 310 can be provided to the predictor (inter-frame predictor 332 and intra-frame predictor 331), and the residual value (i.e., quantized transform coefficients and related parameter information) entropy decoded by the entropy decoder 310 can be input to the residual processor 320. The residual processor 320 can derive the residual signal (residual block, residual sample, residual sample array). In addition, the information about filtering among the information decoded by the entropy decoder 310 can be provided to the filter 350. In addition, a receiver (not illustrated) for receiving a signal output from the encoding device may be further configured as an internal / external element of the decoding device 300, or the receiver may also be a component of the entropy decoder 310. In addition, the decoding device according to this document may be referred to as a video / image / picture decoding device, and the decoding device may also be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoder 310, and the sample decoder may include at least one of the inverse quantizer 321, the inverse transformer 322, the adder 340, the filter 350, the memory 360, the inter-frame predictor 332, and the intra-frame predictor 331.

[0084] The inverse quantizer 321 may inversely quantize the quantized transform coefficients to output the transform coefficients. The inverse quantizer 321 may rearrange the quantized transform coefficients in a two-dimensional block format. In this case, the rearrangement may be performed based on the coefficient scanning order performed by the encoding device. The inverse quantizer 321 may perform inverse quantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain the transform coefficients.

[0085] The inverse transformer 322 inversely transforms the transform coefficients to obtain a residual signal (residual block, residual sample array).

[0086] The predictor 330 may perform prediction of the current block and generate a prediction block including prediction samples of the current block. The predictor may determine whether to apply intra prediction or inter prediction to the current block based on the information on prediction output from the entropy decoder 310 and may determine a specific intra / inter prediction mode.

[0087] The predictor can generate a prediction signal based on various prediction methods to be described later. For example, the predictor can not only apply intra prediction or inter prediction to predict a block, but also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). In addition, the predictor can be based on an intra-block copy (IBC) prediction mode or a palette mode to perform prediction on the block. The IBC prediction mode or palette mode can be used for content image / video coding such as games such as screen content coding (SCC). IBC basically performs prediction in the current picture, but in terms of deriving the reference block in the current picture, it is performed similarly to inter prediction. That is, IBC can use at least one of the inter prediction techniques described in this document. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, information about the palette table and palette index can be included in the video / image information and signaled.

[0088] The intra-frame predictor 331 can predict the current block by referencing samples in the current picture. Depending on the prediction mode, the referenced samples may be located in the neighborhood of the current block or may be located separately from the current block. In intra-frame prediction, prediction modes can include multiple non-directional modes and multiple directional modes. The intra-frame predictor 331 can also determine the prediction mode to be applied to the current block by using the prediction modes applied to neighboring blocks.

[0089] The inter-frame predictor 332 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can also include information about the inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, the neighboring blocks can include spatial neighboring blocks in the current picture and temporal neighboring blocks in the reference picture. For example, the inter-frame predictor 332 can construct a motion information candidate list based on the neighboring blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and the prediction information can include information indicating the inter-frame prediction mode for the current block.

[0090] The adder 340 may add the obtained residual signal to the prediction signal (prediction block or prediction sample array) output from the predictor 330 (including the intra-frame predictor 331 and the inter-frame predictor 332) to generate a reconstructed signal (reconstructed picture, reconstructed block or reconstructed sample array). For example, in the case of applying the skip mode, if there is no residual of the block to be processed, the prediction block can be used as the reconstructed block.

[0091] The adder 340 may be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal may be used for intra-frame prediction of the next block to be processed in the current picture, and as described later, may also be output through filtering or may also be used for inter-frame prediction of the next picture.

[0092] In addition, luma mapping with chroma scaling (LMCS) can also be applied to the picture decoding process.

[0093] The filter 350 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 350 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image, and store the modified reconstructed image in the memory 360, specifically, in the DPB of the memory 360. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.

[0094] The (modified) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter-frame predictor 332. The memory 360 can store the motion information of the block from which the motion information in the current picture is derived (or decoded) and / or the motion information of the block in the reconstructed picture. The stored motion information can be transmitted to the inter-frame predictor 332 to be used as the motion information of the spatially adjacent block or the motion information of the temporally adjacent block. The memory 360 can store the reconstructed samples of the reconstructed block in the current picture and transmit the reconstructed samples to the intra-frame predictor 331.

[0095] In this document, the exemplary embodiments described in the filter 260, the inter-frame predictor 221, and the intra-frame predictor 222 of the encoding device 200 may be equally or correspondingly applied to the filter 350, the inter-frame predictor 332, and the intra-frame predictor 331 of the decoding device 300, respectively.

[0096] In addition, as described above, when performing video encoding, prediction is performed to enhance compression efficiency. Accordingly, a prediction block including prediction samples of a current block to be encoded (i.e., a target coding block) can be generated. Here, the prediction block includes prediction samples in the spatial domain (or pixel domain). The prediction block is derived identically in the encoding device and the decoding device, and the encoding device can signal information (residual information) about the residual between the original block and the prediction block (rather than the original sample values ​​of the original block themselves) to the decoding device to enhance image coding efficiency. The decoding device can derive a residual block including residual samples based on the residual information, add the residual block and the prediction block to generate a reconstructed block including reconstructed samples, and generate a reconstructed picture including the reconstructed block.

[0097] Residual information can be generated through a transform process and a quantization process. For example, the encoding device may derive a residual block between the original block and the prediction block, may perform a transform process on the residual samples (residual sample array) included in the residual block to derive transform coefficients, may perform a quantization process on the transform coefficients to derive quantized transform coefficients, and may signal the relevant residual information to the decoding device (via a bitstream). In this case, the residual information may include value information such as the quantized transform coefficients, position information, transform scheme, transform kernel, and quantization parameter. The decoding device may perform an inverse quantization / inverse transform process based on the residual information and may derive residual samples (or residual blocks). The decoding device may generate a reconstructed picture based on the prediction block and the residual block. In addition, for inter-frame prediction reference of subsequent pictures, the encoding device may also perform inverse quantization / inverse transform on the quantized transform coefficients to derive a residual block, and may generate a reconstructed picture based on this.

[0098] Furthermore, as described above, when prediction is performed on the current block, intra-frame prediction or inter-frame prediction can be applied. In an embodiment, when inter-frame prediction is applied to the current block, a predictor (more specifically, an inter-frame predictor) of the encoding / decoding device can derive prediction samples by performing inter-frame prediction on a block-by-block basis. Inter-frame prediction may refer to a prediction derived using a method that depends on data elements (e.g., sample values ​​or motion information) of a picture other than the current picture. When inter-frame prediction is applied to the current block, the prediction block (prediction sample array) of the current block can be derived based on a reference block (reference sample array) specified by a motion vector in a reference picture indicated by a reference picture index. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, the motion information of the current block can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may also include information on the inter-frame prediction type (L0 prediction, L1 prediction, Bi prediction, etc.). When inter-frame prediction is applied, the neighboring blocks may include spatially neighboring blocks in the current picture and temporally neighboring blocks in the reference picture. The reference picture including the reference block and the reference picture including the temporally neighboring block may be the same as or different from each other. The temporally neighboring block may be referred to as a collocated reference block, collocated CU (colCU), etc., and the reference picture including the temporally neighboring block may be referred to as a collocated picture (colPic). For example, a motion information candidate list may be configured based on the neighboring blocks of the current block, and a flag or index information indicating which candidate is selected (used) to derive the motion vector and / or reference picture index of the current block may be signaled. Inter-frame prediction may be performed based on various prediction modes, and for example, in skip mode and merge mode, the motion information of the current block may be the same as the motion information of the selected neighboring block. In skip mode, a residual signal may not be transmitted as in merge mode. In motion vector prediction (MVP) mode, the motion vector of the selected neighboring block may be used as a motion vector predictor, and the motion vector difference may be signaled. In this case, the motion vector of the current block may be derived using the sum of the motion vector predictor and the motion vector difference.

[0099] Depending on the inter-frame prediction type (L0 prediction, L1 prediction, Bi prediction, etc.), the motion information may also include L0 motion information and / or L1 motion information. The L0-direction motion vector may be referred to as the L0 motion vector or MVL0, and the L1-direction motion vector may be referred to as the L1 motion vector or MVL1. Prediction based on the L0 motion vector may be referred to as L0 prediction, prediction based on the L1 motion vector may be referred to as L1 prediction, and prediction based on both the L0 motion vector and the L1 motion vector may be referred to as bidirectional prediction. Here, the L0 motion vector may indicate a motion vector associated with reference picture list L0, and the L1 motion vector may indicate a motion vector associated with reference picture list L1. Reference picture list L0 may include pictures preceding the current picture in output order, and reference picture list L1 may include pictures following the current picture in output order as reference pictures. The preceding picture may be referred to as a forward (reference) picture, and the following picture may be referred to as a backward (reference) picture. Reference picture list L0 may also include pictures following the current picture in output order as reference pictures. In this case, the previous picture in reference picture list L0 may be indexed first, followed by the next picture. Reference picture list L1 may also include pictures preceding the current picture in output order as reference pictures. In this case, the next picture in reference picture list L1 may be indexed first, followed by the previous picture. Here, the output order may correspond to the picture order count (POC) order.

[0100] Figure 4 An example of an exemplary video / image encoding process to which one or more embodiments of this document may be applied is shown. Figure 4 In the above, S400 can be Figure 2 S400 may include the inter / intra prediction process described in this document; S410 may include the residual processing process described in this document; and S420 may include the information encoding process described in this document.

[0101] Reference Figure 4 The video / image encoding process may illustratively include a process of generating a reconstructed picture for the current picture and a process of applying loop filtering to the reconstructed picture (optional), and encoding information for reconstructing the picture (e.g., prediction information, residual information, or segmentation information) to generate a reconstructed picture as shown in FIG. Figure 2The process of outputting the encoded information in the form of a bitstream is described. The encoding device can derive (modified) residual samples from the quantized transform coefficients through the inverse quantizer 234 and the inverse transformer 235, and generate a reconstructed picture based on the prediction samples and the (modified) residual samples output as S400. The reconstructed picture thus generated can be the same as the reconstructed picture mentioned above generated by the decoding device. The modified reconstructed picture can be generated by a loop filtering process for the reconstructed picture and can be stored in the decoded picture buffer or memory 270 and, as in the case of the decoding device, used as a reference picture in the inter-frame prediction process when the picture is subsequently encoded. As described above, in some cases, some or all of the loop filtering process can be omitted. If a loop filtering process is performed, the (loop) filtering related information (parameters) can be encoded by the entropy encoder 240 and output in the form of a bitstream, and the decoding device can perform a loop filtering process based on the filtering related information in the same way as the encoding device.

[0102] Through loop filtering, noise such as blocking artifacts and ringing artifacts generated when encoding images / videos can be reduced, and subjective and objective visual quality can be enhanced. In addition, by performing loop filtering in both the encoding and decoding devices, the encoding and decoding devices can derive the same prediction results, increasing the reliability of picture encoding and reducing the amount of data transmitted for picture encoding.

[0103] As described above, picture reconstruction processing can be performed in an encoding device as well as in a decoding device. A reconstructed block can be generated based on intra prediction / inter prediction in units of each block, and a reconstructed picture including the reconstructed block can be generated. If the current picture / slice / tile group is an I picture / slice / tile group, the blocks included in the current picture / slice / tile group can be reconstructed based only on intra prediction. In addition, if the current picture / slice / tile group is a P or B picture / slice / tile group, the blocks included in the current picture / slice / tile group can be reconstructed based on intra prediction or inter prediction. In this case, inter prediction can be applied to some blocks in the current picture / slice / tile group, and intra prediction can also be applied to other blocks. The color components of a picture may include a luminance component and a chrominance component, and unless otherwise expressly limited in this document, the method and exemplary embodiments proposed in this document may be applied to the luminance component and the chrominance component.

[0104] Figure 5 An example of an exemplary video / image decoding process to which one or more embodiments of this document may be applied is shown. Figure 5 In the above, S500 can be Figure 3S500 may include the information decoding process described in this document; S510 may include the inter / intra prediction process described in this document; S520 may include the residual processing process described in this document; S530 may include the block / picture reconstruction process described in this document; and S540 may include the loop filtering process described in this document.

[0105] Reference Figure 5 , such as in the case of Figure 3 As shown in the description, the picture decoding process can illustratively include image / video information acquisition processing S500 from the bitstream (through decoding), picture reconstruction processing S510 to S530, and loop filtering processing S540 for reconstructing the picture. The picture reconstruction process can be performed based on the residual samples and prediction samples obtained through the inter-frame / intra-frame prediction S510 and residual processing S520 (dequantization and inverse transformation of quantized transform coefficients) described in this document. By loop filtering the reconstructed picture generated by the picture reconstruction process, a modified reconstructed picture can be generated. The modified reconstructed picture can be output as a decoded picture, or stored in the memory 360 or decoded picture buffer of the decoding device and used as a reference picture in the inter-frame prediction process of subsequent picture decoding.

[0106] Depending on the situation, the loop filtering process can be skipped, and in this case, the reconstructed picture can be output as a decoded picture, or it can be stored in the memory 360 or decoded picture buffer of the decoding device and used as a reference picture in the inter-frame prediction process of subsequent picture decoding. The loop filtering process S540 may include the deblocking filtering process, sample adaptive offset (SAO) process, adaptive loop filter (ALF) process and / or bidirectional filter process as described above, and all or some of them can be skipped. In addition, one or some of the deblocking filtering process, sample adaptive offset (SAO) process, adaptive loop filter (ALF) process and bidirectional filter process can be applied sequentially, or all of them can be applied sequentially. For example, after applying the deblocking filtering process to the reconstructed picture, the SAO process can be performed on it. Alternatively, for example, after applying the deblocking filtering process to the reconstructed picture, the ALF process can be performed on it. This can also be performed in the encoding device.

[0107] In addition, as described above, the encoding device can perform entropy encoding based on various encoding methods such as exponential Golomb coding, context-adaptive variable length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). In addition, the decoding device can perform entropy decoding based on encoding methods such as exponential Golomb coding, CAVLC, or CABAC. Hereinafter, the entropy encoding / decoding process will be described.

[0108] Figure 6 An example of an entropy coding method to which the embodiments of this document are applicable is schematically illustrated, and Figure 7 An entropy encoder in an encoding device is schematically illustrated. Figure 7 The entropy encoder in the encoding device can be applied equally or even correspondingly to the above-mentioned Figure 2 The entropy encoder 240 of the encoding device 200.

[0109] Reference Figure 6 and Figure 7 , the encoding device (entropy encoder) can perform entropy encoding processing on the image / video information. The image / video information may include partition related information, prediction related information (for example, inter / intra prediction classification information, intra prediction mode information, inter prediction mode information, etc.), residual information and loop filter related information, and may also include various syntax elements thereof. Entropy encoding can be performed in units of syntax elements. S600 to S610 may be performed as described above. Figure 2 The entropy encoder 240 of the encoding device 200 is performed.

[0110] The encoding device may perform binarization on the target syntax element (S600). Here, the binarization may be based on various binarization methods such as truncated Rice binarization or fixed-length binarization, and the binarization method used for the target syntax element may be predefined. The binarization process may be performed by the binarizer 242 in the entropy encoder 240.

[0111] The encoding device may perform entropy encoding on the target syntax element (S610). The encoding device may perform encoding of the bin string of the target syntax element based on normal encoding (based on context) or bypass encoding based on an entropy encoding technique such as context adaptive arithmetic coding (CABAC) or context adaptive variable length coding (CAVLC), and its output may be included in the bitstream. The entropy encoding process may be performed by the entropy encoding processor 243 in the entropy encoder 240. The bitstream may be transmitted to the decoding device via a (digital) storage medium or a network as described above.

[0112] Figure 8 An example of an entropy decoding method to which the embodiments of this document are applicable is schematically illustrated, and Figure 9 An entropy decoder in a decoding device is schematically illustrated. Figure 9 The entropy decoder in the decoding device can be applied equally or even correspondingly to the above-mentioned Figure 3 The entropy decoder 310 of the decoding device 300.

[0113] Reference Figure 8 and Figure 9 , the decoding device (entropy decoder) can decode the encoded image / video information. The image / video information may include partition related information, prediction related information (e.g., inter / intra prediction classification information, intra prediction mode information, inter prediction mode information, etc.), residual information and loop filter related information, and may also include various syntax elements thereof. Entropy coding may be performed in units of syntax elements. S800 to S810 may be performed as described above. Figure 3 The entropy decoder 310 of the decoding device 300 is executed.

[0114] The decoding device may perform binarization on the target syntax element (S800). Here, the binarization may be based on various binarization methods such as truncated Rice binarization or fixed-length binarization, and the binarization method used for the target syntax element may be predefined. The decoding device may derive an enabled bin string (bin string candidate) of an enabled value of the target syntax element through the binarization process. The binarization process may be performed by the binarizer 312 in the entropy decoder 310.

[0115] The decoding device can perform entropy decoding on the target syntax element (S810). While sequentially decoding and parsing the various bins of the target syntax element from the input bits in the bitstream, the decoding device compares the derived bin string with the enabled bin string of the corresponding syntax element. If the derived bin string is equal to one of the enabled bin strings, the value corresponding to the corresponding bin string can be derived as the value of the corresponding syntax element. Otherwise, the above process can be performed again after further parsing the next bit in the bitstream. Through such processing, variable length bits can be used to signal the corresponding information, even without using the start bit or end bit of specific information (specific syntax element) in the bitstream. In this way, a relatively small number of bits can be allocated for low values, so the overall coding efficiency can be enhanced.

[0116] The decoding device can perform context-based or bypass-based decoding of the corresponding bin in the bin string from the bitstream based on an entropy coding technique such as CABAC or CAVLC. Here, the bitstream may include various information for image / video decoding as described above. The bitstream may be transmitted to the decoding device via a (digital) storage medium or network as described above.

[0117] Figure 10 The hierarchical structure of the coded image / video is shown as an example.

[0118] Reference Figure 10 The coded image / video is divided into the VCL (Video Coding Layer) that handles the image / video decoding process and itself, the subsystem that sends and stores the coded information, and the Network Abstraction Layer (NAL) that exists between the VCL and the subsystem and is responsible for the network adaptation function.

[0119] VCL can generate VCL data including compressed image data (slice data), or generate parameter sets including picture parameter sets (picture parameter set: PPS), sequence parameter sets (sequence parameter set: SPS), video parameter sets (video parameter set: VPS), etc., or supplementary enhancement information (SEI) messages necessary for the decoding processing of images.

[0120] In NAL, a NAL unit can be generated by adding header information (NAL unit header) to the raw byte sequence payload (RBSP) generated in the VCL. In this case, RBSP refers to slice data, parameter sets, SEI messages, etc. generated in the VCL. The NAL unit header may include NAL unit type information specified according to the RBSP data included in the corresponding NAL unit.

[0121] In addition, according to the RBSP generated in the VCL, the NAL unit can be divided into a VCL NAL unit and a non-VCL NAL unit. A VCL NAL unit may refer to a NAL unit including information about an image (slice data), and a non-VCL-NAL unit may refer to a NAL unit including information required for decoding an image (parameter set or SEI message).

[0122] VCL NAL units and non-VCL NAL units can be sent over a network by attaching header information according to the data standard of the subsystem. For example, NAL units can be converted into a predetermined standard data format such as H.266 / VVC file format, Real-time Transport Protocol (RTP), and Transport Stream (TS), and sent over various networks.

[0123] As described above, in a NAL unit, a NAL unit type may be specified according to an RBSP data structure included in a corresponding NAL unit, and information about the NAL unit type may be stored in a NAL unit header and signaled.

[0124] For example, NAL units can be roughly divided into VCL NAL unit types and non-VCL NAL unit types according to whether the NAL unit includes information about the image (slice data). VCL NAL unit types can be classified according to the nature and type of the picture included in the VCL NAL unit, and non-VCL NAL unit types can be classified according to the type of parameter set.

[0125] The following are examples of NAL unit types specified according to the type of parameter sets included in non-VCL NAL unit types.

[0126] -APS (Adaptation Parameter Set) NAL unit: type of NAL unit including APS

[0127] -DPS (Decoding Parameter Set) NAL unit: the type of NAL unit that includes DPS

[0128] -VPS (Video Parameter Set) NAL unit: Type of NAL unit containing VPS

[0129] -SPS (Sequence Parameter Set) NAL unit: the type of NAL unit that includes the SPS

[0130] -PPS (Picture Parameter Set) NAL unit: Type of NAL unit including PPS

[0131] -PH (Picture Header) NAL unit: the type of NAL unit including PH

[0132] The above-mentioned NAL unit type has syntax information for the NAL unit type, and the syntax information can be stored in the NAL unit header and signaled. For example, the syntax information can be nal_unit_type, and the NAL unit type can be specified by the nal_unit_type value.

[0133] In addition, as described above, a picture may include multiple slices, and a slice may include a slice header and slice data. In this case, a picture header may be further added to the multiple slices (slice header and slice data set) in a picture. The picture header (picture header syntax) may include information / parameters common to the picture. In this document, a tile group may be mixed with a slice or a picture or replaced with a slice or a picture. In addition, in this document, a tile group header may be mixed with a slice header or a picture header or replaced with a slice header or a picture header.

[0134] A slice header (slice header syntax) may include information / parameters common to slices. APS (APS syntax) or PPS (PPS syntax) may include information / parameters common to one or more slices or pictures. SPS (SPS syntax) may include information / parameters common to one or more sequences. VPS (VPS syntax) may include information / parameters common to multiple layers. DPS (DPS syntax) may include information / parameters common to the entire video. DPS may include information / parameters related to the concatenation of coded video sequences (CVS). In this document, a high-level syntax (HLS) may include at least one of APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, picture header syntax, and slice header syntax.

[0135] In this document, the image / video information encoded in the encoding device and signaled to the decoding device in the form of a bitstream may include not only picture segmentation related information, intra / inter prediction information, residual information, loop filtering information, etc. in the picture, but also information included in the slice header, information included in the picture header, information included in the APS, information included in the PPS, information included in the SPS, information included in the VPS, and / or information included in the DPS. In addition, the image / video information may also include information in the NAL unit header.

[0136] As described above, high-level syntax (HLS) can be encoded / signaled for video / image coding. In this document, the video / image information may include HLS. For example, a coded picture may consist of one or more slices. Parameters describing the coded picture may be signaled in a picture header (PH), and parameters describing the slice may be signaled in a slice header (SH). The PH may be sent as its own NAL unit type. The SH may be present at the beginning of a NAL unit including the payload of the slice (i.e., slice data). The details of the syntax and semantics of the PH and SH may be as disclosed in the VVC standard. Each picture may be associated with a PH. A picture may consist of different types of slices: intra-coded slices (i.e., I slices) and inter-coded slices (i.e., P slices and B slices). As a result, the PH may include syntax elements necessary for intra slices of a picture and inter slices of a picture.

[0137] In addition, generally, one NAL unit type can be set for one picture. The NAL unit type can be signaled by nal_unit_type in the NAL unit header of the NAL unit containing the slice. nal_unit_type is syntax information for specifying the NAL unit type, that is, as shown in Table 1 or Table 2 below, it can specify the type of RBSP data structure contained in the NAL unit.

[0138] Table 1 below shows examples of NAL unit type codes and NAL unit type categories.

[0139] [Table 1]

[0140]

[0141]

[0142]

[0143] Alternatively, as an example, the NAL unit type code and the NAL unit type category may be defined as shown in Table 2 below.

[0144] [Table 2]

[0145]

[0146]

[0147]

[0148] As shown in Table 1 or Table 2, the name of the NAL unit type and its value can be specified according to the RBSP data structure included in the NAL unit, and can be divided into a VCL NAL unit type and a non-VCL NAL unit type according to whether the NAL unit includes information about the image (slice data). The VCL NAL unit type can be classified according to the nature and type of the picture, and the non-VCL NAL unit type can be classified according to the type of the parameter set, etc. For example, the NAL unit type can be specified according to the nature and type of the picture included in the VCL NAL unit as follows.

[0149] TRAIL: This indicates the type of the NAL unit including the coded slice data of the trailing picture / sub-picture. For example, nal_unit_type may be defined as TRAIL_NUT, and the value of nal_unit_type may be specified as 0.

[0150] Here, a trailing picture refers to a picture that follows a picture that can be randomly accessed in both output order and decoding order. A trailing picture can be a non-IRAP picture that follows the associated IRAP picture in output order and is not an STSA picture. For example, a trailing picture associated with an IRAP picture follows the IRAP picture in decoding order. Pictures that follow the associated IRAP picture in output order and precede the associated IRAP picture in decoding order are not permitted.

[0151] STSA (Step-by-Step Temporal Sub-Layer Access): This indicates the type of the NAL unit including the coded slice data of the STSA picture / sub-picture. For example, nal_unit_type may be defined as STSA_NUT, and the value of nal_unit_type may be specified as 1.

[0152] Here, the STSA picture is a picture that can be switched between temporal sub-layers in a bitstream that supports temporal scalability, and is a picture that indicates a position where up-switching from a lower sub-layer to an upper sub-layer that is one step higher than the lower sub-layer is possible. The STSA picture does not use pictures with the same TemporalId as the STSA picture and in the same layer as the STSA picture for inter-frame prediction. Pictures that follow the STSA picture with the same TemporalId as the STSA picture and in the same layer in decoding order do not use pictures that precede the STSA picture with the same TemporalId as the STSA picture and in the same layer in decoding order for inter-frame prediction reference. The STSA picture enables up-switching from the immediately next sub-layer to the sub-layer including the STSA picture in the STSA picture. In this case, the picture being encoded must not belong to the lowest sub-layer. That is, the STSA picture must always have a TemporalId greater than 0.

[0153] RADL (Random Access Decodable Preamble (Picture)): This indicates the type of the NAL unit including slice data of the RADL picture / sub-picture to be encoded. For example, nal_unit_type may be defined as RADL_NUT, and the value of nal_unit_type may be specified as 2.

[0154] Here, all RADL pictures are leading pictures. RADL pictures are not used as reference pictures in the decoding process of the trailing pictures of the same associated IRAP picture. Specifically, a RADL picture with nuh_layer_id equal to layerId is a picture that follows the IRAP picture associated with the RADL picture in output order and is not used as a reference picture in the decoding process of pictures with nuh_layer_id equal to layerId. When field_seq_flag (i.e., sps_field_seq_flag) is 0, all RADL pictures precede all non-leading pictures of the same associated IRAP picture in decoding order (i.e., if a RADL picture exists). Furthermore, a leading picture refers to a picture that precedes the associated IRAP picture in output order.

[0155] RASL (Random Access Skip Preamble (Picture)): This indicates the type of the NAL unit including slice data of the RASL picture / sub-picture to be encoded. For example, nal_unit_type may be defined as RASL_NUT, and the value of nal_unit_type may be specified as 3.

[0156] Here, all RASL pictures are leading pictures of the associated CRA picture. When the associated CRA picture has NoOutputBeforeRecoveryFlag whose value is 1, the RASL picture may not be output nor decoded correctly because the RASL picture may include references to pictures that do not exist in the bitstream. RASL pictures are not used as reference pictures for the decoding process of non-RASL pictures of the same layer. However, RADL sub-pictures in RASL pictures of the same layer can be used for inter-frame prediction of collocated RADL sub-pictures in RADL pictures associated with the same CRA picture as the RASL picture. When field_seq_flag (i.e., sps_field_seq_flag) is 0, all RASL pictures precede all non-leading pictures of the same associated CRA picture in decoding order (i.e., if a RASL picture exists).

[0157] For non-IRAP VCL NAL unit types, there may be a reserved nal_unit_type. For example, nal_unit_type may be defined as RSV_VCL_4 and RSV_VCL_6, and the value of nal_unit_type may be specified as 4 to 6, respectively.

[0158] Here, an intra random access point (IRAP) is information indicating a NAL unit for a picture that can be randomly accessed. An IRAP picture can be a CRA picture or an IDR picture. For example, an IRAP picture refers to a picture having a NAL unit type defined as IDR_W_RADL, IDR_N_LP, and CRA_NUT as in Table 1 or Table 2 above, and the value of nal_unit_type can be specified as 7 to 9, respectively.

[0159] IRAP pictures do not use any reference pictures in the same layer for inter prediction during the decoding process. In other words, an IRAP picture does not reference any pictures other than itself for inter prediction during the decoding process. The first picture in the bitstream in decoding order is called an IRAP or GDR picture. For a single-layer bitstream, if the necessary parameter set is available when it is needed to reference it, all subsequent non-RASL pictures and IRAP pictures in the coded layer video sequence (CLVS) in decoding order can be accurately decoded without performing the decoding process for pictures that precede the IRAP picture in decoding order.

[0160] The value of mixed_nalu_types_in_pic_flag for an IRAP picture is 0. When the value of mixed_nalu_types_in_pic_flag for a picture is 0, one slice in the picture may have a NAL unit type (nal_unit_type) in the range from IDR_W_RADL to CRA_NUT (for example, the NAL unit type values ​​in Table 1 or Table 2 are 7 to 9), and all other slices in the picture may have the same NAL unit type (nal_unit_type). In this case, the picture can be regarded as an IRAP picture.

[0161] Instantaneous Decoding Refresh (IDR): This indicates the type of NAL unit containing slice data of the IDR picture / sub-picture to be encoded. For example, the nal_unit_type of the IDR picture / sub-picture can be defined as IDR_W_RADL or IDR_N_LP, and the value of nal_unit_type can be specified as 7 or 8, respectively.

[0162] Here, an IDR picture may not use inter-frame prediction in the decoding process (i.e., it does not refer to pictures other than itself for inter-frame prediction), but may be the first picture in the bitstream in decoding order, or may appear later in the bitstream (i.e., not first, but later). Each IDR picture is the first picture of a coded video sequence (CVS) in decoding order. For example, when an IDR picture is associated with a decodable leading picture, the NAL unit type of the IDR picture may be represented as IDR_W_RADL, and when the IDR picture is not associated with a leading picture, the NAL unit type of the IDR picture may be represented as IDR_N_LP. That is, an IDR picture whose NAL unit type is IDR_W_RADL may not have an associated RADL picture present in the bitstream, but may have an associated RADL picture in the bitstream. An IDR picture whose NAL unit type is IDR_N_LP does not have an associated leading picture present in the bitstream.

[0163] Clean Random Access (CRA): This indicates the type of the NAL unit including the slice data of the CRA picture / sub-picture to be encoded. For example, nal_unit_type may be defined as CRA_NUT, and the value of nal_unit_type may be specified as 9.

[0164] Here, a CRA picture may not use inter-frame prediction in the decoding process (i.e., it does not refer to pictures other than itself for inter-frame prediction), but may be the first picture in the decoding order in the bitstream, or may appear later in the bitstream (i.e., not first, but later). A CRA picture may have an associated RADL or RASL picture present in the bitstream. For a CRA picture in which the value of NoOutputBeforeRecoveryFlag is 1, the associated RASL picture may not be output by the decoder. This is because decoding is not possible in this case because a reference to a picture that does not exist in the bitstream is included.

[0165] Gradual Decoding Refresh (GDR): This indicates the type of the NAL unit including slice data of the GDR picture / sub-picture to be encoded. For example, nal_unit_type may be defined as GDR_NUT, and the value of nal_unit_type may be specified as 10.

[0166] Here, the pps_mixed_nalu_types_in_pic_flag value of the GDR picture may be 0. When the value of pps_mixed_nalu_types_in_pic_flag of the picture is 0 and one slice in the picture has the NAL unit type of GDR_NUT, all other slices in the picture have the same NAL unit type (nal_unit_type) value, and in this case, the picture may become a GDR picture after the first slice is received.

[0167] In addition, for example, the NAL unit type may be specified according to the type of parameters included in the non-VCL NAL unit, and, as shown in Table 1 or Table 2 above, the following NAL unit types (nal_unit_type) may be included: for example, VPS_NUT indicating the type of NAL unit including a video parameter set, SPS_NUT indicating the type of NAL unit including a sequence parameter set, PPS_NUT indicating the type of NAL unit including a picture parameter set, and PH_NUT indicating the type of NAL unit including a picture header.

[0168] In addition, a bitstream that supports temporal scalability (or a temporally scalable bitstream) includes information about a temporal layer for temporal scaling. The information about the temporal layer may be identification information of the temporal layer specified according to the temporal scalability of the NAL unit. For example, the identification information of the temporal layer may use temporal_id syntax information, and the temporal_id syntax information may be stored in the NAL unit header in the encoding device and signaled to the decoding device. Hereinafter, in this specification, a temporal layer may be referred to as a sublayer, a temporal sublayer, a temporal scalable layer, etc.

[0169] Figure 11 is a diagram showing a temporal layer structure of a NAL unit in a bitstream supporting temporal scalability.

[0170] When the bitstream supports temporal scalability, the NAL units included in the bitstream have identification information of the temporal layer (e.g., temporal_id). As an example, a temporal layer consisting of NAL units whose temporal_id value is 0 can provide the lowest temporal scalability, and a temporal layer consisting of NAL units whose temporal_id value is 2 can provide the highest temporal scalability.

[0171] exist Figure 11 , a block marked with I refers to an I picture, and a block marked with B refers to a B picture. In addition, an arrow indicates a reference relationship regarding whether a picture refers to another picture.

[0172] like Figure 11 As shown in , the NAL unit of the temporal layer whose temporal_id value is 0 is a reference picture that can be referenced by the NAL unit of the temporal layer whose temporal_id value is 0, 1, or 2. The NAL unit of the temporal layer whose temporal_id value is 1 is a reference picture that can be referenced by the NAL unit of the temporal layer whose temporal_id value is 1 or 2. The NAL unit of the temporal layer whose temporal_id value is 2 can be a reference picture that can be referenced by the NAL unit of the same temporal layer (i.e., the temporal layer whose temporal_id value is 2), or can be a non-reference picture that is not referenced by other pictures.

[0173] If Figure 11 If the NAL units of the temporal layer (ie, the highest temporal layer) whose temporal_id value is 2 shown in are non-reference pictures, these NAL units are extracted (or removed) from the bitstream in the decoding process without affecting other pictures.

[0174] Among the NAL unit types described above, the IDR and CRA types are information indicating that the NAL unit includes a picture that can be randomly accessed (or spliced), that is, a random access point (RAP) or an intra random access point (IRAP) picture used as a random access point. In other words, an IRAP picture can be an IDR or CRA picture and can include only I slices. In the bitstream, the first picture in decoding order becomes an IRAP picture.

[0175] If an IRAP picture (IDR, CRA picture) is included in the bitstream, there may be pictures that precede the IRAP picture in output order but follow it in decoding order. These pictures are called leading pictures (LP).

[0176] Figure 12 A diagram used to describe a picture that can be randomly accessed.

[0177] A picture that can be randomly accessed (ie, a RAP or IRAP picture used as a random access point) is the first picture in the bitstream in decoding order during random access and includes only I slices.

[0178] Figure 12 The output order (or display order) and decoding order of pictures are shown. As illustrated, the output order and decoding order of pictures may be different from each other. For convenience, the description is made while dividing the pictures into predetermined groups.

[0179] The pictures belonging to the first group (I) indicate pictures that precede the IRAP picture in both output order and decoding order, and the pictures belonging to the second group (II) indicate pictures that precede the IRAP picture in output order but follow it in decoding order. The pictures of the third group (III) follow the IRAP picture in both output order and decoding order.

[0180] The first group (I) of pictures may be decoded and output regardless of IRAP pictures.

[0181] A picture belonging to the second group (II) output before an IRAP picture is called a leading picture, and when an IRAP picture is used as a random access point, the leading picture may become a problem in a decoding process.

[0182] The pictures belonging to the third group (III) following the IRAP picture in output order and decoding order are called normal pictures. Normal pictures are not used as reference pictures for leading pictures.

[0183] A random access point where random access occurs in a bitstream becomes an IRAP picture, and random access starts as the first picture of the second group (II) is output.

[0184] Figure 13 This is a diagram used to describe an IDR picture.

[0185] An IDR picture is a picture that becomes a random access point when a group of pictures has a closed structure. As described above, since an IDR picture is an IRAP picture, it only includes I slices and can be the first picture in the bitstream in decoding order, or it can appear in the middle of the bitstream. When an IDR picture is decoded, all reference pictures stored in the decoded picture buffer (DPB) are marked as "unused for reference."

[0186] Figure 13 The bar shown in indicates a picture, and the arrow indicates a reference relationship related to whether the picture can use another picture as a reference picture. An x mark on the arrow indicates that the picture cannot refer to the picture indicated by the arrow.

[0187] As shown, a picture whose POC is 32 is an IDR picture. A picture whose POC is 25 to 31 and output before the IDR picture is a leading picture 1310. A picture whose POC is 33 or greater corresponds to a normal picture 1320.

[0188] The leading picture 1310 preceding the IDR picture in output order can use a leading picture other than the IDR picture as a reference picture, but cannot use a past picture 1330 preceding the leading picture 1310 in output order and decoding order as a reference picture.

[0189] The normal picture 1320 following the IDR picture in output order and decoding order can be decoded with reference to the IDR picture, the leading picture, and other normal pictures.

[0190] Figure 14 A diagram used to describe a CRA picture.

[0191] A CRA picture is a picture that becomes a random access point when a group of pictures has an open structure. As described above, since a CRA picture is also an IRAP picture, it only includes I slices and can be the first picture in the decoding order of the bitstream, or can appear in the middle of the bitstream for normal display.

[0192] Figure 14 The bars shown in indicate pictures, and the arrows indicate reference relationships related to whether a picture can use another picture as a reference picture. An x mark on an arrow indicates that one or more pictures cannot refer to the picture indicated by the arrow.

[0193] A leading picture 1410 that precedes the CRA picture in output order may use all of the CRA picture, other leading pictures, and past pictures 1430 that precede the leading picture 1410 in output order and decoding order as reference pictures.

[0194] Conversely, a normal picture 1420 following a CRA picture in output order and decoding order may be decoded with reference to a normal picture different from the CRA picture. The normal picture 1420 may not use the leading picture 1410 as a reference picture.

[0195] In addition, in the VVC standard, it is possible to allow the coded picture (ie, the current picture) to include slices of different NAL unit types. Whether the current picture includes slices of different NAL unit types can be indicated based on the syntax element mixed_nalu_types_in_pic_flag. For example, when the current picture includes slices of different NAL unit types, the value of the syntax element mixed_nalu_types_in_pic_flag can be represented as 1. In this case, the current picture must refer to a PPS including mixed_nalu_types_in_pic_flag with a value of 1. The semantics of the flag (mixed_nalu_types_in_pic_flag) are as follows:

[0196] When the value of the syntax element mixed_nalu_types_in_pic_flag is 1, it may indicate that each picture of the reference PPS has one or more VCL NAL units, the VCL NAL units do not have the same NAL unit type (nal_unit_type), and the picture is not an IRAP picture.

[0197] When the value of the syntax element mixed_nalu_types_in_pic_flag is 0, it may indicate that each picture of the referenced PPS has one or more VCL NAL units, and the VCL NAL units of each picture of the referenced PPS have the same value of the NAL unit type (nal_unit_type).

[0198] When the value of no_mixed_nalu_types_in_pic_constraint_flag is 1, the value of mixed_nalu_types_in_pic_flag must be 0. The no_mixed_nalu_types_in_pic_constraint_flag syntax element indicates a constraint on whether the value of mixed_nalu_types_in_pic_flag of a picture must be 0. For example, based on no_mixed_nalu_types_in_pic_constraint_flag information signaled from a higher-level syntax (e.g., PPS) or a syntax including information about constraints (e.g., GCI; general constraint information), it can be determined whether the value of mixed_nalu_types_in_pic_flag must be 0.

[0199] In a picture picA that also includes one or more slices with different values ​​of NAL unit types (i.e., when the value of mixed_nalu_types_in_pic_flag of picture picA is 1), for each slice with a NAL unit type value nalUnitTypeA in the range from IDR_W_RADL to CRA_NUT (e.g., in Table 1 or Table 2, the values ​​of NAL unit type are 7 to 9), the following may apply.

[0200] - The slice must belong to a subpicA in which the value of the corresponding subpic_treated_as_pic_flag[i] is 1. Here, subpic_treated_as_pic_flag[i] is information about whether the i-th subpicture of each picture encoded in the CLVS is treated as a picture in the decoding process other than the loop filtering operation. For example, when the value of subpic_treated_as_pic_flag[i] is 1, it can indicate that the i-th subpicture is treated as a picture in the decoding process other than the loop filtering operation. Alternatively, when the value of subpic_treated_as_pic_flag[i] is 0, it can indicate that the i-th subpicture is not treated as a picture in the decoding process other than the loop filtering operation.

[0201] - A slice shall not belong to a sub-picture of picA that includes a VCL NAL unit with a NAL unit type (nal_unit_type) not equal to nalUnitTypeA.

[0202] - For all PUs that follow CLVS in decoding order, the RefPicList[0] or RefPicList[1] of the slice in subpicA shall not include pictures that precede picA in decoding order in the active entry.

[0203] To implement the above-mentioned concepts, the following can be specified. For example, the following can be applied to the VCL NAL unit of a specific picture.

[0204] - When the value of mixed_nalu_types_in_pic_flag is 0, the value of the NAL unit type (nal_unit_type) must be the same for all slice NAL units coded in a picture. A picture or PU can be considered to have the same NAL unit type as the slice NAL units coded in the picture or PU.

[0205] Otherwise (when the value of mixed_nalu_types_in_pic_flag is 1), one or more VCL NAL units must have a NAL unit type of a specific value in the range from IDR_W_RADL to CRA_NUT (e.g., the values ​​of NAL unit type in Table 1 or Table 2 are 7 to 9), and all other VCL NAL units must have the same NAL unit type as GRA_NUT or a NAL unit type of a specific value in the range from TRAIL_NUT to RSV_VCL_6 (e.g., the values ​​of NAL unit type in Table 1 or Table 2 are 0 to 6).

[0206] In the current VVC standard, in the case of a picture with mixed NAL unit types, there may be at least the following problems.

[0207] 1. When a picture includes IDR and non-IRAP NAL units, and when signaling for a reference picture list (RPL) is present in the slice header, the signaling must also be present in the header of the IDR slice. When the value of sps_idr_rpl_present_flag is 1, RPL signaling is present in the slice header of the IDR slice. Currently, the value of this flag (sps_idr_rpl_present_flag) can be 0 even when there are one or more pictures with mixed NAL unit types. Here, the sps_idr_rpl_present_flag syntax element can indicate whether RPL syntax elements can be present in the slice header of a slice with a NAL unit type such as IDR_N_LP or IDR_W_RADL. For example, when the value of sps_idr_rpl_present_flag is 1, it can indicate that RPL syntax elements can be present in the slice header of a slice with a NAL unit type such as IDR_N_LP or IDR_W_RADL. Alternatively, when the value of sps_idr_rpl_present_flag is 0, it may indicate that the RPL syntax element is not present in the slice header of a slice with a NAL unit type such as IDR_N_LP or IDR_W_RADL.

[0208] 2.- When the current picture references a PPS with the value of mixed_nalu_types_in_pic_flag set to 1, one or more of the VCL NAL units of the current picture must have a NAL unit type with a specific value in the range from IDR_W_RADL to CRA_NUT (e.g., NAL unit type values ​​7 to 9 in Table 1 or Table 2 above), and all other VCL NAL units must have a NAL unit type that is the same as GRA_NUT or a NAL unit type with a specific value in the range from TRAIL_NUT to RSV_VCL_6 (e.g., NAL unit type values ​​0 to 6 in Table 1 or Table 2 above). This constraint applies only to current pictures that include a mix of IRAP and non-IRAP NAL unit types. However, it does not yet correctly apply to pictures that include a mix of RASL / RADL and non-IRAP NAL unit types.

[0209] This document provides a solution to the above-mentioned problem. That is, as described above, a picture (i.e., a current picture) including two or more sub-pictures can have a mixed NAL unit type. In the case of the current VVC standard, a picture with a mixed NAL unit type can have a mixed form of an IRAP NAL unit type and a non-IRAP NAL unit type. However, a leading picture associated with a CRA NAL unit type can also have a mixed form with a non-IRAP NAL unit type, and pictures with such a mixed NAL unit type are not supported under the current standard. Therefore, a solution is needed for pictures with CRA NAL unit types and non-IRAP NAL unit types in a mixed form.

[0210] Therefore, this document provides a method for allowing pictures to include a mixed form of leading picture NAL unit types (e.g., RASL_NUT, RADL_NUT) and other non-IRAP NAL unit types (e.g., TRAIL_NUT, STSA, NUT). In addition, this document defines constraints that allow the presence or signaling of reference picture lists when IDR sub-pictures and other non-IRAP sub-pictures are mixed. Therefore, pictures with mixed NAL unit types are provided with a form that not only mixes IRAP but also CRA NAL units, providing more flexible characteristics.

[0211] For example, it can be applied in the following embodiments, thereby solving the above-mentioned problems. The following embodiments can be applied independently or in combination.

[0212] In one embodiment, when pictures are allowed to have mixed NAL unit types (when the value of mixed_nal_types_in_pic_flag is 1), the signaling of the reference picture list allows it to exist even for slices with IDR-type NAL unit types (e.g., IDR_W_RADL or IDR_N_LP). This constraint can be expressed as follows.

[0213] - When there is at least one PPS that references an SPS whose mixed_nal_types_in_pic_flag has the value 1, the value of sps_idr_rpl_present_flag must become 1. This constraint may be a requirement for bitstream conformance.

[0214] Alternatively, in one embodiment, for a picture with mixed NAL unit types, the picture is allowed to include slices with a specific NAL unit type (e.g., RADL or RASL) for the leading picture and a specific NAL unit type (non-IRAP) for the non-leading picture. This can be expressed as follows.

[0215] For the VCL NAL units of a particular picture, the following may apply.

[0216] - When the value of mixed_nalu_types_in_pic_flag is 0, the value of the NAL unit type (nal_unit_type) must be the same for all slice NAL units coded in a picture. A picture or PU can be considered to have the same NAL unit type as the slice NAL units coded in the picture or PU.

[0217] - Otherwise (when the value of mixed_nalu_types_in_pic_flag is 1), one of the following must be satisfied (ie, one of the following may have a value of true).

[0218] 1) One or more VCL NAL units must have a NAL unit type (nal_unit_type) of a specific value in the range from IDR_W_RADL to CRA_NUT (e.g., the values ​​of NAL unit type in Table 1 or Table 2 above are 7 to 9), and all other VCL NAL units must have the same NAL unit type as GRA_NUT or a NAL unit type of a specific value in the range from TRAIL_NUT to RSV_VCL_6 (e.g., the values ​​of NAL unit type in Table 1 or Table 2 above are 0 to 6).

[0219] 2) One or more VCL NAL units must all have the same specific value of NAL unit type as RADL_NUT (e.g., the value of NAL unit type is 2 in Table 1 or Table 2 above) or RASL_NUT (e.g., the value of NAL unit type is 3 in Table 1 or Table 2 above), and all other VCL NAL units must have the same specific value of NAL unit type as TRAIL_NUT (e.g., the value of NAL unit type is 0 in Table 1 or Table 2 above), STSA_NUT (e.g., the value of NAL unit type is 1 in Table 1 or Table 2 above), RSV_VCL_4 (e.g., the value of NAL unit type is 4 in Table 1 or Table 2 above), RSV_VCL_5 (e.g., the value of NAL unit type is 5 in Table 1 or Table 2 above), RSV_VCL_6 (e.g., the value of NAL unit type is 6 in Table 1 or Table 2 above), or GRA_NUT.

[0220] Furthermore, this document proposes a method for providing pictures having the above-mentioned mixed NAL unit types even for a single-layer bitstream. As an embodiment, in the case of a single-layer bitstream, the following constraints may be applied.

[0221] - In the bitstream, every picture except the first picture in decoding order is considered to be associated with the previous IRAP picture in decoding order.

[0222] - If the picture is the leading picture of an IRAP picture, it must be a RADL or RASL picture.

[0223] - If the picture is the trailing picture of an IRAP picture, it must be neither a RADL nor a RASL picture.

[0224] - RASL pictures shall not be present in the bitstream associated with an IDR picture.

[0225] - A RADL picture shall not be present in the bitstream associated with an IDR picture whose NAL unit type (nal_unit_type) is IDR_N_LP.

[0226] When referenced, and when each parameter set is available, random access can be performed at the location of the IRAP PU by discarding all PUs before the IRAP PU (and the IRAP picture and all subsequent non-RASL pictures can be correctly decoded in decoding order).

[0227] - Pictures that precede an IRAP picture in decoding order must precede the IRAP picture in output order, and must precede the RADL picture associated with the IRAP picture in output order.

[0228] - RASL pictures associated with a CRA picture must precede RADL pictures associated with the CRA picture in output order.

[0229] - RASL pictures associated with a CRA picture must follow, in output order, the IRAP pictures that precede the CRA picture in decoding order.

[0230] - If the value of field_seq_flag is 0 and the current picture is a leading picture associated with an IRAP picture, then it must precede all non-leading pictures associated with the same IRAP picture in decoding order. Otherwise, when pictures picA and picB are the first and last leading pictures associated with an IRAP picture in decoding order, respectively, there can be at most one non-leading picture before picA in decoding order, and there must be no non-leading pictures between picA and picB in decoding order.

[0231] The following figures are prepared to illustrate specific examples of this document. Since specific terms or names or names of specific devices (e.g., names of grammar / grammar elements, etc.) described in the figures are presented as examples, the technical features of this document are not limited to the specific names used in the following figures.

[0232] Figure 15 Schematically shows an example of a video / image encoding method to which the embodiments of this document are applicable. Figure 2 The encoding device 200 disclosed in Figure 15 The method disclosed in .

[0233] Reference Figure 15 , the encoding device can determine the NAL unit type of the slice in the picture (S1500).

[0234] For example, the encoding device can determine the NAL unit type according to the properties, types, etc. of the picture or sub-picture as described in Tables 1 and 2 above, and based on the NAL unit type of the picture or sub-picture, the NAL unit type of each slice can be determined.

[0235] For example, when the value of mixed_nalu_types_in_pic_flag is 0, the slices in the picture associated with the PPS can be determined to be of the same NAL unit type. That is, when the value of mixed_nalu_types_in_pic_flag is 0, the NAL unit type defined in the first NAL unit header of the first NAL unit including information about the first slice of the picture is the same as the NAL unit type defined in the second NAL unit header of the second NAL unit including information about the second slice of the same picture. Alternatively, for the case where the value of mixed_nalu_types_in_pic_flag is 1, the slices in the picture associated with the PPS can be determined to be of different NAL unit types. Here, the NAL unit type of the slice in the picture can be determined based on the method proposed in the above embodiment.

[0236] The encoding device may generate NAL unit type related information (S1510). The NAL unit type related information may include information / syntax elements related to the NAL unit types described in the above embodiments and / or Tables 1 and 2 above. For example, the information related to the NAL unit type may include the mixed_nalu_types_in_pic_flag syntax element included in the PPS. In addition, the information related to the NAL unit type may include the nal_unit_type syntax element in the NAL unit header of the NAL unit including information about the coded slice.

[0237] The encoding device may generate a bitstream (S1520). The bitstream may include at least one NAL unit including image information about the coded slice. In addition, the bitstream may include a PPS.

[0238] Figure 16Schematically shows an example of a video / image decoding method applicable to the embodiment of this document. Figure 3 The decoding device 300 disclosed in Figure 16 The method disclosed in .

[0239] Reference Figure 16 , the decoding device may receive a bitstream (S1600). Here, the bitstream may include at least one NAL unit including image information about a coded slice. In addition, the bitstream may include a PPS.

[0240] The decoding device may obtain NAL unit type related information (S1610). The NAL unit type related information may include information / syntax elements related to the NAL unit types described in the above embodiments and / or Tables 1 and 2 above. For example, the information related to the NAL unit type may include the mixed_nalu_types_in_pic_flag syntax element included in the PPS. In addition, the information related to the NAL unit type may include the nal_unit_type syntax element in the NAL unit header of the NAL unit including information about the coded slice.

[0241] The decoding apparatus may determine a NAL unit type of a slice in a picture ( S1620 ).

[0242] For example, when the value of mixed_nalu_types_in_pic_flag is 0, the slices in the picture associated with the PPS use the same NAL unit type. That is, when the value of mixed_nalu_types_in_pic_flag is 0, the NAL unit type defined in the first NAL unit header of the first NAL unit including information about the first slice of the picture is the same as the NAL unit type defined in the second NAL unit header of the second NAL unit including information about the second slice of the same picture. Alternatively, when the value of mixed_nalu_types_in_pic_flag is 1, the slices in the picture associated with the PPS use different NAL unit types. Here, the NAL unit type of the slice in the picture can be determined based on the method proposed in the above embodiment.

[0243] The decoding device may decode / reconstruct the sample / block / slice based on the NAL unit type of the slice (S1630). The sample / block in the slice may be decoded / reconstructed based on the NAL unit type of the slice.

[0244] For example, when a first NAL unit type is set for a first slice of a current picture and a second NAL unit type (different from the first NAL unit type) is set for a second slice of the current picture, samples / blocks in the first slice or the first slice itself can be decoded / reconstructed based on the first NAL unit type, and samples / blocks in the second slice or the second slice itself can be decoded / reconstructed based on the second NAL unit type.

[0245] Figure 17 and Figure 18 An example of a video / image encoding method and associated components according to an embodiment of this document is schematically shown.

[0246] Figure 17 The method disclosed in Figure 2 or Figure 18 Here, Figure 18 The encoding device 200 disclosed in Figure 2 Specifically, Figure 17 Steps S1700 to S1720 can be performed by Figure 2 The entropy encoder 240 disclosed in the embodiment is executed. In addition, according to the embodiment, each step can be performed by Figure 2 The image segmenter 210, the predictor 220, the residual processor 230, the adder 340, etc. disclosed in the embodiment of the present invention are executed. In addition, the embodiments described above in this document may be executed. Figure 17 Therefore, in Figure 17 In the embodiment, detailed descriptions of contents corresponding to the repetitions of the above-mentioned embodiments will be omitted or simplified.

[0247] Reference Figure 17 , the encoding device can determine the NAL unit type of the NAL unit in the current picture (S1700).

[0248] The current picture may include multiple slices, and one slice may include a slice header and slice data. In addition, a NAL unit may be generated by adding a NAL unit header to a slice (slice header and slice data). The NAL unit header may include NAL unit type information specified according to the slice data included in the corresponding NAL unit.

[0249] As an embodiment, the encoding device may generate a first NAL unit for a first slice in the current picture and a second NAL unit for a second slice in the current picture. In addition, the encoding device may determine the first NAL unit type for the first slice and the second NAL unit type for the second slice based on the types of the first slice and the second slice.

[0250] For example, based on the type of slice data included in the NAL unit as shown in Table 1 or Table 2 above, the NAL unit type may include TRAIL_NUT, STSA_NUT, RADL_NUT, RASL_NUT, IDR_W_RADL, IDR_N_LP, CRA_NUT, etc. In addition, the NAL unit type may be signaled based on the nal_unit_type syntax element in the NAL unit header. The nal_unit_type syntax element is syntax information for specifying the NAL unit type and, as shown in Table 1 or Table 2 above, may be represented as a specific value corresponding to a specific NAL unit type.

[0251] The encoding apparatus may generate NAL unit type related information of a NAL unit type ( S1710 ).

[0252] NAL unit type related information may include information / syntax elements related to the NAL unit types described in the above embodiments and / or Tables 1 and 2 above. For example, NAL unit type related information may be information about whether the current picture has mixed NAL unit types, and may be represented by the mixed_nalu_types_in_pic_flag syntax element included in the PPS. For example, when the value of the mixed_nalu_types_in_pic_flag syntax element is 0, it may indicate that the NAL units in the current picture have the same NAL unit type. Alternatively, when the value of the mixed_nalu_types_in_pic_flag syntax element is 1, it may indicate that the NAL units in the current picture have different NAL unit types.

[0253] As an embodiment, when all NAL unit types of NAL units in the current picture are the same, the encoding device may generate NAL unit type related information (e.g., mixed_nalu_types_in_pic_flag) having a value of 0. Alternatively, when the NAL unit types of NAL units in the current picture are different, the encoding device may generate NAL unit type related information (e.g., mixed_nalu_types_in_pic_flag) having a value of 1.

[0254] That is, based on the NAL unit type related information about the current picture with a mixed NAL unit type (for example, the value of mixed_nalu_types_in_pic_flag is 1), the first NAL unit of the first slice of the current picture and the second NAL unit of the second slice of the current picture may have different NAL unit types. Alternatively, based on the NAL unit type related information about the current picture without a mixed NAL unit type (for example, the value of mixed_nalu_types_in_pic_flag is 0), the first NAL unit of the first slice of the current picture and the second NAL unit of the second slice of the current picture may have the same NAL unit type.

[0255] As an example, based on the NAL unit type related information about the current picture having a mixed NAL unit type (for example, the value of mixed_nalu_types_in_pic_flag is 1), the first NAL unit of the first slice may have a leading picture NAL unit type, and the second NAL unit of the second slice may have a non-IRAP NAL unit type or a non-leading picture NAL unit type. Here, the leading picture NAL unit type may include a RADL NAL unit type or a RASL NAL unit type, and the non-IRAP NAL unit type or the non-leading picture NAL unit type may include a trailing NAL unit type or an STSA NAL unit type.

[0256] Alternatively, as an example, based on NAL unit type-related information about the current picture having a mixed NAL unit type (e.g., the value of mixed_nalu_types_in_pic_flag is 1), the first NAL unit of the first slice may have an IRAP NAL unit type, and the second NAL unit of the second slice may have a non-IRAP NAL unit type or a non-leading picture NAL unit type. Here, the IRAP NAL unit type may include an IDR NAL unit type (i.e., IDR_N_LPNAL or IDR_W_RADL NAL unit type) or a CRA NAL unit type, and the non-IRAP NAL unit type or the non-leading picture NAL unit type may include a tracking NAL unit type or an STSA NAL unit type. In addition, according to an embodiment, the non-IRAP NAL unit type or the non-leading picture NAL unit type may refer to only the tracking NAL unit type.

[0257] According to an embodiment, based on the case where the current picture is allowed to have mixed NAL unit types, for a slice having an IDR NAL unit type (e.g., IDR_W_RADL or IDR_N_LP) in the current picture, information related to signaling a reference picture list must be present. The information related to signaling a reference picture list may indicate whether a syntax element for signaling a reference picture list is present in the slice header of the slice. That is, if the value of the information related to signaling a reference picture list is 1, the syntax element for signaling a reference picture list may be present in the slice header of the slice having the IDR NAL unit type. Alternatively, if the value of the information related to signaling a reference picture list is 0, the syntax element for signaling a reference picture list may not be present in the slice header of the slice having the IDR NAL unit type.

[0258] For example, the information related to signaling a reference picture list may be the sps_idr_rpl_present_flag syntax element described above. When the value of sps_idr_rpl_present_flag is 1, it may indicate that the syntax element for signaling a reference picture list may be present in the slice header of a slice having a NAL unit type such as IDR_N_LP or IDR_W_RADL. Alternatively, when the value of sps_idr_rpl_present_flag is 0, it may indicate that the syntax element for signaling a reference picture list may not be present in the slice header of a slice having a NAL unit type such as IDR_N_LP or IDR_W_RADL.

[0259] The encoding apparatus may encode image / video information including NAL unit type related information ( S1720 ).

[0260] For example, when the first NAL unit of the first slice in the current picture and the second NAL unit of the second slice in the current picture have different NAL unit types, the encoding device may encode the image / video information including NAL unit type related information (e.g., mixed_nalu_types_in_pic_flag) having a value of 1. Alternatively, when the first NAL unit of the first slice in the current picture and the second NAL unit of the second slice in the current picture have the same NAL unit type, the encoding device may encode the image / video information including NAL unit type related information (e.g., mixed_nalu_types_in_pic_flag) having a value of 0.

[0261] In addition, for example, the encoding device may encode image / video information including nal_unit_type information indicating each NAL unit type of a slice in the current picture.

[0262] In addition, for example, the encoding apparatus may encode image / video information including information related to signaling of a reference picture list (eg, sps_idr_rpl_present_flag).

[0263] In addition, for example, the encoding device may encode image / video information of a NAL unit including a slice in the current picture.

[0264] The image / video information including the various information as described above can be encoded and output in the form of a bit stream. The bit stream can be sent to a decoding device via a network or a (digital) storage medium. Here, the network can include a broadcast network, a communication network and / or the like, and the digital storage medium can include various storage media such as a universal serial bus (USB), secure digital (SD), compact disc (CD), digital video disc (DVD), Blu-ray, hard disk drive (HDD), solid state drive (SSD), etc.

[0265] Figure 19 and Figure 20 An example of a video / image decoding method and associated components according to an embodiment of this document is schematically shown.

[0266] Figure 19 The method disclosed in Figure 3 or Figure 20 The decoding device 300 disclosed in the embodiment is executed. Here, Figure 20 The decoding device 300 disclosed in Figure 3 Specifically, Figure 19 Steps S1900 to S1920 can be performed by Figure 3 The entropy decoder 310 disclosed in the embodiment is executed. In addition, according to the embodiment, each step can be performed by Figure 3 The residual processor 320, the predictor 330, the adder 340, etc. disclosed in the embodiment of the present invention are executed. In addition, the embodiments including those described above in this document may be executed. Figure 19 Therefore, in Figure 19 In the embodiment, detailed descriptions of contents corresponding to the repetitions of the above-mentioned embodiments will be omitted or simplified.

[0267] Reference Figure 19 , the decoding device can obtain image / video information including NAL unit type related information from the bitstream (S1900).

[0268] For example, the decoding device can parse the bitstream and derive information required for image reconstruction (or picture reconstruction) (e.g., video / image information). In this case, the image information may include the above-mentioned NAL unit type related information (e.g., mixed_nalu_types_in_pic_flag), nal_unit_type information indicating each NAL unit type of the slice in the current picture, information related to signaling the reference picture list (e.g., sps_idr_rpl_present_flag), NAL units of slices in the current picture, etc. That is, the image information may include various information required in the decoding process and may be decoded based on a coding method such as exponential Golomb coding, CAVLC, or CABAC.

[0269] As described above, the NAL unit type related information may include information / syntax elements related to the NAL unit types described in the above embodiments and / or Tables 1 and 2 above. For example, the NAL unit type related information may be information about whether the current picture has mixed NAL unit types, and may be represented by the mixed_nalu_types_in_pic_flag syntax element included in the PPS. For example, when the value of the mixed_nalu_types_in_pic_flag syntax element is 0, it may indicate that the NAL units in the current picture have the same NAL unit type. Alternatively, when the value of the mixed_nalu_types_in_pic_flag syntax element is 1, it may indicate that the NAL units in the current picture have different NAL unit types.

[0270] The decoding apparatus may determine a NAL unit type of a NAL unit in a current picture based on NAL unit type related information ( S1910 ).

[0271] The current picture may include multiple slices, and one slice may include a slice header and slice data. In addition, a NAL unit may be generated by adding a NAL unit header to a slice (slice header and slice data). The NAL unit header may include NAL unit type information specified according to the slice data included in the corresponding NAL unit.

[0272] For example, based on the type of slice data included in the NAL unit as shown in Table 1 or Table 2 above, the NAL unit type may include TRAIL_NUT, STSA_NUT, RADL_NUT, RASL_NUT, IDR_W_RADL, IDR_N_LP, CRA_NUT, etc. In addition, the NAL unit type may be signaled based on the nal_unit_type syntax element in the NAL unit header. The nal_unit_type syntax element is syntax information for specifying the NAL unit type and, as shown in Table 1 or Table 2 above, may be represented as a specific value corresponding to a specific NAL unit type.

[0273] In an embodiment, the decoding device may determine that the first NAL unit of the first slice of the current picture and the second NAL unit of the second slice of the current picture may have different NAL unit types based on the NAL unit type related information about the current picture with a mixed NAL unit type (for example, the value of mixed_nalu_types_in_pic_flag is 1). Alternatively, the decoding device may determine that the first NAL unit of the first slice of the current picture and the second NAL unit of the second slice of the current picture may have the same NAL unit type based on the NAL unit type related information about the current picture without a mixed NAL unit type (for example, the value of mixed_nalu_types_in_pic_flag is 0).

[0274] As an example, based on the NAL unit type related information about the current picture having a mixed NAL unit type (for example, the value of mixed_nalu_types_in_pic_flag is 1), the first NAL unit of the first slice may have a leading picture NAL unit type, and the second NAL unit of the second slice may have a non-IRAP NAL unit type or a non-leading picture NAL unit type. Here, the leading picture NAL unit type may include a RADL NAL unit type or a RASL NAL unit type, and the non-IRAP NAL unit type or the non-leading picture NAL unit type may include a tracking NAL unit type or an STSA NAL unit type.

[0275] Alternatively, as an example, based on NAL unit type-related information about the current picture having a mixed NAL unit type (e.g., the value of mixed_nalu_types_in_pic_flag is 1), the first NAL unit of the first slice may have an IRAP NAL unit type, and the second NAL unit of the second slice may have a non-IRAP NAL unit type or a non-leading picture NAL unit type. Here, the IRAP NAL unit type may include an IDR NAL unit type (i.e., IDR_N_LPNAL or IDR_W_RADL NAL unit type) or a CRA NAL unit type, and the non-IRAP NAL unit type or the non-leading picture NAL unit type may include a tracking NAL unit type or an STSA NAL unit type. In addition, according to an embodiment, the non-IRAP NAL unit type or the non-leading picture NAL unit type may refer to only the tracking NAL unit type.

[0276] According to an embodiment, based on the case where the current picture is allowed to have mixed NAL unit types, for a slice having an IDR NAL unit type (e.g., IDR_W_RADL or IDR_N_LP) in the current picture, information related to signaling a reference picture list must be present. The information related to signaling a reference picture list may indicate whether a syntax element for signaling a reference picture list is present in the slice header of the slice. That is, if the value of the information related to signaling a reference picture list is 1, the syntax element for signaling a reference picture list may be present in the slice header of the slice having the IDR NAL unit type. Alternatively, if the value of the information related to signaling a reference picture list is 0, the syntax element for signaling a reference picture list may not be present in the slice header of the slice having the IDR NAL unit type.

[0277] For example, the information related to signaling a reference picture list may be the sps_idr_rpl_present_flag syntax element described above. When the value of sps_idr_rpl_present_flag is 1, it may indicate that the syntax element for signaling a reference picture list may be present in the slice header of a slice having a NAL unit type such as IDR_N_LP or IDR_W_RADL. Alternatively, when the value of sps_idr_rpl_present_flag is 0, it may indicate that the syntax element for signaling a reference picture list may not be present in the slice header of a slice having a NAL unit type such as IDR_N_LP or IDR_W_RADL.

[0278] The decoding device may decode / reconstruct the current picture based on the NAL unit type ( S1920 ).

[0279] For example, for a first slice in a current picture determined to be of a first NAL unit type and a second slice in the current picture determined to be of a second NAL unit type, the decoding device may decode / restore the first slice based on the first NAL unit type and decode / reconstruct the second slice based on the second NAL unit type. In addition, the decoding device may decode / reconstruct the sample / block in the first slice based on the first NAL unit type and decode / reconstruct the sample / block in the second slice based on the second NAL unit type.

[0280] Although the method has been described based on a flowchart that lists steps or blocks in sequence in the above embodiments, the steps of this document are not limited to a specific order, and specific steps can be performed in different steps or in a different order or simultaneously with respect to the above steps. In addition, it will be understood by those skilled in the art that the steps in the flowchart are not exclusive, and another step can be included therein, or one or more steps in the flowchart can be deleted without affecting the scope of the present disclosure.

[0281] The above-mentioned method according to the present disclosure may be in the form of software, and the encoding device and / or decoding device according to the present disclosure may be included in an apparatus for performing image processing (e.g., TV, computer, smart phone, set-top box, display device, etc.).

[0282] When the embodiments of the present disclosure are implemented with software, the above-mentioned methods can be implemented with modules (processing or functions) that perform the above-mentioned functions. The modules can be stored in a memory and executed by a processor. The memory can be installed inside or outside the processor and can be connected to the processor via various well-known devices. The processor may include an application-specific integrated circuit (ASIC), other chipsets, logic circuits and / or data processing devices. The memory may include a read-only memory (ROM), a random access memory (RAM), a flash memory, a memory card, a storage medium and / or other storage devices. In other words, according to the embodiments of the present disclosure, it can be implemented and executed on a processor, a microprocessor, a controller or a chip. For example, the functional units illustrated in the corresponding figures can be implemented and executed on a computer, a processor, a microprocessor, a controller or a chip. In this case, information about the implementation (for example, information about instructions) or the algorithm can be stored in a digital storage medium.

[0283] In addition, the decoding device and encoding device of the embodiment of the present document can be included in a multimedia broadcast transceiver, a mobile communication terminal, a home theater video device, a digital theater video device, a surveillance camera, a video chat device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a portable camera, a video on demand (VoD) service provider, an over-the-top (OTT) video device, an Internet streaming service provider, a 3D video device, a virtual reality (VR) device, an augmented reality (AR) device, an image phone video device, a vehicle terminal (e.g., a vehicle (including an autonomous vehicle) terminal, an aircraft terminal, or a ship terminal) and a medical video device; and can be used to process image signals or data. For example, the OTT video device may include a game console, a Blueray player, a networked TV, a home theater system, a smartphone, a tablet PC, and a digital video recorder (DVR).

[0284] In addition, the processing method of the embodiment of the application of this document can be generated in the form of a program executed by a computer and can be stored in a computer-readable recording medium. The multimedia data with a data structure according to the embodiment of this document can also be stored in a computer-readable recording medium. The computer-readable recording medium includes all kinds of storage devices and distributed storage devices stored with computer-readable data. The computer-readable recording medium may include, for example, a Blu-ray disc (BD), a universal serial bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. The computer-readable recording medium also includes a medium implemented in the form of a carrier wave (for example, transmission on the Internet). In addition, the bit stream generated by the encoding method can be stored in a computer-readable recording medium, or can be transmitted through a wired or wireless communication network.

[0285] In addition, the embodiments of this document can be implemented as a computer program product based on a program code, and the program code can be executed on a computer according to the embodiments of this document. The program code can be stored on a computer-readable carrier.

[0286] Figure 21 This shows an example of a content streaming system to which the embodiments of this document can be applied.

[0287] Reference Figure 21 A content streaming system to which embodiments of this document are applied may generally include an encoding server, a streaming server, a network server, a media storage, a user device, and a multimedia input device.

[0288] The encoding server is used to compress content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data, generate a bitstream, and transmit it to the streaming server. As another example, if the multimedia input device such as smartphones, cameras, and camcorders directly generates the bitstream, the encoding server can be omitted.

[0289] The bitstream may be generated by the encoding method or the bitstream generation method to which the embodiments of this document are applied, and the streaming server may temporarily store the bitstream in the process of transmitting or receiving the bitstream.

[0290] The streaming server transmits multimedia data to the user device via a network server based on the user's request. The network server serves as a tool for notifying the user of available services. When the user requests a desired service, the network server transfers the request to the streaming server, which then transmits the multimedia data to the user. In this regard, the content streaming system may include a separate control server, in which case the control server is used to control commands and responses between the various devices in the content streaming system.

[0291] The streaming server may receive content from a media storage device and / or an encoding server. For example, when receiving content from an encoding server, the content may be received in real time. In this case, the streaming server may store the bitstream for a predetermined period of time to smoothly provide a streaming service.

[0292] For example, user devices may include mobile phones, smart phones, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation, tablet PCs, tablet PCs, ultrabooks, wearable devices (e.g., watch-type terminals (smart watches), glasses-type terminals (smart glasses), head-mounted displays (HMDs)), digital TVs, desktop computers, digital signage, etc.

[0293] Each server in the content streaming system may be operated as a distributed server, and in such case, data received by each server may be processed in a distributed manner.

[0294] The claims in this specification may be combined in various ways. For example, the technical features in the method claims of this specification may be combined to be implemented or performed in a device, and the technical features in the device claims may be combined to be implemented or performed in a method. Furthermore, the technical features in method and device claims may be combined to be implemented or performed in a device. Furthermore, the technical features in method and device claims may be combined to be implemented or performed in a method.

Claims

1. An image decoding method performed by a decoding device, the image decoding method comprising the following steps: Obtain image information including network abstraction layer NAL unit type related information from the bitstream; Determine a NAL unit type for a NAL unit in a current picture based on the NAL unit type related information; as well as Decoding the current picture based on the NAL unit type, The NAL unit type related information is information about whether the current picture has a mixed NAL unit type. wherein, based on the NAL unit type related information for the current picture having the mixed NAL unit type, a first NAL unit type for a first slice of the current picture is different from a second NAL unit type for a second slice of the current picture, wherein, based on the first NAL unit type being an instantaneous decoding refresh (IDR) NAL unit type, a value of information related to signaling a reference picture list is equal to 1, and Wherein, based on the value of the information related to signaling the reference picture list being equal to 1, a syntax element for signaling the reference picture list is present in a slice header of the first slice having the IDR NAL unit type.

2. An image encoding method performed by an encoding device, the image encoding method comprising the following steps: Determine the NAL unit type for the NAL unit in the current picture; generating NAL unit type related information based on the NAL unit type; as well as Encode the image information including the NAL unit type related information, The NAL unit type related information is information about whether the current picture has a mixed NAL unit type. wherein, based on the NAL unit type related information for the current picture having the mixed NAL unit type, a first NAL unit type for a first slice of the current picture is different from a second NAL unit type for a second slice of the current picture, and Wherein, based on the first NAL unit type being an instantaneous decoding refresh (IDR) NAL unit type, a syntax element for signaling a reference picture list is present in a slice header of the first slice having the IDR NAL unit type.

3. A method for transmitting image data, the method comprising the following steps: Obtaining a bitstream for the image, wherein the bitstream is generated based on the following operations: determining a NAL unit type for a NAL unit in a current picture, generating NAL unit type related information based on the NAL unit type, and encoding image information including the NAL unit type related information; and sending said data comprising said bitstream, The NAL unit type related information is information about whether the current picture has a mixed NAL unit type. wherein, based on the NAL unit type related information for the current picture having the mixed NAL unit type, a first NAL unit type for a first slice of the current picture is different from a second NAL unit type for a second slice of the current picture, and Wherein, based on the first NAL unit type being an instantaneous decoding refresh (IDR) NAL unit type, a syntax element for signaling a reference picture list is present in a slice header of the first slice having the IDR NAL unit type.