Nal unit type based picture or video coding for slices or pictures
By encoding images/videos based on NAL unit type information, the problem of efficient compression of high-resolution, high-quality image/video data is solved, encoding efficiency is improved, and transmission and storage costs are reduced.
Patent Information
- Application Number
- CN202080096777.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-23
- Filing Date
- 2020-12-10
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2040-12-10
AI Technical Summary
Existing technologies have difficulty in effectively compressing and encoding high-resolution, high-quality image/video data, especially pictures with mixed NAL unit types, resulting in increased transmission and storage costs.
By determining whether a picture has a mixed NAL unit type based on NAL unit type related information, and signaling and encoding the slice type for slices of the picture having the mixed NAL unit type, efficient encoding of reference picture list related information is allowed.
Improves image/video coding efficiency, especially for pictures with mixed NAL unit types, enables flexible encoding and decoding, and reduces transmission and storage costs.
Smart Images

Figure CN115104316B_ABST
Abstract
Description
Technical Field
[0001] The present technology relates to video or image coding, for example, to an image or video coding technology based on a Network Abstraction Layer (NAL) unit type for a slice or a picture. Background Art
[0002] Recently, there has been an increasing demand for high-resolution, high-quality images / videos, such as 4K or 8K ultra-high-definition (UHD) images / videos, in various fields. As image / video resolution or quality becomes higher, relatively more information or bits are transmitted compared to conventional image / video data. Therefore, if image / video data is transmitted via a medium such as an existing wired / wireless broadband line or stored in a conventional storage medium, the cost of transmission and storage is likely to increase.
[0003] In addition, there is growing interest and demand for virtual reality (VR) and artificial reality (AR) content and immersive media such as holograms; and there is also growing broadcasting of images / videos that exhibit image / video characteristics that are different from actual images / videos (e.g., game images / videos).
[0004] Therefore, highly efficient image / video compression technology is required to effectively compress and transmit, store, or play high-resolution, high-quality images / videos exhibiting various characteristics as described above.
[0005] In addition, a method for improving image / video encoding efficiency is required, and to this end, a method for efficiently signaling and encoding information related to a Network Abstraction Layer (NAL) unit is necessary. Summary of the Invention
[0006] Technical issues
[0007] This document provides methods and devices for improving video / image coding efficiency.
[0008] This document also provides a method and apparatus for improving video / image coding efficiency based on NAL unit related information.
[0009] This document also provides methods and apparatus for improving video / image coding efficiency for pictures with mixed (hybrid) NAL unit types.
[0010] This document also provides methods and apparatus for allowing the signaling or presence of reference picture list related information for slices in a picture having a specific NAL unit type relative to a picture having a mixed NAL unit type.
[0011] Technical Solution
[0012] According to an embodiment of the present document, whether a picture has a mixed NAL unit type can be determined based on NAL unit type related information, and the NAL unit type can be determined for a slice of a picture with a mixed NAL unit type. For example, based on a value of 1 for the NAL unit type related information, the NAL unit type of the first slice in the picture may have a leading picture NAL unit type, and the NAL unit type of the second slice in the picture may have a non-intra random access point (IRAP) NAL unit type or a non-leading picture NAL unit type. According to an embodiment of the present document, based on the case where a picture is allowed to have a mixed NAL unit type, for a slice with a specific NAL unit type in the picture, information related to signaling a reference picture list may exist.
[0013] According to an embodiment of this document, based on the case where a picture is allowed to have mixed NAL unit types, information related to signaling a reference picture list may exist for slices with a specific NAL unit type in the picture.
[0014] According to an embodiment of the present document, a video / image decoding method performed by a decoding device is provided. The video / image decoding method may include the method disclosed in the embodiment of the present document.
[0015] According to an embodiment of the present document, a decoding device for performing video / image decoding is provided. The decoding device can execute the method disclosed in the embodiment of the present document.
[0016] According to an embodiment of the present document, a video / image encoding method performed by an encoding device is provided. The video / image encoding method may include the method disclosed in the embodiment of the present document.
[0017] According to an embodiment of the present document, an encoding device for performing video / image encoding is provided. The encoding device can execute the method disclosed in the embodiment of the present document.
[0018] According to an embodiment of this document, a computer-readable digital storage medium is provided that stores encoded video / image information generated according to the video / image encoding method disclosed in at least one of the embodiments of this document.
[0019] According to an embodiment of this document, there is provided a computer-readable digital storage medium storing encoded information or encoded video / image information that causes a decoding device to perform the video / image decoding method disclosed in at least one of the embodiments of this document.
[0020] Technical Effects
[0021] This document can have various effects. For example, according to the embodiments of this document, the overall image / video compression efficiency can be improved. In addition, according to the embodiments of this document, the video / image coding efficiency can be improved based on the NAL unit related information. In addition, according to the embodiments of this document, the video / image coding efficiency of pictures with mixed NAL unit types can be improved. In addition, according to the embodiments of this document, for pictures with mixed NAL unit types, reference picture list related information can be efficiently signaled and encoded. In addition, according to the embodiments of this document, by allowing pictures to include leading picture NAL unit types (e.g., RASL_NUT, RADL_NUT) and other non-IRAP NAL unit types (e.g., TRAIL_NUT, STSA, NUT) in a mixed form, for pictures with mixed NAL unit types, a form mixed not only with IRAP but also with other types of NAL units can be provided, and accordingly, more flexible characteristics can be provided.
[0022] The effects that can be obtained through the detailed examples of this document are not limited to the effects listed above. For example, there may be various technical effects that a person of ordinary skill in the relevant field can understand or deduce from this document. Therefore, the detailed effects of this document are not limited to the effects explicitly described in this document, but may include various effects that can be understood or derived from the technical features of this document. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 An example of a video / image encoding system to which the embodiments of this document are applicable is schematically illustrated.
[0024] Figure 2 is a diagram schematically illustrating a configuration of a video / image encoding device to which an embodiment of this document can be applied.
[0025] Figure 3 is a schematic diagram illustrating a configuration of a video / image decoding device to which the embodiments of this document can be applied.
[0026] Figure 4 An example of an exemplary video / image encoding method to which the embodiments of this document are applicable is shown.
[0027] Figure 5 An example of an exemplary video / image decoding method to which the embodiments of this document are applied is shown.
[0028] Figure 6 The hierarchical structure of the coded image / video is shown as an example.
[0029] Figure 7 Schematically shows an example of an entropy coding method applicable to the embodiments of this document, Figure 8An entropy encoder in an encoding device is schematically shown.
[0030] Figure 9 An example of an entropy decoding method applicable to the embodiments of this document is schematically shown. Figure 10 An entropy decoder in a decoding device is schematically shown.
[0031] Figure 11 is a diagram showing a temporal layer structure of a NAL unit in a bitstream supporting temporal scalability.
[0032] Figure 12 A diagram used to describe pictures that can be randomly accessed.
[0033] Figure 13 This is a diagram used to describe an IDR picture.
[0034] Figure 14 A diagram used to describe a CRA picture.
[0035] Figure 15 An example of a video / image encoding method to which the embodiments of this document are applied is schematically shown.
[0036] Figure 16 An example of a video / image decoding method to which the embodiments of this document are applied is schematically shown.
[0037] Figure 17 and Figure 18 An example of a video / image encoding method and related components according to an embodiment of this document is schematically illustrated.
[0038] Figure 19 and Figure 20 An example of a video / image decoding method and related components according to an embodiment of this document is schematically illustrated.
[0039] Figure 21 An example of a content streaming system to which the embodiments disclosed in this document are applicable is illustrated. DETAILED DESCRIPTION
[0040] The present disclosure can be modified in various forms, and its specific embodiments will be described and illustrated in the accompanying drawings. However, these embodiments are not intended to limit the present disclosure. The terms used in the following description are only used to describe specific embodiments and are not intended to limit the present disclosure. Singular expressions include plural expressions as long as they are clearly interpreted differently. Terms such as "including" and "having" are intended to indicate the presence of features, quantities, steps, operations, elements, parts, or combinations thereof used in the following description, so it should be understood that the possibility of the presence or addition of one or more different features, quantities, steps, operations, elements, parts, or combinations thereof is not excluded.
[0041] In addition, the various configurations of the drawings described in this document are independent illustrations for the purpose of illustrating the functions of different features, and do not imply that the various configurations are implemented by different hardware or different software. For example, two or more of the configurations may be combined to form a single configuration, and a single configuration may also be divided into multiple configurations. Without departing from the subject matter of this document, embodiments in which the configurations are combined and / or separated are included within the scope of the claims.
[0042] In this document, the term "A or B" may mean "only A," "only B," or "both A and B." In other words, in this document, the term "A or B" may be interpreted to mean "A and / or B." For example, in this document, the term "A, B, or C" may mean "only A," "only B," "only C," or "any combination of A, B, and C."
[0043] As used in this document, a slash ( / ) or a comma may mean "and / or." For example, "A / B" may mean "A and / or B." Thus, "A / B" may mean "only A," "only B," or "both A and B." For example, "A, B, C" may mean "A, B, or C."
[0044] In this document, "at least one of A and B" may mean "only A", "only B", or "both A and B". In addition, in this document, the expression "at least one of A or B" or "at least one of A and / or B" may be interpreted as the same as "at least one of A and B".
[0045] In addition, in this document, “at least one of A, B, and C” may mean “only A,” “only B,” “only C,” or “any combination of A, B, and C.” In addition, “at least one of A, B, or C” or “at least one of A, B, and / or C” may mean “at least one of A, B, and C.”
[0046] In addition, the parentheses used in this document may mean "for example." Specifically, when the expression "prediction (intra-frame prediction)" is used, it can indicate an example in which "intra-frame prediction" is proposed as "prediction." In other words, the term "prediction" in this document is not limited to "intra-frame prediction" and can indicate an example in which "intra-frame prediction" is proposed as "prediction." In addition, even when the expression "prediction (i.e., intra-frame prediction)" is used, it can indicate an example in which "intra-frame prediction" is proposed as "prediction."
[0047] This document relates to video / image coding. For example, the methods / embodiments disclosed in this document can be applied to methods disclosed in the Versatile Video Coding (VVC) standard. In addition, the methods / embodiments disclosed in this document can be applied to methods disclosed in the Essential Video Coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the second generation of Audio Video Coding standard (AVS2), or the next generation video / image coding standard (e.g., H.267 or H.268, etc.).
[0048] This document proposes various embodiments of video / image coding, and the above embodiments can also be executed in combination with each other unless otherwise mentioned.
[0049] In this document, a video can mean a set of a series of images over time. A picture generally means a unit representing one image in a specific time period, and a slice / tile is a unit constituting a part of a picture at the time of encoding. A slice / tile can include one or more coding tree units (CTUs). One picture can be composed of one or more slices / tiles. A tile is a rectangular region of CTUs within a specific tile column and a specific tile row in a picture. A tile column is a rectangular region of CTUs whose height is equal to the height of a picture and whose width is specified by a syntax element in a picture parameter set. A tile row is a rectangular region of CTUs whose width is specified by a syntax element in a picture parameter set and whose height is equal to the height of a picture. Tile scanning is a specific order of partitioning CTUs of a picture in which the CTUs are continuously ordered in a CTU raster scan in a tile, and the tiles in a picture are continuously ordered in a raster scan of tiles of the picture. A slice includes an integer number of complete tiles of a picture or an integer number of consecutive complete rows of CTUs within a tile that can be exclusively contained in a single NAL unit.
[0050] In addition, one picture can be divided into two or more sub-pictures. A sub-picture can be a rectangular region of one or more slices within a picture.
[0051] A pixel or pel can mean the smallest unit constituting one picture (or image). In addition, a "sample" can be used as a term corresponding to a pixel. A sample can generally represent a pixel or a pixel value, and can represent only a pixel / pixel value of a luminance component or only a pixel / pixel value of a chrominance component. Alternatively, a sample can mean a pixel value in a spatial domain, or can mean a transform coefficient in a frequency domain when the pixel value is transformed into the frequency domain.
[0052] A unit may represent a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to the region. A unit may include a luma block and two chroma (e.g., CB, CR) blocks. In some cases, a unit may be used interchangeably with terms such as block or region. In general, an M×N block may include M columns and N rows of samples (or sample arrays) or a set (or array) of transform coefficients.
[0053] In addition, in this document, at least one of quantization / dequantization and / or transform / inverse transform may be omitted. When quantization / dequantization is omitted, the quantized transform coefficient may be referred to as a transform coefficient. When transform / inverse transform is omitted, the transform coefficient may be referred to as a coefficient or a residual coefficient, or may still be referred to as a transform coefficient for the sake of consistency.
[0054] In this document, quantized transform coefficients and transform coefficients may be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, residual information may include information about the transform coefficients, and information about the transform coefficients may be signaled via residual coding syntax. The transform coefficients may be derived based on the residual information (or information about the transform coefficients), and the scaled transform coefficients may be derived by inverse transforming (scaling) the transform coefficients. Residual samples may be derived based on the inverse transform (transform) of the scaled transform coefficients. This may also be applied / expressed in other parts of this document.
[0055] In this document, technical features described individually in one drawing may be implemented individually or simultaneously.
[0056] Hereinafter, preferred embodiments of the present document will be described in more detail with reference to the accompanying drawings. Hereinafter, in the accompanying drawings, the same reference numerals are used for the same elements, and redundant descriptions of the same elements may be omitted.
[0057] Figure 1 An example of a video / image encoding system to which the embodiments of this document can be applied is illustrated.
[0058] Reference Figure 1 The video / image coding system may include a source device and a receiving device. The source device may send the coded video / image information or data to the receiving device in the form of a file or stream transmission via a digital storage medium or a network.
[0059] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.
[0060] A video source may acquire video / images through a process of capturing, synthesizing, or generating video / images. A video source may include a video / image capture device and / or a video / image generation device. For example, a video / image capture device may include one or more cameras, a video / image archive including previously captured video / images, etc. A video / image generation device may include, for example, a computer, a tablet computer, and a smartphone, and may (electronically) generate video / images. For example, a virtual video / image may be generated by a computer, etc. In this case, the video / image capture process may be replaced by a process of generating relevant data.
[0061] An encoding device encodes input video / images. For compression and coding efficiency, the encoding device performs a series of processes such as prediction, transformation, and quantization. The encoded data (encoded video / image information) is output as a bitstream.
[0062] The transmitter can transmit the encoded image / image information or data, output as a bitstream, to a receiver of a receiving device in the form of a file or stream transmission via a digital storage medium or network. Digital storage media may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter may include components for generating a media file in a predetermined file format and may also include components for transmission via a broadcast / communication network. The receiver may receive / extract the bitstream and transmit the received bitstream to a decoding device.
[0063] The decoding device may decode a video / image by performing a series of processes such as inverse quantization, inverse transformation, and prediction corresponding to the operations of the encoding device.
[0064] The renderer may render the decoded video / image, and the rendered video / image may be displayed on a display.
[0065] Figure 2 Schematically illustrates a configuration of a video / image encoding device to which embodiments of the present document can be applied. Hereinafter, the so-called encoding device may include an image encoding device and / or a video encoding device.
[0066] Reference Figure 2, the encoding device 200 may include and be configured with an image segmenter 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, an inverse quantizer 234, and an inverse transformer 235. The residual processor 230 may also include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstructed block generator. According to an embodiment, the image segmenter 210, the predictor 220, the residual processor 230, the entropy encoder 240, the adder 250, and the filter 260 described above may be configured by one or more hardware components (e.g., an encoder chipset or processor). In addition, the memory 270 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware components may also include the memory 270 as an internal / external component.
[0067] The image splitter 210 may split the input image (or picture or frame) input to the encoding device 200 into one or more processing units. For example, a processing unit may be referred to as a coding unit (CU). In this case, the coding unit may be recursively split from a coding tree unit (CTU) or a maximum coding unit (LCU) based on a quadtree, binary tree, and / or ternary tree (QTBTTT) structure. For example, a coding unit may be split into multiple coding units of increasing depth based on a quadtree, binary tree, and / or ternary tree structure. In this case, for example, the quadtree structure may be applied first, followed by the binary tree and / or ternary tree structure. Alternatively, the binary tree structure may be applied first. The encoding process according to the present disclosure may be performed based on the final coding unit that is no longer split. In this case, the maximum coding unit may be used as the final coding unit based on image characteristics, coding efficiency, and the like. Alternatively, if necessary, the coding unit may be recursively split into coding units of increasing depth so that a coding unit of optimal size is used as the final coding unit. The encoding process may include processes such as prediction, transformation, and reconstruction (described later). As another example, a processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, each of the prediction unit and the transform unit may be split or partitioned from the final coding unit. A prediction unit may be a unit for sample prediction, and a transform unit may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.
[0068] In some cases, the term "unit" may be used interchangeably with terms such as "block" or "region." Typically, an M×N block may represent a set of samples or transform coefficients consisting of M columns and N rows. A sample may generally represent a pixel or pixel value, and may represent only the pixel / pixel value of the luma component or only the pixel / pixel value of the chroma component. A sample may be used as a term corresponding to a pixel or picture element that constitutes a picture (or image).
[0069] The encoding device 200 can generate a residual signal (residual block, residual sample array) by subtracting the prediction signal (prediction block, prediction sample array) output from the inter predictor 221 or the intra predictor 222 from the input image signal (original block, original sample array), and the generated residual signal is transmitted to the transformer 232. In this case, as illustrated, the unit for subtracting the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) within the encoding device 200 can be referred to as a subtractor 231. The predictor can perform prediction on a block to be processed (hereinafter referred to as a current block) and generate a prediction block including prediction samples of the current block. The predictor can determine whether to apply intra prediction or inter prediction in units of the current block or CU. The predictor can generate various prediction-related information, such as prediction mode information, and transmit the generated information to the entropy encoder 240, as described later when describing each prediction mode. The prediction-related information can be encoded by the entropy encoder 240 and output in the form of a bitstream.
[0070] The intra-frame predictor 222 can predict the current block with reference to samples in the current picture. Depending on the prediction mode, the referenced samples can be located near the current block or far away from the current block. In intra-frame prediction, the prediction mode may include multiple non-directional modes and multiple directional modes. For example, the non-directional mode may include a DC mode and a planar mode. For example, depending on the degree of refinement of the prediction direction, the directional mode may include 33 directional prediction modes or 65 directional prediction modes. However, this is only an example, and more or fewer directional prediction modes may be used depending on the settings. The intra-frame predictor 222 can use the prediction mode applied to the neighboring blocks to determine the prediction mode applied to the current block.
[0071] The inter-frame predictor 221 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector in a reference picture. To reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. It can also include information about the inter-frame prediction direction (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, neighboring blocks can include spatially neighboring blocks in the current picture and temporally neighboring blocks in a reference picture. The reference picture containing the reference block and the reference picture containing the temporally neighboring block can be the same or different. Temporally neighboring blocks can be referred to as collocated reference blocks, collocated CUs (colCUs), etc., and reference pictures containing temporally neighboring blocks can be referred to as collocated pictures (colPics). For example, the inter-frame predictor 221 can configure a motion information candidate list based on the neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index for the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the inter-frame predictor 221 can use the motion information of the neighboring block as the motion information of the current block. In skip mode, unlike merge mode, a residual signal may not be transmitted. The motion vector prediction (MVP) mode can indicate the motion vector of the current block by using the motion vector of the neighboring block as a motion vector predictor and signaling the motion vector difference.
[0072] The predictor 220 can generate a prediction signal based on various prediction methods to be described later. For example, the predictor can not only apply intra prediction or inter prediction to predict a block, but also apply intra prediction and inter prediction simultaneously. This can be referred to as combined inter and intra prediction (CIIP). In addition, the predictor can perform prediction on the block based on an intra-block copy (IBC) prediction mode or a palette mode. The IBC prediction mode or palette mode can be used for content image / video coding such as games such as screen content coding (SCC). IBC basically performs prediction in the current picture, but is similar to inter prediction in deriving a reference block in the current picture. That is, IBC can use at least one of the inter prediction techniques described in this document. The palette mode can be considered an example of intra coding or intra prediction. When the palette mode is applied, the sample values in the picture can be signaled based on information about the palette index and the palette table.
[0073] The prediction signal generated by the predictor (including the inter-predictor 221 and / or the intra-predictor 222) can be used to generate a reconstructed signal or to generate a residual signal. The transformer 232 can generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique can include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loève transform (KLT), a graph-based transform (GBT), or a conditional nonlinear transform (CNT). Here, when the relationship information between pixels is exemplified as a graph, the GBT means a transform obtained from the graph. The CNT means a transform obtained based on a prediction signal generated by using all previously reconstructed pixels. In addition, the transform process can also be applied to a pixel block of the same size as a square, and can also be applied to a block of a variable size that is not a square.
[0074] The quantizer 233 can quantize the transform coefficients and transmit the quantized transform coefficients to the entropy encoder 240. The entropy encoder 240 can encode the quantized signal (information about the quantized transform coefficients) into an encoded quantized signal and output it as a bitstream. The information about the quantized transform coefficients can be referred to as residual information. The quantizer 233 can rearrange the quantized transform coefficients in block form into a one-dimensional vector form based on the coefficient scanning order, and also generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. The entropy encoder 240 can perform various encoding methods such as exponential Golomb coding, context-adaptive variable length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy encoder 240 can also encode information necessary for video / image reconstruction (e.g., syntax element values, etc.) in addition to the quantized transform coefficients, either together or separately. The encoded information (e.g., encoded video / image information) can be transmitted in the form of a bitstream or stored in units of network abstraction layer (NAL) units. The video / image information may also include information about various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may also include conventional constraint information. The signaled / sent information and / or syntax elements described later in this document may be encoded by the encoding process mentioned above and thus included in the bitstream. The bitstream may be sent over a network or stored in a digital storage medium. Here, the network may include a broadcast network and / or a communication network, etc., and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. A transmitter (not shown) for sending a signal output from the entropy encoder 240 and / or a memory (not shown) for storing the signal may be configured as an internal / external element of the encoding device 200, or the transmitter may also be included in the entropy encoder 240.
[0075] The quantized transform coefficients output from the quantizer 233 can be used to generate a prediction signal. For example, the inverse quantizer 234 and the inverse transformer 235 apply inverse quantization and inverse transformation to the quantized transform coefficients so that the residual signal (residual block or residual sample) can be reconstructed. The adder 250 adds the reconstructed residual signal to the prediction signal output from the inter-frame predictor 221 or the intra-frame predictor 222 so that a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) can be generated. For example, in the case of applying the skip mode, if there is no residual in the block to be processed, the prediction block can be used as a reconstructed block. The adder 250 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current picture, and as described later, it is also used for inter-frame prediction of the next picture by filtering.
[0076] Furthermore, luma mapping with chroma scaling (LMCS) may be applied in the picture encoding and / or reconstruction process.
[0077] The filter 260 can apply filtering to the reconstructed signal to improve the subjective / objective image quality. For example, the filter 260 can apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture, and store the modified reconstructed picture in the memory 270, specifically, in the DPB of the memory 270. Various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filter 260 can generate various information related to filtering to transmit the generated information to the entropy encoder 240, as described subsequently in the description of each filtering method. The information related to filtering can be encoded by the entropy encoder 240 and output in the form of a bitstream.
[0078] The modified reconstructed picture transmitted to the memory 270 may be used as a reference picture in the inter predictor 221. If inter prediction is applied through the inter predictor, prediction mismatch between the encoding apparatus 200 and the decoding apparatus may be avoided, and encoding efficiency may be improved.
[0079] The DPB of the memory 270 can store the modified reconstructed picture for use as a reference picture in the inter-frame predictor 221. The memory 270 can store the motion information of the block from which the motion information in the current picture is derived (or encoded) and / or the motion information of the block in the previously reconstructed picture. The stored motion information can be transmitted to the inter-frame predictor 221 to be used as the motion information of the spatially adjacent blocks or the motion information of the temporally adjacent blocks. The memory 270 can store the reconstructed samples of the reconstructed blocks in the current picture and can transmit the reconstructed samples to the intra-frame predictor 222.
[0080] Figure 3 FIG2 is a diagram schematically illustrating the configuration of a video / image decoding device to which this document is applicable.
[0081] Reference Figure 3, the decoding device 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-frame predictor 332 and an intra-frame predictor 331. The residual processor 320 may include an inverse quantizer 321 and an inverse transformer 322. According to an embodiment, the entropy decoder 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 described above may be configured by hardware components (e.g., a decoder chipset or processor). In addition, the memory 360 may include a decoded picture buffer (DPB) or may be configured by a digital storage medium. The hardware components may also include the memory 360 as an internal / external component.
[0082] When a bit stream including video / image information is input, the decoding apparatus 300 may generate a signal in response to the bit stream received in the video / image processing. Figure 2 , the image is reconstructed by processing the video / image information in the encoding device illustrated in . For example, the decoding device 300 may derive a unit / block based on block segmentation related information obtained from the bit stream. The decoding device 300 may perform decoding using a processing unit applied to the encoding device. Thus, for example, the processing unit of decoding may be a coding unit, and the coding unit may be split from a coding tree unit or a maximum coding unit according to a quadtree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units may be derived from the coding unit. In addition, the reconstructed image signal decoded and output by the decoding device 300 may be reproduced by a reproduction device.
[0083] The decoding device 300 may receive the data from the Figure 2The signal output by the encoding device illustrated in the example, and the received signal can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can derive the information necessary for image reconstruction (or picture reconstruction) (e.g., video / image information) by parsing the bitstream. The video / image information may also include information about various parameter sets such as the Adaptive Parameter Set (APS), the Picture Parameter Set (PPS), the Sequence Parameter Set (SPS), and the Video Parameter Set (VPS). In addition, the video / image information may also include conventional constraint information. The decoding device can further decode the picture based on the information about the parameter sets and / or the conventional constraint information. The signaled / received information and / or syntax elements described later in this document can be decoded and obtained from the bitstream through a decoding process. For example, the entropy decoder 310 can decode the information in the bitstream based on a coding method such as Exponential Golomb coding, CAVLC, or CABAC, and output the values of the syntax elements necessary for image reconstruction and the quantized values of the residual-related transform coefficients. More specifically, the CABAC entropy decoding method can receive the bin corresponding to each syntax element from the bitstream, use the syntax element information to be decoded and the decoded information of the neighboring blocks and the block to be decoded or the information of the symbol / bin decoded in the previous stage to determine the context model, and generate the symbol corresponding to the value of each syntax element by predicting the bin generation probability according to the determined context model to perform arithmetic decoding of the bin. At this time, the CABAC entropy decoding method can determine the context model and then update the context model using the decoded symbol / bin information of the context model for the next symbol / bin. The information about prediction among the information decoded by the entropy decoder 310 can be provided to the predictor (inter-frame predictor 332 and intra-frame predictor 331), and the residual value (i.e., quantized transform coefficients and related parameter information) entropy decoded by the entropy decoder 310 can be input to the residual processor 320. The residual processor 320 can derive the residual signal (residual block, residual sample, residual sample array). In addition, the information about filtering among the information decoded by the entropy decoder 310 can be provided to the filter 350. In addition, a receiver (not illustrated) for receiving a signal output from the encoding device may be further configured as an internal / external element of the decoding device 300, or the receiver may also be a component of the entropy decoder 310. In addition, the decoding device according to this document may be referred to as a video / image / picture decoding device, and the decoding device may also be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoder 310, and the sample decoder may include at least one of the inverse quantizer 321, the inverse transformer 322, the adder 340, the filter 350, the memory 360, the inter-frame predictor 332, and the intra-frame predictor 331.
[0084] The dequantizer 321 can dequantize the quantized transform coefficients to output transform coefficients. The dequantizer 321 can rearrange the quantized transform coefficients in a two-dimensional block form. In this case, the rearrangement can be performed based on a coefficient scan order performed by the encoding device. The dequantizer 321 can perform dequantization on the quantized transform coefficients using a quantization parameter (e.g., quantization step information) and obtain transform coefficients.
[0085] The inverse transformer 322 inverse-transforms the transform coefficients to obtain a residual signal (a residual block, a residual sample array).
[0086] The predictor 330 can perform prediction for a current block and generate a prediction block including prediction samples of the current block. The predictor can determine whether to apply intra prediction or inter prediction to the current block based on information on prediction output from the entropy decoder 310 and can determine a specific intra / inter prediction mode.
[0087] The predictor can generate a prediction signal based on various prediction methods that will be described later. For example, the predictor can not only apply intra prediction or inter prediction to predict one block, but also simultaneously apply intra prediction and inter prediction. This can be referred to as combined inter and intra prediction (CIIP). In addition, the predictor can predict a block based on an intra block copy (IBC) prediction mode or a palette mode. The IBC prediction mode or the palette mode can be used for content image / video encoding of a game or the like such as screen content coding (SCC). The IBC basically performs prediction in a current picture, but in deriving a reference block in the current picture, it is similar to inter prediction in that it uses at least one of the inter prediction techniques described in this document. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, information on a palette table and a palette index can be included in video / image information and signaled.
[0088] The intra predictor 331 can predict a current block by referring to samples in the current picture. Depending on the prediction mode, the referred samples can be located in the neighbors of the current block, or their positions can be separated from the current block. In intra prediction, the prediction mode can include a variety of non-directional modes and a variety of directional modes. The intra predictor 331 can also determine a prediction mode to be applied to the current block by using a prediction mode applied to a neighboring block.
[0089] The inter-frame predictor 332 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can also include information about the inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, the neighboring blocks can include spatial neighboring blocks in the current picture and temporal neighboring blocks in the reference picture. For example, the inter-frame predictor 332 can construct a motion information candidate list based on the neighboring blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and the prediction information can include information indicating the inter-frame prediction mode for the current block.
[0090] The adder 340 may add the obtained residual signal to the prediction signal (prediction block or prediction sample array) output from the predictor 330 (including the intra-frame predictor 331 and the inter-frame predictor 332) to generate a reconstructed signal (reconstructed picture, reconstructed block or reconstructed sample array). For example, in the case of applying the skip mode, if there is no residual of the block to be processed, the prediction block can be used as the reconstructed block.
[0091] The adder 340 may be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal may be used for intra-frame prediction of the next block to be processed in the current picture, and as described later, may also be output through filtering or may also be used for inter-frame prediction of the next picture.
[0092] In addition, luma mapping with chroma scaling (LMCS) can also be applied to the picture decoding process.
[0093] The filter 350 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 350 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image, and store the modified reconstructed image in the memory 360, specifically, in the DPB of the memory 360. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0094] The (modified) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter prediction 332. The memory 360 can store motion information of a block from which motion information in the current picture is derived (or decoded) and / or motion information of a block in the already reconstructed picture. The stored motion information can be transferred to the inter prediction 332 to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory 360 can store reconstructed samples of a reconstructed block in the current picture and transfer the reconstructed samples to the intra prediction 331.
[0095] In this document, the exemplary embodiments described in the filter 260, the inter prediction 221, and the intra prediction 222 of the encoding apparatus 200 can be equally or correspondingly applied to the filter 350, the inter prediction 332, and the intra prediction 331 of the decoding apparatus 300, respectively.
[0096] Further, as described above, in performing video encoding, prediction is performed to enhance compression efficiency. Thereby, a prediction block including prediction samples of a current block (i.e., a target encoding block) to be encoded can be generated. Here, the prediction block includes prediction samples in a spatial domain (or pixel domain). The prediction block is equally derived in the encoding apparatus and the decoding apparatus, and the encoding apparatus can signal information (residual information) about a residual between the original block and the prediction block (rather than original sample values of the original block itself) to the decoding apparatus to enhance image encoding efficiency. The decoding apparatus can derive a residual block including residual samples based on the residual information, can add the residual block and the prediction block to generate a reconstructed block including reconstructed samples, and can generate a reconstructed picture including the reconstructed block.
[0097] The residual information can be generated through a transform process and a quantization process. For example, the encoding apparatus can derive a residual block between the original block and the prediction block, can perform a transform process on residual samples (a residual sample array) included in the residual block to derive transform coefficients, can perform a quantization process on the transform coefficients to derive quantized transform coefficients, and can signal the relevant residual information (through a bitstream) to the decoding apparatus. In this case, the residual information can include value information of the quantized transform coefficients, position information, a transform scheme, a transform kernel, and a quantization parameter, etc. The decoding apparatus can perform a dequantization / inverse transform process based on the residual information and can derive residual samples (or a residual block). The decoding apparatus can generate a reconstructed picture based on the prediction block and the residual block. Further, for inter prediction reference of a subsequent picture, the encoding apparatus can also dequantize / inverse transform the quantized transform coefficients to derive a residual block, and can generate a reconstructed picture based thereon.
[0098] Furthermore, as described above, when prediction is performed on the current block, intra-frame prediction or inter-frame prediction can be applied. In an embodiment, when inter-frame prediction is applied to the current block, a predictor (more specifically, an inter-frame predictor) of the encoding / decoding device can derive prediction samples by performing inter-frame prediction on a block-by-block basis. Inter-frame prediction may refer to prediction derived using a method that depends on data elements (e.g., sample values or motion information) of a picture other than the current picture. When inter-frame prediction is applied to the current block, the prediction block (prediction sample array) of the current block can be derived based on a reference block (reference sample array) specified by a motion vector in a reference picture indicated by a reference picture index. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, the motion information of the current block can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may also include information on the inter-frame prediction type (L0 prediction, L1 prediction, Bi prediction, etc.). When inter-frame prediction is applied, the neighboring blocks may include spatially neighboring blocks in the current picture and temporally neighboring blocks in the reference picture. The reference picture including the reference block and the reference picture including the temporally neighboring block may be the same as or different from each other. The temporally neighboring block may be referred to as a collocated reference block, collocated CU (colCU), etc., and the reference picture including the temporally neighboring block may be referred to as a collocated picture (colPic). For example, a motion information candidate list may be configured based on the neighboring blocks of the current block, and a flag or index information indicating which candidate is selected (used) to derive the motion vector and / or reference picture index of the current block may be signaled. Inter-frame prediction may be performed based on various prediction modes, and for example, in skip mode and merge mode, the motion information of the current block may be the same as the motion information of the selected neighboring block. In skip mode, a residual signal may not be transmitted as in merge mode. In motion vector prediction (MVP) mode, the motion vector of the selected neighboring block may be used as a motion vector predictor, and the motion vector difference may be signaled. In this case, the motion vector of the current block may be derived using the sum of the motion vector predictor and the motion vector difference.
[0099] Depending on the inter-frame prediction type (L0 prediction, L1 prediction, Bi prediction, etc.), the motion information may also include L0 motion information and / or L1 motion information. The L0-direction motion vector may be referred to as the L0 motion vector or MVL0, and the L1-direction motion vector may be referred to as the L1 motion vector or MVL1. Prediction based on the L0 motion vector may be referred to as L0 prediction, prediction based on the L1 motion vector may be referred to as L1 prediction, and prediction based on both the L0 motion vector and the L1 motion vector may be referred to as bidirectional prediction. Here, the L0 motion vector may indicate a motion vector associated with reference picture list L0, and the L1 motion vector may indicate a motion vector associated with reference picture list L1. Reference picture list L0 may include pictures preceding the current picture in output order, and reference picture list L1 may include pictures following the current picture in output order as reference pictures. The preceding picture may be referred to as a forward (reference) picture, and the following picture may be referred to as a backward (reference) picture. Reference picture list L0 may also include pictures following the current picture in output order as reference pictures. In this case, the previous picture in reference picture list L0 may be indexed first, followed by the next picture. Reference picture list L1 may also include pictures preceding the current picture in output order as reference pictures. In this case, the next picture in reference picture list L1 may be indexed first, followed by the previous picture. Here, the output order may correspond to the picture order count (POC) order.
[0100] Figure 4 An example of an exemplary video / image encoding method to which the embodiments of this document are applicable is shown.
[0101] Figure 4 The method disclosed in Figure 2 Specifically, S400 may be performed by the inter-frame predictor 221 or the intra-frame predictor 222 of the encoding device 200, and S410, S420, S430, and S440 may be performed by the subtractor 231, the transformer 232, the quantizer 233, and the entropy encoder 240 of the encoding device 200, respectively.
[0102] Reference Figure 4 , the encoding device may derive a prediction sample by predicting the current block (S400). The encoding device may determine whether to perform inter-frame prediction or intra-frame prediction on the current block, and determine a specific inter-frame prediction mode or a specific intra-frame prediction mode based on the RD cost. According to the determined mode, the encoding device may derive a prediction sample for the current block.
[0103] The encoding apparatus may compare the predicted sample of the current block with the original sample and derive a residual sample ( S410 ).
[0104] The encoding apparatus may derive a transform coefficient through a transform process on the residual sample ( S420 ), and may derive a quantized transform coefficient by quantizing the derived transform coefficient ( S430 ).
[0105] Quantization may be performed based on a quantization parameter. The transform process and / or the quantization process may be skipped. When the transform process is skipped, the (quantized) (residual) coefficients of the residual samples may be encoded according to the residual coding technique described later. For the sake of terminology, the (quantized) (residual) coefficients may also be referred to as (quantized) transform coefficients.
[0106] The encoding device may encode image information including residual information and prediction information, and may output the encoded image information in the form of a bitstream (S440). The prediction information may include information about motion information (for example, when inter-frame prediction is applied) and prediction mode information as a plurality of information related to the prediction process. The residual information may include information about quantized transform coefficients. The residual information may be entropy-encoded. Alternatively, the residual information may include information about (quantized) (residual) coefficients.
[0107] The output bitstream can be transmitted to a decoding device via a storage medium or a network.
[0108] Figure 5 An example of an exemplary video / image decoding method to which the embodiments of this document are applied is shown.
[0109] Figure 5 The method disclosed in Figure 3 The decoding device 300 described above is executed. Specifically, S500 may be performed by the inter-frame predictor 332 or the intra-frame predictor 331 of the decoding device 300. In S500, the entropy decoder 310 of the decoding device 300 may decode the prediction information included in the bitstream and derive the value of the related syntax element. S510, S520, S530, and S540 may be performed by the entropy decoder 310, the inverse quantizer 321, the inverse transformer 322, and the adder 340 of the decoding device 300, respectively.
[0110] Reference Figure 5 The decoding device may perform an operation corresponding to the operation performed in the encoding device. The decoding device may perform inter-frame prediction or intra-frame prediction on the current block and derive a prediction sample based on the received prediction information (S500).
[0111] The decoding apparatus may derive a quantized transform coefficient of the current block based on the received residual information (S510). The decoding apparatus may derive the quantized transform coefficient from the residual information through entropy decoding.
[0112] The decoding apparatus may inverse quantize the quantized transform coefficient and derive the transform coefficient (S520). The inverse quantization may be performed based on a quantization parameter.
[0113] The decoding apparatus may induce residual samples through an inverse transform process on the transform coefficients ( S530 ).
[0114] The inverse transform process and / or the inverse quantization process may be skipped. When the inverse transform process is skipped, (quantized) (residual) coefficients may be derived from the residual information, and residual samples may be derived based on the (quantized) (residual) coefficients.
[0115] The decoding apparatus may generate reconstructed samples of the current block based on the residual samples and the prediction samples, and generate a reconstructed picture based on the reconstructed samples (S540).Thereafter, a loop filtering process may be further applied to the reconstructed picture as described above.
[0116] Figure 6 The hierarchical structure of the encoded image / video is shown as an example.
[0117] Reference Figure 6 The encoded image / video is divided into the VCL (Video Coding Layer) that manipulates the image / video decoding process and its own, the subsystem that sends and stores the encoded information, and the Network Abstraction Layer (NAL) that exists between the VCL and the subsystem and is responsible for the network adaptation function.
[0118] VCL can generate VCL data including compressed image data (slice data), or generate parameter sets including picture parameter sets (picture parameter set: PPS), sequence parameter sets (sequence parameter set: SPS), video parameter sets (video parameter set: VPS), etc., or supplementary enhancement information (SEI) messages necessary for the decoding processing of images.
[0119] In NAL, a NAL unit can be generated by adding header information (NAL unit header) to the raw byte sequence payload (RBSP) generated in the VCL. In this case, RBSP refers to slice data, parameter sets, SEI messages, etc. generated in the VCL. The NAL unit header may include NAL unit type information specified according to the RBSP data included in the corresponding NAL unit.
[0120] In addition, according to the RBSP generated in the VCL, the NAL unit can be divided into a VCL NAL unit and a non-VCL NAL unit. A VCL NAL unit may refer to a NAL unit including information about an image (slice data), and a non-VCL-NAL unit may refer to a NAL unit including information required for decoding an image (parameter set or SEI message).
[0121] VCL NAL units and non-VCL NAL units can be transmitted over a network by attaching header information according to the data standard of the subsystem. For example, NAL units can be converted into a predetermined standard data format such as H.266 / VVC file format, Real-time Transport Protocol (RTP), and Transport Stream (TS), and transmitted over various networks.
[0122] As described above, in a NAL unit, a NAL unit type may be specified according to an RBSP data structure included in a corresponding NAL unit, and information about the NAL unit type may be stored in a NAL unit header and signaled.
[0123] For example, NAL units can be roughly divided into VCL NAL unit types and non-VCL NAL unit types according to whether the NAL unit includes information about the image (slice data). VCL NAL unit types can be classified according to the nature and type of the picture included in the VCL NAL unit, and non-VCL NAL unit types can be classified according to the type of parameter set.
[0124] The following are examples of NAL unit types specified according to the type of parameter sets included in non-VCL NAL unit types.
[0125] -APS (Adaptation Parameter Set) NAL unit: type of NAL unit including APS
[0126] -DPS (Decoding Parameter Set) NAL unit: the type of NAL unit that includes DPS
[0127] -VPS (Video Parameter Set) NAL unit: Type of NAL unit containing VPS
[0128] -SPS (Sequence Parameter Set) NAL unit: the type of NAL unit that includes the SPS
[0129] -PPS (Picture Parameter Set) NAL unit: Type of NAL unit including PPS
[0130] -PH (Picture Header) NAL unit: the type of NAL unit including PH
[0131] The above-mentioned NAL unit type has syntax information for the NAL unit type, and the syntax information can be stored in the NAL unit header and signaled. For example, the syntax information can be nal_unit_type, and the NAL unit type can be specified by the nal_unit_type value.
[0132] In addition, as described above, a picture may include multiple slices, and a slice may include a slice header and slice data. In this case, a picture header may be further added to the multiple slices (slice header and slice data set) in a picture. The picture header (picture header syntax) may include information / parameters common to the picture. In this document, a tile group may be mixed with a slice or a picture or replaced with a slice or a picture. In addition, in this document, a tile group header may be mixed with a slice header or a picture header or replaced with a slice header or a picture header.
[0133] A slice header (slice header syntax) may include information / parameters common to slices. APS (APS syntax) or PPS (PPS syntax) may include information / parameters common to one or more slices or pictures. SPS (SPS syntax) may include information / parameters common to one or more sequences. VPS (VPS syntax) may include information / parameters common to multiple layers. DPS (DPS syntax) may include information / parameters common to the entire video. DPS may include information / parameters related to the concatenation of coded video sequences (CVS). In this document, a high-level syntax (HLS) may include at least one of APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, picture header syntax, and slice header syntax.
[0134] In this document, the image / video information encoded in the encoding device and signaled to the decoding device in the form of a bitstream may include not only picture segmentation related information, intra-frame / inter-frame prediction information, residual information, loop filtering information, etc. in the picture, but also information included in the slice header, information included in the picture header, information included in the APS, information included in the PPS, information included in the SPS, information included in the VPS, and / or information included in the DPS. In addition, the image / video information may also include information in the NAL unit header.
[0135] As described above, high-level syntax (HLS) can be encoded / signaled for video / image coding. In this document, the video / image information may include HLS. For example, a coded picture may consist of one or more slices. Parameters describing the coded picture may be signaled in a picture header (PH), and parameters describing the slice may be signaled in a slice header (SH). The PH may be sent as its own NAL unit type. The SH may be present at the beginning of a NAL unit including the payload of the slice (i.e., slice data). The details of the syntax and semantics of the PH and SH may be as disclosed in the VVC standard. Each picture may be associated with a PH. A picture may consist of different types of slices: intra-coded slices (i.e., I slices) and inter-coded slices (i.e., P slices and B slices). As a result, the PH may include syntax elements necessary for intra slices of a picture and inter slices of a picture.
[0136] In addition, as described above, the encoding device performs entropy encoding based on various encoding methods such as exponential Golomb coding, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. In addition, the decoding device can perform entropy decoding based on encoding methods such as exponential Golomb coding, CAVLC, or CABAC. Hereinafter, the entropy encoding / decoding process will be described.
[0137] Figure 7 An example of an entropy coding method to which the embodiments of this document are applicable is schematically shown, and Figure 8 An entropy encoder in an encoding device is schematically shown. Figure 8 The entropy encoder in the encoding device can also be applied to the above-mentioned Figure 2 The entropy encoder 240 of the encoding device 200.
[0138] Reference Figure 7 and Figure 8 , the encoding device (entropy encoder) performs entropy encoding processing on the image / video information. The image / video information may include segmentation related information, prediction related information (for example, inter-frame / intra-frame prediction distinction information, intra-frame prediction mode information, inter-frame prediction mode information, etc.), residual information, loop filter related information, or may include various syntax elements related thereto. Entropy encoding can be performed in syntax element units. S700 and S710 may be performed by Figure 2 The above-mentioned entropy encoder 240 of the encoding device 200 is performed.
[0139] The encoding device may perform binarization on the target syntax element (S700). Here, the binarization may be based on various binarization methods such as truncated Rice binarization, fixed-length binarization, etc., and the binarization method for the target syntax element may be predefined. The binarization process may be performed by the binarizer 242 in the entropy encoder 240.
[0140] The encoding device may perform entropy encoding on the target syntax element (S710). The encoding device may perform rule-based encoding (context-based) or bypass-based encoding on the bin string of the target syntax element based on an entropy encoding scheme such as context-adaptive arithmetic coding (CABAC) or context-adaptive variable length coding (CAVLC), and its output may be incorporated into the bitstream. The entropy encoding process may be performed by the entropy encoding processor 243 in the entropy encoder 240. As described above, the bitstream may be transmitted to the decoding device via a (digital) storage medium or a network.
[0141] Figure 9 An example of an entropy decoding method to which the embodiments of this document are applicable is schematically shown, and Figure 10 An entropy decoder in a decoding device is schematically shown. Figure 10 The entropy decoder in the decoding device can also be used with the above Figure 3 The entropy decoder 310 of the decoding device 300 is equivalent to or corresponds to the entropy decoder 310 of the decoding device 300.
[0142] refer to Figure 9 and Figure 10 , the decoding device (entropy decoder) can decode the encoded image / video information. The image / video information may include segmentation related information, prediction related information (e.g., inter-frame / intra-frame prediction distinction information, intra-frame prediction mode information, inter-frame prediction mode information, etc.), residual information, loop filter related information, or may include various syntax elements related thereto. Entropy coding can be performed in syntax element units. S900 and S910 may be performed by Figure 3 The above-mentioned entropy decoder 310 of the decoding device 300 is executed.
[0143] The decoding device may perform binarization on the target syntax element (S900). Here, the binarization may be based on various binarization methods such as truncated Rice binarization, fixed-length binarization, etc., and the binarization method for the target syntax element may be predefined. The decoding device may derive an enabled bin string (bin string candidate) of an enabled value of the target syntax element through the binarization process. The binarization process may be performed by the binarizer 312 in the entropy decoder 310.
[0144] The decoding device may perform entropy decoding on the target syntax element (S910). When sequentially decoding and parsing each bin of the target syntax element from the input bits in the bitstream, the decoding device compares the derived bin string with the enabled bin string of the corresponding syntax element. When the derived bin string is the same as one of the enabled bin strings, the value corresponding to the bin string may be derived as the value of the syntax element. If not, the above process may be performed again after further parsing the next bit in the bitstream. Through these processes, even if the start bit or end bit is not used for specific information (specific syntax element) in the bitstream, the decoding device may use variable length bits to signal the information. Accordingly, relatively fewer bits may be allocated to low values, thereby improving overall coding efficiency.
[0145] The decoding device may perform context-based or bypass-based decoding on each bin in the bin string from the bitstream based on an entropy coding technique such as CABAC, CAVLC, etc. In this regard, the bitstream may include various information used for image / video decoding as described above. As described above, the bitstream may be transmitted to the decoding device via a (digital) storage medium or a network.
[0146] In addition, generally, one NAL unit type can be set for one picture. The NAL unit type can be signaled by nal_unit_type in the NAL unit header of the NAL unit containing the slice. nal_unit_type is syntax information for specifying the NAL unit type, that is, as shown in Table 1 or Table 2 below, it can specify the type of RBSP data structure contained in the NAL unit.
[0147] Table 1 below shows examples of NAL unit type codes and NAL unit type categories.
[0148] [Table 1]
[0149]
[0150] Alternatively, as an example, the NAL unit type code and the NAL unit type category may be defined as shown in Table 2 below.
[0151] [Table 2]
[0152]
[0153] As shown in Table 1 or Table 2, the name of the NAL unit type and its value can be specified according to the RBSP data structure included in the NAL unit, and can be classified into a VCL NAL unit type and a non-VCL NAL unit type according to whether the NAL unit includes information (slice data) on a picture. The VCL NAL unit type can be classified according to the properties, types, and the like of a picture, and the non-VCL NAL unit type can be classified according to the types of parameter sets and the like. For example, the NAL unit type can be specified according to the properties and types of a picture included in the VCL NAL unit as follows.
[0154] TRAIL: It indicates the type of a NAL unit including coded slice data of a trailing picture / sub-picture. For example, nal_unit_type can be defined as TRAIL_NUT, and the value of nal_unit_type can be specified as 0.
[0155] Here, a trailing picture refers to a picture that follows a picture that can be randomly accessed in output order and decoding order. The trailing picture can be a non-IRAP picture that follows the associated IRAP picture in output order, and is not an STSA picture. For example, the trailing picture associated with the IRAP picture follows the IRAP picture in decoding order. A picture that follows the associated IRAP picture in output order and precedes the associated IRAP picture in decoding order is not allowed.
[0156] STSA (Stepwise Temporal Sub-Layer Access): It indicates the type of a NAL unit including coded slice data of an STSA picture / sub-picture. For example, nal_unit_type can be defined as STSA_NUT, and the value of nal_unit_type can be specified as 1.
[0157] Here, an STSA picture is a picture that can switch between temporal sub-layers in a bitstream supporting temporal scalability, and refers to a picture indicating a position where an up-switch from a lower sub-layer to an upper sub-layer higher than the lower sub-layer is possible. The STSA picture does not use a picture in the same layer as the STSA picture Temporalld and the same as the STSA picture for inter prediction. A picture in the same layer as the STSA picture Temporalld and following the STSA picture in decoding order does not use a picture in the same layer as the STSA picture Temporalld and preceding the STSA picture in decoding order for inter prediction reference. The STSA picture enables an up-switch from an immediately next sub-layer to a sub-layer including the STSA picture in the STSA picture. In this case, the picture to be coded must not belong to the lowest sub-layer. That is, the STSA picture must always have a Temporalld greater than 0.
[0158] RADL (Random Access Decodable Preamble (Picture)): This indicates the type of the NAL unit including slice data of the RADL picture / sub-picture to be encoded. For example, nal_unit_type may be defined as RADL_NUT, and the value of nal_unit_type may be specified as 2.
[0159] Here, all RADL pictures are leading pictures. RADL pictures are not used as reference pictures in the decoding process of the trailing pictures of the same associated IRAP picture. Specifically, a RADL picture with nuh_layer_id equal to layerId is a picture that follows the IRAP picture associated with the RADL picture in output order and is not used as a reference picture in the decoding process of pictures with nuh_layer_id equal to layerId. When field_seq_flag (i.e., sps_field_seq_flag) is 0, all RADL pictures precede all non-leading pictures of the same associated IRAP picture in decoding order (i.e., if a RADL picture exists). Furthermore, a leading picture refers to a picture that precedes the associated IRAP picture in output order.
[0160] RASL (Random Access Skip Preamble (Picture)): This indicates the type of the NAL unit including slice data of the RASL picture / sub-picture to be encoded. For example, nal_unit_type may be defined as RASL_NUT, and the value of nal_unit_type may be specified as 3.
[0161] Here, all RASL pictures are leading pictures of the associated CRA picture. When the associated CRA picture has NoOutputBeforeRecoveryFlag whose value is 1, the RASL picture may not be output nor decoded correctly because the RASL picture may include references to pictures that do not exist in the bitstream. RASL pictures are not used as reference pictures for the decoding process of non-RASL pictures of the same layer. However, RADL sub-pictures in RASL pictures of the same layer can be used for inter-frame prediction of collocated RADL sub-pictures in RADL pictures associated with the same CRA picture as the RASL picture. When field_seq_flag (i.e., sps_field_seq_flag) is 0, all RASL pictures precede all non-leading pictures of the same associated CRA picture in decoding order (i.e., if a RASL picture exists).
[0162] For non-IRAP VCL NAL unit types, there may be a reserved nal_unit_type. For example, nal_unit_type may be defined as RSV_VCL_4 and RSV_VCL_6, and the value of nal_unit_type may be specified as 4 to 6, respectively.
[0163] Here, an intra random access point (IRAP) is information indicating a NAL unit for a picture that can be randomly accessed. An IRAP picture can be a CRA picture or an IDR picture. For example, an IRAP picture refers to a picture having a NAL unit type defined as IDR_W_RADL, IDR_N_LP, and CRA_NUT as in Table 1 or Table 2 above, and the value of nal_unit_type can be specified as 7 to 9, respectively.
[0164] IRAP pictures do not use any reference pictures in the same layer for inter prediction during the decoding process. In other words, an IRAP picture does not reference any pictures other than itself for inter prediction during the decoding process. The first picture in the bitstream in decoding order is called an IRAP or GDR picture. For a single-layer bitstream, if the necessary parameter set is available when it is needed to reference it, all subsequent non-RASL pictures and IRAP pictures in the coded layer video sequence (CLVS) in decoding order can be accurately decoded without performing the decoding process for pictures that precede the IRAP picture in decoding order.
[0165] The value of mixed_nalu_types_in_pic_flag for an IRAP picture is 0. When the value of mixed_nalu_types_in_pic_flag for a picture is 0, one slice in the picture may have a NAL unit type (nal_unit_type) in the range from IDR_W_RADL to CRA_NUT (for example, the NAL unit type values in Table 1 or Table 2 are 7 to 9), and all other slices in the picture may have the same NAL unit type (nal_unit_type). In this case, the picture can be regarded as an IRAP picture.
[0166] Instantaneous Decoding Refresh (IDR): This indicates the type of NAL unit containing slice data of the IDR picture / sub-picture to be encoded. For example, the nal_unit_type of the IDR picture / sub-picture can be defined as IDR_W_RADL or IDR_N_LP, and the value of nal_unit_type can be specified as 7 or 8, respectively.
[0167] Here, an IDR picture may not use inter-frame prediction in the decoding process (i.e., it does not refer to pictures other than itself for inter-frame prediction), but may be the first picture in the bitstream in decoding order, or may appear later in the bitstream (i.e., not first, but later). Each IDR picture is the first picture of a coded video sequence (CVS) in decoding order. For example, when an IDR picture is associated with a decodable leading picture, the NAL unit type of the IDR picture may be represented as IDR_W_RADL, and when the IDR picture is not associated with a leading picture, the NAL unit type of the IDR picture may be represented as IDR_N_LP. That is, an IDR picture whose NAL unit type is IDR_W_RADL may not have an associated RADL picture present in the bitstream, but may have an associated RADL picture in the bitstream. An IDR picture whose NAL unit type is IDR_N_LP does not have an associated leading picture present in the bitstream.
[0168] Clean Random Access (CRA): This indicates the type of the NAL unit including the slice data of the CRA picture / sub-picture to be encoded. For example, nal_unit_type may be defined as CRA_NUT, and the value of nal_unit_type may be specified as 9.
[0169] Here, a CRA picture may not use inter-frame prediction in the decoding process (i.e., it does not refer to pictures other than itself for inter-frame prediction), but may be the first picture in the decoding order in the bitstream, or may appear later in the bitstream (i.e., not first, but later). A CRA picture may have an associated RADL or RASL picture present in the bitstream. For a CRA picture in which the value of NoOutputBeforeRecoveryFlag is 1, the associated RASL picture may not be output by the decoder. This is because decoding is not possible in this case because a reference to a picture that does not exist in the bitstream is included.
[0170] Gradual Decoding Refresh (GDR): This indicates the type of the NAL unit including slice data of the GDR picture / sub-picture to be encoded. For example, nal_unit_type may be defined as GDR_NUT, and the value of nal_unit_type may be specified as 10.
[0171] Here, the pps_mixed_nalu_types_in_pic_flag value of the GDR picture may be 0. When the value of pps_mixed_nalu_types_in_pic_flag of the picture is 0 and one slice in the picture has the NAL unit type of GDR_NUT, all other slices in the picture have the same NAL unit type (nal_unit_type) value, and in this case, the picture may become a GDR picture after the first slice is received.
[0172] In addition, for example, the NAL unit type may be specified according to the type of parameters included in the non-VCL NAL unit, and, as shown in Table 1 or Table 2 above, the following NAL unit types (nal_unit_type) may be included: for example, VPS_NUT indicating the type of NAL unit including a video parameter set, SPS_NUT indicating the type of NAL unit including a sequence parameter set, PPS_NUT indicating the type of NAL unit including a picture parameter set, and PH_NUT indicating the type of NAL unit including a picture header.
[0173] In addition, a bitstream that supports temporal scalability (or a temporally scalable bitstream) includes information about a temporal layer for temporal scaling. The information about the temporal layer may be identification information of the temporal layer specified according to the temporal scalability of the NAL unit. For example, the identification information of the temporal layer may use temporal_id syntax information, and the temporal_id syntax information may be stored in the NAL unit header in the encoding device and signaled to the decoding device. Hereinafter, in this specification, a temporal layer may be referred to as a sublayer, a temporal sublayer, a temporal scalable layer, etc.
[0174] Figure 11 is a diagram showing a temporal layer structure of a NAL unit in a bitstream supporting temporal scalability.
[0175] When the bitstream supports temporal scalability, the NAL units included in the bitstream have identification information of the temporal layer (e.g., temporal_id). As an example, a temporal layer consisting of NAL units whose temporal_id value is 0 can provide the lowest temporal scalability, and a temporal layer consisting of NAL units whose temporal_id value is 2 can provide the highest temporal scalability.
[0176] exist Figure 11 , a block marked with I refers to an I picture, and a block marked with B refers to a B picture. In addition, an arrow indicates a reference relationship regarding whether a picture refers to another picture.
[0177] like Figure 11As shown in , the NAL unit of the temporal layer whose temporal_id value is 0 is a reference picture that can be referenced by the NAL unit of the temporal layer whose temporal_id value is 0, 1, or 2. The NAL unit of the temporal layer whose temporal_id value is 1 is a reference picture that can be referenced by the NAL unit of the temporal layer whose temporal_id value is 1 or 2. The NAL unit of the temporal layer whose temporal_id value is 2 can be a reference picture that can be referenced by the NAL unit of the same temporal layer (i.e., the temporal layer whose temporal_id value is 2), or can be a non-reference picture that is not referenced by other pictures.
[0178] If Figure 11 If the NAL units of the temporal layer (ie, the highest temporal layer) whose temporal_id value is 2 shown in are non-reference pictures, these NAL units are extracted (or removed) from the bitstream in the decoding process without affecting other pictures.
[0179] Among the NAL unit types described above, the IDR and CRA types are information indicating that the NAL unit includes a picture that can be randomly accessed (or spliced), that is, a random access point (RAP) or an intra random access point (IRAP) picture used as a random access point. In other words, an IRAP picture can be an IDR or CRA picture and can include only I slices. In the bitstream, the first picture in decoding order becomes an IRAP picture.
[0180] If an IRAP picture (IDR, CRA picture) is included in the bitstream, there may be pictures that precede the IRAP picture in output order but follow it in decoding order. These pictures are called leading pictures (LP).
[0181] Figure 12 A diagram used to describe a picture that can be randomly accessed.
[0182] A picture that can be randomly accessed (ie, a RAP or IRAP picture used as a random access point) is the first picture in the bitstream in decoding order during random access and includes only I slices.
[0183] Figure 12 The output order (or display order) and decoding order of pictures are shown. As illustrated, the output order and decoding order of pictures may be different from each other. For convenience, the description is made while dividing the pictures into predetermined groups.
[0184] The pictures belonging to the first group (I) indicate pictures that precede the IRAP picture in both output order and decoding order, and the pictures belonging to the second group (II) indicate pictures that precede the IRAP picture in output order but follow it in decoding order. The pictures of the third group (III) follow the IRAP picture in both output order and decoding order.
[0185] The first group (I) of pictures may be decoded and output regardless of IRAP pictures.
[0186] A picture belonging to the second group (II) output before an IRAP picture is called a leading picture, and when an IRAP picture is used as a random access point, the leading picture may become a problem in a decoding process.
[0187] The pictures belonging to the third group (III) following the IRAP picture in output order and decoding order are called normal pictures. Normal pictures are not used as reference pictures for leading pictures.
[0188] A random access point where random access occurs in a bitstream becomes an IRAP picture, and random access starts as the first picture of the second group (II) is output.
[0189] Figure 13 This is a diagram used to describe an IDR picture.
[0190] An IDR picture is a picture that becomes a random access point when a group of pictures has a closed structure. As described above, since an IDR picture is an IRAP picture, it only includes I slices and can be the first picture in the bitstream in decoding order, or it can appear in the middle of the bitstream. When an IDR picture is decoded, all reference pictures stored in the decoded picture buffer (DPB) are marked as "unused for reference."
[0191] Figure 13 The bar shown in indicates a picture, and the arrow indicates a reference relationship related to whether the picture can use another picture as a reference picture. An x mark on the arrow indicates that the picture cannot refer to the picture indicated by the arrow.
[0192] As shown, a picture whose POC is 32 is an IDR picture. A picture whose POC is 25 to 31 and output before the IDR picture is a leading picture 1310. A picture whose POC is 33 or greater corresponds to a normal picture 1320.
[0193] The leading picture 1310 preceding the IDR picture in output order can use a leading picture other than the IDR picture as a reference picture, but cannot use a past picture 1330 preceding the leading picture 1310 in output order and decoding order as a reference picture.
[0194] The normal picture 1320 following the IDR picture in output order and decoding order can be decoded with reference to the IDR picture, the leading picture, and other normal pictures.
[0195] Figure 14 A diagram used to describe a CRA picture.
[0196] A CRA picture is a picture that becomes a random access point when a group of pictures has an open structure. As described above, since a CRA picture is also an IRAP picture, it only includes I slices and can be the first picture in the decoding order of the bitstream, or can appear in the middle of the bitstream for normal display.
[0197] Figure 14 The bars shown in indicate pictures, and the arrows indicate reference relationships related to whether a picture can use another picture as a reference picture. An x mark on an arrow indicates that one or more pictures cannot refer to the picture indicated by the arrow.
[0198] A leading picture 1410 that precedes the CRA picture in output order may use all of the CRA picture, other leading pictures, and past pictures 1430 that precede the leading picture 1410 in output order and decoding order as reference pictures.
[0199] Conversely, a normal picture 1420 following a CRA picture in output order and decoding order may be decoded with reference to a normal picture different from the CRA picture. The normal picture 1420 may not use the leading picture 1410 as a reference picture.
[0200] In addition, in the VVC standard, it is possible to allow the coded picture (ie, the current picture) to include slices of different NAL unit types. Whether the current picture includes slices of different NAL unit types can be indicated based on the syntax element mixed_nalu_types_in_pic_flag. For example, when the current picture includes slices of different NAL unit types, the value of the syntax element mixed_nalu_types_in_pic_flag can be represented as 1. In this case, the current picture must refer to a PPS including mixed_nalu_types_in_pic_flag with a value of 1. The semantics of the flag (mixed_nalu_types_in_pic_flag) are as follows:
[0201] When the value of the syntax element mixed_nalu_types_in_pic_flag is 1, it may indicate that each picture of the reference PPS has one or more VCL NAL units, the VCL NAL units do not have the same NAL unit type (nal_unit_type), and the picture is not an IRAP picture.
[0202] When the value of the syntax element mixed_nalu_types_in_pic_flag is 0, it may indicate that each picture of the referenced PPS has one or more VCL NAL units, and the VCL NAL units of each picture of the referenced PPS have the same value of the NAL unit type (nal_unit_type).
[0203] When the value of no_mixed_nalu_types_in_pic_constraint_flag is 1, the value of mixed_nalu_types_in_pic_flag must be 0. The no_mixed_nalu_types_in_pic_constraint_flag syntax element indicates a constraint on whether the value of mixed_nalu_types_in_pic_flag of a picture must be 0. For example, based on no_mixed_nalu_types_in_pic_constraint_flag information signaled from a higher-level syntax (e.g., PPS) or a syntax including information about constraints (e.g., GCI; general constraint information), it can be determined whether the value of mixed_nalu_types_in_pic_flag must be 0.
[0204] In a picture picA that also includes one or more slices with different values of NAL unit types (i.e., when the value of mixed_nalu_types_in_pic_flag of picture picA is 1), for each slice with a NAL unit type value nalUnitTypeA in the range from IDR_W_RADL to CRA_NUT (e.g., in Table 1 or Table 2, the values of NAL unit type are 7 to 9), the following may apply.
[0205] - The slice must belong to a subpicA in which the corresponding subpic_treated_as_pic_flag has a value of 1. Here, subpic_treated_as_pic_flag is information about whether the i-th subpicture of each picture encoded in the CLVS is treated as a picture in the decoding process other than the loop filtering operation. For example, when the value of subpic_treated_as_pic_flag is 1, it can indicate that the i-th subpicture is treated as a picture in the decoding process other than the loop filtering operation. Alternatively, when the value of subpic_treated_as_pic_flag is 0, it can indicate that the i-th subpicture is not treated as a picture in the decoding process other than the loop filtering operation.
[0206] - A slice shall not belong to a sub-picture of picA that includes a VCL NAL unit with a NAL unit type (nal_unit_type) not equal to nalUnitTypeA.
[0207] - For all PUs that follow CLVS in decoding order, the RefPicList or RefPicList of the slice in subpicA shall not include pictures that precede picA in decoding order in the active entry.
[0208] To operate the concept described above, the following can be specified. For example, the following can be applied to the VCL NAL unit of a specific picture.
[0209] - When the value of mixed_nalu_types_in_pic_flag is 0, the value of the NAL unit type (nal_unit_type) must be the same for all slice NAL units coded in a picture. A picture or PU can be considered to have the same NAL unit type as the slice NAL units coded in the picture or PU.
[0210] Otherwise (when the value of mixed_nalu_types_in_pic_flag is 1), one or more VCL NAL units must have a NAL unit type of a specific value in the range from IDR_W_RADL to CRA_NUT (e.g., the values of NAL unit type in Table 1 or Table 2 are 7 to 9), and all other VCL NAL units must have the same NAL unit type as GRA_NUT or a NAL unit type of a specific value in the range from TRAIL_NUT to RSV_VCL_6 (e.g., the values of NAL unit type in Table 1 or Table 2 are 0 to 6).
[0211] In the current VVC standard, in the case of a picture with mixed NAL unit types, there may be at least the following problems.
[0212] 1. When a picture includes IDR and non-IRAP NAL units, and when signaling for a reference picture list (RPL) is present in the slice header, the signaling must also be present in the header of the IDR slice. When the value of sps_idr_rpl_present_flag is 1, RPL signaling is present in the slice header of the IDR slice. Currently, the value of this flag (sps_idr_rpl_present_flag) can be 0 even when there are one or more pictures with mixed NAL unit types. Here, the sps_idr_rpl_present_flag syntax element can indicate whether RPL syntax elements can be present in the slice header of a slice with a NAL unit type such as IDR_N_LP or IDR_W_RADL. For example, when the value of sps_idr_rpl_present_flag is 1, it can indicate that RPL syntax elements can be present in the slice header of a slice with a NAL unit type such as IDR_N_LP or IDR_W_RADL. Alternatively, when the value of sps_idr_rpl_present_flag is 0, it may indicate that the RPL syntax element is not present in the slice header of a slice with a NAL unit type such as IDR_N_LP or IDR_W_RADL.
[0213] 2. When the current picture references a PPS with the value of mixed_nalu_types_in_pic_flag set to 1, one or more of the VCL NAL units of the current picture must have a NAL unit type with a specific value in the range from IDR_W_RADL to CRA_NUT (e.g., the NAL unit type values in Table 1 or Table 2 above are 7 to 9), and all other VCL NAL units must have the same NAL unit type as GRA_NUT or a NAL unit type with a specific value in the range from TRAIL_NUT to RSV_VCL_6 (e.g., the NAL unit type values in Table 1 or Table 2 above are 0 to 6). This constraint applies only to current pictures that include a mix of IRAP and non-IRAP NAL unit types. However, it does not yet apply correctly to pictures that include a mix of RASL / RADL and non-IRAP NAL unit types.
[0214] This document provides a solution to the above-mentioned problem. That is, as described above, a picture (i.e., a current picture) including two or more sub-pictures can have a mixed NAL unit type. In the case of the current VVC standard, a picture with a mixed NAL unit type can have a mixed form of an IRAP NAL unit type and a non-IRAP NAL unit type. However, a leading picture associated with a CRA NAL unit type can also have a mixed form with a non-IRAP NAL unit type, and pictures with such a mixed NAL unit type are not supported under the current standard. Therefore, a solution is needed for pictures with CRA NAL unit types and non-IRAP NAL unit types in a mixed form.
[0215] Therefore, this document provides a method for allowing pictures to include a mixed form of leading picture NAL unit types (e.g., RASL_NUT, RADL_NUT) and other non-IRAP NAL unit types (e.g., TRAIL_NUT, STSA, NUT). In addition, this document defines constraints that allow the presence or signaling of reference picture lists when IDR sub-pictures and other non-IRAP sub-pictures are mixed. Therefore, pictures with mixed NAL unit types are provided with a form that mixes not only IRAP but also CRANAL units, providing more flexible characteristics.
[0216] For example, it can be applied in the following embodiments, thereby solving the above-mentioned problems. The following embodiments can be applied independently or in combination.
[0217] In one embodiment, when pictures are allowed to have mixed NAL unit types (when the value of mixed_nal_types_in_pic_flag is 1), the signaling of the reference picture list allows it to exist even for slices with IDR-type NAL unit types (e.g., IDR_W_RADL or IDR_N_LP). This constraint can be expressed as follows.
[0218] - When there is at least one PPS that references an SPS whose mixed_nal_types_in_pic_flag has the value 1, the value of sps_idr_rpl_present_flag must become 1. This constraint may be a requirement for bitstream conformance.
[0219] Alternatively, in one embodiment, for a picture with mixed NAL unit types, the picture is allowed to include slices with a specific NAL unit type (e.g., RADL or RASL) for the leading picture and a specific NAL unit type (non-IRAP) for the non-leading picture. This can be expressed as follows.
[0220] For the VCL NAL units of a particular picture, the following may apply.
[0221] - When the value of mixed_nalu_types_in_pic_flag is 0, the value of the NAL unit type (nal_unit_type) must be the same for all slice NAL units coded in a picture. A picture or PU can be considered to have the same NAL unit type as the slice NAL units coded in the picture or PU.
[0222] - Otherwise (when the value of mixed_nalu_types_in_pic_flag is 1), one of the following must be satisfied (ie, one of the following may have a value of true).
[0223] 1) One or more VCL NAL units must have a NAL unit type (nal_unit_type) of a specific value in the range from IDR_W_RADL to CRA_NUT (e.g., the values of NAL unit type in Table 1 or Table 2 above are 7 to 9), and all other VCL NAL units must have the same NAL unit type as GRA_NUT or a NAL unit type of a specific value in the range from TRAIL_NUT to RSV_VCL_6 (e.g., the values of NAL unit type in Table 1 or Table 2 above are 0 to 6).
[0224] 2) One or more VCL NAL units must all have the same specific value of NAL unit type as RADL_NUT (e.g., the value of NAL unit type is 2 in Table 1 or Table 2 above) or RASL_NUT (e.g., the value of NAL unit type is 3 in Table 1 or Table 2 above), and all other VCL NAL units must have the same specific value of NAL unit type as TRAIL_NUT (e.g., the value of NAL unit type is 0 in Table 1 or Table 2 above), STSA_NUT (e.g., the value of NAL unit type is 1 in Table 1 or Table 2 above), RSV_VCL_4 (e.g., the value of NAL unit type is 4 in Table 1 or Table 2 above), RSV_VCL_5 (e.g., the value of NAL unit type is 5 in Table 1 or Table 2 above), RSV_VCL_6 (e.g., the value of NAL unit type is 6 in Table 1 or Table 2 above), or GRA_NUT.
[0225] Furthermore, this document proposes a method for providing pictures having the above-mentioned mixed NAL unit types even for a single-layer bitstream. As an embodiment, in the case of a single-layer bitstream, the following constraints may be applied.
[0226] - In the bitstream, every picture except the first picture in decoding order is considered to be associated with the previous IRAP picture in decoding order.
[0227] - If the picture is the leading picture of an IRAP picture, it must be a RADL or RASL picture.
[0228] - If the picture is the trailing picture of an IRAP picture, it must be neither a RADL nor a RASL picture.
[0229] - RASL pictures shall not be present in the bitstream associated with an IDR picture.
[0230] - A RADL picture shall not be present in the bitstream associated with an IDR picture whose NAL unit type (nal_unit_type) is IDR_N_LP.
[0231] When referenced, and when each parameter set is available, random access can be performed at the location of the IRAP PU by discarding all PUs before the IRAP PU (and the IRAP picture and all subsequent non-RASL pictures can be correctly decoded in decoding order).
[0232] - Pictures that precede an IRAP picture in decoding order must precede the IRAP picture in output order, and must precede the RADL picture associated with the IRAP picture in output order.
[0233] - RASL pictures associated with a CRA picture must precede RADL pictures associated with the CRA picture in output order.
[0234] - RASL pictures associated with a CRA picture must follow, in output order, the IRAP pictures that precede the CRA picture in decoding order.
[0235] - If the value of field_seq_flag is 0 and the current picture is a leading picture associated with an IRAP picture, then it must precede all non-leading pictures associated with the same IRAP picture in decoding order. Otherwise, when pictures picA and picB are the first and last leading pictures associated with an IRAP picture in decoding order, respectively, there can be at most one non-leading picture before picA in decoding order, and there must be no non-leading pictures between picA and picB in decoding order.
[0236] The following figures are prepared to illustrate specific examples of this document. Since specific terms or names or names of specific devices (e.g., names of grammar / grammar elements, etc.) described in the figures are presented as examples, the technical features of this document are not limited to the specific names used in the following figures.
[0237] Figure 15 Schematically shows an example of a video / image encoding method to which the embodiments of this document are applicable. Figure 2 The encoding device 200 disclosed in Figure 15 The method disclosed in .
[0238] Reference Figure 15 , the encoding device can determine the NAL unit type of the slice in the picture (S1500).
[0239] For example, the encoding device can determine the NAL unit type according to the properties, types, etc. of the picture or sub-picture as described in Tables 1 and 2 above, and based on the NAL unit type of the picture or sub-picture, the NAL unit type of each slice can be determined.
[0240] For example, when the value of mixed_nalu_types_in_pic_flag is 0, the slices in the picture associated with the PPS can be determined to be of the same NAL unit type. That is, when the value of mixed_nalu_types_in_pic_flag is 0, the NAL unit type defined in the first NAL unit header of the first NAL unit including information about the first slice of the picture is the same as the NAL unit type defined in the second NAL unit header of the second NAL unit including information about the second slice of the same picture. Alternatively, for the case where the value of mixed_nalu_types_in_pic_flag is 1, the slices in the picture associated with the PPS can be determined to be of different NAL unit types. Here, the NAL unit type of the slice in the picture can be determined based on the method proposed in the above embodiment.
[0241] The encoding device may generate NAL unit type related information (S1510). The NAL unit type related information may include information / syntax elements related to the NAL unit types described in the above embodiments and / or Tables 1 and 2 above. For example, the information related to the NAL unit type may include the mixed_nalu_types_in_pic_flag syntax element included in the PPS. In addition, the information related to the NAL unit type may include the nal_unit_type syntax element in the NAL unit header of the NAL unit including information about the coded slice.
[0242] The encoding device may generate a bitstream (S1520). The bitstream may include at least one NAL unit including image information about the coded slice. In addition, the bitstream may include a PPS.
[0243] Figure 16Schematically shows an example of a video / image decoding method applicable to the embodiment of this document. Figure 3 The decoding device 300 disclosed in Figure 16 The method disclosed in .
[0244] Reference Figure 16 , the decoding device may receive a bitstream (S1600). Here, the bitstream may include at least one NAL unit including image information about a coded slice. In addition, the bitstream may include a PPS.
[0245] The decoding device may obtain NAL unit type related information (S1610). The NAL unit type related information may include information / syntax elements related to the NAL unit types described in the above embodiments and / or Tables 1 and 2 above. For example, the information related to the NAL unit type may include the mixed_nalu_types_in_pic_flag syntax element included in the PPS. In addition, the information related to the NAL unit type may include the nal_unit_type syntax element in the NAL unit header of the NAL unit including information about the coded slice.
[0246] The decoding apparatus may determine a NAL unit type of a slice in a picture ( S1620 ).
[0247] For example, when the value of mixed_nalu_types_in_pic_flag is 0, the slices in the picture associated with the PPS use the same NAL unit type. That is, when the value of mixed_nalu_types_in_pic_flag is 0, the NAL unit type defined in the first NAL unit header of the first NAL unit including information about the first slice of the picture is the same as the NAL unit type defined in the second NAL unit header of the second NAL unit including information about the second slice of the same picture. Alternatively, when the value of mixed_nalu_types_in_pic_flag is 1, the slices in the picture associated with the PPS use different NAL unit types. Here, the NAL unit type of the slice in the picture can be determined based on the method proposed in the above embodiment.
[0248] The decoding device may decode / reconstruct the sample / block / slice based on the NAL unit type of the slice (S1630). The sample / block in the slice may be decoded / reconstructed based on the NAL unit type of the slice.
[0249] For example, when a first NAL unit type is set for a first slice of a current picture and a second NAL unit type (different from the first NAL unit type) is set for a second slice of the current picture, samples / blocks in the first slice or the first slice itself can be decoded / reconstructed based on the first NAL unit type, and samples / blocks in the second slice or the second slice itself can be decoded / reconstructed based on the second NAL unit type.
[0250] Figure 17 and Figure 18 An example of a video / image encoding method and associated components according to an embodiment of this document is schematically shown.
[0251] Figure 17 The method disclosed in Figure 2 or Figure 18 Here, Figure 18 The encoding device 200 disclosed in Figure 2 Specifically, Figure 17 Steps S1700 to S1730 can be performed by Figure 2 The entropy encoder 240 disclosed in the embodiment is executed. In addition, according to the embodiment, each step can be performed by Figure 2 The image segmenter 210, the predictor 220, the residual processor 230, the adder 340, etc. disclosed in the embodiment of the present invention are executed. In addition, the embodiment including the embodiment described in this document can be executed. Figure 17 Therefore, in Figure 17 In the embodiment, the detailed description corresponding to the repetition of the above-mentioned embodiments will be omitted or simplified.
[0252] Reference Figure 17 , the encoding device can determine the NAL unit type of the slice in the current picture (S1700).
[0253] The current picture may include multiple slices, and one slice may include a slice header and slice data. In addition, a NAL unit may be generated by adding a NAL unit header to a slice (slice header and slice data). The NAL unit header may include NAL unit type information specified according to the slice data included in the corresponding NAL unit.
[0254] As an embodiment, the encoding device may generate a first NAL unit for a first slice in the current picture and a second NAL unit for a second slice in the current picture. In addition, the encoding device may determine the first NAL unit type of the first slice and the second NAL unit type of the second slice based on the types of the first and second slices.
[0255] For example, based on the type of slice data included in the NAL units shown in Table 1 or Table 2 above, the NAL unit types can include TRAIL_NUT, STSA_NUT, RADL_NUT, RASL_NUT, IDR_W_RADL, IDR_N_LP, CRA_NUT, and the like. Further, the NAL unit types can be signaled based on a nal_unit_type syntax element in the NAL unit header. The nal_unit_type syntax element is syntax information for specifying the NAL unit type, and as shown in Table 1 or Table 2 above, can be represented as a specific value corresponding to a specific NAL unit type.
[0256] The encoding device can determine whether the current picture has mixed NAL unit types based on the NAL unit types (S1710).
[0257] For example, when all NAL unit types of the slices in the current picture are the same, the encoding device can determine that the current picture does not have mixed NAL unit types. Alternatively, when all NAL unit types of the slices in the current picture are not the same, the encoding device can determine that the current picture has mixed NAL unit types.
[0258] The encoding device can generate NAL unit type-related information based on whether the current picture has mixed NAL unit types (S1720).
[0259] The NAL unit type-related information can include information / syntax elements related to the NAL unit types described in the above embodiments and / or Table 1 and Table 2 above. For example, the NAL unit type-related information can be information about whether the current picture has mixed NAL unit types, and can be represented by a mixed_nalu_types_in_pic_flag syntax element included in the PPS. For example, when the value of the mixed_nalu_types_in_pic_flag syntax element is 0, it can be indicated that the NAL units in the current picture have the same NAL unit type. Alternatively, when the value of the mixed_nalu_types_in_pic_flag syntax element is 1, it can be indicated that the NAL units in the current picture have different NAL unit types.
[0260] In one embodiment, when all NAL unit types of slices in a current picture are the same, the encoding device can determine that the current picture does not have mixed NAL unit types, and can generate NAL unit type related information (e.g., mixed_nalu_type_in_pic_flag). In this case, the value of the NAL unit type related information (e.g., mixed_nalu_type_in_pic_flag) can be determined as 0. Alternatively, when NAL unit types of slices in a current picture are not the same, the encoding device can determine that the current picture has mixed NAL unit types, and can generate NAL unit type related information (e.g., mixed_nalu_type_in_pic_flag). In this case, the value of the NAL unit type related information (e.g., mixed_nalu_type_in_pic_flag) can be determined as 1.
[0261] That is, based on the NAL unit type related information about the current picture having mixed NAL unit types (e.g., the value of mixed_nalu_types_in_pic_flag is 1), a first NAL unit of a first slice of the current picture and a second NAL unit of a second slice of the current picture can have different NAL unit types. Alternatively, based on the NAL unit type related information about the current picture not having mixed NAL unit types (e.g., the value of mixed_nalu_types_in_pic_flag is 0), the first NAL unit of the first slice of the current picture and the second NAL unit of the second slice of the current picture can have the same NAL unit type.
[0262] As an example, based on the NAL unit type related information about the current picture having mixed NAL unit types (e.g., the value of mixed_nalu_types_in_pic_flag is 1), the first NAL unit of the first slice can have a leading picture NAL unit type, and the second NAL unit of the second slice can have a non-IRAP NAL unit type or a non-leading picture NAL unit type. Here, the leading picture NAL unit type can include a RADL NAL unit type or a RASL NAL unit type, and the non-IRAP NAL unit type or the non-leading picture NAL unit type can include a trail NAL unit type or a STSA NAL unit type.
[0263] Alternatively, as an example, based on NAL unit type-related information about the current picture having a mixed NAL unit type (e.g., the value of mixed_nalu_types_in_pic_flag is 1), the first NAL unit of the first slice may have an IRAP NAL unit type, and the second NAL unit of the second slice may have a non-IRAP NAL unit type or a non-leading picture NAL unit type. Here, the IRAP NAL unit type may include an IDR NAL unit type (i.e., IDR_N_LPNAL or IDR_W_RADL NAL unit type) or a CRAN NAL unit type, and the non-IRAP NAL unit type or the non-leading picture NAL unit type may include a tracking NAL unit type or an STS NAL unit type. In addition, according to an embodiment, the non-IRAP NAL unit type or the non-leading picture NAL unit type may refer to only the tracking NAL unit type.
[0264] According to an embodiment, based on the case where the current picture is allowed to have mixed NAL unit types, for a slice having an IDR NAL unit type (e.g., IDR_W_RADL or IDR_N_LP) in the current picture, information related to signaling a reference picture list must be present. The information related to signaling a reference picture list may indicate whether a syntax element for signaling a reference picture list is present in the slice header of the slice. That is, if the value of the information related to signaling a reference picture list is 1, the syntax element for signaling a reference picture list may be present in the slice header of the slice having the IDR NAL unit type. Alternatively, if the value of the information related to signaling a reference picture list is 0, the syntax element for signaling a reference picture list may not be present in the slice header of the slice having the IDR NAL unit type.
[0265] For example, the information related to signaling a reference picture list may be the sps_idr_rpl_present_flag syntax element described above. When the value of sps_idr_rpl_present_flag is 1, it may indicate that the syntax element for signaling a reference picture list may be present in the slice header of a slice having a NAL unit type such as IDR_N_LP or IDR_W_RADL. Alternatively, when the value of sps_idr_rpl_present_flag is 0, it may indicate that the syntax element for signaling a reference picture list may not be present in the slice header of a slice having a NAL unit type such as IDR_N_LP or IDR_W_RADL.
[0266] The encoding apparatus may encode image / video information including NAL unit type related information ( S1730 ).
[0267] For example, when the first NAL unit of the first slice in the current picture and the second NAL unit of the second slice in the current picture have different NAL unit types, the encoding device may encode the image / video information including NAL unit type related information (e.g., mixed_nalu_types_in_pic_flag) having a value of 1. Alternatively, when the first NAL unit of the first slice in the current picture and the second NAL unit of the second slice in the current picture have the same NAL unit type, the encoding device may encode the image / video information including NAL unit type related information (e.g., mixed_nalu_types_in_pic_flag) having a value of 0.
[0268] In addition, for example, the encoding device may encode image / video information including nal_unit_type information indicating each NAL unit type of a slice in the current picture.
[0269] In addition, for example, the encoding apparatus may encode image / video information including information related to signaling of a reference picture list (eg, sps_idr_rpl_present_flag).
[0270] In addition, for example, the encoding device may encode image / video information of a NAL unit including a slice in the current picture.
[0271] The image / video information including the various information as described above can be encoded and output in the form of a bit stream. The bit stream can be sent to a decoding device via a network or a (digital) storage medium. Here, the network can include a broadcast network, a communication network and / or the like, and the digital storage medium can include various storage media such as a universal serial bus (USB), secure digital (SD), compact disc (CD), digital video disc (DVD), Blu-ray, hard disk drive (HDD), solid state drive (SSD), etc.
[0272] Figure 19 and Figure 20 An example of a video / image decoding method and associated components according to an embodiment of this document is schematically shown.
[0273] Figure 19 The method disclosed in Figure 3 or Figure 20 The decoding device 300 disclosed in the embodiment is executed. Here, Figure 20 The decoding device 300 disclosed in Figure 3Specifically, Figure 19 Steps S1900 to S1930 can be performed by Figure 3 The entropy decoder 310 disclosed in the embodiment is executed. In addition, according to the embodiment, each step can be performed by Figure 3 The residual processor 320, the predictor 330, the adder 340, etc. disclosed in the embodiment of the present invention are executed. In addition, the embodiments including those described above in this document may be executed. Figure 19 Therefore, in Figure 19 In the embodiment, detailed descriptions of contents corresponding to the repetitions in the above-mentioned embodiments will be omitted or simplified.
[0274] Reference Figure 19 , the decoding device may obtain image / video information including NAL unit type related information from the bitstream (S1900).
[0275] For example, the decoding device can parse the bitstream and derive information required for image reconstruction (or picture reconstruction) (e.g., video / image information). In this case, the image information may include the above-mentioned NAL unit type related information (e.g., mixed_nalu_types_in_pic_flag), nal_unit_type information indicating each NAL unit type of the slice in the current picture, information related to signaling the reference picture list (e.g., sps_idr_rpl_present_flag), NAL units of slices in the current picture, etc. That is, the image information may include various information required in the decoding process and may be decoded based on a coding method such as exponential Golomb coding, CAVLC, or CABAC.
[0276] As described above, the NAL unit type related information may include information / syntax elements related to the NAL unit types described in the above embodiments and / or Tables 1 and 2 above. For example, the NAL unit type related information may be information about whether the current picture has mixed NAL unit types, and may be represented by the mixed_nalu_types_in_pic_flag syntax element included in the PPS. For example, when the value of the mixed_nalu_types_in_pic_flag syntax element is 0, it may indicate that the NAL units in the current picture have the same NAL unit type. Alternatively, when the value of the mixed_nalu_types_in_pic_flag syntax element is 1, it may indicate that the NAL units in the current picture have different NAL unit types.
[0277] The decoding apparatus may determine whether the current picture has a mixed NAL unit type based on NAL unit type related information ( S1910 ).
[0278] For example, the decoding device can determine that the current picture has mixed NAL unit types based on a value of the NAL unit type related information (e.g., mixed_nalu_type_in_pic_flag) being 1. Alternatively, the decoding device can determine that the current picture does not have mixed NAL unit types based on a value of the NAL unit type related information (e.g., mixed_nalu_type_in_pic_flag) being 0.
[0279] The decoding device can determine a NAL unit type of a slice in the current picture based on the NAL unit type related information regarding whether the current picture has mixed NAL unit types (S1920).
[0280] The current picture can include a plurality of slices, and one slice can include a slice header and slice data. In addition, a NAL unit can be generated by adding a NAL unit header to the slice (slice header and slice data). The NAL unit header can include NAL unit type information designated according to slice data included in the corresponding NAL unit.
[0281] For example, based on a type of slice data included in a NAL unit as shown in Table 1 or Table 2 above, the NAL unit type can include TRAIL_NUT, STSA_NUT, RADL_NUT, RASL_NUT, IDR_W_RADL, IDR_N_LP, CRA_NUT, etc. In addition, the NAL unit type can be signaled based on a nal_unit_type syntax element in the NAL unit header. The nal_unit_type syntax element is syntax information for designating the NAL unit type, and as shown in Table 1 or Table 2 above, can be expressed as a specific value corresponding to a specific NAL unit type.
[0282] In an embodiment, the decoding device can determine that a first NAL unit of a first slice of the current picture and a second NAL unit of a second slice of the current picture can have different NAL unit types based on the NAL unit type related information regarding the current picture having mixed NAL unit types (e.g., a value of mixed_nalu_types_in_pic_flag being 1). Alternatively, the decoding device can determine that the first NAL unit of the first slice of the current picture and the second NAL unit of the second slice of the current picture can have the same NAL unit type based on the NAL unit type related information regarding the current picture not having mixed NAL unit types (e.g., a value of mixed_nalu_types_in_pic_flag being 0).
[0283] As an example, based on the NAL unit type related information about the current picture having a mixed NAL unit type (for example, the value of mixed_nalu_types_in_pic_flag is 1), the first NAL unit of the first slice may have a leading picture NAL unit type, and the second NAL unit of the second slice may have a non-IRAP NAL unit type or a non-leading picture NAL unit type. Here, the leading picture NAL unit type may include a RADL NAL unit type or a RASL NAL unit type, and the non-IRAP NAL unit type or the non-leading picture NAL unit type may include a tracking NAL unit type or an STS NAL unit type.
[0284] Alternatively, as an example, based on NAL unit type-related information about the current picture having a mixed NAL unit type (e.g., the value of mixed_nalu_types_in_pic_flag is 1), the first NAL unit of the first slice may have an IRAP NAL unit type, and the second NAL unit of the second slice may have a non-IRAP NAL unit type or a non-leading picture NAL unit type. Here, the IRAP NAL unit type may include an IDR NAL unit type (i.e., IDR_N_LPNAL or IDR_W_RADL NAL unit type) or a CRAN NAL unit type, and the non-IRAP NAL unit type or the non-leading picture NAL unit type may include a tracking NAL unit type or an STS NAL unit type. In addition, according to an embodiment, the non-IRAP NAL unit type or the non-leading picture NAL unit type may refer to only the tracking NAL unit type.
[0285] According to an embodiment, based on the case where the current picture is allowed to have mixed NAL unit types, for a slice having an IDR NAL unit type (e.g., IDR_W_RADL or IDR_N_LP) in the current picture, information related to signaling a reference picture list must be present. The information related to signaling a reference picture list may indicate whether a syntax element for signaling a reference picture list is present in the slice header of the slice. That is, if the value of the information related to signaling a reference picture list is 1, the syntax element for signaling a reference picture list may be present in the slice header of the slice having the IDR NAL unit type. Alternatively, if the value of the information related to signaling a reference picture list is 0, the syntax element for signaling a reference picture list may not be present in the slice header of the slice having the IDR NAL unit type.
[0286] For example, the information related to signaling a reference picture list may be the sps_idr_rpl_present_flag syntax element described above. When the value of sps_idr_rpl_present_flag is 1, it may indicate that the syntax element for signaling a reference picture list may be present in the slice header of a slice having a NAL unit type such as IDR_N_LP or IDR_W_RADL. Alternatively, when the value of sps_idr_rpl_present_flag is 0, it may indicate that the syntax element for signaling a reference picture list may not be present in the slice header of a slice having a NAL unit type such as IDR_N_LP or IDR_W_RADL.
[0287] The decoding device may decode / reconstruct the current picture based on the NAL unit type ( S1930 ).
[0288] For example, for a first slice in a current picture determined to be of a first NAL unit type and a second slice in the current picture determined to be of a second NAL unit type, the decoding device may decode / restore the first slice based on the first NAL unit type and decode / reconstruct the second slice based on the second NAL unit type. In addition, the decoding device may decode / reconstruct the sample / block in the first slice based on the first NAL unit type and decode / reconstruct the sample / block in the second slice based on the second NAL unit type.
[0289] Although the method has been described based on a flowchart that lists steps or blocks in sequence in the above embodiments, the steps of this document are not limited to a specific order, and specific steps can be performed in different steps or in a different order or simultaneously with respect to the above steps. In addition, it will be understood by those skilled in the art that the steps in the flowchart are not exclusive, and another step can be included therein, or one or more steps in the flowchart can be deleted without affecting the scope of the present disclosure.
[0290] The above-mentioned method according to the present disclosure may be in the form of software, and the encoding device and / or decoding device according to the present disclosure may be included in an apparatus for performing image processing (e.g., TV, computer, smart phone, set-top box, display device, etc.).
[0291] When the embodiments of the present disclosure are implemented with software, the above-mentioned methods can be implemented with modules (processing or functions) that perform the above-mentioned functions. The modules can be stored in a memory and executed by a processor. The memory can be installed inside or outside the processor and can be connected to the processor via various well-known devices. The processor may include an application-specific integrated circuit (ASIC), other chipsets, logic circuits and / or data processing devices. The memory may include a read-only memory (ROM), a random access memory (RAM), a flash memory, a memory card, a storage medium and / or other storage devices. In other words, according to the embodiments of the present disclosure, it can be implemented and executed on a processor, a microprocessor, a controller or a chip. For example, the functional units illustrated in the corresponding figures can be implemented and executed on a computer, a processor, a microprocessor, a controller or a chip. In this case, information about the implementation (for example, information about instructions) or the algorithm can be stored in a digital storage medium.
[0292] In addition, the decoding device and encoding device of the embodiment of the present document can be included in a multimedia broadcast transceiver, a mobile communication terminal, a home theater video device, a digital theater video device, a surveillance camera, a video chat device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a portable camera, a video on demand (VoD) service provider, an over-the-top (OTT) video device, an Internet streaming service provider, a 3D video device, a virtual reality (VR) device, an augmented reality (AR) device, an image phone video device, a vehicle terminal (e.g., a vehicle (including an autonomous vehicle) terminal, an aircraft terminal, or a ship terminal) and a medical video device; and can be used to process image signals or data. For example, the OTT video device may include a game console, a Blueray player, a networked TV, a home theater system, a smartphone, a tablet PC, and a digital video recorder (DVR).
[0293] In addition, the processing method of the embodiment of the application of this document can be generated in the form of a program executed by a computer and can be stored in a computer-readable recording medium. The multimedia data with a data structure according to the embodiment of this document can also be stored in a computer-readable recording medium. The computer-readable recording medium includes all kinds of storage devices and distributed storage devices stored with computer-readable data. The computer-readable recording medium may include, for example, a Blu-ray disc (BD), a universal serial bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. The computer-readable recording medium also includes a medium implemented in the form of a carrier wave (for example, transmission on the Internet). In addition, the bit stream generated by the encoding method can be stored in a computer-readable recording medium, or can be transmitted through a wired or wireless communication network.
[0294] In addition, the embodiments of this document can be implemented as a computer program product based on a program code, and the program code can be executed on a computer according to the embodiments of this document. The program code can be stored on a computer-readable carrier.
[0295] Figure 21 This shows an example of a content streaming system to which the embodiments of this document can be applied.
[0296] Reference Figure 21 A content streaming system to which embodiments of this document are applied may generally include an encoding server, a streaming server, a network server, a media storage, a user device, and a multimedia input device.
[0297] The encoding server is used to compress content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data, generate a bitstream, and transmit it to the streaming server. As another example, if the multimedia input device such as smartphones, cameras, and camcorders directly generates the bitstream, the encoding server can be omitted.
[0298] The bitstream may be generated by the encoding method or the bitstream generation method to which the embodiments of this document are applied, and the streaming server may temporarily store the bitstream in the process of transmitting or receiving the bitstream.
[0299] The streaming server transmits multimedia data to the user device via a network server based on the user's request. The network server serves as a tool for notifying the user of available services. When the user requests a desired service, the network server transfers the request to the streaming server, which then transmits the multimedia data to the user. In this regard, the content streaming system may include a separate control server, in which case the control server is used to control commands and responses between the various devices in the content streaming system.
[0300] The streaming server may receive content from a media storage device and / or an encoding server. For example, when receiving content from an encoding server, the content may be received in real time. In this case, the streaming server may store the bitstream for a predetermined period of time to smoothly provide a streaming service.
[0301] For example, user devices may include mobile phones, smart phones, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation, tablet PCs, tablet PCs, ultrabooks, wearable devices (e.g., watch-type terminals (smart watches), glasses-type terminals (smart glasses), head-mounted displays (HMDs)), digital TVs, desktop computers, digital signage, etc.
[0302] Each server in the content streaming system can be operated as a distributed server, and in this case, data received by each server can be processed in a distributed manner.
[0303] The claims in this specification can be combined in various ways. For example, the technical features in the method claims of this specification can be combined to be implemented or performed in an apparatus, and the technical features in the apparatus claims can be combined to be implemented or performed in a method. Also, the technical features in the method claims and the apparatus claims can be combined to be implemented or performed in an apparatus. Also, the technical features in the method claims and the apparatus claims can be combined to be implemented or performed in a method.
Claims
1. An image decoding method performed by a decoding device, the image decoding method comprising: Obtain image information including network abstraction layer NAL unit type related information from the bitstream; Determining whether the current picture has a mixed NAL unit type based on the NAL unit type related information; determining a NAL unit type for a slice in the current picture based on the NAL unit type related information for the current picture having the mixed NAL unit type; as well as decoding the current picture based on the NAL unit type, Wherein, based on the value of the NAL unit type related information being equal to 1, it is determined that the current picture has the mixed NAL unit type, wherein, based on the NAL unit type for the first slice in the current picture being an instantaneous decoding refresh (IDR) NAL unit type, for the first slice in the current picture, the value of the information related to signaling a reference picture list is equal to 1, and Wherein, based on the value of the NAL unit type related information being equal to 1, the NAL unit type for the first slice in the current picture is different from the NAL unit type for the second slice in the current picture.
2. An image encoding method performed by an encoding device, the image encoding method comprising: Determine the NAL unit type for the slice in the current picture; determining, based on the NAL unit type, whether the current picture has a mixed NAL unit type; generating NAL unit type related information based on whether the current picture has the mixed NAL unit type; as well as Encode the image information including the NAL unit type related information, Wherein, based on the current picture having the mixed NAL unit type, the value of the NAL unit type related information is determined to be 1, wherein, based on the NAL unit type for the first slice in the current picture being an instantaneous decoding refresh (IDR) NAL unit type, for the first slice in the current picture, the value of the information related to signaling a reference picture list is equal to 1, and Wherein, based on the value of the NAL unit type related information being equal to 1, the NAL unit type for the first slice in the current picture is different from the NAL unit type for the second slice in the current picture.
3. A method for transmitting image data, the method comprising the following steps: Obtaining a bitstream for the image, wherein the bitstream is generated based on the following operations: determining a NAL unit type for a slice in a current picture, determining whether the current picture has a mixed NAL unit type based on the NAL unit type, generating NAL unit type related information based on whether the current picture has the mixed NAL unit type, and encoding image information including the NAL unit type related information; and sending said data comprising said bitstream, Wherein, based on the current picture having the mixed NAL unit type, the value of the NAL unit type related information is determined to be 1, wherein, based on the NAL unit type for the first slice in the current picture being an instantaneous decoding refresh (IDR) NAL unit type, for the first slice in the current picture, the value of the information related to signaling a reference picture list is equal to 1, and Wherein, based on the value of the NAL unit type related information being equal to 1, the NAL unit type for the first slice in the current picture is different from the NAL unit type for the second slice in the current picture.