Video or video coding based on NAL unit-related information.
By determining NAL unit types for video slices and signaling reference picture lists based on NAL unit type-related information, the method enhances video compression efficiency, especially for mixed NAL unit types, addressing the challenges of high-resolution and high-quality video coding.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- LG ELECTRONICS INC
- Filing Date
- 2025-05-15
- Publication Date
- 2026-05-15
AI Technical Summary
The increasing demand for high-resolution and high-quality video, including VR and AR content, has led to a need for more efficient video compression technologies, particularly in signaling and coding information related to the Network Abstraction Layer (NAL) units, to manage the increased data volume effectively.
The method determines NAL unit types for slices within a picture based on NAL unit type-related information, allowing for flexible signaling of reference picture lists, especially for pictures with mixed NAL unit types, and includes encoding and decoding devices to enhance video coding efficiency.
This approach improves overall video compression efficiency, particularly for pictures with mixed NAL unit types, enabling effective signaling and coding of reference picture lists, thus providing more flexible characteristics.
Smart Images

Figure 0007860310000004 
Figure 0007860310000005 
Figure 0007860310000006
Abstract
Description
Technical Field
[0001] This technology relates to video or video coding, for example, to coding techniques based on NAL (network abstraction layer) unit related information.
Background Art
[0002] In recent years, the demand for high-resolution and high-quality video / video such as 4K or UHD (Ultra High Definition) video / video of 8K or higher has been increasing in various fields. As the video / video data becomes higher in resolution and quality, the amount of information or bits transmitted relatively increases compared to the existing video / video data. Therefore, when transmitting video data using a medium such as an existing wired or wireless broadband line or storing video / video data using an existing storage medium, the transmission cost and storage cost increase.
[0003] Also, in recent years, the interest and demand for immersive media such as VR (Virtual Reality), AR (Artificial Reality) content, and holograms have been increasing, and the broadcast of video / video having video characteristics different from real-world video, such as game video, has been increasing.
[0004] Thus, there is a need for a highly efficient video / video compression technology to effectively compress, transmit, store, and reproduce the information of high-resolution and high-quality video / video having various characteristics as described above.
[0005] Also, a solution for improving the efficiency of video / video coding is necessary. Therefore, a solution for effectively signaling and coding information related to the NAL (network abstraction layer) unit is necessary.
Summary of the Invention
Problems to be Solved by the Invention
[0006] The technical objective of this document is to provide a method and apparatus for improving the efficiency of video / image coding.
[0007] Another technical objective of this document is to provide a method and apparatus for improving the efficiency of video / image coding based on NAL unit-related information.
[0008] Another technical challenge of this paper is to provide a method and apparatus for improving the efficiency of video / image coding for pictures having mixed NAL unit types.
[0009] Another technical challenge of this document is to provide a method and apparatus for making reference picture list-related information present or signaled for slices in a picture having a specific NAL unit type, for pictures having mixed NAL unit types. [Means for solving the problem]
[0010] According to one embodiment of this document, the NAL unit type for a slice within a picture can be determined based on NAL unit type-related information regarding whether the picture has a mixed NAL unit type. For example, based on NAL unit type-related information regarding whether the picture has a mixed NAL unit type, the first NAL unit for the first slice of the picture and the second NAL unit for the second slice of the picture may have different NAL unit types. Alternatively, based on NAL unit type-related information regarding whether the picture does not have a mixed NAL unit type, the first NAL unit for the first slice of the picture and the second NAL unit for the second slice of the picture may have the same NAL unit type.
[0011] According to one embodiment of this document, based on the case where a picture is allowed to have mixed NAL unit types, signaling-related information of a reference picture list can be made available for slices having a specific NAL unit type within a picture.
[0012] According to one embodiment of this document, a method for decoding video / images performed by a decoding device is provided. The video / image decoding method may include the methods disclosed in the embodiments of this document.
[0013] According to one embodiment of this document, a decoding device is provided for performing video / image decoding. The decoding device can perform the methods disclosed in the embodiments of this document.
[0014] According to one embodiment of this document, a method for encoding video / images performed by an encoding device is provided. The video / image encoding method may include the methods disclosed in the embodiments of this document.
[0015] According to one embodiment of this document, an encoding device for performing video / image encoding is provided. The encoding device can perform the method disclosed in the embodiments of this document.
[0016] According to one embodiment of this document, a computer-readable digital storage medium is provided which stores encoded video / image information generated by a video / image encoding method disclosed in at least one embodiment of this document.
[0017] According to one embodiment of this document, a computer-readable digital storage medium is provided which stores encoded information or encoded video / image information that causes a decoding device to perform a video / image decoding method disclosed in at least one embodiment of this document. [Effects of the Invention]
[0018] This document can have a variety of effects. For example, according to one embodiment of this document, the overall video compression efficiency can be improved. Also, according to one embodiment of this document, the efficiency of video coding can be improved based on NAL unit-related information. Furthermore, according to one embodiment of this document, the efficiency of video coding for pictures having mixed NAL unit types can be improved. Also, according to one embodiment of this document, reference picture list-related information can be effectively signaled and coded for pictures having mixed NAL unit types. Furthermore, according to one embodiment of this document, by allowing pictures containing a mixed form of leading picture NAL unit types (e.g., RASL_NUT, RADL_NUT) and other non-IRAP NAL unit types (e.g., TRAIL_NUT, STSA, NUT), pictures having mixed NAL unit types can be provided with a mixed form not only with IRAP but also with other types of NAL units, thereby providing more flexible characteristics.
[0019] The effects obtained through the specific embodiments of this document are not limited to those listed above. For example, there may be a variety of technical effects that a person having ordinary skill in the related art can understand or derive from this document. Thus, the specific effects of this document are not limited to those explicitly stated herein, but may include a variety of effects that can be understood or derive from the technical features of this document. [Brief explanation of the drawing]
[0020] [Figure 1] An example of a video / image coding system that can be applied to the embodiments described herein is schematically shown. [Figure 2] This diagram schematically illustrates the configuration of a video / image encoding device that can be applied to the embodiments described herein. [Figure 3] A drawing schematically explaining the configuration of a video / video decoding apparatus applicable to an embodiment of this document. [Figure 4] An example of a schematic video / video encoding procedure applicable to an embodiment of this document is shown. [Figure 5] An example of a schematic video / video decoding procedure applicable to an embodiment of this document is shown. [Figure 6] An example of an entropy encoding method applicable to an embodiment of this document is schematically shown. [Figure 7] The entropy encoding section in the encoding apparatus is schematically shown. [Figure 8] An example of an entropy decoding method applicable to an embodiment of this document is schematically shown. [Figure 9] The entropy decoding section in the decoding apparatus is schematically shown. [Figure 10] A hierarchical structure for the coded video / video is exemplarily shown. [Figure 11] A diagram showing the temporal layer structure for NAL units in a bitstream supporting temporal scalability. [Figure 12] A diagram for explaining a picture capable of random access. [Figure 13] A diagram for explaining an IDR picture. [Figure 14] A diagram for explaining a CRA picture. [Figure 15] An example of a video / video encoding method applicable to an embodiment of this document is schematically shown. [Figure 16] An example of a video / video decoding method applicable to an embodiment of this document is schematically shown. [Figure 17] An example of a video / video encoding method and related components according to an embodiment(s) of this document is schematically shown. [Figure 18] An example of a video / video encoding method and related components according to an embodiment(s) of this document is schematically shown. [Figure 19] An example of a video decoding method and related components according to the embodiments (e) of this document is schematically shown. [Figure 20] An example of a video decoding method and related components according to the embodiments (e) of this document is schematically shown. [Figure 21] Examples of content streaming systems to which the embodiments disclosed in this document can be applied are shown. [Modes for carrying out the invention]
[0021] This document may be modified in various ways and may have various embodiments, and specific embodiments are illustrated in the drawings and described in detail. However, this is not intended to limit this document to any particular embodiment. Terms used herein are used solely to describe specific embodiments and are not intended to limit the technical ideas of this document. Singular expressions include plural expressions unless the context clearly indicates otherwise. Terms such as “includes” or “has” herein are intended to specify the existence of features, figures, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood not to preemptively exclude the existence or possibility of adding one or more other features, figures, steps, actions, components, parts, or combinations thereof.
[0022] On the other hand, each configuration shown in the diagrams described in this document is illustrated independently for the purpose of explaining its distinct characteristic functions, and does not mean that each configuration is embodied in separate hardware or separate software. For example, two or more of the configurations may be combined to form a single configuration, and one configuration may be divided into multiple configurations. Embodiments in which each configuration is integrated and / or separated are also included in the scope of the rights of this document, as long as they do not deviate from the essence of this document.
[0023] In this document, "A or B" may mean "just A," "just B," or "both A and B." In other words, in this document, "A or B" may be interpreted as "A and / or B." For example, in this document, "A, B or C" may mean "just A," "just B," "just C," or "any combination of A, B and C."
[0024] The slashes ( / ) and commas used in this document can mean "and / or". For example, "A / B" can mean "A and / or B". Thus, "A / B" can mean "just A", "just B", or "both A and B". For example, "A, B, C" can mean "A, B or C".
[0025] In this document, "at least one of A and B" may mean "just A," "just B," or "both A and B." Furthermore, in this document, the expressions "at least one of A or B" and "at least one of A and / or B" may be interpreted similarly to "at least one of A and B."
[0026] Furthermore, in this document, "at least one of A, B and C" may mean "just A," "just B," "just C," or "any combination of A, B and C." Also, "at least one of A, B or C" or "at least one of A, B and / or C" may mean "at least one of A, B and C."
[0027] Furthermore, parentheses used in this document may mean "for example." Specifically, when "prediction (intra prediction)" is displayed, "intra prediction" may be proposed as an example of "prediction." In other words, "prediction" in this document is not limited to "intra prediction," and "intra prediction" may be proposed as an example of "prediction." Also, when "prediction (i.e., intra prediction)" is displayed, "intra prediction" may be proposed as an example of "prediction."
[0028] This document relates to video / image coding. For example, the methods / examples disclosed in this document can be applied to methods disclosed in the VVC (versatile video coding) standard. Furthermore, the methods / examples disclosed in this document can be applied to methods disclosed in the EVC (essential video coding) standard, AV1 (AOMedia Video 1) standard, AVS2 (2nd generation of audio video coding standard), or next-generation video / image coding standards (e.g., 267 or H.268).
[0029] This document presents various examples of video / image coding, and unless otherwise noted, these examples may be combined with each other.
[0030] In this document, video may mean a collection of images over time. Picture generally refers to a unit representing a single image at a specific time point in time, and slice / tile is a unit that constitutes part of a picture in coding. A slice / tile may contain one or more CTUs (coding tree units). A single picture may consist of one or more slices / tiles. A tile is a row of tiles within a picture and a rectangular area of CTUs within a row of tiles. The row of tiles is a rectangular area of CTUs, and the rectangular area has the same height as the height of the picture, and its width may be specified by syntax elements in the picture parameter set. The row of tiles is a rectangular area of CTUs, and the rectangular area has the width specified by syntax elements in the picture parameter set, and its height may be the same as the height of the picture. A tile scan may represent a specific sequential ordering of CTUs partitioning a picture, where the CTUs may be aligned consecutively by a raster scan of the CTUs within the tile, and the tiles within a picture may be aligned consecutively by a raster scan of the tiles within the picture. A slice may exclusively contain an integer number of complete tiles or an integer number of consecutive complete rows of CTUs within the tiles of a picture.
[0031] On the other hand, a single picture can be divided into two or more subpictures. A subpicture can be a rectangular region of one or more slices within the picture.
[0032] A pixel or pel can refer to the smallest unit that makes up a picture (or image). Alternatively, the term "sample" may be used as a counterpart to pixel. A sample generally refers to a pixel or a pixel value, sometimes representing only the luma component pixel / pixel value, or sometimes only the chroma component pixel / pixel value. Alternatively, a sample may refer to a pixel value in the spatial domain, and if such a pixel value is converted to the frequency domain, it may refer to the conversion coefficient in the frequency domain.
[0033] A unit can represent a basic unit of image processing. A unit may include a specific region of a picture and at least one piece of information about that region. A unit may include one luma block and two chroma (e.g., cb, cr) blocks. The term unit may sometimes be used interchangeably with terms such as block or area. In general, an M×N block may include a set (or array) of samples (or sample arrays) or transform coefficients consisting of M columns and N rows.
[0034] Furthermore, in this document, at least one of quantization / inverse quantization and / or transformation / inverse transformation may be omitted. If quantization / inverse quantization is omitted, the quantized transformation coefficients may be called transformation coefficients. If transformation / inverse transformation is omitted, the transformation coefficients may be called coefficients or residual coefficients, or for consistency of expression, they may still be called transformation coefficients.
[0035] In this document, quantized transformation coefficients and transformation coefficients may be referred to as transformation coefficients and scaled transformation coefficients, respectively. In this case, residual information may include information about the transformation coefficients, and information about the transformation coefficients may be signaled via residual coding syntax. Transformation coefficients may be derived based on residual information (or information about the transformation coefficients), and scaled transformation coefficients may be derived via inverse transformation (scaling) of the transformation coefficients. Residual samples may be derived based on inverse transformation (transformation) of the scaled transformation coefficients. This may be applied / expressed similarly in other parts of this document.
[0036] In this document, technical features described individually within a single drawing may be implemented individually or simultaneously.
[0037] The preferred embodiments of this document will be described in more detail below with reference to the attached drawings. The same reference numerals will be used for the same components in the drawings, and redundant descriptions of the same components may be omitted.
[0038] Figure 1 schematically shows an example of a video / image coding system that can be applied to the embodiments described herein.
[0039] As shown in Figure 1, a video / image coding system may comprise a first device (source device) and a second device (receiving device). The source device can transmit encoded video / image information or data to the receiving device in file or streaming form via a digital storage medium or network.
[0040] The source device may comprise a video source, an encoding device, and a transmitter. The receiving device may comprise a receiver, a decoding device, and a renderer. The encoding device may be called a video / image encoding device, and the decoding device may be called a video / image decoding device. The transmitter may be provided in the encoding device. The receiver may be provided in the decoding device. The renderer may comprise a display unit, which may consist of a separate device or external component.
[0041] A video source can acquire video / images through processes such as video / image capture, synthesis, or generation. A video source may include video / image capture devices and / or video / image generation devices. Video / image capture devices may include, for example, one or more cameras, or video / image archives containing previously captured video / images. Video / image generation devices may include, for example, computers, tablets, and smartphones, and can generate video / images (electronically). For example, virtual video / images may be generated via a computer, in which case the video / image capture process may be replaced by the process of generating the associated data.
[0042] An encoding device can encode input video / image data. For compression and coding efficiency, the encoding device can perform a series of steps, including prediction, transformation, and quantization. The encoded data (encoded video / image information) can be output in bitstream format.
[0043] The transmitting unit can transmit encoded video / image information or data output in bitstream format to the receiving unit of a receiving device via a digital storage medium or network in file or streaming format. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit may include elements for generating media files via a predetermined file format and may include elements for transmission via a broadcast / communication network. The receiving unit can receive / extract the bitstream and transmit it to a decoding device.
[0044] A decoding device can decode video / images by performing a series of steps, such as inverse quantization, inverse transformation, and prediction, corresponding to the operation of an encoding device.
[0045] The renderer can render the decoded video / image. The rendered video / image can be displayed via the display unit.
[0046] Figure 2 is a schematic diagram illustrating the configuration of a video / image encoding device to which the embodiments described in this document can be applied. Hereinafter, the term "encoding device" may include an image encoding device and / or a video encoding device.
[0047] As shown in Figure 2, the encoding device 200 can be configured to include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-predictor 221 and an intra-predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be called a reconstructor or a reconstructed block generator. The aforementioned video splitting unit 210, prediction unit 220, residual processing unit 230, entropy encoding unit 240, addition unit 250, and filtering unit 260 can be configured by one or more hardware components (e.g., an encoder chipset or processor) depending on the embodiment. The memory 270 may also include a DPB (decoded picture buffer) and may be configured by a digital storage medium. The hardware components may further include the memory 270 as an internal / external component.
[0048] The video splitting unit 210 can split the input video (or picture, frame) input to the encoding device 200 into one or more processing units. For example, the processing units may be called coding units (CUs). In this case, the coding units can be recursively split from a coding tree unit (CTU) or the largest coding unit (LCU) using a QTBTTT (Quad-tree binary-tree ternary-tree) structure. For example, one coding unit can be split into multiple coding units of deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first, followed by the binary-tree structure and / or the ternary structure. Alternatively, the binary-tree structure may be applied first. The coding procedure according to this disclosure may be performed based on the final coding unit that is not further split. In this case, based on coding efficiency due to video characteristics, the largest coding unit can be used as the final coding unit, or, if necessary, the coding unit can be recursively divided into lower-depth coding units so that the optimally sized coding unit is used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration, which will be described later. As another example, the processing unit may further include a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit can each be separated or partitioned from the final coding unit described above.The prediction unit may be a unit of sample prediction, and the conversion unit may be a unit for deriving conversion coefficients and / or a unit for deriving a residual signal from conversion coefficients.
[0049] The term "unit" can sometimes be used interchangeably with terms such as "block" or "area." Generally, an M×N block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and may represent only the luminance (luma) component pixel / pixel value, or only the chroma component pixel / pixel value. A sample can be used to refer to a single picture (or image) as a pixel or pel.
[0050] The encoding device 200 can generate a residual signal (residual block, residual sample array) by subtracting the prediction signal (predicted block, predicted sample array) output from the inter-prediction unit 221 or intra-prediction unit 222 from the input video signal (original block, original sample array), and the generated residual signal is transmitted to the conversion unit 232. In this case, as shown in the figure, the unit that subtracts the prediction signal (predicted block, predicted sample array) from the input video signal (original block, original sample array) within the encoder 200 can be called the subtraction unit 231. The prediction unit can make predictions for the block to be processed (hereinafter referred to as the current block) and generate a predicted block that includes predicted samples for the current block. The prediction unit can determine whether intra-prediction or inter-prediction is applied on a current block or CU basis. The prediction unit can generate various prediction-related information, such as prediction mode information, and transmit it to the entropy encoding unit 240, as will be described later in the explanation of each prediction mode. Prediction information can be encoded by the entropy encoding unit 240 and output in bitstream format.
[0051] The intra-prediction unit 222 can predict the current block by referring to a sample in the current picture. The referenced sample may be located adjacent to the current block or at a distance, depending on the prediction mode. The prediction mode in intra-prediction may include multiple non-directional modes and multiple directional modes. Non-directional modes may include, for example, DC mode and planar mode. Directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes, depending on the degree of fineness of the prediction direction. However, this is illustrative, and more or fewer directional prediction modes may be used depending on the settings. The intra-prediction unit 222 may also determine the prediction mode to apply to the current block using the prediction modes applied to adjacent blocks.
[0052] The interprediction unit 221 can derive a predicted block relative to the current block based on a reference block (reference sample array) identified by motion vectors on the reference picture. In this case, in order to reduce the amount of motion information transmitted in interprediction mode, motion information can be predicted in units of blocks, subblocks, or samples based on the correlation of motion information between adjacent blocks and the current block. The motion information may include motion vectors and reference picture indices. The motion information may further include interprediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of interprediction, adjacent blocks may include spatial neighboring blocks that exist in the current picture and temporal neighboring blocks that exist in the reference picture. The reference picture containing the reference block and the reference picture containing the temporal neighboring block may be the same or different. The temporal neighboring block may be called a collocated reference block, col CU, etc., and the reference picture containing the temporal neighboring block may be called a collocated picture (colPic). For example, the inter-prediction unit 221 can construct a motion information candidate list based on adjacent blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter-prediction can be performed based on various prediction modes; for example, in skip mode and merge mode, the inter-prediction unit 221 can use the motion information of adjacent blocks as the motion information of the current block. In skip mode, unlike merge mode, a residual signal may not be transmitted.In motion vector prediction (MVP) mode, the motion vector of an adjacent block is used as a motion vector predictor, and the motion vector difference is signaled to indicate the motion vector of the current block.
[0053] The prediction unit 220 can generate prediction signals based on various prediction methods described later. For example, the prediction unit can apply intra-prediction or inter-prediction for predictions on a single block, and can also apply intra-prediction and inter-prediction simultaneously. This can be called combined inter and intra prediction (CIIP). The prediction unit can also base its predictions on an intra-block copy (IBC) prediction mode or a palette mode for predictions on a block. The IBC prediction mode or palette mode can be used for content video / video coding such as in games, for example, in SCC (screen content coding). IBC basically performs predictions within the current picture, but can be performed similarly to inter-prediction in that it derives reference blocks within the current picture. That is, IBC can use at least one of the inter-prediction techniques described in this document. Palette mode can be considered an example of intra-coding or intra-prediction. When palette mode is applied, sample values within the picture can be signaled based on information about the palette table and palette index.
[0054] The prediction signal generated via the prediction unit (including the inter-prediction unit 221 and / or the intra-prediction unit 222) may be used to generate a reconstructed signal or to generate a residual signal. The transformation unit 232 can apply a transformation technique to the residual signal to generate transformation coefficients. For example, the transformation technique may include at least one of DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), KLT (Karhunen-Loeve Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT means a transformation obtained from a graph when relational information between pixels is represented by a graph. CNT means a transformation obtained by generating a prediction signal using all previously reconstructed pixels and obtaining a transformation based on it. The transformation process may also be applied to pixel blocks of the same size that are square, or to non-square blocks of variable size.
[0055] The quantization unit 233 quantizes the conversion coefficients and transmits them to the entropy encoding unit 240, which can encode the quantized signal (information about the quantized conversion coefficients) and output it as a bitstream. The information about the quantized conversion coefficients can be called residual information. The quantization unit 233 can rearrange the block-form quantized conversion coefficients into a one-dimensional vector form based on the coefficient scan order, and can also generate information about the quantized conversion coefficients based on the one-dimensional vector form of the quantized conversion coefficients. The entropy encoding unit 240 can perform various encoding methods, such as exponential Golomb, CAVLC (context-adaptive variable length coding), and CABAC (context-adaptive binary arithmetic coding). In addition to the quantized conversion coefficients, the entropy encoding unit 240 can also encode information necessary for video / image restoration (e.g., the values of syntax elements) together with or separately from the quantized conversion coefficients. Encoded information (e.g., encoded video / image information) can be transmitted or stored in bitstream form in units of NAL (network abstraction layer) units. The video / image information may further include information about various parameter sets, such as the adaptation parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS). The video / image information may also further include general constraint information. In this document, information and / or syntax elements transmitted / signaled from the encoding device to the decoding device may be included in the video / image information.The video / image information can be encoded via the encoding procedure described above and included in the bitstream. The bitstream can be transmitted over a network or stored in a digital storage medium. Here, the network can include broadcast networks and / or communication networks, and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The signal output from the entropy encoding unit 240 can be transmitted by a transmitting unit (not shown) and / or stored by a storage unit (not shown) which are configured as internal / external elements of the encoding device 200, or the transmitting unit may be included in the entropy encoding unit 240.
[0056] The quantized conversion coefficients output from the quantization unit 233 can be used to generate a prediction signal. For example, a residual signal (residual block or residual sample) can be reconstructed by applying inverse quantization and inverse transformation to the quantized conversion coefficients via the inverse quantization unit 234 and the inverse transformation unit 235. The adder 155 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter-prediction unit 221 or the intra-prediction unit 222. If there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as the reconstructed block. The adder 250 can be called the reconstruction unit or reconstructed block generation unit. The generated reconstructed signal can be used for intra-prediction of the next block to be processed in the current picture, or, as described later, for inter-prediction of the next picture after filtering.
[0057] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied during the picture encoding and / or restoration process.
[0058] The filtering unit 260 can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit 260 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be stored in the memory 270, specifically in the DPB of the memory 270. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, and bilateral filter. The filtering unit 260 can generate various filtering-related information and transmit it to the entropy encoding unit 240, as will be described later in the explanation of each filtering method. The filtering-related information can be encoded by the entropy encoding unit 240 and output in bitstream format.
[0059] The corrected restored picture sent to memory 270 can be used as a reference picture in the interpretation unit 221. When interpretation is applied via this, the encoding device can avoid prediction mismatches between the encoding device 100 and the decoding device, and can also improve encoding efficiency.
[0060] The DPB in memory 270 can store the corrected restored picture for use as a reference picture in the inter-prediction unit 221. Memory 270 can store motion information of blocks from which motion information in the current picture has been derived (or encoded) and / or motion information of blocks in the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit 221 for use as motion information of spatially adjacent blocks or motion information of temporally adjacent blocks. Memory 270 can store restored samples of restored blocks in the current picture and transmit them to the intra-prediction unit 222.
[0061] Figure 3 is a schematic diagram illustrating the configuration of a video / image decoding device to which this document can be applied.
[0062] As shown in Figure 3, the decoding device 300 can be configured to include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-prediction unit 331 and an intra-prediction unit 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 321. The aforementioned entropy decoder 310, residual processor 320, predictor 330, adder 340, and filtering unit 350 can be configured by a single hardware component (e.g., a decoder chipset or processor) depending on the embodiment. The memory 360 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The aforementioned hardware component may also further include memory 360 as an internal / external component.
[0063] When a bitstream containing video / image information is input, the decoding device 300 can reconstruct the image in accordance with the process by which the video / image information was processed in the encoding device shown in Figure 3. For example, the decoding device 300 can derive units / blocks based on block division-related information obtained from the bitstream. The decoding device 300 can perform decoding using the processing units applied in the encoding device. Therefore, the decoding processing unit can be, for example, a coding unit, which can be divided from a coding tree unit or a maximum coding unit according to a quad-tree structure, a binary tree structure, and / or a terminally tree structure. One or more conversion units can be derived from the coding unit. The reconstructed video signal decoded and output via the decoding device 300 can then be played back via a playback device.
[0064] The decoding device 300 can receive the signal output from the encoding device shown in Figure 3 in bitstream form, and the received signal can be decoded via the entropy decoding unit 310. For example, the entropy decoding unit 310 can parse the bitstream to derive information necessary for video restoration (or picture restoration) (e.g., video / image information). The video / image information may further include information about various parameter sets, such as the adaptation parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS). The video / image information may also further include general constraint information. The decoding device can further decode the picture based on the parameter set information and / or the general constraint information. The signaling / received information and / or syntax elements described later in this document can be decoded via the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 310 can decode information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the values of syntax elements necessary for image restoration and the quantized values of conversion coefficients related to the residual. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element in the bitstream, determines a context model using the syntax element information to be decoded and the decoding information of adjacent and decoded blocks or symbol / bin information decoded in a previous step, predicts the probability of bin occurrence based on the determined context model, performs arithmetic decoding of the bins, and generates symbols corresponding to the values of each syntax element.In this case, the CABAC entropy decoding method can update the context model after determining the context model by utilizing the decoded symbol / bin information for the context model of the next symbol / bin. Of the information decoded by the entropy decoding unit 310, information related to prediction is provided to the prediction unit (inter-prediction unit 332 and intra-prediction unit 331), and the residual values that have been entropy decoded by the entropy decoding unit 310, i.e., quantized conversion coefficients and related parameter information, can be input to the residual processing unit 320. The residual processing unit 320 can derive residual signals (residual blocks, residual samples, residual sample arrays). In addition, of the information decoded by the entropy decoding unit 310, information related to filtering can be provided to the filtering unit 350. On the other hand, a receiving unit (not shown) that receives signals output from the encoding device can be further configured as an internal / external element of the decoding device 300, or the receiving unit can be a component of the entropy decoding unit 310. On the other hand, the decoding device relating to this document may be called a video / image / picture decoding device, and the decoding device may also be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoding unit 310, and the sample decoder may include at least one of the inverse quantization unit 321, inverse transformation unit 322, addition unit 340, filtering unit 350, memory 360, interpretation unit 332, and intraprediction unit 331.
[0065] The inverse quantization unit 321 can inverse quantize the quantized transformation coefficients and output the transformation coefficients. The inverse quantization unit 321 can rearrange the quantized transformation coefficients in a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scan order performed by the encoding device. The inverse quantization unit 321 can perform inverse quantization on the quantized transformation coefficients using quantization parameters (e.g., quantization step size information) and obtain the transformation coefficients.
[0066] In the inverse conversion unit 322, the conversion coefficients are inversely converted to obtain a residual signal (residual block, residual sample array).
[0067] The prediction unit can make predictions for the current block and generate a predicted block containing prediction samples for the current block. Based on the prediction information output from the entropy decoding unit 310, the prediction unit can determine whether intra-prediction or inter-prediction is applied to the current block and can determine a specific intra / inter-prediction mode.
[0068] The prediction unit 320 can generate prediction signals based on various prediction methods described later. For example, the prediction unit can apply intra-prediction or inter-prediction for prediction of a single block, and can also apply intra-prediction and inter-prediction simultaneously. This can be called combined inter and intra prediction (CIIP). The prediction unit can also base its prediction on an intra-block copy (IBC) prediction mode or on a palette mode for prediction of a block. The IBC prediction mode or palette mode can be used for content video / movie coding such as games, for example, as in SCC (screen content coding). IBC basically performs prediction within the current picture, but can be performed similarly to inter-prediction in that it derives a reference block within the current picture. That is, IBC can utilize at least one of the inter-prediction techniques described in this document. Palette mode can be considered an example of intra-coding or intra-prediction. When palette mode is applied, information about the palette table and palette index can be included in the video / movie information and signaled.
[0069] The intra-prediction unit 331 can predict the current block by referring to a sample in the current picture. The referenced sample can be located adjacent to or far from the current block depending on the prediction mode. In intra-prediction, the prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The intra-prediction unit 331 can also determine the prediction mode to be applied to the current block using the prediction modes applied to adjacent blocks.
[0070] The interprediction unit 332 can derive a predicted block for the current block based on a reference block (reference sample array) identified by motion vectors on the reference picture. In this case, in order to reduce the amount of motion information transmitted in interprediction mode, motion information can be predicted in blocks, subblocks, or samples based on the correlation of motion information between adjacent blocks and the current block. The motion information may include motion vectors and reference picture indices. The motion information may further include interprediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of interprediction, adjacent blocks may include spatially adjacent blocks that exist in the current picture and temporally adjacent blocks that exist in the reference picture. For example, the interprediction unit 332 can construct a motion information candidate list based on adjacent blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Interprediction can be performed based on various prediction modes, and the prediction information may include information indicating the mode of interprediction for the current block.
[0071] The summing unit 340 can generate a restored signal (restored picture, restored block, restored sample array) by adding the acquired residual signal to the predicted signal (predicted block, predicted sample array) output from the prediction unit (which comprises an inter-prediction unit 332 and / or an intra-prediction unit 331). If there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as the restored block.
[0072] The addition unit 340 may be called the restoration unit or restoration block generation unit. The generated restoration signal can be used for intra-prediction of the next block to be processed in the current picture, and can be output after filtering as described later, or it can be used for intra-prediction of the next picture.
[0073] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied during the picture decoding process.
[0074] The filtering unit 350 can apply filtering to the restored signal to improve subjective / objective image quality. For example, the filtering unit 350 can apply various filtering methods to the restored picture to generate a modified restored picture, and can transmit the modified restored picture to the memory 360, specifically to the DPB of the memory 360. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, and bilateral filter.
[0075] The (modified) restored picture stored in the DPB of memory 360 can be used as a reference picture by the inter-prediction unit 332. Memory 360 can store motion information of blocks from which motion information in the current picture has been derived (or decoded) and / or motion information of blocks in the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit 260 for use as motion information of spatially adjacent blocks or motion information of temporally adjacent blocks. Memory 360 can store restored samples of restored blocks in the current picture and transmit them to the intra-prediction unit 331.
[0076] In this specification, embodiments described for the filtering unit 260, inter-prediction unit 221, and intra-prediction unit 222 of the encoding device 200 can also be applied identically or in a corresponding manner to the filtering unit 350, inter-prediction unit 332, and intra-prediction unit 331 of the decoding device 300, respectively.
[0077] As mentioned above, prediction is performed to improve compression efficiency when performing video coding. Through this, a predicted block containing predicted samples for the current block, which is the block to be coded, can be generated. Here, the predicted block contains predicted samples in the spatial domain (or pixel domain). The predicted block is derived in both the encoding and decoding devices, and the encoding device can improve video coding efficiency by signaling the decoding device with information (residual information) about the residual between the original block and the predicted block, which is not the original sample value of the original block itself. The decoding device can derive a residual block containing residual samples based on the residual information, and can generate a restored block containing restored samples by combining the residual block and the predicted block, and can generate a restored picture containing the restored block.
[0078] The residual information can be generated through transformation and quantization procedures. For example, an encoding device can signal the relevant residual information (via a bitstream) to a decoding device by deriving a residual block between the original block and the predicted block, performing a transformation procedure on the residual samples (residual sample array) contained in the residual block to derive transformation coefficients, and performing a quantization procedure on the transformation coefficients to derive quantized transformation coefficients. Here, the residual information may include information such as the value information, position information, transformation technique, transformation kernel, and quantization parameters of the quantized transformation coefficients. The decoding device can derive a residual sample (or residual block) by performing an inverse quantization / inverse transformation procedure based on the residual information. The decoding device can generate a reconstructed picture based on the predicted block and the residual block. The encoding device can also derive a residual block by inverse quantization / inverse transformation of the quantized transformation coefficients for reference for subsequent interpretation of the picture, and generate a reconstructed picture based on this.
[0079] On the other hand, as mentioned above, intra-prediction or inter-prediction can be applied when making predictions for the current block. In one embodiment, when inter-prediction is applied to the current block, the prediction unit (more specifically, the inter-prediction unit) of the encode / decode device can perform inter-prediction on a block-by-block basis to derive predicted samples. Inter-prediction can indicate predictions derived in a manner dependent on data elements of pictures other than the current picture (e.g., sample values or motion information). When inter-prediction is applied to the current block, a predicted block (predicted sample array) for the current block can be derived based on a reference block (reference sample array) identified by motion vectors on the reference picture pointed to by the reference picture index. In this case, in order to reduce the amount of motion information transmitted in inter-prediction mode, the motion information of the current block can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between surrounding blocks and the current block. Motion information may include motion vectors and reference picture indexes. Motion information may further include inter-prediction type (L0 prediction, L1 prediction, Bi prediction, etc.) information. When interpretation is applied, a neighboring block may include a spatial neighboring block currently present in the picture and a temporal neighboring block present in the reference picture. The reference picture containing the reference block and the reference picture containing the temporal neighboring block may be the same or different. The temporal neighboring block may be called a collocated reference block, colCU, etc., and the reference picture containing the temporal neighboring block may also be called a collocated picture (colPic). For example, a list of motion information candidates may be constructed based on the neighboring blocks of the current block, and flags or index information may be signaled to indicate which candidate is selected (used) to derive the motion vector and / or reference picture index of the current block.Interpretation is performed based on various prediction modes. For example, in skip mode and merge mode, the motion information of the current block may be the same as the motion information of the selected surrounding block. In skip mode, unlike merge mode, a residual signal may not be transmitted. In motion vector prediction (MVP) mode, the motion vector of the selected surrounding block can be used as a motion vector predictor, and the motion vector difference can be signaled. In this case, the motion vector of the current block can be derived using the sum of the motion vector predictor and the motion vector difference.
[0080] The motion information may include L0 motion information and / or L1 motion information depending on the interpretation type (L0 prediction, L1 prediction, Bi prediction, etc.). A motion vector in the L0 direction may be called an L0 motion vector or MVL0, and a motion vector in the L1 direction may be called an L1 motion vector or MVL1. A prediction based on an L0 motion vector may be called an L0 prediction, a prediction based on an L1 motion vector may be called an L1 prediction, and a prediction based on both the L0 motion vector and the L1 motion vector may be called a bi (Bi) prediction. Here, an L0 motion vector may represent a motion vector associated with a reference picture list L0 (L0), and an L1 motion vector may represent a motion vector associated with a reference picture list L1 (L1). The reference picture list L0 may include pictures earlier in the output order than the current picture, and the reference picture list L1 may include pictures later in the output order than the current picture. The aforementioned earlier picture may be called a forward (reference) picture, and the aforementioned later picture may be called a reverse (reference) picture. The reference picture list L0 may include further reference pictures that are later in the output order than the current picture. In this case, the earlier picture may be indexed first in the reference picture list L0, and the later picture may be indexed afterward. The reference picture list L1 may include further reference pictures that are earlier in the output order than the current picture. In this case, the later picture may be indexed first in the reference picture list L1, and the earlier picture may be indexed afterward. Here, the output order may correspond to the POC (picture order count) order.
[0081] Figure 4 shows an example of a schematic video / image encoding procedure to which the embodiments described herein can be applied. In Figure 4, S400 can be performed in the prediction unit 220 of the encoding device detailed in Figure 2, S410 can be performed in the residual processing unit 230, and S420 can be performed in the entropy encoding unit 240. S400 may include the inter / intra prediction procedure described herein, S410 may include the residual processing procedure described herein, and S420 may include the information encoding procedure described herein.
[0082] Referring to Figure 4, the video / image encoding procedure may include not only a procedure for encoding information for picture reconstruction (e.g., prediction information, residual information, partitioning information, etc.) and outputting it in the form of a bitstream, as shown in the description relating to Figure 2, but also a procedure for generating a reconstruction picture for the current picture, and an optional procedure for applying in-loop filtering to the reconstruction picture. The encoding device can derive (corrected) residual samples from the quantized conversion coefficients via the inverse quantization unit 234 and the inverse conversion unit 235, and can generate a reconstruction picture based on the prediction samples, which are the output of S400, and the (corrected) residual samples. The reconstruction picture thus generated may be identical to the reconstruction picture generated by the decoding device described above. A corrected reconstruction picture is generated via the in-loop filtering procedure on the reconstruction picture, which is stored in the decoded picture buffer or memory 270 and can be used as a reference picture in the inter-prediction procedure when encoding subsequent pictures, as in the case of the decoding device. As described above, some or all of the in-loop filtering procedure may be omitted depending on the circumstances. When the in-loop filtering procedure is performed, the (in-loop) filtering-related information (parameters) can be encoded by the entropy encoding unit 240 and output in the form of a bitstream, and the decoding device can perform the in-loop filtering procedure in the same way as the encoding device based on the filtering-related information.
[0083] Through such in-loop filtering procedures, noise that occurs during image / video coding, such as blocking artifacts and ringing artifacts, can be reduced, improving the subjective and objective visual quality. Furthermore, by having both the encoding and decoding devices perform in-loop filtering, they can derive the same prediction results, increasing the reliability of picture coding and reducing the amount of data that must be transmitted for picture coding.
[0084] As mentioned above, the picture restoration procedure is performed not only in the decoding device but also in the encoding device. A restored block is generated for each block based on intra-prediction / inter-prediction, and a restored picture containing the restored block may be generated. If the current picture / slice / tile group is an I picture / slice / tile group, the blocks included in the current picture / slice / tile group can be restored based only on intra-prediction. On the other hand, if the current picture / slice / tile group is a P or B picture / slice / tile group, the blocks included in the current picture / slice / tile group can be restored based on intra-prediction or inter-prediction. In this case, inter-prediction may be applied to some blocks in the current picture / slice / tile group, and intra-prediction may be applied to the remaining blocks. The color components of a picture may include luminous and chroma components, and unless expressly limited in this document, the methods and embodiments proposed in this document may be applied to luminous and chroma components.
[0085] Figure 5 shows an example of a schematic video / image decoding procedure to which the embodiments described herein can be applied. In Figure 5, S500 may be performed in the entropy decoding unit 310 of the decoding device detailed in Figure 3, S510 may be performed in the prediction unit 330, S520 may be performed in the residual processing unit 320, S530 may be performed in the addition unit 340, and S540 may be performed in the filtering unit 350. S500 may include the information decoding procedure described herein, S510 may include the inter / intra prediction procedure described herein, S520 may include the residual processing procedure described herein, S530 may include the block / picture restoration procedure described herein, and S540 may include the in-loop filtering procedure described herein.
[0086] Referring to Figure 5, the picture decoding procedure can, in general terms, include a procedure for acquiring video information (via decoding) from a bitstream (S500), a picture restoration procedure (S510-S530), and an in-loop filtering procedure (S540) for the restored picture, as shown in the explanation relating to Figure 3. The picture restoration procedure can be performed based on predicted samples and residual samples obtained through the inter / intra prediction (S510) and residual processing (S520, inverse quantization and inverse transformation of quantized transformation coefficients) processes described in this document. A modified restored picture is generated through an in-loop filtering procedure on the restored picture generated through the picture restoration procedure, and the modified restored picture is output as a decoded picture, and is also stored in the decoded picture buffer or memory 360 of the decoding device, and can be used as a reference picture in the inter-prediction procedure when decoding subsequent pictures.
[0087] In some cases, the in-loop filtering procedure may be omitted. In this case, the restored picture is output as a decoded picture and stored in the decoding picture buffer or memory 360 of the decoding device, and can be used as a reference picture in the inter-prediction procedure when decoding subsequent pictures. The in-loop filtering procedure (S540) may include, as described above, a deblocking filtering procedure, an SAO (sample adaptive offset) procedure, an ALF (adaptive loop filter) procedure, and / or a bi-lateral filter procedure, and some or all of these may be omitted. Also, one or some of the deblocking filtering procedure, the SAO (sample adaptive offset) procedure, the ALF (adaptive loop filter) procedure, and the bi-lateral filter procedure may be applied sequentially, or all of them may be applied sequentially. For example, the SAO procedure may be executed after the deblocking filtering procedure is applied to the restored picture. Alternatively, for example, the ALF procedure may be executed after the deblocking filtering procedure is applied to the restored picture. This can also be done in an encoding device.
[0088] On the other hand, as mentioned above, encoding devices can perform entropy encoding based on various encoding methods such as exponential Golomb, CAVLC (context-adaptive variable length coding), and CABAC (context-adaptive binary arithmetic coding). Similarly, decoding devices can perform entropy decoding based on coding methods such as exponential Golomb coding, CAVLC, or CABAC. The entropy encoding / decoding procedures will be described below.
[0089] Figure 6 schematically shows an example of an entropy encoding method to which the embodiments described in this document can be applied, and Figure 7 schematically shows the entropy encoding unit in an encoding device. The entropy encoding unit in the encoding device in Figure 7 can be applied to the same or corresponding entropy encoding unit 240 of the encoding device 200 in Figure 2 described above.
[0090] Referring to Figures 6 and 7, the encoding device (entropy encoding unit) can perform entropy coding procedures for video information. Video information may include partitioning-related information, prediction-related information (e.g., inter / intra prediction classification information, intra prediction mode information, inter prediction mode information, etc.), residual information, in-loop filtering-related information, etc., or may include various syntax elements related thereto. Entropy coding can be performed on a unit of syntax elements. Steps S600 to S610 can be performed by the entropy encoding unit 240 of the encoding device 200 shown in Figure 2 above.
[0091] The encoding device can perform binary encoding on the target syntax element (S600). Here, the binary encoding method for the target syntax element can be predefined based on various binary encoding methods such as the Truncated Rice binarization process and the Fixed-length binarization process. The binary encoding procedure can be performed by the binary encoding unit 242 within the entropy encoding unit 240.
[0092] The encoding device can perform entropy encoding on the target syntax element (S610). Based on entropy coding techniques such as CABAC (context-adaptive arithmetic coding) or CAVLC (context-adaptive variable length coding), the encoding device can encode the bin string of the target syntax element based on regular coding (context-based) or bypass coding, and the output may be included in a bitstream. The entropy encoding procedure can be performed by the entropy encoding processing unit 243 within the entropy encoding unit 240. As previously mentioned, the bitstream can be transmitted to the decoding device via a (digital) storage medium or a network.
[0093] Figure 8 schematically shows an example of an entropy decoding method to which the embodiments described in this document can be applied, and Figure 9 schematically shows the entropy decoding section in a decoding device. The entropy decoding section in the decoding device in Figure 9 can be applied in the same way as or corresponding to the entropy decoding section 310 of the decoding device 300 in Figure 3 described above.
[0094] Referring to Figures 8 and 9, the decoding device (entropy decoding unit) can decode the encoded video information. The video information may include partitioning-related information, prediction-related information (e.g., inter / intra prediction classification information, intra prediction mode information, inter prediction mode information, etc.), residual information, in-loop filtering-related information, etc., or may include various syntax elements related thereto. Entropy coding can be performed on a unit of syntax elements. Steps S800 to S810 can be performed by the entropy decoding unit 310 of the decoding device 300 shown in Figure 3 above.
[0095] The decoding device can perform binary conversion on the target syntax element (S800). Here, the binary conversion method for the target syntax element can be predefined based on various binary conversion methods such as the Truncated Rice binarization process and the Fixed-length binarization process. The decoding device can derive available binstrings (candidate binstrings) for the available values of the target syntax element through the binary conversion procedure. The binary conversion procedure can be performed by the binary conversion unit 312 within the entropy decoding unit 310.
[0096] The decoding device can perform entropy decoding on the target syntax element (S810). The decoding device sequentially decodes and parses each bin for the target syntax element from the input bits in the bitstream, and compares the derived binstring with the available binstrings for the syntax element. If the derived binstring is identical to one of the available binstrings, the value corresponding to that binstring can be derived as the value of the syntax element. Otherwise, the next bit in the bitstream can be further parsed, and the above procedure can be repeated. Through this process, information can be signaled using bits of variable length without using start bits or end bits for specific information (specific syntax elements) in the bitstream. Through this, fewer bits can be allocated to lower values, thereby improving overall coding efficiency.
[0097] A decoding device can decode each bin in a bin string from a bitstream based on context or by bypass, using entropy coding techniques such as CABAC or CAVLC. Here, the bitstream may contain various information for video decoding, as mentioned above. As mentioned above, the bitstream can be transmitted to the decoding device via a (digital) storage medium or a network.
[0098] Figure 10 illustrates the hierarchical structure for coded video.
[0099] Referring to Figure 10, coded video is divided into the VCL (video coding layer), which handles the video decoding process and the video itself; a lower-level system that transmits and stores the coded information; and the NAL (network abstraction layer), which exists between the VCL and the lower-level system and is responsible for network adaptation functions.
[0100] VCL can generate VCL data containing compressed video data (slice data), or generate parameter sets containing information such as Picture Parameter Set (PPS), Sequence Parameter Set (SPS), and Video Parameter Set (VPS), or SEI (Supplemental Enhancement Information) messages that are additionally necessary during the video decoding process.
[0101] In NAL, a NAL unit can be generated by adding header information (NAL unit header) to an RBSP (Raw Byte Sequence Payload) generated by VCL. In this case, the RBSP refers to slice data, parameter sets, SEI messages, etc., generated by VCL. The NAL unit header may include NAL unit type information identified by the RBSP data contained in the NAL unit.
[0102] Furthermore, NAL units can be classified into VCL NAL units and Non-VCL NAL units based on the RBSP generated by VCL. A VCL NAL unit may represent a NAL unit containing information about the video (slice data), while a Non-VCL NAL unit may represent a NAL unit containing information necessary for decoding the video (parameter set or SEI message).
[0103] VCL NAL units and Non-VCL NAL units can be transmitted over a network with header information according to the data standards of the underlying system. For example, NAL units can be transformed into data formats of predetermined standards such as H.266 / VVC file format, RTP (Real-time Transport Protocol), and TS (Transport Stream) and transmitted over various networks.
[0104] As mentioned above, the NAL unit type can be identified by the RBSP data structure contained within the NAL unit, and information about such NAL unit types can be stored in the NAL unit header and signaled.
[0105] For example, NAL units can be broadly classified into VCL NAL unit types and Non-VCL NAL unit types depending on whether or not they contain information (slice data) related to the video. VCL NAL unit types can be further classified by the nature and type of picture they contain, while Non-VCL NAL unit types can be further classified by the type of parameter set.
[0106] The following is an example of a NAL unit type identified by the type of parameter set included in the Non-VCL NAL unit type.
[0107] -APS (Adaptation Parameter Set) NAL unit: Type for NAL units that include APS
[0108] -DPS (Decoding Parameter Set) NAL unit: Type for NAL units including DPS
[0109] -VPS (Video Parameter Set) NAL unit: Type for NAL unit including VPS
[0110] -SPS (Sequence Parameter Set) NAL unit: Type for NAL units that include SPS
[0111] -PPS (Picture Parameter Set) NAL unit: Type for NAL units that include PPS
[0112] -PH (Picture header) NAL unit: Type for NAL units that include PH
[0113] The aforementioned NAL unit type has syntax information for the NAL unit type, and this syntax information can be stored in the NAL unit header and signaled. For example, the syntax information is nal_unit_type, and the NAL unit type can be identified by the nal_unit_type value.
[0114] On the other hand, as mentioned above, a single picture may contain multiple slices, and a single slice may contain a slice header and slice data. In this case, an additional picture header may be added to each of the multiple slices (sets of slice headers and slice data) within a single picture. The picture header (picture header syntax) may contain information / parameters that are applicable to pictures in common. In this document, tile groups may be used interchangeably with or substituted for slices or pictures. Also, in this document, tile group headers may be used interchangeably with or substituted for slice headers or picture headers.
[0115] A slice header (slice header syntax) may contain information / parameters that are applicable to slices in common. An APS (APS syntax) or PPS (PPS syntax) may contain information / parameters that are applicable to one or more slices or pictures in common. An SPS (SPS syntax) may contain information / parameters that are applicable to one or more sequences in common. A VPS (VPS syntax) may contain information / parameters that are applicable to multiple layers in common. A DPS (DPS syntax) may contain information / parameters that are applicable to video in general. A DPS may contain information / parameters related to the concatenation of CVS (coded video sequence). In this document, High-level syntax (HLS) may include at least one of the APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, picture header syntax, and slice header syntax.
[0116] In this document, the video information encoded from the encoding device to the decoding device and signaled in the form of a bitstream includes not only partitioning-related information, intra / inter prediction information, residual information, and in-loop filtering information within the picture, but may also include information contained in the slice header, picture header, APS, PPS, SPS, VPS, and / or DPS. Furthermore, the video information may further include information in the NAL unit header.
[0117] As mentioned above, HLS (High-level syntax) can be coded / signaled for video / image coding. In this document, video / image information may include HLS. For example, a coded picture may consist of one or more slices. Parameters describing the coded picture may be signaled in the picture header (PH), and parameters describing the slices may be signaled in the slice header (SH). The PH may be sent to its own NAL unit type. The SH may be present at the beginning of the NAL unit containing the payload of the slice (i.e., slice data). Detailed information regarding the syntax and semantics of the PH and SH is disclosed in the VVC standard. Each picture may be associated with a PH. A picture may consist of different types of slices, which are intra-coded slices (i.e., I-slices) and inter-coded slices (i.e., P-slices and B-slices). As a result, the PH may contain the syntax elements required for the intra-slice and inter-slice of the picture.
[0118] On the other hand, generally, one NAL unit type can be set for one picture. The NAL unit type can be signaled via nal_unit_type in the NAL unit header of a NAL unit containing a slice. nal_unit_type is syntax information for identifying the NAL unit type, that is, it can identify the type of RBSP data structure contained in the NAL unit, as shown in Table 1 or Table 2 below.
[0119] Table 1 shows examples of NAL unit type codes and NAL unit type classes.
[0120] [Table 1]
[0121] Alternatively, as an example, the codes and classes of the NAL unit type can be defined as shown in Table 2 below.
[0122] [Table 2-1]
[0123] [Table 2-2]
[0124] As shown in Table 1 or Table 2 above, the name and value of the NAL unit type can be identified by the RBSP data structure contained in the NAL unit, and the NAL unit can be classified into VCL NAL unit types and Non-VCL NAL unit types depending on whether or not the NAL unit contains information (slice data) for the video. VCL NAL unit types can be classified according to the nature and type of the picture, and Non-VCL NAL unit types can be classified according to the type of parameter set, etc. For example, the NAL unit type can be identified according to the nature and type of the picture contained in the VCL NAL unit as follows.
[0125] TRAIL: Indicates the type for the NAL unit containing the coded slice data for the trailing picture / subpicture. For example, nal_unit_type can be defined as TRAIL_NUT, and the value of nal_unit_type can be specified as 0.
[0126] Here, a trailing picture is a picture that follows a randomly accessible picture in both the output order and the decoding order. A trailing picture may be a non-IRAP picture that follows an associated IRAP picture in the output order and is not an STSA picture. For example, an associated trailing picture follows an IRAP picture in the decoding order. A picture that follows an associated IRAP picture in the output order and precedes an associated IRAP picture in the decoding order is not permitted.
[0127] STSA (Step-wise Temporal Sub-layer Access): Indicates the type of NAL unit containing the coded slice data for STSA pictures / subpictures. For example, nal_unit_type can be defined as STSA_NUT, and the value of nal_unit_type can be specified as 1.
[0128] Here, an STSA picture is a picture that supports time scalability and is switchable between time sublayers, indicating a position in a lower sublayer where up-switching to a higher sublayer one level above the lower sublayer is possible. An STSA picture does not use pictures with the same TemporalId and on the same layer as the STSA picture for interpretation reference. In the decoding order, a picture that follows an STSA picture with the same TemporalId and on the same layer as the STSA picture does not use pictures that precede the STSA picture with the same TemporalId and on the same layer in the decoding order for interpretation reference. An STSA picture immediately enables up-switching in a lower sublayer to the sublayer containing the STSA picture. In this case, the coded picture must not belong to the lowest sublayer; that is, an STSA picture must always have a TemporalId greater than 0.
[0129] RADL (random access decodable leading (picture)): Indicates the type for a NAL unit containing the coded slice data of a RADL picture / subpicture. For example, nal_unit_type can be defined as RADL_NUT, and the value of nal_unit_type can be specified as 2.
[0130] Here, all RADL pictures are leading pictures. RADL pictures are not used as reference pictures for decoding the trailing pictures of the same associated IRAP pictures. Specifically, RADL pictures with a nuh_layer_id, such as layerId, are not used as reference pictures for decoding pictures with a nuh_layer_id, such as layerId, as they follow the IRAP pictures associated with the RADL picture in the output order. If field_seq_flag (i.e., sps_field_seq_flag) is 0, all RADL pictures (i.e., if RADL pictures exist) precede all non-leading pictures of the same associated IRAP pictures in the decoding order. On the other hand, a leading picture is a picture that precedes the associated IRAP picture in the output order.
[0131] RASL (random access skipped leading(picture)): Indicates the type for a NAL unit containing the coded slice data of a RASL picture / subpicture. For example, nal_unit_type can be defined as RASL_NUT, and the value of nal_unit_type can be specified as 3.
[0132] Here, every RASL picture is the leading picture of its associated CRA picture. If the associated CRA picture has a NoOutputBeforeRecoveryFlag with a value of 1, the RASL picture may not be output and may not be decoded correctly, as it may contain references to pictures that do not exist in the bitstream. RASL pictures are not used as reference pictures for the decoding process of non-RASL pictures of the same layer. However, RADL subpictures in a RASL picture of the same layer can be used for interpretation of collocated RADL subpictures in a RADL picture associated with the same CRA picture as the RASL picture. If field_seq_flag (i.e., sps_field_seq_flag) is 0, every RASL picture (i.e., if a RASL picture exists) precedes all non-leading pictures of the same associated CRA picture in the decoding order.
[0133] There may be reserved nal_unit_types for non-IRAP VCL NAL unit types. For example, nal_unit_type can be defined as RSV_VCL_4 through RSV_VCL_6, and the values of nal_unit_type can be specified as 4 through 6, respectively.
[0134] Here, IRAP (intra random access point) is information indicating the NAL unit for a picture that is randomly accessible. An IRAP picture can be a CRA picture or an IDR picture. For example, an IRAP picture is a picture that has NAL unit types defined as IDR_W_RADL, IDR_N_LP, or CRA_NU, as shown in Table 1 or Table 2 above, and the value of nal_unit_type can be specified as 7 to 9, respectively.
[0135] IRAP pictures do not use reference pictures on the same layer for interpretation during the decoding process. In other words, IRAP pictures do not reference any other pictures for interpretation during the decoding process. In the decoding order, the first picture in the bitstream will be an IRAP or GDR picture. For a single-layer bitstream, if the necessary parameter set is available where references are needed, all subsequent non-RASL and IRAP pictures in the CLVS (coded layer video sequence) in the decoding order can be decoded accurately without having to decode the pictures that precede the IRAP picture in the decoding order.
[0136] The value of mixed_nalu_types_in_pic_flag for an IRAP picture is 0. When the value of mixed_nalu_types_in_pic_flag for a picture is 0, one slice in the picture may have an NAL unit type (nal_unit_type) within the range of IDR_W_RADL to CRA_NUT (for example, a NAL unit type value of 7 to 9 in Table 1 or Table 2), and all other slices in the picture may have the same NAL unit type (nal_unit_type). In this case, the picture may be considered an IRAP picture.
[0137] IDR (instantaneous decoding refresh): Indicates the type for the NAL unit containing the coded slice data of the IDR picture / subpicture. For example, the nal_unit_type for an IDR picture / subpicture can be defined as IDR_W_RADL or IDR_N_LP, and the value of nal_unit_type can be specified as 7 or 8, respectively.
[0138] Here, an IDR picture does not use interpretation during the decoding process (i.e., it does not refer to other pictures for interpretation) and may be the first picture in the bitstream in the decoding order, or it may appear later in the bitstream (i.e., not the first). Each IDR picture is the first picture in the CVS (coded video sequence) in the decoding order. For example, if an IDR picture is associated with a decodeable reading picture, the NAL unit type of the IDR picture may be IDR_W_RADL, and if an IDR picture is not associated with a reading picture, the NAL unit type of the IDR picture may be IDR_N_LP. That is, an IDR picture with NAL unit type IDR_W_RADL may not have an associated RASL picture in the bitstream, but may have an associated RADL picture in the bitstream. An IDR picture with NAL unit type IDR_N_LP does not have an associated reading picture in the bitstream.
[0139] CRA (Clean Random Access): Indicates the type of NAL unit containing the coded slice data for CRA pictures / subpictures. For example, nal_unit_type can be defined as CRA_NUT, and the value of nal_unit_type can be specified as 9.
[0140] Here, the CRA picture does not use interpretation during the decoding process (i.e., it does not refer to other pictures for interpretation) and may be the first picture in the bitstream in terms of the decoding order, or it may appear later in the bitstream (i.e., not the first). The CRA picture may have associated RADL or RASL pictures present in the bitstream. For a CRA picture with a NoOutputBeforeRecoveryFlag value of 1, the associated RASL picture may not be output by the decoder because it may contain references to pictures that do not exist in the bitstream, and therefore cannot be decoded in this case.
[0141] GDR (gradual decoder refresh): Indicates the type for the NAL unit containing the coded slice data for the GDR picture / subpicture. For example, nal_unit_type can be defined as GDR_NUT, and the value of nal_unit_type can be specified as 10.
[0142] Here, the value of pps_mixed_nalu_types_in_pic_flag for a GDR picture can be 0. If the value of pps_mixed_nalu_types_in_pic_flag for a picture is 0, and one slice in the picture has the NAL unit type GDR_NUT, then all other slices in the picture will have the same NAL unit type (nal_unit_type), and in this case, the picture can become a GDR picture after receiving the first slice.
[0143] Furthermore, for example, the NAL unit type can be identified depending on the type of parameters included in the Non-VCL NAL unit, and may include NAL unit types (nal_unit_type) such as VPS_NUT indicating the type for NAL units containing video parameter sets, SPS_NUT indicating the type for NAL units containing sequence parameter sets, PPS_NUT indicating the type for NAL units containing picture parameter sets, and PH_NUT indicating the type for NAL units containing picture headers, as shown in Table 1 or Table 2.
[0144] On the other hand, a bitstream that supports temporal scalability (or a temporally scalable bitstream) includes information about a temporally scalable temporal layer. This information about the temporal layer may be identification information for the temporal layer, which is determined by the temporal scalability of the NAL unit. For example, the identification information for the temporal layer can be temporal_id syntax information, which is stored in the NAL unit header by the encoding device and signaled to the decoding device. Hereinafter, the temporal layer may be referred to as a sub-layer, temporal sub-layer, or temporally scalable layer, etc.
[0145] Figure 11 shows the time layer structure for NAL units in a bitstream that supports time scalability.
[0146] If a bitstream supports temporal scalability, the NAL units contained in the bitstream have temporal layer identification information (e.g., temporal_id). For example, a temporal layer composed of NAL units with a temporal_id value of 0 can provide the lowest temporal scalability, while a temporal layer composed of NAL units with a temporal_id value of 2 can provide the highest temporal scalability.
[0147] In Figure 11, boxes labeled "I" are called I-pictures, and boxes labeled "B" are called B-pictures. Arrows indicate reference relationships, showing whether a picture references another picture.
[0148] As shown in Figure 11, a NAL unit of a time layer with a temporal_id value of 0 is a reference picture that can be referenced by NAL units of time layers with a temporal_id value of 0, 1, or 2. A NAL unit of a time layer with a temporal_id value of 1 is a reference picture that can be referenced by NAL units of time layers with a temporal_id value of 1 or 2. A NAL unit of a time layer with a temporal_id value of 2 may be a reference picture that can be referenced by NAL units of the same time layer, i.e., NAL units of time layers with a temporal_id value of 2, or it may be an unreferenced picture that is not referenced by other pictures.
[0149] If, as shown in Figure 11, the NAL unit of the time layer with a temporal_id value of 2, i.e., the topmost time layer, is a non-referenced picture, then such a NAL unit can be extracted (or removed) from the bitstream during the decoding process without affecting other pictures.
[0150] On the other hand, among the aforementioned NAL unit types, the IDR and CRA types indicate an NAL unit containing a picture that is capable of random access (or splicing), i.e., a RAP (Random Access Point) or IRAP (Intra Random Access Point) picture that becomes a random access point. In other words, an IRAP picture can be an IDR or a CRA picture, and can only contain I slices. In the bitstream, the first picture in the decoding order will be an IRAP picture.
[0151] If an IRAP picture (IDR, CRA picture) is included in the bitstream, there may be a picture that is output before the IRAP picture but decoded later. Such a picture is called a leading picture (LP).
[0152] Figure 12 is a diagram illustrating a picture that can be accessed randomly.
[0153] A randomly accessible picture, i.e., a RAP or IRAP picture that becomes a random access point, is the first picture in the bitstream's decoding order during random access and contains only I slices.
[0154] Figure 12 illustrates the output order (or display order) and decoding order of pictures. As shown, the output order and decoding order of pictures may differ from each other. For convenience, the pictures are described by dividing them into predetermined groups.
[0155] The first group (I) consists of pictures that precede the IRAP pictures in both output and decoding order. The second group (II) consists of pictures that precede the IRAP pictures in output order but decode later. The third group (III) consists of pictures that all follow the IRAP pictures in both output and decoding order.
[0156] Pictures in Group 1 (I) can be decoded and output independently of IRAP pictures.
[0157] Pictures belonging to the second group (II), which are output before the IRAP picture, are called leading pictures, and leading pictures can cause problems during the decoding process when the IRAP picture is used as a random access point.
[0158] Pictures belonging to the third group (III), whose output and decoding order follows that of IRAP pictures, are called normal pictures. Normal pictures are not used as reference pictures for reading pictures.
[0159] Random access points in the bitstream where random access occurs become IRAP pictures, and random access begins while the first picture of the second group (II) is output.
[0160] Figure 13 is a diagram illustrating the IDR picture.
[0161] An IDR picture is a picture that becomes a random access point when the group of pictures has a closed structure. As mentioned above, an IDR picture is an IRAP picture, so it contains only I slices and may be the first picture in the bitstream's decoding order, or it may appear in the middle of the bitstream. When an IDR picture is decoded, all reference pictures stored in the DPB (decoded picture buffer) are displayed as "unused for reference".
[0162] In Figure 13, the bars represent pictures, and the arrows indicate the reference relationship, showing whether a picture can use another picture as a reference picture. The "x" mark on the arrow indicates that the picture cannot reference the picture pointed to by the arrow.
[0163] As shown, a picture with a POC of 32 is an IDR picture. A picture with a POC of 25 to 31 that is output before an IDR picture is a leading picture 1310. A picture with a POC of 33 or higher is a normal picture 1320.
[0164] In terms of output order, the reading picture 1310, which precedes the IDR picture, can use the IDR picture and other reading pictures as reference pictures. However, in terms of output order and decoding order, the past picture 1330, which precedes the reading picture 1310, cannot be used as a reference picture.
[0165] In terms of output and decoding order, the normal picture 1320, which follows the IDR picture, can be decoded by referencing the IDR picture, the reading picture, and other normal pictures.
[0166] Figure 14 is a diagram illustrating the CRA picture.
[0167] A CRA picture is a picture that becomes a random access point when the group of pictures has an open structure. As mentioned earlier, since a CRA picture is also an IRA picture, it contains only I slices and may be the first picture in the bitstream in terms of decoding order, or it may be in the middle of the bitstream for normal play.
[0168] The bars shown in Figure 14 represent pictures, and the arrows indicate the reference relationship, showing whether a picture can use another picture as a reference picture. The x mark displayed on the arrow indicates that the picture in question, or the picture pointed to by the arrow, cannot reference the picture.
[0169] In terms of output order, the reading picture 1410 that precedes the CRA picture can use the CRA picture, other reading pictures, and all of the past pictures 1430 that precede the reading picture 1410 in terms of output order and decoding order as reference pictures.
[0170] On the other hand, in terms of output and decoding order, the normal picture 1420, which follows the CRA picture, can be decoded by referencing the CRA picture and other normal pictures. The normal picture 1420 may not use the leading picture 1410 as a reference picture.
[0171] On the other hand, the VVC standard allows the coded picture (i.e., the current picture) to contain slices of other NAL unit types. Whether the current picture contains slices of other NAL unit types can be indicated by the syntax element mixed_nalu_types_in_pic_flag. For example, if the current picture contains slices of other NAL unit types, the value of the syntax element mixed_nalu_types_in_pic_flag may be 1. In this case, the current picture must refer to a PPS that contains a mixed_nalu_types_in_pic_flag with a value of 1. The semantics of the aforementioned flag (mixed_nalu_types_in_pic_flag) are as follows:
[0172] If the value of the syntax element mixed_nalu_types_in_pic_flag is 1, it may indicate that each picture referencing a PPS has one or more VCL NAL units, that the VCL NAL units do not have NAL unit types (nal_unit_type) with the same value, and that the picture is not an IRAP picture.
[0173] If the value of the syntax element mixed_nalu_types_in_pic_flag is 0, it may indicate that each picture referencing the PPS has one or more VCL NAL units, and that the VCL NAL units of each picture referencing the PPS have the same NAL unit type (nal_unit_type).
[0174] If the value of no_mixed_nalu_types_in_pic_constraint_flag is 1, then the value of mixed_nalu_types_in_pic_flag must be 0. The no_mixed_nalu_types_in_pic_constraint_flag syntax element indicates a constraint on whether the value of mixed_nalu_types_in_pic_flag for a picture must be 0. For example, the no_mixed_nalu_types_in_pic_constraint_flag information can be signaled from higher-level syntax (e.g., PPS) or syntax containing constraint information (e.g., GCI; general constraints information) to determine whether the value of mixed_nalu_types_in_pic_flag must be 0.
[0175] In picture picA, which also includes one or more slices with other NAL unit type values (i.e., the value of mixed_nalu_types_in_pic_flag for picture picA is 1), the following can be applied to each slice having a NAL unit type value nalUnitTypeA in the range from IDR_W_RADL to CRA_NUT (for example, the NAL unit type value being 7 to 9 in Table 1 or Table 2):
[0176] - A slice must belong to the subpicture subpicA whose corresponding subpic_treated_as_pic_flag[i] value is 1, where subpic_treated_as_pic_flag[i] is information about whether the i-th subpicture of each picture coded in CLVS is treated as a picture during the decoding process excluding in-loop filtering. For example, a value of subpic_treated_as_pic_flag[i] of 1 may indicate that the i-th subpicture is treated as a picture during the decoding process excluding in-loop filtering. Alternatively, a value of subpic_treated_as_pic_flag[i] of 0 may indicate that the i-th subpicture is not treated as a picture during the decoding process excluding in-loop filtering.
[0177] - A slice must not belong to a subpicture of picA that contains a VCL NAL unit having a NAL unit type (nal_unit_type) that is not identical to nalUnitTypeA.
[0178] - In the decoding order, for all subsequent PUs in the CLVS, RefPicList[0] or RefPicList[1] of the slice in subpicA should not contain any active entries that precede picA in the decoding order.
[0179] To make the aforementioned concept work, you can specify it as follows. For example, you can apply it to the VCL NAL unit of a particular picture as follows:
[0180] - If the value of mixed_nalu_types_in_pic_flag is 0, the value of NAL unit type (nal_unit_type) must be the same for all coded slice NAL units in the picture. A picture or PU may be considered to have the same NAL unit type as the coded slice NAL units within the picture or PU.
[0181] - Otherwise (if the value of mixed_nalu_types_in_pic_flag is 1), one or more VCL NAL units must each have a specific NAL unit type within the range of IDR_W_RADL to CRA_NUT (for example, NAL unit type values of 7 to 9 in Table 1 or Table 2), and each other VCL NAL units must each have a specific NAL unit type within the range of TRAIL_NUT to RSV_VCL_6 (for example, NAL unit type values of 0 to 6 in Table 1 or Table 2), or have the same NAL unit type as GRA_NUT.
[0182] The current VVC standard may have at least the following problems when dealing with pictures that have mixed NAL unit types:
[0183] 1. If a picture contains IDR and non-IRAP NAL units, and signaling for the RPL (reference picture list) exists in the slice header, then such signaling must also exist in the header of the IDR slice. RPL signaling exists in the slice header of an IDR slice if the value of sps_idr_rpl_present_flag is 1. Currently, the value of such a flag (sps_idr_rpl_present_flag) can be 0 even if there is one or more pictures with mixed NAL unit types. Here, the sps_idr_rpl_present_flag syntax element may indicate whether the RPL syntax element may exist in the slice header of a slice with NAL unit types such as IDR_N_LP or IDR_W_RADL. For example, if the value of sps_idr_rpl_present_flag is 1, it may indicate that the RPL syntax element may exist in the slice header of a slice with NAL unit types such as IDR_N_LP or IDR_W_RADL. Alternatively, if the value of sps_idr_rpl_present_flag is 0, it may indicate that the RPL syntax element is not present in the slice header of a slice with a NAL unit type such as IDR_N_LP or IDR_W_RADL.
[0184] 2. If the current picture references a PPS where the value of mixed_nalu_types_in_pic_flag is 1, then one or more of the VCL NAL units in the current picture must each have a specific NAL unit type within the range of IDR_W_RADL to CRA_NUT (for example, NAL unit type values from 7 to 9 in Table 1 or Table 2), and each of the other VCL NAL units must each have a specific NAL unit type within the range of TRAIL_NUT to RSV_VCL_6 (for example, NAL unit type values from 0 to 6 in Table 1 or Table 2), or have the same NAL unit type as GRA_NUT. Such constraints apply only to current pictures that include a mixture of IRAP and non-IRAP NAL unit types. However, they are not yet properly applied to pictures that include a mixture of RASL / RADL and non-IRAP NAL unit types.
[0185] This document provides solutions to the problems described above. Specifically, as previously stated, a picture containing two or more subpictures (i.e., a current picture) may have a mixed NAL unit type. In the current VVC standard, a picture with a mixed NAL unit type may have a mixed form of IRAP NAL unit types and non-IRAP NAL unit types. However, a leading picture associated with a CRA NAL unit type may also have a mixed form with non-IRAP NAL unit types, but such pictures with mixed NAL unit types are not currently supported in the standard. Therefore, a solution is needed for pictures that have a mixed form of CRA NAL unit types and non-IRAP NAL unit types.
[0186] Therefore, this document provides a method for allowing pictures that contain a mixture of leading picture NAL unit types (e.g., RASL_NUT, RADL_NUT) and other non-IRAP NAL unit types (e.g., TRAIL_NUT, STSA, NUT). Furthermore, this document defines constraints that ensure a reference picture list exists or is signaled when IDR subpictures are mixed with other non-IRAP subpictures. Thus, pictures with mixed NAL unit types are given more flexible characteristics by providing a mixed form that includes not only IRAP but also CRA NAL units.
[0187] For example, the following embodiments can be applied, thereby solving the aforementioned problems. The following embodiments may be applied individually or in combination.
[0188] In one embodiment, when allowing pictures to have mixed NAL unit types (when the value of mixed_nal_types_in_pic_flag is 1), signaling for the reference picture list is also allowed to exist for slices that have IDR type NAL unit types (e.g., IDR_W_RADL or IDR_N_LP). This constraint may be expressed as follows:
[0189] - If there is at least one PPS that references an SPS for which the value of mixed_nal_types_in_pic_flag eqaul is 1, then the value of sps_idr_rpl_present_flag must be 1. Such constraints may be requirements for bitstream conformance.
[0190] Alternatively, in one embodiment, a picture having mixed NAL unit types may be allowed to include slices having specific NAL unit types for the leading picture (e.g., RADL or RASL) and specific NAL unit types for the non-leading picture, non-IRAP. This can be shown as follows:
[0191] For the VCL NAL units of a specific picture, the following can be applied:
[0192] - If the value of mixed_nalu_types_in_pic_flag is 0, the value of NAL unit type (nal_unit_type) must be the same for all coded slice NAL units in the picture. A picture or PU may be considered to have the same NAL unit type as the coded slice NAL units in the picture or PU.
[0193] - Otherwise (if the value of mixed_nalu_types_in_pic_flag is 1), one of the following must be satisfied (i.e., one of the following may have a true value):
[0194] 1) Each of the VCL NAL units must have a specific NAL unit type (nal_unit_type) within the range of IDR_W_RADL to CRA_NUT (for example, NAL unit type values from 7 to 9 in Table 1 or Table 2), and each of the other VCL NAL units must have a specific NAL unit type within the range of TRAIL_NUT to RSV_VCL_6 (for example, NAL unit type values from 0 to 6 in Table 1 or Table 2), or have the same NAL unit type as GRA_NUT.
[0195] 2) One or more VCL NAL units must each have the same NAL unit type specific value as RADL_NUT (e.g., a NAL unit type value of 2 in Table 1 or Table 2) or RASL_NUT (e.g., a NAL unit type value of 3 in Table 1 or Table 2), and each other VCL NAL units must each have the same NAL unit type specific value as TRAIL_NUT (e.g., a NAL unit type value of 0 in Table 1 or Table 2), STSA_NUT (e.g., a NAL unit type value of 1 in Table 1 or Table 2), RSV_VCL_4 (e.g., a NAL unit type value of 4 in Table 1 or Table 2), RSV_VCL_5 (e.g., a NAL unit type value of 5 in Table 1 or Table 2), RSV_VCL_6 (e.g., a NAL unit type value of 6 in Table 1 or Table 2), or GRA_NUT.
[0196] On the other hand, this document proposes a method for providing a picture with the mixed NAL unit types described above, even for a single-layer bitstream. In one embodiment, the following constraints may apply to a single-layer bitstream.
[0197] - In the bitstream, each picture except the first one is considered to be related to the previous IRAP picture in the decoding order.
[0198] - If the picture is the leading picture of an IRAP picture, it must be a RADL or RASL picture.
[0199] - If the picture is a trailing picture of an IRAP picture, it must not be a RADL or RASL picture.
[0200] - RASL pictures must not be present within the bitstream associated with an IDR picture.
[0201] - A RADL picture must not exist within a bitstream associated with an IDR picture whose NAL unit type (nal_unit_type) is IDR_N_LP.
[0202] If referenced, and each parameter set is available, random access can be performed at the IRAP PU location by discarding all PUs prior to the IRAP PU (and also by accurately decoding the IRAP picture and all subsequent non-RASL pictures in the decoding order).
[0203] - Pictures that precede an IRAP picture in the decoding order must precede the IRAP picture in the output order, and must precede any RADL pictures associated with the IRAP picture in the output order.
[0204] - The RASL picture associated with the CRA picture must precede the RADL picture associated with the CRA picture in the output order.
[0205] - RASL pictures associated with CRA pictures must be output after IRAP pictures that precede CRA pictures in the decoding order.
[0206] - If the value of field_seq_flag is 0 and the picture is currently a reading picture associated with an IRAP picture, it must precede all non-reading pictures associated with the same IRAP picture in the decoding order. Otherwise, if pictures picA and picB are the first and last reading pictures associated with an IRAP picture in the decoding order, there can be at most one non-reading picture preceding picA in the decoding order, and there can be no non-reading pictures between picA and picB in the decoding order.
[0207] The following drawings were created to illustrate a specific example of this document. The names of specific devices, terms, and names (e.g., syntax / syntax element names) shown in the drawings are presented illustratively, and the technical features of this document are not limited to the specific names used in the following drawings.
[0208] Figure 15 schematically shows an example of a video / image encoding method to which embodiments of this document can be applied. The method disclosed in Figure 15 can be performed by the encoding device 200 disclosed in Figure 2.
[0209] Referring to Figure 15, the encoding device can determine the NAL unit type for a slice in the picture (S1500).
[0210] For example, as described in Tables 1 and 2 above, the encoding device can determine the NAL unit type according to the properties and type of the picture or subpicture, and can determine the NAL unit type for each slice based on the NAL unit type of the picture or subpicture.
[0211] For example, if the value of mixed_nalu_types_in_pic_flag is 0, the slices in the picture associated with the PPS can be determined to be of the same NAL unit type. That is, when the value of mixed_nalu_types_in_pic_flag is 0, the NAL unit type defined in the first NAL unit header of the first NAL unit containing information for the first slice of the picture is the same as the NAL unit type defined in the second NAL unit header of the second NAL unit containing information for the second slice of the same picture. Alternatively, if the value of mixed_nalu_types_in_pic_flag is 1, the slices in the picture associated with the PPS can be determined to be of a different NAL unit type. Here, the NAL unit type for the slices in the picture can be determined based on the method proposed in the embodiments described above.
[0212] The encoding device can generate NAL unit type-related information (S1510). The NAL unit type-related information may include information / syntax elements associated with the NAL unit types described in the embodiments described above and / or in Tables 1 and 2. For example, the information associated with the NAL unit type may include the mixed_nalu_types_in_pic_flag syntax element included in the PPS. The information associated with the NAL unit type may also include the nal_unit_type syntax element in the NAL unit header of the NAL unit, which contains information for the coded slice.
[0213] The encoding device can generate a bitstream (S1520). The bitstream may include at least one NAL unit containing video information for the coded slice. The bitstream may also include a PPS.
[0214] Figure 16 schematically shows an example of a video / image decoding method to which the embodiments of this document can be applied. The method disclosed in Figure 16 can be performed by the decoding device 300 disclosed in Figure 3.
[0215] Referring to Figure 16, the decoding device can receive a bitstream (S1600). Here, the bitstream may include at least one NAL unit containing video information for the coded slice. The bitstream may also include a PPS.
[0216] The decoding device can obtain NAL unit type-related information (S1610). The NAL unit type-related information may include information / syntax elements related to the NAL unit types described in the embodiments described above and / or in Tables 1 and 2. For example, the information related to the NAL unit type may include the mixed_nalu_types_in_pic_flag syntax element included in the PPS. The information related to the NAL unit type may also include the nal_unit_type syntax element in the NAL unit header of the NAL unit, which contains information for the coded slice.
[0217] The decoding device can determine the NAL unit type for a slice in the picture (S1620).
[0218] For example, if the value of mixed_nalu_types_in_pic_flag is 0, slices in a picture associated with a PPS use the same NAL unit type. That is, if the value of mixed_nalu_types_in_pic_flag is 0, the NAL unit type defined in the first NAL unit header of the first NAL unit containing information for the first slice of the picture is the same as the NAL unit type defined in the second NAL unit header of the second NAL unit containing information for the second slice of the same picture. Alternatively, if the value of mixed_nalu_types_in_pic_flag is 1, slices in a picture associated with a PPS use a different NAL unit type. Here, the NAL unit type for slices in a picture can be determined based on the method proposed in the embodiments described above.
[0219] The decoding device can decode / restore samples / blocks / slices based on the NAL unit type of the slice (S1630). Samples / blocks within a slice can be decoded / restored based on the NAL unit type of that slice.
[0220] For example, if a first NAL unit type is currently set for the first slice of a picture, and a second NAL unit type (different from the first NAL unit type) is currently set for the second slice of the picture, then the samples / blocks within the first slice or the first slice itself can be decoded / restored based on the first NAL unit type, and the samples / blocks within the second slice or the second slice itself can be decoded / restored based on the second NAL unit type.
[0221] Figures 17 and 18 schematically illustrate an example of a video / image encoding method and related components according to the embodiments described in this document.
[0222] The method disclosed in Figure 17 can be performed by the encoding device 200 disclosed in Figure 2 or Figure 18. Here, the encoding device 200 disclosed in Figure 18 is a simplified representation of the encoding device 200 disclosed in Figure 2. Specifically, steps S1700 to S1720 in Figure 17 can be performed by the entropy encoding unit 240 disclosed in Figure 2, and depending on the embodiment, each step can also be performed by the video splitting unit 210, prediction unit 220, residual processing unit 230, adder unit 340, etc., disclosed in Figure 2. Furthermore, the method disclosed in Figure 17 can be performed including the embodiments described above in this document. Accordingly, in Figure 17, specific explanations of content that overlaps with the embodiments described above will be omitted or simplified.
[0223] Referring to Figure 17, the encoding device can now determine the NAL unit type for the NAL units in the picture (S1700).
[0224] Currently, a picture can contain multiple slices, and a single slice can contain a slice header and slice data. Furthermore, a NAL unit can be generated by adding a NAL unit header to a slice (slice header and slice data). The NAL unit header may contain NAL unit type information identified by the slice data contained within that NAL unit.
[0225] In one embodiment, the encoding device can generate a first NAL unit for a first slice in the current picture and a second NAL unit for a second slice in the current picture. The encoding device can also determine the type of the first NAL unit for the first slice and the type of the second NAL unit for the second slice based on the types of the first and second slices.
[0226] For example, NAL unit types may include TRAIL_NUT, STSA_NUT, RADL_NUT, RASL_NUT, IDR_W_RADL, IDR_N_LP, CRA_NUT, etc., based on the type of slice data contained in the NAL unit, as shown in Table 1 or Table 2. Furthermore, NAL unit types may be signaled in the NAL unit header based on the nal_unit_type syntax element. The nal_unit_type syntax element is syntax information for identifying the NAL unit type and may be represented by a specific value corresponding to a particular NAL unit type, as shown in Table 1 or Table 2.
[0227] The encoding device can generate NAL unit type-related information for the NAL unit type (S1710).
[0228] NAL unit type-related information may include information / syntax elements related to the NAL unit types described in the embodiments and / or in Tables 1 and 2. For example, NAL unit type-related information may be information about whether the current picture has mixed NAL unit types, and may be indicated by the mixed_nalu_types_in_pic_flag syntax element included in the PPS. For example, if the value of the mixed_nalu_types_in_pic_flag syntax element is 0, it may indicate that the NAL units in the current picture have the same NAL unit type. Alternatively, if the value of the mixed_nalu_types_in_pic_flag syntax element is 1, it may indicate that the NAL units in the current picture have other NAL unit types.
[0229] In one embodiment, if all NAL unit types for the NAL units in the picture are currently the same, the encoding device can generate NAL unit type-related information (e.g., mixed_nalu_types_in_pic_flag) with a value of 0. Alternatively, if the NAL unit types for the NAL units in the picture are not currently the same, the encoding device can generate NAL unit type-related information (e.g., mixed_nalu_types_in_pic_flag) with a value of 1.
[0230] In other words, based on NAL unit type-related information indicating that the current picture has mixed NAL unit types (for example, a value of 1 for mixed_nalu_types_in_pic_flag), the first NAL unit for the first slice of the current picture and the second NAL unit for the second slice of the current picture may have different NAL unit types. Alternatively, based on NAL unit type-related information indicating that the current picture does not have mixed NAL unit types (for example, a value of 0 for mixed_nalu_types_in_pic_flag), the first NAL unit for the first slice of the current picture and the second NAL unit for the second slice of the current picture may have the same NAL unit type.
[0231] For example, based on NAL unit type-related information regarding the current picture having mixed NAL unit types (e.g., the value of mixed_nalu_types_in_pic_flag is 1), the first NAL unit for the first slice may have a leading picture NAL unit type, and the second NAL unit for the second slice may have a non-IRAP NAL unit type or a non-leading picture NAL unit type. Here, the leading picture NAL unit type may include the RADL NAL unit type or the RASL NAL unit type, and the non-IRAP NAL unit type or non-leading picture NAL unit type may include the trail NAL unit type or the STSA NAL unit type.
[0232] Alternatively, for example, based on NAL unit type-related information regarding the current picture having mixed NAL unit types (e.g., a value of mixed_nalu_types_in_pic_flag of 1), the first NAL unit for the first slice may have an IRAP NAL unit type, and the second NAL unit for the second slice may have a non-IRAP NAL unit type or a non-leading picture NAL unit type. Here, the IRAP NAL unit type may include an IDR NAL unit type (i.e., an IDR_N_LP NAL or IDR_W_RADL NAL unit type) or a CRA NAL unit type, and the non-IRAP NAL unit type or non-leading picture NAL unit type may include a trail NAL unit type or a STSA NAL unit type. Also, depending on the embodiment, the non-IRAP NAL unit type or non-leading picture NAL unit type may refer only to the trail NAL unit type.
[0233] Depending on the embodiment, if the current picture is allowed to have mixed NAL unit types, then for slices having IDR NAL unit types (e.g., IDR_W_RADL or IDR_N_LP) in the current picture, signaling-related information for the reference picture list must exist. The signaling-related information for the reference picture list may indicate whether a syntax element for the signaling of the reference picture list exists in the slice header of the slice. That is, based on a value of 1 for the signaling-related information for the reference picture list, a syntax element for the signaling of the reference picture list may exist in the slice header of a slice having IDR NAL unit types. Alternatively, based on a value of 0 for the signaling-related information for the reference picture list, a syntax element for the signaling of the reference picture list may not exist in the slice header of a slice having IDR NAL unit types.
[0234] For example, signaling-related information for a reference picture list can be the aforementioned sps_idr_rpl_present_flag syntax element. If the value of sps_idr_rpl_present_flag is 1, it may indicate that the syntax element for the signaling of the reference picture list may be present in the slice header of a slice with a NAL unit type such as IDR_N_LP or IDR_W_RADL. Alternatively, if the value of sps_idr_rpl_present_flag is 0, it may indicate that the syntax element for the signaling of the reference picture list is not present in the slice header of a slice with a NAL unit type such as IDR_N_LP or IDR_W_RADL.
[0235] The encoding device can encode video information including NAL unit type-related information (S1720).
[0236] For example, if the first NAL unit for the first slice in the current picture and the second NAL unit for the second slice in the current picture have different NAL unit types, the encoding device can encode video information that includes NAL unit type-related information with a value of 1 (e.g., mixed_nalu_types_in_pic_flag). Alternatively, if the first NAL unit for the first slice in the current picture and the second NAL unit for the second slice in the current picture have the same NAL unit type, the encoding device can encode video information that includes NAL unit type-related information with a value of 0 (e.g., mixed_nalu_types_in_pic_flag).
[0237] Furthermore, for example, an encoding device can encode video information that includes nal_unit_type information indicating the NAL unit type for each slice in the picture.
[0238] Furthermore, for example, an encoding device can encode video information that includes signaling-related information from a reference picture list (e.g., sps_idr_rpl_present_flag).
[0239] Furthermore, for example, an encoding device can now encode video information, including NAL units, for slices within a picture.
[0240] Video information containing the diverse information described above can be encoded and output in the form of a bitstream. The bitstream can be transmitted to a decoding device via a network or (digital) storage medium. Here, the network may include broadcast networks and / or communication networks, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.
[0241] Figures 19 and 20 schematically illustrate an example of a video / image decoding method and related components according to the embodiments described herein.
[0242] The method disclosed in Figure 19 can be performed by the decoding device 300 disclosed in Figure 3 or Figure 20. Here, the decoding device 300 disclosed in Figure 20 is a simplified representation of the decoding device 300 disclosed in Figure 3. Specifically, steps S1900 to S1920 in Figure 19 can be performed by the entropy decoding unit 310 disclosed in Figure 3, and depending on the embodiment, each step can also be performed by the residual processing unit 320, prediction unit 330, adder unit 340, etc., disclosed in Figure 3. Furthermore, the method disclosed in Figure 19 can be performed in this document including the embodiments described above. Accordingly, in Figure 19, specific explanations of content that overlaps with the embodiments described above will be omitted or simplified.
[0243] Referring to Figure 19, the decoding device can obtain video information, including NAL unit type-related information, from the bitstream (S1900).
[0244] In one embodiment, the decoding device can parse the bitstream to derive information necessary for video restoration (or picture restoration) (e.g., video / image information). In this case, the video information may include the aforementioned NAL unit type-related information (e.g., mixed_nalu_types_in_pic_flag), nal_unit_type information indicating each NAL unit type for slices in the current picture, signaling-related information for the reference picture list (e.g., sps_idr_rpl_present_flag), and NAL units for slices in the current picture. In other words, the video information may include various information necessary for the decoding process and can be decoded based on coding methods such as exponential Golomb coding, CAVLC, or CABAC.
[0245] As described above, NAL unit type-related information may include information / syntax elements related to the NAL unit types described in the embodiments and / or in Tables 1 and 2. For example, NAL unit type-related information may be information about whether the current picture has mixed NAL unit types, and may be indicated by the mixed_nalu_types_in_pic_flag syntax element included in the PPS. For example, if the value of the mixed_nalu_types_in_pic_flag syntax element is 0, it may indicate that the NAL units in the current picture have the same NAL unit type. Alternatively, if the value of the mixed_nalu_types_in_pic_flag syntax element is 1, it may indicate that the NAL units in the current picture have other NAL unit types.
[0246] The decoding device can determine the NAL unit type for the NAL unit currently in the picture based on the NAL unit type-related information (S1910).
[0247] Currently, a picture can contain multiple slices, and one slice can contain a slice header and slice data. Furthermore, a NAL unit can be generated by adding a NAL unit header to a slice (slice header and slice data). The NAL unit header may contain NAL unit type information identified by the slice data contained within that NAL unit.
[0248] For example, NAL unit types may include TRAIL_NUT, STSA_NUT, RADL_NUT, RASL_NUT, IDR_W_RADL, IDR_N_LP, CRA_NUT, etc., based on the type of slice data contained in the NAL unit, as shown in Table 1 or Table 2. Furthermore, NAL unit types may be signaled in the NAL unit header based on the nal_unit_type syntax element. The nal_unit_type syntax element is syntax information for identifying the NAL unit type and may be represented by a specific value corresponding to a particular NAL unit type, as shown in Table 1 or Table 2.
[0249] In one embodiment, the decoding device can determine that the first NAL unit for the first slice of the current picture and the second NAL unit for the second slice of the current picture have different NAL unit types, based on NAL unit type-related information (e.g., a value of mixed_nalu_types_in_pic_flag of 1) indicating that the current picture has a mixed NAL unit type. Alternatively, the decoding device can determine that the first NAL unit for the first slice of the current picture and the second NAL unit for the second slice of the current picture have the same NAL unit type, based on NAL unit type-related information (e.g., a value of mixed_nalu_types_in_pic_flag of 0) indicating that the current picture does not have a mixed NAL unit type.
[0250] For example, based on NAL unit type-related information regarding the current picture having mixed NAL unit types (e.g., the value of mixed_nalu_types_in_pic_flag is 1), the first NAL unit for the first slice may have a leading picture NAL unit type, and the second NAL unit for the second slice may have a non-IRAP NAL unit type or a non-leading picture NAL unit type. Here, the leading picture NAL unit type may include the RADL NAL unit type or the RASL NAL unit type, and the non-IRAP NAL unit type or non-leading picture NAL unit type may include the trail NAL unit type or the STSA NAL unit type.
[0251] Alternatively, for example, based on NAL unit type-related information regarding the current picture having mixed NAL unit types (e.g., a value of 1 for mixed_nalu_types_in_pic_flag), the first NAL unit for the first slice may have an IRAP NAL unit type, and the second NAL unit for the second slice may have a non-IRAP NAL unit type or a non-leading picture NAL unit type. Here, the IRAP NAL unit type may include an IDR NAL unit type (i.e., an IDR_N_LP NAL or IDR_W_RADL NAL unit type) or a CRA NAL unit type, and the non-IRAP NAL unit type or non-leading picture NAL unit type may include a trail NAL unit type or a STSA NAL unit type. Also, depending on the embodiment, the non-IRAP NAL unit type or non-leading picture NAL unit type may refer only to the trail NAL unit type.
[0252] Depending on the embodiment, if the current picture is allowed to have mixed NAL unit types, then for slices having IDR NAL unit types (e.g., IDR_W_RADL or IDR_N_LP) in the current picture, signaling-related information for the reference picture list must exist. The signaling-related information for the reference picture list may indicate whether a syntax element for the signaling of the reference picture list exists in the slice header of the slice. That is, based on a value of 1 for the signaling-related information for the reference picture list, a syntax element for the signaling of the reference picture list may exist in the slice header of a slice having IDR NAL unit types. Alternatively, based on a value of 0 for the signaling-related information for the reference picture list, a syntax element for the signaling of the reference picture list may not exist in the slice header of a slice having IDR NAL unit types.
[0253] For example, signaling-related information for a reference picture list can be the aforementioned sps_idr_rpl_present_flag syntax element. If the value of sps_idr_rpl_present_flag is 1, it may indicate that the syntax element for the signaling of the reference picture list may be present in the slice header of a slice with a NAL unit type such as IDR_N_LP or IDR_W_RADL. Alternatively, if the value of sps_idr_rpl_present_flag is 0, it may indicate that the syntax element for the signaling of the reference picture list is not present in the slice header of a slice with a NAL unit type such as IDR_N_LP or IDR_W_RADL.
[0254] The decoding device can now decode / restore the picture based on the NAL unit type (S1920).
[0255] For example, with respect to a first slice in the current picture determined to be a first NAL unit type and a second slice in the current picture determined to be a second NAL unit type, the decoding device can decode / restore the first slice based on the first NAL unit type and decode / restore the second slice based on the second NAL unit type. Furthermore, the decoding device can decode / restore samples / blocks in the first slice based on the first NAL unit type and decode / restore samples / blocks in the second slice based on the second NAL unit type.
[0256] In the embodiments described above, the method is explained based on a flowchart as a series of steps or blocks; however, the embodiments in this document are not limited to the order of the steps, and some steps may occur in a different order or simultaneously with other steps than those described above. Furthermore, those skilled in the art will understand that the steps shown in the flowchart are not exclusive, and other steps may be included, or one or more steps in the flowchart may be deleted without affecting the scope of this document.
[0257] The method described in this document can be implemented in software form, and the encoding and / or decoding devices described in this document may be included in, for example, video processing devices such as TVs, computers, smartphones, set-top boxes, and display devices.
[0258] In this document, when embodiments are implemented in software, the methods described above may be implemented by modules (processes, functions, etc.) that perform the functions described above. These modules may be stored in memory and executed by a processor. The memory may be internal or external to the processor and may be connected to the processor by various well-known means. The processor may include an ASIC (application-specific integrated circuit), other chipsets, logic circuits, and / or data processing devices. The memory may include ROM (read-only memory), RAM (random access memory), flash memory, memory cards, storage media, and / or other storage devices. That is, the embodiments described in this document may be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units shown in each drawing may be implemented and executed on a computer, processor, microprocessor, controller, or chip. In this case, information on instructions or algorithms for implementation may be stored on a digital storage medium.
[0259] Furthermore, decoding and encoding devices to which this document applies may include multimedia broadcasting transceivers, mobile communication terminals, home cinema video equipment, digital cinema video equipment, surveillance cameras, video interaction devices, real-time communication devices such as video communications, mobile streaming devices, storage media, camcorders, customized video (VoD) service providers, OTT video (Over the Top Video) devices, internet streaming service providers, 3D video devices, VR (virtual reality) devices, AR (argumente reality) devices, video telephone video devices, transportation terminals (e.g., vehicle terminals (including autonomous vehicles), airplane terminals, ship terminals, etc.), and medical video equipment, etc., and may be used to process video signals or data signals. For example, OTT video (Over the Top Video) devices may include game consoles, Blu-ray players, internet access TVs, home theater systems, smartphones, tablet PCs, DVRs (Digital Video Recorders), etc.
[0260] Furthermore, the processing methods to which the embodiments of this document apply can be produced in the form of programs executed on a computer and stored on a computer-readable recording medium. Multimedia data having a data structure according to the embodiments of this document can also be stored on a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices that store data to be read by a computer. The computer-readable recording medium may include, for example, Blu-ray discs (BDs), general-purpose serial buses (USBs), ROMs, PROMs, EPROMs, EEPROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, and optical data storage devices. The computer-readable recording medium also includes media embodied in the form of carrier waves (e.g., transmission over the Internet). Furthermore, a bitstream generated by an encoding method can be stored on a computer-readable recording medium or transmitted over a wired wireless network.
[0261] Furthermore, embodiments of this document can be embodied in computer program products comprising program code, and said program code can be executed on a computer according to the embodiments of this document. The said program code can be stored on a computer-readable carrier.
[0262] Figure 21 shows an example of a content streaming system to which the embodiments disclosed in this document may be applied.
[0263] Referring to Figure 21, the content streaming system to which the embodiments described in this document apply can broadly include an encoding server, a streaming server, a web server, media storage, user equipment, and multimedia input devices.
[0264] The encoding server is responsible for compressing content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream, and then transmitting this bitstream to the streaming server. As an alternative, if a multimedia input device such as a smartphone, camera, or camcorder directly generates the bitstream, the encoding server may be omitted.
[0265] The bitstream can be generated by an encoding method or a bitstream generation method to which an embodiment of this document applies, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.
[0266] The streaming server transmits multimedia data to user devices based on user requests via a web server, and the web server acts as an intermediary to inform users about available services. When a user requests a desired service from the web server, the web server transmits this to the streaming server, and the streaming server transmits multimedia data to the user. In this case, the content streaming system may include a separate control server, in which case the control server controls the commands and responses between the devices within the content streaming system.
[0267] The streaming server can receive content from a media storage and / or encoding server. For example, if it starts receiving content from the encoding server, it can receive the content in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.
[0268] Examples of user devices include mobile phones, smartphones, laptop computers, digital broadcasting terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (such as smartwatches, smart glasses, and HMDs), digital TVs, desktop computers, and digital signage.
[0269] Each server within the aforementioned content streaming system can be operated as a distributed server, in which case the data received by each server can be processed in a distributed manner.
[0270] The claims described herein can be combined in various ways. For example, the technical features of the method claims herein can be combined to embody an apparatus, and the technical features of the apparatus claims herein can be combined to embody a method. Furthermore, the technical features of the method claims and the technical features of the apparatus claims herein can be combined to embody an apparatus, and the technical features of the method claims and the technical features of the apparatus claims herein can be combined to embody a method.
Claims
1. In a video decoding method performed by a decoding device, The steps include obtaining video information from a bitstream, including network abstraction layer (NAL) unit type-related information and information related to signaling a reference picture list, The steps include: performing interpretation for the current block based on the aforementioned reference picture list; The aforementioned NAL unit type-related information is information regarding whether the picture currently has a mixed NAL unit type. The aforementioned video information further includes NAL unit type restriction information for limiting the value of the NAL unit type related information, Based on the fact that the NAL unit type-related information for the current picture has the mixed NAL unit type, the first NAL unit type for the first slice of the current picture is different from the second NAL unit type for the second slice of the current picture. The first NAL unit type is an IDR (instantaneous decoding refresh) NAL unit type, and the value of the information related to signaling the reference picture list is equal to 1. A method wherein, based on the value of the information relating to signaling the reference picture list being equal to 1, the syntax element for signaling the reference picture list is present in the slice header of the first slice having the IDR NAL unit type.
2. In a video encoding method performed by an encoding device, A step of generating NAL unit type-related information based on the NAL unit type, Steps to generate information related to signaling a list of reference pictures, The step of encoding video information including the NAL unit type-related information and the information related to signaling a reference picture list, The aforementioned NAL unit type-related information is information regarding whether the picture currently has a mixed NAL unit type. The aforementioned video information further includes NAL unit type restriction information for limiting the value of the NAL unit type related information, Based on the fact that the NAL unit type-related information for the current picture has the mixed NAL unit type, the first NAL unit type for the first slice of the current picture is different from the second NAL unit type for the second slice of the current picture. The first NAL unit type is an IDR (instantaneous decoding refresh) NAL unit type, and the syntax elements for signaling a reference picture list are present in the slice header of the first slice having the IDR NAL unit type.
3. A method for transmitting video data, A step of generating a bitstream of the aforementioned video, wherein the bitstream is: A step of generating NAL unit type-related information based on the NAL unit type, Steps to generate information related to signaling a list of reference pictures, A step of encoding video information including the NAL unit type related information and the information related to signaling the reference picture list, and a step of generating based on this, The step of transmitting the data, which includes the bitstream, The aforementioned NAL unit type-related information is information regarding whether the picture currently has a mixed NAL unit type. The aforementioned video information further includes NAL unit type restriction information for limiting the value of the NAL unit type related information, Based on the fact that the NAL unit type-related information for the current picture has the mixed NAL unit type, the first NAL unit type for the first slice of the current picture is different from the second NAL unit type for the second slice of the current picture. The first NAL unit type is an IDR (instantaneous decoding refresh) NAL unit type, and the syntax elements for signaling a reference picture list are present in the slice header of the first slice having the IDR NAL unit type.