Image information encoding / decoding apparatus and apparatus for bitstream
By determining image slices with mixed NAL unit types based on NAL unit type information, and performing signal notification and encoding for slices of specific NAL unit types, the problem of encoding efficiency for high-resolution, high-quality image/video data is solved, achieving more efficient image/video compression and encoding.
Patent Information
- Application Number
- CN202511336078.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-23
- Filing Date
- 2020-12-10
- Publication Date
- 2025-11-14
AI Technical Summary
Existing technologies struggle to effectively compress and encode high-resolution, high-quality image/video data, especially images with mixed NAL unit types, leading to increased transmission and storage costs.
By determining whether an image has a mixed NAL unit type based on NAL unit type related information, and by signaling and encoding slices of images with mixed NAL unit types for specific NAL unit types, effective encoding of information related to the reference image list is allowed.
It improves image/video encoding efficiency, especially for images with mixed NAL unit types, and can flexibly send signals to notify and encode reference image list information, thereby improving overall compression efficiency.
Smart Images

Figure CN120956932A_ABST
Abstract
Description
[0001] This application is a divisional application of the original invention patent application No. 202080096777.6 (International Application No.: PCT / KR2020 / 018061, Application Date: December 10, 2020, Invention Title: Image or Video Coding Based on NAL Unit Type for Slices or Pictures). Technical Field
[0002] This technology relates to video or image coding, for example, to image or video coding techniques based on Network Abstraction Layer (NAL) unit types for slicing or images. Background Technology
[0003] Recently, there has been a growing demand for high-resolution, high-quality images / videos, such as 4K or 8K Ultra High Definition (UHD) images / videos, across various fields. As image / video resolution or quality increases, relatively more information or bits are transmitted compared to traditional image / video data. Therefore, if image / video data is transmitted via media such as existing wired / wireless broadband lines or stored in traditional storage media, the costs of transmission and storage can easily increase.
[0004] In addition, there is growing interest and demand for virtual reality (VR) and artificial reality (AR) content, as well as immersive media such as holograms; and the broadcasting of images / videos that exhibit characteristics different from actual images / videos (e.g., game images / videos) is also increasing.
[0005] Therefore, highly efficient image / video compression technology is needed to effectively compress and send, store, or play high-resolution, high-quality images / videos that exhibit the various characteristics described above.
[0006] In addition, methods are needed to improve image / video coding efficiency, and for this purpose, methods for effectively signaling and encoding information related to Network Abstraction Layer (NAL) units are necessary. Summary of the Invention
[0007] Technical issues
[0008] This document provides methods and devices for improving the efficiency of video / image coding.
[0009] This document also provides methods and devices for improving video / image coding efficiency based on NAL unit-related information.
[0010] This document also provides methods and apparatus for improving the efficiency of video / image coding for images with mixed (hybrid) NAL unit types.
[0011] This document also provides methods and apparatus for allowing signaling notifications or information related to the existence of a list of reference images for slices in an image with a specific NAL unit type relative to an image with mixed NAL unit types.
[0012] Technical solution
[0013] According to the embodiments of this document, it is possible to determine whether an image has a mixed NAL unit type based on NAL unit type related information, and the NAL unit type can be determined for slices of an image with mixed NAL unit types. For example, based on a value of 1 for the NAL unit type related information, the NAL unit type of the first slice in the image may have a preceding image NAL unit type, and the NAL unit type of the second slice in the image may have a non-Intra-Random Access Point (IRAP) NAL unit type or a non-preamble image NAL unit type. According to the embodiments of this document, based on the case where an image is allowed to have mixed NAL unit types, for slices in an image with a specific NAL unit type, information related to signaling a reference image list can exist.
[0014] According to the implementation of this document, based on the case where images are allowed to have mixed NAL unit types, for slices in an image that have a specific NAL unit type, information related to signaling a list of reference images may exist.
[0015] According to embodiments of this document, a video / image decoding method performed by a decoding device is provided. The video / image decoding method may include the methods disclosed in the embodiments of this document.
[0016] According to embodiments of this document, a decoding apparatus is provided for performing video / image decoding. The decoding apparatus can execute the methods disclosed in the embodiments of this document.
[0017] According to embodiments of this document, a video / image coding method performed by an encoding device is provided. The video / image coding method may include the methods disclosed in embodiments of this document.
[0018] According to embodiments of this document, an encoding apparatus for performing video / image encoding is provided. The encoding apparatus can perform the methods disclosed in the embodiments of this document.
[0019] According to embodiments of this document, a computer-readable digital storage medium is provided for storing encoded video / image information generated by a video / image encoding method disclosed in at least one of the embodiments of this document.
[0020] According to embodiments of this document, a computer-readable digital storage medium is provided that stores encoded information or encoded video / image information that enables a decoding device to perform at least one of the video / image decoding methods disclosed in embodiments of this document.
[0021] Technical effect
[0022] This document can have various effects. For example, according to the embodiments of this document, the overall image / video compression efficiency can be improved. Additionally, according to the embodiments of this document, video / image coding efficiency can be improved based on NAL unit-related information. Furthermore, according to the embodiments of this document, the video / image coding efficiency of images with mixed NAL unit types can be improved. Furthermore, according to the embodiments of this document, for images with mixed NAL unit types, signaling and encoding of reference image list-related information can be effectively achieved. Furthermore, according to the embodiments of this document, by allowing images to include leading image NAL unit types (e.g., RASL_NUT, RADL_NUT) and other non-IRAP NAL unit types (e.g., TRAIL_NUT, STSA, NUT) in a mixed form, images with mixed NAL unit types can be provided with a form that mixes not only IRAP but also other types of NAL units, thereby enabling more flexible characteristics.
[0023] The effects achievable through the detailed examples in this document are not limited to those listed above. For example, there may be various technical effects that can be understood or derived from this document by a person skilled in the art. Therefore, the detailed effects of this document are not limited to those explicitly stated herein, but may include various effects that can be understood or derived from the technical features of this document. Attached Figure Description
[0024] Figure 1 Examples of video / image coding systems applicable to the embodiments of this document are illustrated.
[0025] Figure 2 This is a schematic illustration of the configuration of a video / image encoding device to which the embodiments of this document can be applied.
[0026] Figure 3 This is a schematic diagram illustrating the configuration of a video / image decoding device to which the embodiments described in this document can be applied.
[0027] Figure 4 Examples of illustrative video / image coding methods applicable to the implementation of this document.
[0028] Figure 5 Examples of illustrative video / image decoding methods applicable to the embodiments described in this document.
[0029] Figure 6 An example is shown to represent the hierarchical structure of an encoded image / video.
[0030] Figure 7 An example of an entropy coding method applicable to the embodiments described in this document is illustrated schematically. Figure 8 An entropy encoder in an encoding device is illustrated schematically.
[0031] Figure 9 An example of an entropy decoding method applicable to the embodiments of this document is illustrated schematically. Figure 10 An entropy decoder in a decoding device is illustrated schematically.
[0032] Figure 11 This is a diagram illustrating the temporal layer structure of NAL cells in a bitstream that supports temporal scalability.
[0033] Figure 12 It is a diagram used to describe an image that may be accessed randomly.
[0034] Figure 13 This is a diagram used to describe IDR images.
[0035] Figure 14 This is a diagram used to describe CRA images.
[0036] Figure 15 Examples of video / image coding methods applicable to the implementation of this document are illustrated.
[0037] Figure 16 Examples of video / image decoding methods applicable to the implementation of this document are illustrated.
[0038] Figure 17 and Figure 18 Examples of video / image coding methods and related components according to embodiments of this document are illustrated.
[0039] Figure 19 and Figure 20 Examples of video / image decoding methods and related components according to embodiments of this document are illustrated schematically.
[0040] Figure 21 Examples of content streaming systems applicable to the implementations disclosed in this document are illustrated. Detailed Implementation
[0041] This disclosure may be modified in various forms, and specific embodiments thereof will be described and illustrated in the accompanying drawings. However, these embodiments are not intended to limit this disclosure. The terminology used in the following description is for the purpose of describing specific embodiments only and is not intended to limit this disclosure. Singular expressions include plural expressions, provided that different interpretations are clear. Terms such as “comprising” and “having” are intended to indicate the presence of the features, quantities, steps, operations, elements, components, or combinations thereof used in the following description, and therefore it should be understood that the possibility of having or adding one or more different features, quantities, steps, operations, elements, components, or combinations thereof is not excluded.
[0042] Furthermore, the various configurations described in the accompanying drawings are illustrated independently to illustrate functions that are distinct from each other, and do not imply that the configurations are implemented using different hardware or different software. For example, two or more configurations may be combined to form one configuration, and a configuration may be divided into multiple configurations. Without departing from the spirit of this document, embodiments in which configurations are combined and / or separated are included within the scope of the claims.
[0043] In this document, the term "A or B" may mean "A only", "B only", or "both A and B". In other words, in this document, the term "A or B" may be interpreted as indicating "A and / or B". For example, in this document, the term "A, B or C" may mean "A only", "B only", "C only", or "any combination of A, B, and C".
[0044] In this document, a forward slash ( / ) or a comma can mean "and / or". For example, "A / B" can mean "A and / or B". Therefore, "A / B" can mean "A only", "B only", or "both A and B". For example, "A, B, C" can mean "A, B, or C".
[0045] In this document, "at least one of A and B" can mean "only A", "only B" or "both A and B". Furthermore, in this document, the expressions "at least one of A or B" or "at least one of A and / or B" can be interpreted as the same as "at least one of A and B".
[0046] Additionally, in this document, "at least one of A, B, and C" may mean "A only", "B only", "C only" or "any combination of A, B, and C". Furthermore, "at least one of A, B, or C" or "at least one of A, B, and / or C" may mean "at least one of A, B, and C".
[0047] Additionally, the parentheses used in this document may mean "for example". Specifically, when expressing "prediction (intra-prediction)", it may indicate an example where "intra-prediction" is proposed as "prediction". In other words, the term "prediction" in this document is not limited to "intra-prediction", and may indicate an example where "intra-prediction" is proposed as "prediction". Furthermore, even when expressing "prediction (i.e., intra-prediction)", it may indicate an example where "intra-prediction" is proposed as "prediction".
[0048] This document relates to video / image coding. For example, the methods / implementations disclosed in this document can be applied to methods disclosed in the Universal Video Coding (VVC) standard. Additionally, the methods / implementations disclosed in this document can be applied to methods disclosed in the Basic Video Coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the second-generation Audio Video Coding (AVS2) standard, or next-generation video / image coding standards (e.g., H.267 or H.268).
[0049] This document proposes various implementations of video / image coding, and unless otherwise mentioned, the above implementations can also be combined with each other.
[0050] In this document, video can refer to a collection of images over time. An image generally refers to a unit representing an image within a specific time period, and a slice / tile is a unit that constitutes part of an image during encoding. A slice / tile can include one or more Code Tree Units (CTUs). An image can consist of one or more slices / tiles. A tile is a rectangular area of CTUs within a specific tile column and a specific tile row in an image. A tile column is a rectangular area of CTUs with a height equal to the height of the image and a width specified by a syntax element in the image parameter set. A tile row is a rectangular area of CTUs with a width specified by a syntax element in the image parameter set and a height equal to the height of the image. A tile scan is a specific ordering of CTUs segmented in an image: CTUs are sequentially ordered in a raster scan of CTUs within a tile, and tiles within an image are sequentially ordered in a raster scan of tiles within the image. A slice includes an integer number of complete tiles of an image that can be exclusively contained within a single NAL unit, or an integer number of consecutive complete CTU rows within a tile.
[0051] Furthermore, an image can be divided into two or more sub-images. A sub-image can be a rectangular region of one or more slices within the image.
[0052] A pixel, or image unit, can refer to the smallest unit that makes up a picture (or image). Alternatively, "sample" can be used as the term corresponding to a pixel. A sample can typically represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component. Alternatively, a sample can refer to a pixel value in the spatial domain, or, when a pixel value is transformed to the frequency domain, to the transform coefficients in the frequency domain.
[0053] A unit can represent a basic unit of image processing. A unit may include a specific region of an image and at least one of the information associated with that region. A unit may include a luminance block and two chrominance (e.g., cb, cr) blocks. In some cases, the term "unit" may be used interchangeably with terms such as "block" or "region". In general, an M×N block may include a set (or array) of samples (or sample arrays) or transform coefficients in M columns and N rows.
[0054] Furthermore, in this document, at least one of quantization / dequantization and / or transform / inverse transform may be omitted. When quantization / dequantization is omitted, the quantization transform coefficients may be referred to as transform coefficients. When transform / inverse transform is omitted, the transform coefficients may be referred to as coefficients or residual coefficients, or, for the sake of consistency, may still be referred to as transform coefficients.
[0055] In this document, quantization transform coefficients and transform coefficients can be referred to as transform coefficients and scaling transform coefficients, respectively. In this case, residual information can include information about the transform coefficients and can be signaled via residual coding syntax. Transform coefficients can be derived based on residual information (or information about the transform coefficients), and scaling transform coefficients can be derived through the inverse transform (scaling) of the transform coefficients. Residual samples can be derived based on the inverse transform (scaling) of the scaling transform coefficients. This can also be applied / expressed in other parts of this document.
[0056] In this document, a technical feature described separately in a single figure may be implemented individually or simultaneously.
[0057] In the following, preferred embodiments of this document are described in more detail with reference to the accompanying drawings. In the drawings, the same reference numerals are used for the same elements, and redundant descriptions of the same elements may be omitted.
[0058] Figure 1 Examples of video / image coding systems to which the implementation methods of this document can be applied are illustrated.
[0059] Reference Figure 1A video / image encoding system may include a source device and a receiving device. The source device may transmit encoded video / image information or data to the receiving device in the form of a file or stream via a digital storage medium or network.
[0060] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.
[0061] Video sources can acquire video / images through processes that capture, synthesize, or generate video / images. Video sources may include video / image capture devices and / or video / image generation devices. For example, a video / image capture device may include one or more cameras, a video / image archive containing previously captured video / images, etc. For example, a video / image generation device may include a computer, tablet computer, and smartphone, and may generate video / images (electronically). For example, virtual video / images may be generated via a computer, etc. In this case, the video / image capture process may be replaced by a process that generates related data.
[0062] Encoding devices can encode input video / images. For compression and encoding efficiency, encoding devices can perform a series of processes such as prediction, transformation, and quantization. The encoded data (encoded video / image information) can be output as a bitstream.
[0063] The transmitter can send encoded images / image information or data, output as a bitstream, to the receiver of the receiving device in the form of a file or stream via a digital storage medium or network. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter may include elements for generating media files according to a predetermined file format and may include elements for transmission over a broadcast / communication network. The receiver can receive / extract the bitstream and send the received bitstream to a decoding device.
[0064] Decoding devices can decode video / images by performing a series of processes such as inverse quantization, inverse transform, and prediction, which correspond to the operations of encoding devices.
[0065] The renderer can render decoded video / images. The rendered video / images can be displayed on a monitor.
[0066] Figure 2This is a schematic illustration of the configuration of a video / image encoding apparatus to which the embodiments of this document can be applied. Hereinafter, the term "encoding apparatus" may include image encoding apparatus and / or video encoding apparatus.
[0067] Reference Figure 2 The encoding device 200 may include and be configured with an image segmenter 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, an inverse quantizer 234, and an inverse transformer 235. The residual processor 230 may also include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstruction block generator. According to embodiments, the image segmenter 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 described above may be configured by one or more hardware components (e.g., an encoder chipset or processor). Additionally, the memory 270 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware components may also include the memory 270 as an internal / external component.
[0068] Image segmenter 210 can segment an input image (or picture or frame) input to encoding device 200 into one or more processing units. For example, a processing unit may be referred to as a coding unit (CU). In this case, the coding unit can be recursively segmented from a coding tree unit (CTU) or a maximum coding unit (LCU) according to a quadtree-binary-tritree (QTBTTT) structure. For example, a coding unit can be segmented into multiple deeper coding units based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. In this case, for example, a quadtree structure may be applied first, followed by a binary tree structure and / or a ternary tree structure. Alternatively, a binary tree structure may be applied first. The encoding process according to this disclosure can be performed based on the final coding unit that is no longer segmented. In this case, the maximum coding unit may be used as the final coding unit based on image characteristics, coding efficiency, etc., or, if necessary, the coding unit may be recursively segmented into deeper coding units such that a coding unit with an optimal size can be used as the final coding unit. Here, the encoding process may include processes such as prediction, transformation, and reconstruction (described later). As another example, the processing unit may also include a prediction unit (PU) or a transform unit (TU). In this case, each of the prediction unit and the transform unit may be split or divided from the aforementioned final encoding unit. The prediction unit may be a unit for predicting samples, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving residual signals from transform coefficients.
[0069] In some cases, a unit can be used interchangeably with terms such as block or region. Typically, an M×N block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can typically represent a pixel or pixel value, and can represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component. A sample can be used as a term corresponding to the pixels or cells that make up a picture (or image).
[0070] Encoding device 200 generates a residual signal (residual block, residual sample array) by subtracting the prediction signal (prediction block, prediction sample array) output from inter-frame predictor 221 or intra-frame predictor 222 from the input image signal (original block, original sample array), and the generated residual signal is sent to converter 232. In this case, as illustrated, the unit for subtracting the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) within encoding device 200 can be referred to as subtractor 231. The predictor can perform prediction on the block to be processed (hereinafter referred to as the current block) and generate a prediction block including the prediction samples of the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction on a unit of the current block or CU. The predictor can generate various information about the prediction, such as prediction mode information, to transmit the generated information to entropy encoder 240, as described later in the description of each prediction mode. The information about the prediction can be encoded by entropy encoder 240 and output as a bitstream.
[0071] Intra-predictor 222 can refer to samples in the current image to predict the current block. Depending on the prediction mode, the referenced samples can be located near or far from the current block. In intra-prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. For example, non-directional modes can include DC mode and planar mode. For example, depending on the fineness of the prediction direction, the directional modes can include 33 directional prediction modes or 65 directional prediction modes. However, this is just an example, and more or fewer directional prediction modes can be used depending on the settings. Intra-predictor 222 can use the prediction modes applied to neighboring blocks to determine the prediction mode applied to the current block.
[0072] Inter-frame predictor 221 can deduce the predicted block of the current block based on a reference block (reference sample array) specified by a motion vector on a reference image. Here, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation between motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference image indices. Motion information may also include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current image and temporally neighboring blocks existing in the reference image. The reference image including the reference block and the reference image including the temporally neighboring block may be the same or different from each other. The temporally neighboring block may be referred to as a juxtaposed reference block, a juxtaposed CU (colCU), etc., and the reference image including the temporally neighboring block may also be referred to as a juxtaposed image (colPic). For example, inter-frame predictor 221 can configure a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to deduce the motion vector and / or reference image index of the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in skip mode and merge mode, the inter-frame predictor 221 can use motion information of neighboring blocks as motion information for the current block. In skip mode, unlike merge mode, residual signals may not be sent. The motion vector prediction (MVP) mode indicates the motion vector of the current block by using the motion vectors of neighboring blocks as motion vector predictors and signaling the motion vector difference.
[0073] Predictor 220 can generate prediction signals based on various prediction methods described below. For example, the predictor can not only apply intra-frame prediction or inter-frame prediction to predict a block, but also apply both intra-frame prediction and inter-frame prediction simultaneously. This can be referred to as combined intra-frame and inter-frame prediction (CIIP). Alternatively, the predictor can perform prediction on blocks based on an intra-block copy (IBC) prediction mode or a palette mode. The IBC prediction mode or palette mode can be used for content image / video coding such as in games, etc. IBC essentially performs prediction in the current image, but it performs similarly to inter-frame prediction in deriving reference blocks in the current image. That is, IBC can use at least one of the inter-frame prediction techniques described in this document. The palette mode can be considered as an example of intra-frame coding or intra-frame prediction. When applying a palette mode, sample values in the image can be signaled based on information about the palette index and the palette table.
[0074] The predicted signal generated by the predictor (including inter-frame predictor 221 and / or intra-frame predictor 222) can be used to generate a reconstructed signal or a residual signal. Transformer 232 can generate transform coefficients by applying transform techniques to the residual signal. For example, the transform technique can include at least one of Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Karhunen-Loève Transform (KLT), Graph-Based Transform (GBT), or Conditional Nonlinear Transform (CNT). Here, when the relationship information between pixels is illustrated as a graph, GBT refers to the transform obtained from that graph. CNT refers to the transform obtained based on the predicted signal generated using all previously reconstructed pixels. Furthermore, the transform processing can also be applied to pixel blocks of the same square size, and can also be applied to blocks of variable size that are not square.
[0075] Quantizer 233 can quantize the transform coefficients to send the quantized transform coefficients to entropy encoder 240, and entropy encoder 240 can encode the quantized signal (information about the quantized transform coefficients) into an encoded quantized signal and output it as a bitstream. The information about the quantized transform coefficients can be referred to as residual information. Quantizer 233 can rearrange the quantized transform coefficients in block form into a one-dimensional vector form based on the coefficient scan order, and also generate information about the quantized transform coefficients based on the one-dimensional vector form of the quantized transform coefficients. Entropy encoder 240 can perform various encoding methods such as Golomb coding, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). Entropy encoder 240 can also encode information necessary for video / image reconstruction other than the quantized transform coefficients (e.g., values of syntax elements, etc.) together or separately. The encoded information (e.g., encoded video / image information) can be sent as a bitstream or stored in units of Network Abstraction Layer (NAL) units. The video / image information may also include information about various parameter sets such as Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). Additionally, the video / image information may also include general constraint information. The information and / or syntax elements to be signaled / sent, as described subsequently in this document, can be encoded using the encoding process mentioned above and thus included in the bitstream. The bitstream can be transmitted over a network or stored in a digital storage medium. Here, the network may include broadcast networks and / or communication networks, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. A transmitter (not shown) for transmitting the signal output from the entropy encoder 240 and / or a memory (not shown) for storing the signal can be configured as internal / external components of the encoding device 200, or the transmitter may also be included within the entropy encoder 240.
[0076] The quantized transform coefficients output from quantizer 233 can be used to generate a prediction signal. For example, dequantizer 234 and inverse transformer 235 vectorize the transform coefficients and apply dequantization and inverse transform to reconstruct the residual signal (residual block or residual sample). Adder 250 adds the reconstructed residual signal to the prediction signal output from inter-frame predictor 221 or intra-frame predictor 222 to generate a reconstructed signal (reconstructed image, reconstructed block, reconstructed sample array). If no residual exists in the block to be processed, such as when a skip mode is applied, the predicted block can be used as the reconstructed block. Adder 250 can be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current image, and, as described later, for inter-frame prediction of the next image by filtering.
[0077] In addition, luminance mapping with chroma scaling (LMCS) can be applied in image encoding and / or reconstruction processing.
[0078] Filter 260 can apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, filter 260 can apply various filtering methods to the reconstructed image to generate a modified reconstructed image, which is then stored in memory 270, specifically in the DPB of memory 270. Various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filtering, bilateral filtering, etc. Filter 260 can generate various filtering-related information to transmit the generated information to entropy encoder 240, as described subsequently in the description of each filtering method. The filtering-related information can be encoded by entropy encoder 240 and output as a bitstream.
[0079] The modified reconstructed image sent to memory 270 can be used as a reference image in inter-frame predictor 221. Applying inter-frame prediction through the inter-frame predictor can avoid prediction mismatch between encoding device 200 and decoding device, and can improve encoding efficiency.
[0080] The DPB of memory 270 can store modified reconstructed images for use as reference images in inter-frame predictor 221. Memory 270 can store motion information of blocks from which motion information within the current image is derived (or encoded) and / or motion information of blocks within previously reconstructed images. The stored motion information can be transmitted to inter-frame predictor 221 to be used as motion information for spatially or temporally neighboring blocks. Memory 270 can store reconstructed samples of reconstructed blocks in the current image and can transmit these reconstructed samples to intra-frame predictor 222.
[0081] Figure 3 This is a schematic diagram illustrating the configuration of a video / image decoding device applicable to this document. In the following text, the term "decoding device" may include image decoding devices and / or video decoding devices.
[0082] Reference Figure 3The decoding device 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-frame predictor 332 and an intra-frame predictor 331. The residual processor 320 may include an inverse quantizer 321 and an inverse transformer 322. According to embodiments, the entropy decoder 310, residual processor 320, predictor 330, adder 340, and filter 350 described above may be configured by hardware components (e.g., a decoder chipset or processor). Additionally, the memory 360 may include a decoded image buffer (DPB) or may be configured by a digital storage medium. The hardware components may also include the memory 360 as an internal / external component.
[0083] When an input bitstream including video / image information is received, the decoding device 300 can respond to... Figure 2 The encoding apparatus illustrated herein processes video / image information to reconstruct the image. For example, decoding apparatus 300 may deduce units / blocks based on block segmentation information obtained from the bitstream. Decoding apparatus 300 may use processing units applied to the encoding apparatus to perform decoding. Thus, for example, the decoding processing unit may be an encoding unit, and the encoding unit may be segmented from encoding tree units or maximum encoding units according to a quadtree structure, binary tree structure, and / or ternary tree structure. One or more transform units may be derived from the encoding units. Furthermore, the reconstructed image signal decoded and output by decoding apparatus 300 may be reproduced by a reproduction apparatus.
[0084] Decoding device 300 can receive data from... in the form of a bitstream. Figure 2The signal output by the encoding device illustrated herein can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can deduce the information necessary for image reconstruction (or picture reconstruction) (e.g., video / image information) by parsing the bitstream. The video / image information may also include information about various parameter sets such as Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), and Video Parameter Set (VPS). In addition, the video / image information may also include general constraint information. The decoding device can further decode the picture based on the information about the parameter sets and / or general constraint information. The transmitted / received information and / or syntax elements, which will be described subsequently in this document, can be decoded and obtained from the bitstream through the decoding process. For example, the entropy decoder 310 can decode the information within the bitstream based on encoding methods such as Exponential Golomb coding, CAVLC, or CABAC, and output the values of the syntax elements necessary for image reconstruction and the quantized values of the residual correlation transform coefficients. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element from the bitstream, determine the context model using information about the syntax element to be decoded, as well as decoding information of neighboring blocks and the block to be decoded, or information about symbols / bins decoded in the previous stage, and generate symbols corresponding to the values of each syntax element by predicting bin generation probabilities based on the determined context model and performing arithmetic decoding of the bins. At this point, the CABAC entropy decoding method can determine the context model and then update the context model using information about decoded symbols / bins for the context model of the next symbol / bin. Prediction information from the information decoded by the entropy decoder 310 can be provided to the predictors (inter-frame predictor 332 and intra-frame predictor 331), and the residual values (i.e., quantization transform coefficients and related parameter information) from the entropy decoding performed by the entropy decoder 310 can be input to the residual processor 320. The residual processor 320 can derive the residual signals (residual blocks, residual samples, residual sample arrays). Additionally, filtering information from the information decoded by the entropy decoder 310 can be provided to the filter 350. Furthermore, the receiver (not illustrated) for receiving the signal output from the encoding device can be further configured as an internal / external component of the decoding device 300, or the receiver can also be a component of the entropy decoder 310. Additionally, the decoding device according to this document can be referred to as a video / image / picture decoding device, and the decoding device can also be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoder 310, and the sample decoder may include at least one of an inverse quantizer 321, an inverse transformer 322, an adder 340, a filter 350, a memory 360, an inter-frame predictor 332, and an intra-frame predictor 331.
[0085] The dequantizer 321 can dequantize the quantized transform coefficients to output transform coefficients. The dequantizer 321 can rearrange the quantized transform coefficients in a two-dimensional block format. In this case, the rearrangement can be performed based on the coefficient scan order performed by the encoding device. The dequantizer 321 can use quantization parameters (e.g., quantization step size information) to perform dequantization on the quantized transform coefficients and obtain the transform coefficients.
[0086] The inverse transformer 322 performs inverse transformation on the transformation coefficients to obtain the residual signal (residual block, residual sample array).
[0087] Predictor 330 can perform prediction for the current block and generate a prediction block that includes prediction samples of the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction to the current block based on information about the prediction output from entropy decoder 310, and can determine a specific intra-frame / inter-frame prediction mode.
[0088] The predictor can generate a prediction signal based on various prediction methods described below. For example, the predictor can not only apply intra-frame prediction or inter-frame prediction to predict a block, but also apply both intra-frame prediction and inter-frame prediction simultaneously. This can be referred to as combined intra-frame and inter-frame prediction (CIIP). Alternatively, the predictor can perform prediction on blocks based on an intra-block copy (IBC) prediction mode or a palette mode. The IBC prediction mode or palette mode can be used for content image / video coding such as in games with screen content coding (SCC). IBC essentially performs prediction in the current image, but it performs similarly to inter-frame prediction in deriving reference blocks in the current image. That is, IBC can use at least one of the inter-frame prediction techniques described in this document. The palette mode can be considered as an example of intra-frame coding or intra-frame prediction. When a palette mode is applied, information about the palette table and palette index can be included in the video / image information and signaled.
[0089] Intra-predictor 331 can predict the current block by referencing samples in the current image. Depending on the prediction mode, the referenced samples may be located among the neighbors of the current block, or their location may be separate from the current block. In intra-prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. Intra-predictor 331 can also determine the prediction mode to be applied to the current block by using prediction modes applied to neighboring blocks.
[0090] Inter-frame predictor 332 can deduce the predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference image. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the motion information correlation between neighboring blocks and the current block. Motion information may include motion vectors and reference image indices. Motion information may also include information about the inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current image and temporally neighboring blocks existing in the reference image. For example, inter-frame predictor 332 can construct a motion information candidate list based on neighboring blocks and deduce the motion vector and / or reference image index for the current block based on the received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and the information about the prediction may include information indicating the mode of inter-frame prediction for the current block.
[0091] Adder 340 can add the obtained residual signal to the prediction signal (prediction block or prediction sample array) output from predictor 330 (including intra-frame predictor 331 and inter-frame predictor 332) to generate a reconstruction signal (reconstructed image, reconstruction block, or reconstruction sample array). If no residual exists for the block to be processed when a jump mode is applied, the prediction block can be used as a reconstruction block.
[0092] Adder 340 can be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current image, and as described later, it can also be output by filtering or used for inter-frame prediction of the next image.
[0093] In addition, Luminance Mapping with Chroma Scaling (LMCS) can also be applied to image decoding processing.
[0094] Filter 350 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 350 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image and store the modified reconstructed image in memory 360, specifically in the DPB of memory 360. Various filtering methods may include, for example, deblocking filtering, adaptive sample shifting, adaptive loop filtering, bilateral filtering, etc.
[0095] The (modified) reconstructed image stored in the DPB of memory 360 can be used as a reference image in inter-frame predictor 332. Memory 360 can store motion information of blocks in the current image from which motion information is derived (or decoded) and / or motion information of blocks in already reconstructed images. The stored motion information can be transmitted to inter-frame predictor 332 to be used as motion information for spatially or temporally neighboring blocks. Memory 360 can store reconstructed samples of reconstructed blocks in the current image and transmit the reconstructed samples to intra-frame predictor 331.
[0096] The exemplary embodiments described in this document in the filter 260, inter-frame predictor 221 and intra-frame predictor 222 of the encoding device 200 can be applied equally or correspondingly to the filter 350, inter-frame predictor 332 and intra-frame predictor 331 of the decoding device 300.
[0097] Furthermore, as described above, prediction is performed during video encoding to enhance compression efficiency. Accordingly, a prediction block can be generated, comprising prediction samples as the current block to be encoded (i.e., the target coding block). Here, the prediction block includes prediction samples in the spatial domain (or pixel domain). The prediction block is derived identically in both the encoding and decoding devices, and the encoding device can signal the decoding device with information about the residual (rather than the original sample values of the original block itself) between the original block and the prediction block (residual information) to enhance image coding efficiency. The decoding device can derive a residual block including residual samples based on the residual information, add the residual block and the prediction block to generate a reconstructed block including reconstructed samples, and generate a reconstructed image including the reconstructed block.
[0098] Residual information can be generated through transformation and quantization processes. For example, the encoding device can derive the residual block between the original block and the prediction block, perform a transformation process on the residual samples (residual sample array) included in the residual block to derive the transform coefficients, perform a quantization process on the transform coefficients to derive the quantized transform coefficients, and signal the relevant residual information (via bitstream) to the decoding device. In this case, the residual information may include the values of the quantized transform coefficients, position information, transform scheme, transform kernel, and quantization parameters. The decoding device can perform inverse quantization / inverse transform based on the residual information and derive residual samples (or residual blocks). The decoding device can generate a reconstructed image based on the prediction block and the residual block. Furthermore, for inter-frame prediction reference of subsequent images, the encoding device can also perform inverse quantization / inverse transform on the quantized transform coefficients to derive the residual block and generate a reconstructed image based on this.
[0099] Furthermore, as described above, when performing prediction on the current block, intra-frame prediction or inter-frame prediction can be applied. In an implementation, when applying inter-frame prediction to the current block, the predictor of the encoding / decoding device (more specifically, the inter-frame predictor) can derive prediction samples by performing inter-frame prediction on a block-by-block basis. Inter-frame prediction can represent a prediction derived by a method that depends on data elements (e.g., sample values or motion information) of images other than the current image. When applying inter-frame prediction to the current block, the prediction block (prediction sample array) of the current block can be derived based on the reference block (reference sample array) specified by the motion vector on the reference image indicated by the reference image index. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, the motion information of the current block can be predicted on a block, sub-block, or sample-by-sample basis based on the correlation of motion information between neighboring blocks and the current block. Motion information can include motion vectors and reference image indices. Motion information can also include inter-frame prediction type (L0 prediction, L1 prediction, Bi prediction, etc.) information. When applying inter-frame prediction, neighboring blocks can include spatially neighboring blocks existing in the current image and temporally neighboring blocks existing in the reference image. The reference picture including the reference block and the reference picture including the temporally neighboring block can be the same as or different from each other. The temporally neighboring block can be called a name such as a juxtaposed reference block, a juxtaposed CU (colCU), etc., and the reference picture including the temporally neighboring block can be called a juxtaposed picture (colPic). For example, a motion information candidate list can be configured based on the neighboring blocks of the current block, and a signal can be sent to indicate which candidate's flag or index information to select (use) in order to derive the motion vector and / or reference picture index of the current block. Inter-frame prediction can be performed based on various prediction modes, and for example, in skip mode and merge mode, the motion information of the current block can be the same as the motion information of the selected neighboring block. In the case of skip mode, residual signaling is not required as in merge mode. In the case of motion vector prediction (MVP) mode, the motion vector of the selected neighboring block can be used as a motion vector predictor, and the motion vector difference can be signaled. In this case, the motion vector of the current block can be derived using the sum of the motion vector predictor and the motion vector difference.
[0100] Depending on the inter-frame prediction type (L0 prediction, L1 prediction, Bi prediction, etc.), motion information may also include L0 motion information and / or L1 motion information. The L0 direction motion vector may be referred to as the L0 motion vector or MVL0, and the L1 direction motion vector may be referred to as the L1 motion vector or MVL1. Prediction based on the L0 motion vector may be called L0 prediction, prediction based on the L1 motion vector may be called L1 prediction, and prediction based on both L0 and L1 motion vectors may be called bidirectional prediction. Here, the L0 motion vector may indicate the motion vector associated with the reference image list L0, and the L1 motion vector may indicate the motion vector associated with the reference image list L1. As reference images, the reference image list L0 may include images preceding the current image in output order, and the reference image list L1 may include images following the current image in output order. The preceding images may be referred to as forward (reference) images, and the following images may be referred to as backward (reference) images. The reference image list L0 may also include images following the current image in output order as reference images. In this scenario, the previous images can be indexed first in the reference image list L0, followed by the subsequent images. The reference image list L1 can also include images preceding the current image in the output order as reference images. In this case, the subsequent images can be indexed first in the reference image list L1, followed by the previous images. Here, the output order can correspond to the Image Order Count (POC) order.
[0101] Figure 4 Examples of illustrative video / image coding methods applicable to the implementation of this document.
[0102] Figure 4 The method disclosed in the article can be derived from Figure 2 The above-mentioned encoding device 200 performs the above-mentioned functions. Specifically, S400 can be performed by the inter-frame predictor 221 or the intra-frame predictor 222 of the encoding device 200, and S410, S420, S430 and S440 can be performed by the subtractor 231, the transformer 232, the quantizer 233 and the entropy encoder 240 of the encoding device 200, respectively.
[0103] Reference Figure 4 The encoding device can derive prediction samples by predicting the current block (S400). The encoding device can determine whether to perform inter-frame prediction or intra-frame prediction for the current block, and determine a specific inter-frame prediction mode or a specific intra-frame prediction mode based on the RD cost. According to the determined mode, the encoding device can derive prediction samples for the current block.
[0104] The encoding device can compare the predicted sample of the current block with the original sample and derive the residual sample (S410).
[0105] The encoding device can derive the transform coefficients by transforming the residual samples (S420), and can derive the quantized transform coefficients by quantizing the derived transform coefficients (S430).
[0106] Quantization can be performed based on quantization parameters. Transformation and / or quantization processes can be skipped. When transform processing is skipped, the (quantized) (residual) coefficients of the residual samples can be encoded according to the residual encoding technique described later. For terminology consistency, the (quantized) (residual) coefficients can also be referred to as (quantized) transform coefficients.
[0107] The encoding device can encode image information including residual information and prediction information, and can output the encoded image information in the form of a bitstream (S440). The prediction information may include information about motion information (e.g., when inter-frame prediction is applied) and prediction mode information as multiple pieces of information related to the prediction process. The residual information may include information about the quantization transform coefficients. The residual information can be entropy encoded. Alternatively, the residual information may include information about the (quantization) (residual) coefficients.
[0108] The output bitstream can be transmitted to the decoding device via storage media or network.
[0109] Figure 5 Examples of illustrative video / image decoding methods applicable to the embodiments described in this document.
[0110] Figure 5 The method disclosed in the article can be derived from Figure 3 The above-mentioned decoding device 300 performs this operation. Specifically, S500 can be performed by the inter-frame predictor 332 or the intra-frame predictor 331 of the decoding device 300. In S500, the entropy decoder 310 of the decoding device 300 performs the process of decoding the prediction information included in the bitstream and deriving the values of the relevant syntax elements. S510, S520, S530 and S540 can be performed by the entropy decoder 310, the dequantizer 321, the inverse transformer 322 and the adder 340 of the decoding device 300, respectively.
[0111] Reference Figure 5 The decoding device can perform operations corresponding to those already performed in the encoding device. The decoding device can perform inter-frame prediction or intra-frame prediction on the current block and derive prediction samples based on the received prediction information (S500).
[0112] The decoding device can derive the quantization transform coefficients of the current block based on the received residual information (S510). The decoding device can derive the quantization transform coefficients from the residual information through entropy decoding.
[0113] The decoding device can dequantize the quantization transform coefficients and derive the transform coefficients (S520). Dequantization can be performed based on the quantization parameters.
[0114] The decoding device can derive the residual sample by performing an inverse transformation of the transform coefficients (S530).
[0115] Inverse transform and / or dequantization can be skipped. When inverse transform is skipped, (quantized) (residual) coefficients can be derived from the residual information, and residual samples can be derived based on the (quantized) (residual) coefficients.
[0116] The decoding device can generate reconstructed samples for the current block based on residual samples and predicted samples, and generate a reconstructed image based on these reconstructed samples (S540). Thereafter, loop filtering can be further applied to the reconstructed image as described above.
[0117] Figure 6 An example is shown of the hierarchical structure of the encoded image / video.
[0118] Reference Figure 6 The encoded image / video is divided into a VCL (Video Coding Layer) that manipulates image / video decoding processing and its own subsystem, a subsystem that sends and stores encoded information, and a Network Abstraction Layer (NAL) that exists between the VCL and the subsystem and is responsible for network adaptation functions.
[0119] VCL can generate VCL data including compressed image data (slice data), or generate parameter sets or image decoding processing including image parameter sets (image parameter set: PPS), sequence parameter sets (sequence parameter set: SPS), video parameter sets (video parameter set: VPS), etc., plus necessary supplementary enhancement information (SEI) messages.
[0120] In NAL, NAL cells can be generated by adding header information (NAL cell header) to the raw byte sequence payload (RBSP) generated in VCL. In this case, RBSP refers to slice data, parameter sets, SEI messages, etc., generated in VCL. The NAL cell header can include NAL cell type information specified according to the RBSP data included in the corresponding NAL cell.
[0121] Furthermore, based on the RBSP generated in the VCL, NAL units can be divided into VCL NAL units and non-VCL NAL units. A VCL NAL unit can refer to a NAL unit that includes information about the image (slice data), while a non-VCL NAL unit can refer to a NAL unit that includes information required for decoding the image (parameter set or SEI message).
[0122] VCL NAL units and non-VCL NAL units can be transmitted over a network according to the data standard appender information of the subsystem. For example, NAL units can be converted into predetermined standard data formats such as H.266 / VVC file format, Real-time Transport Protocol (RTP), and Transport Stream (TS) and transmitted over various networks.
[0123] As described above, in a NAL cell, the NAL cell type can be specified according to the RBSP data structure included in the corresponding NAL cell, and information about the NAL cell type can be stored in the NAL cell header and signaled.
[0124] For example, based on whether NAL units include information about the image (slice data), NAL units can be broadly classified into VCL NAL unit types and non-VCL NAL unit types. VCL NAL unit types can be classified according to the nature and type of the image included in the VCL NAL unit, while non-VCL NAL unit types can be classified according to the type of parameter set.
[0125] The following is an example of a NAL cell type specified based on the type of the parameter set included in a non-VCL NAL cell type.
[0126] -APS (Adaptive Parameter Set) NAL Unit: The type of NAL unit including APS.
[0127] -DPS (Decoding Parameter Set) NAL Unit: The type of NAL unit including DPS.
[0128] -VPS (Video Parameter Set) NAL Unit: Includes the type of NAL unit for the VPS.
[0129] -SPS (Sequence Parameter Set) NAL Unit: The type of NAL unit that includes SPS.
[0130] -PPS (Image Parameter Set) NAL Unit: The type of NAL unit including PPS.
[0131] -PH (Image Header) NAL Unit: The type of NAL unit including PH.
[0132] The aforementioned NAL unit types have syntax information specific to the NAL unit type, and this syntax information can be stored in the NAL unit header and signaled. For example, the syntax information can be `nal_unit_type`, and the NAL unit type can be specified through the `nal_unit_type` value.
[0133] Furthermore, as mentioned above, an image can include multiple slices, and a slice can include a slice header and slice data. In this case, an image header can be further added to the multiple slices (slice header and slice dataset) in an image. The image header (image header syntax) can include information / parameters shared for the image. In this document, a tile group can be mixed with or replaced by a slice or image. Additionally, in this document, a tile group header can be mixed with or replaced by a slice header or image header.
[0134] A slice header (slice header syntax) may include information / parameters shared for slices. An APS (APS syntax) or PPS (PPS syntax) may include information / parameters shared for one or more slices or images. An SPS (SPS syntax) may include information / parameters shared for one or more sequences. A VPS (VPS syntax) may include information / parameters shared for multiple layers. A DPS (DPS syntax) may include information / parameters shared for the entire video. A DPS may include information / parameters related to the concatenation of encoded video sequences (CVS). In this document, the High-Level Syntax (HLS) may include at least one of the following: APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, image header syntax, and slice header syntax.
[0135] In this document, the image / video information encoded in the encoding device and transmitted to the decoding device in the form of a bitstream can include not only image segmentation information, intra / inter-frame prediction information, residual information, and loop filtering information in the image, but also information included in the slice header, image header, APS, PPS, SPS, VPS, and / or DPS. Additionally, the image / video information may also include information from the NAL unit header.
[0136] As described above, High-Level Syntax (HLS) can be encoded / signaled for use in video / image coding. In this document, video / image information may include HLS. For example, an encoded picture may consist of one or more slices. Parameters describing the encoded picture may be signaled in the Picture Header (PH), and parameters describing the slices may be signaled in the Slice Header (SH). The PH may be sent according to its own NAL unit type. The SH may be present at the beginning of the NAL unit, which includes the slice's payload (i.e., slice data). Details of the syntax and semantics of the PH and SH may be as disclosed in the VVC standard. Each picture may be associated with a PH. Pictures may consist of different types of slices: intra-frame coded slices (i.e., I-slices) and inter-frame coded slices (i.e., P-slices and B-slices). As a result, the PH may include the syntax elements necessary for intra-frame and inter-frame slices of the picture.
[0137] Furthermore, as mentioned above, the encoding device performs entropy encoding based on various encoding methods such as Exponential Golomb coding, Context Adaptive Variable Length Coding (CAVLC), and Context Adaptive Binary Arithmetic Coding (CABAC). Similarly, the decoding device can perform entropy decoding based on encoding methods such as Exponential Golomb coding, CAVLC, or CABAC. The entropy encoding / decoding process will be described below.
[0138] Figure 7 An example of an entropy coding method applicable to the embodiments of this document is illustrated schematically, and Figure 8 An entropy encoder in an encoding device is illustrated schematically. Figure 8 The entropy encoder in the encoding device can also be applied equivalently or accordingly to the above. Figure 2 Entropy encoder 240 of encoding device 200.
[0139] Reference Figure 7 and Figure 8 The encoding device (entropy encoder) performs entropy coding processing on image / video information. Image / video information may include segmentation-related information, prediction-related information (e.g., inter-frame / intra-frame prediction differentiation information, intra-frame prediction mode information, inter-frame prediction mode information, etc.), residual information, loop filtering-related information, or may include various related syntax elements. Entropy coding can be performed within syntax element units. S700 and S710 can be... Figure 2 The above-mentioned entropy encoder 240 of the encoding device 200 is executed.
[0140] The encoding device can perform binary conversion on the target syntax element (S700). Here, binary conversion can be based on various binary conversion methods such as truncated Rice binary conversion, fixed-length binary conversion, etc., and the binary conversion method for the target syntax element can be predefined. The binary conversion process can be performed by the binary converter 242 in the entropy encoder 240.
[0141] The encoding device can perform entropy encoding on the target syntax element (S710). The encoding device can encode the bin string of the target syntax element using a rule-based (context-based) or bypass-based encoding scheme such as context-adaptive arithmetic coding (CABAC) or context-adaptive variable-length coding (CAVLC), and its output can be merged into the bitstream. The entropy encoding process can be performed by the entropy encoding processor 243 in the entropy encoder 240. As described above, the bitstream can be sent to the decoding device via a (digital) storage medium or network.
[0142] Figure 9 An example of an entropy decoding method applicable to the embodiments of this document is illustrated schematically, and Figure 10 An entropy decoder in a decoding device is illustrated schematically. Figure 10 The entropy decoder in the decoding device can also be used with the above. Figure 3 The decoding device 300 is equivalent to or corresponds to the entropy decoder 310.
[0143] refer to Figure 9 and Figure 10 The decoding device (entropy decoder) can decode the encoded image / video information. The image / video information may include segmentation-related information, prediction-related information (e.g., inter-frame / intra-frame prediction differentiation information, intra-frame prediction mode information, inter-frame prediction mode information, etc.), residual information, loop filtering-related information, or may include various related syntax elements. Entropy coding can be performed in syntax element units. S900 and S910 can be... Figure 3 The aforementioned entropy decoder 310 of the decoding device 300 is executed.
[0144] The decoding device can perform binary conversion on the target syntax element (S900). Here, binary conversion can be based on various binary conversion methods such as truncated Rice binary conversion, fixed-length binary conversion, etc., and a predefined binary conversion method for the target syntax element can be used. The decoding device can derive the enable bin string (bin string candidate) of the enable value of the target syntax element through the binary conversion process. The binary conversion process can be performed by the binary converter 312 in the entropy decoder 310.
[0145] The decoding device can perform entropy decoding (S910) on the target syntax element. When decoding and parsing each bin of the target syntax element sequentially from the input bits in the bitstream, the decoding device compares the derived bin string with the enabled bin string of the corresponding syntax element. When the derived bin string is the same as one of the enabled bin strings, the value corresponding to the bin string can be deduced as the value of the syntax element. If not, the above process can be repeated after further parsing the next bit in the bitstream. Through these processes, even if a start bit or end bit is not used for specific information (a specific syntax element) in the bitstream, the decoding device can use variable-length bits to signal that information. Accordingly, relatively fewer bits can be allocated to low values, thereby improving overall encoding efficiency.
[0146] Decoding devices can perform context-based or bypass-based decoding on individual bins in a bin string from a bitstream based on entropy coding techniques such as CABAC and CAVLC. In this regard, the bitstream can include various information for image / video decoding as described above. As mentioned above, the bitstream can be transmitted to the decoding device via (digital) storage media or a network.
[0147] In addition, a NAL unit type can typically be set for an image. The NAL unit type can be signaled through the nal_unit_type in the NAL unit header that includes the slice. nal_unit_type is syntax information used to specify the NAL unit type; that is, as shown in Table 1 or Table 2 below, it can specify the type of RBSP data structures contained in the NAL unit.
[0148] Table 1 below shows examples of NAL unit type codes and NAL unit type categories.
[0149] [Table 1]
[0150]
[0151]
[0152] Alternatively, as an example, NAL unit type codes and NAL unit type categories can be defined as shown in Table 2 below.
[0153] [Table 2]
[0154]
[0155]
[0156]
[0157] As shown in Table 1 or Table 2, the name and value of the NAL unit type can be specified based on the RBSP data structure included in the NAL unit, and can be divided into VCLNAL unit type and non-VCL NAL unit type based on whether the NAL unit includes information about the image (slice data). VCL NAL unit types can be classified according to the nature and type of the image, and non-VCL NAL unit types can be classified according to the type of parameter set. For example, the NAL unit type can be specified based on the nature and type of the image included in the VCL NAL unit.
[0158] TRAIL: This indicates the type of NAL unit that includes coded slice data of trailing images / sub-images. For example, nal_unit_type can be defined as TRAIL_NUT, and the value of nal_unit_type can be specified as 0.
[0159] Here, a trailing image refers to an image that follows another image and can be accessed randomly in both output and decoding order. A trailing image can be a non-IRAP image that follows its associated IRAP image in output order, and it is not an STSA image. For example, a trailing image associated with an IRAP image follows the IRAP image in decoding order. Images that are both following their associated IRAP image in output order and preceding it in decoding order are not allowed.
[0160] STSA (Step-by-Step Temporal Sublayer Access): This indicates the type of NAL unit that contains the coded slice data of the STSA image / subimage. For example, nal_unit_type can be defined as STSA_NUT, and the value of nal_unit_type can be specified as 1.
[0161] Here, an STSA image is an image that can be switched between temporal sublayers in a bitstream that supports temporal scalability, and it indicates the position where an upswitching from a lower sublayer to an upper sublayer one step above it is possible. An STSA image does not use images with the same TemporalId as the STSA image or images in the same layer as the STSA image for inter-frame prediction. Images following the STSA image with the same TemporalId in the same layer in decoding order do not use images preceding the STSA image with the same TemporalId in the same layer in decoding order for inter-frame prediction reference. STSA images enable upswitching from the immediately following sublayer to a sublayer that includes the STSA image. In this case, the encoded image must not belong to the lowest sublayer. That is, the STSA image must always have a TemporalId greater than 0.
[0162] RADL (Random Access Decodeable Preamble (Picture)): This indicates the type of NAL unit that contains the slice data of the RADL picture / subpicture to be encoded. For example, nal_unit_type can be defined as RADL_NUT, and the value of nal_unit_type can be specified as 2.
[0163] Here, all RADL images are leading images. RADL images are not used as reference images for decoding trailing images of the same associated IRAP image. Specifically, a RADL image with a nuh_layer_id equal to its layerId is the image that follows the IRAP image associated with it in output order, and is not used as a reference image for decoding images with a nuh_layer_id equal to their layerId. When field_seq_flag (i.e., sps_field_seq_flag) is 0, all RADL images precede all non-leading images of the same associated IRAP image in decoding order (i.e., if a RADL image exists). Furthermore, a leading image is the image that precedes the associated IRAP image in output order.
[0164] RASL (Random Access Skip Preamble (Picture)): This indicates the type of NAL unit that includes the slice data of the RASL picture / subpicture to be encoded. For example, nal_unit_type can be defined as RASL_NUT, and the value of nal_unit_type can be specified as 3.
[0165] Here, all RASL images are preceding images of their associated CRA images. When an associated CRA image has a NoOutputBeforeRecoveryFlag value of 1, the RASL image may neither be output nor correctly decoded because it may contain references to images not present in the bitstream. RASL images are not used as reference images for decoding non-RASL images at the same layer. However, RADL sub-images in RASL images at the same layer can be used for inter-frame prediction of juxtaposed RADL sub-images in RADL images associated with the same CRA image as the RASL image. When field_seq_flag (i.e., sps_field_seq_flag) is 0, all RASL images are decoded before all non-preceding images of the same associated CRA image (i.e., if a RASL image exists).
[0166] For non-IRAP VCL NAL unit types, a reserved nal_unit_type can exist. For example, nal_unit_type can be defined as RSV_VCL_4 and RSV_VCL_6, and the value of nal_unit_type can be specified as 4 to 6 respectively.
[0167] Here, Intra-Frame Random Access Point (IRAP) refers to information indicating the NAL units of a picture that can be randomly accessed. IRAP pictures can be CRA pictures or IDR pictures. For example, an IRAP picture is a picture with NAL unit types defined as IDR_W_RADL, IDR_N_LP, and CRA_NUT as shown in Table 1 or Table 2 above, and the value of nal_unit_type can be specified as 7 to 9 respectively.
[0168] IRAP pictures do not use any reference pictures in the same layer for inter-frame prediction during decoding. In other words, IRAP pictures do not reference any pictures other than themselves for inter-frame prediction during decoding. The first picture in the bitstream in decoding order is called the IRAP or GDR picture. For a single-layer bitstream, if the necessary set of parameters is available when it needs to be referenced, all subsequent non-RASL pictures and IRAP pictures in the coded layer video sequence (CLVS) in decoding order can be decoded accurately without performing decoding processing on pictures that precede the IRAP picture in decoding order.
[0169] The value of `mixed_nalu_types_in_pic_flag` for an IRAP image is 0. When the value of `mixed_nalu_types_in_pic_flag` for an image is 0, one slice in the image can have a NAL unit type (nal_unit_type) in the range from `IDR_W_RADL` to `CRA_NUT` (e.g., NAL unit type values of 7 to 9 in Table 1 or Table 2), and all other slices in the image can have the same NAL unit type (nal_unit_type). In this case, the image can be considered an IRAP image.
[0170] Instantaneous Decoding Refresh (IDR): This indicates the type of NAL unit that includes the slice data of the IDR image / subimage to be encoded. For example, the nal_unit_type of the IDR image / subimage can be defined as IDR_W_RADL or IDR_N_LP, and the value of nal_unit_type can be specified as 7 or 8 respectively.
[0171] Here, an IDR image may not use inter-frame prediction during decoding (i.e., it does not reference images other than itself for inter-frame prediction), but it may be the first image in the bitstream in decoding order, or it may appear later in the bitstream (i.e., not first, but later). Each IDR image is the first image in the encoded video sequence (CVS) in decoding order. For example, when an IDR image is associated with a decodeable preamble image, its NAL unit type can be represented as IDR_W_RADL, while when an IDR image is not associated with a preamble image, its NAL unit type can be represented as IDR_N_LP. That is, an IDR image with NAL unit type IDR_W_RADL may not have an associated RASL image present in the bitstream, but it may have an associated RADL image present in the bitstream. An IDR image with NAL unit type IDR_N_LP does not have an associated preamble image present in the bitstream.
[0172] Clean Random Access (CRA): This indicates the type of NAL unit that includes the slice data of the CRA image / subimage to be encoded. For example, nal_unit_type can be defined as CRA_NUT, and the value of nal_unit_type can be specified as 9.
[0173] Here, a CRA image may not use inter-frame prediction during decoding (i.e., it does not reference images other than itself for inter-frame prediction), but it can be the first image in the bitstream in decoding order, or it may appear later in the bitstream (i.e., not first, but later). A CRA image may have associated RADL or RASL images present in the bitstream. For a CRA image where the NoOutputBeforeRecoveryFlag value is 1, the associated RASL image may not be output by the decoder. This is because decoding is impossible in this case due to the inclusion of references to images not present in the bitstream.
[0174] Gradual Decoding Refresh (GDR): This indicates the type of NAL unit that includes the slice data of the GDR image / sub-image to be encoded. For example, nal_unit_type can be defined as GDR_NUT, and the value of nal_unit_type can be specified as 10.
[0175] Here, the pps_mixed_nalu_types_in_pic_flag value of the GDR image can be 0. When the value of pps_mixed_nalu_types_in_pic_flag of the image is 0 and one slice in the image has a GDR_NUT NAL unit type, all other slices in the image have the same NAL unit type (nal_unit_type) value, and in this case, the image can become a GDR image after receiving the first slice.
[0176] Additionally, for example, the NAL unit type can be specified based on the types of parameters included in the non-VCL NAL unit, and as shown in Table 1 or Table 2 above, the following NAL unit types (nal_unit_type) can be included: for example, VPS_NUT indicating the type of NAL unit including a video parameter set, SPS_NUT indicating the type of NAL unit including a sequence parameter set, PPS_NUT indicating the type of NAL unit including a picture parameter set, and PH_NUT indicating the type of NAL unit including a picture header.
[0177] Furthermore, the bitstream supporting time scalability (or time-scalable bitstream) includes information about the time layer of the time scaling. This time layer information can be identification information for the time layer specified according to the time scalability of the NAL unit. For example, the time layer identification information can use the temporal_id syntax information, and this temporal_id syntax information can be stored in the NAL unit header in the encoding device and signaled to the decoding device. Hereinafter, in this specification, the time layer may be referred to as a sublayer, time sublayer, time-scalable layer, etc.
[0178] Figure 11 This is a diagram illustrating the temporal layer structure of NAL cells in a bitstream that supports temporal scalability.
[0179] When a bitstream supports temporal scalability, the NAL units included in the bitstream have temporal layer identification information (e.g., temporal_id). As an example, a temporal layer consisting of NAL units with a temporal_id value of 0 can provide the lowest temporal scalability, while a temporal layer consisting of NAL units with a temporal_id value of 2 can provide the highest temporal scalability.
[0180] exist Figure 11 In the diagram, blocks marked with "I" refer to image "I", and blocks marked with "B" refer to image "B". Additionally, arrows indicate reference relationships between images, specifying whether an image references another image.
[0181] like Figure 11As shown, a NAL unit in a time layer with a temporal_id value of 0 is a reference image that can be referenced by NAL units in time layers with temporal_id values of 0, 1, or 2. A NAL unit in a time layer with a temporal_id value of 1 is a reference image that can be referenced by NAL units in time layers with temporal_id values of 1 or 2. A NAL unit in a time layer with a temporal_id value of 2 can be a reference image that can be referenced by NAL units in the same time layer (i.e., the time layer with a temporal_id value of 2), or it can be a non-reference image that is not referenced by other images.
[0182] If as Figure 11 If the NAL unit of the time layer (i.e., the highest time layer) with a temporal_id value of 2 is a non-reference image, then these NAL units are extracted (or removed) from the bitstream during the decoding process without affecting other images.
[0183] Furthermore, among the aforementioned NAL unit types, IDR and CRA types indicate information about NAL units that include pictures capable of random access (or splicing) (i.e., random access point (RAP) pictures or intra-frame random access point (IRAP) pictures used as random access points). In other words, an IRAP picture can be an IDR or CRA picture and can consist only of I-slices. In the bitstream, the first picture in decoding order becomes the IRAP picture.
[0184] If the bitstream includes IRAP images (IDR, CRA images), there can be images that appear before the IRAP images in output order but follow them in decoding order. These images are called preamble images (LPs).
[0185] Figure 12 It is an illustration used to describe images that can be accessed randomly.
[0186] The randomly accessible picture (i.e., the RAP or IRAP picture used as a random access point) is the first picture in the bitstream in decoding order during random access and includes only I slices.
[0187] Figure 12 The output (or display) order and decoding order of the images are shown. As illustrated, the output and decoding orders of the images may differ from each other. For convenience, the images are described while being divided into predetermined groups.
[0188] Images belonging to Group 1 (I) are images that precede the IRAP images in both output and decoding order. Images belonging to Group 2 (II) are images that precede the IRAP images in output order but follow them in decoding order. Images belonging to Group 3 (III) are images that follow the IRAP images in both output and decoding order.
[0189] The images in the first group (I) can be decoded and output, regardless of the IRAP images.
[0190] The image that is output before the IRAP image and belongs to the second group (II) is called the preamble image, and the preamble image may become a problem in the decoding process when the IRAP image is used as a random access point.
[0191] Images belonging to group III, following the IRAP images in both output and decoding order, are called normal images. Normal images are not used as reference images for leading images.
[0192] The random access point that appears in the bitstream is called the IRAP picture, and random access begins as the first picture of the second group (II) is output.
[0193] Figure 13 This is a diagram used to describe IDR images.
[0194] An IDR image is an image that becomes a random access point when a set of images has a closed structure. As mentioned above, since an IDR image is an IRAP image, it only consists of I slices and can be the first image in the bitstream in decoding order, or it can appear in the middle of the bitstream. When an IDR image is decoded, all reference images stored in the Decoded Image Buffer (DPB) are marked as "not used for reference".
[0195] Figure 13 The bars shown indicate the images, and the arrows indicate the reference relationship between the images and whether another image can be used as a reference image. An 'x' mark on the arrow indicates that the image cannot be referenced by the image indicated by that arrow.
[0196] As shown, the image with a POC of 32 is the IDR image. Images with POCs of 25 to 31, and output before the IDR image, are the leading image 1310. Images with a POC of 33 or greater correspond to the normal image 1320.
[0197] The preceding image 1310, which precedes the IDR image in the output order, can use a preceding image different from the IDR image as a reference image, but the previous image 1330, which precedes the preceding image 1310 in both the output and decoding order, cannot be used as a reference image.
[0198] The normal image 1320, which follows the IDR image in the output and decoding order, can be decoded by referring to the IDR image, the preceding image, and other normal images.
[0199] Figure 14 This is a diagram used to describe CRA images.
[0200] A CRA image is an image that becomes a random access point when a set of images has an open structure. As mentioned above, since a CRA image is also an IRAP image, it only consists of I-slices and can be the first image in the bitstream in decoding order, or it can appear in the middle of the bitstream for normal display.
[0201] Figure 14 The bars shown indicate the images, and the arrows indicate the reference relationships between images and whether another image can be used as a reference image. An 'x' mark on the arrow indicates that one or more images cannot reference the image indicated by that arrow.
[0202] The preceding picture 1410, which precedes the CRA picture in the output order, can use the CRA picture, other preceding pictures, and all of the previous pictures 1430 that precede the preceding picture 1410 in the output and decoding order as reference pictures.
[0203] Conversely, the normal image 1420, which follows the CRA image in both output and decoding order, can be decoded by referencing a different normal image than the CRA image. The normal image 1420 can be decoded without using the preceding image 1410 as a reference image.
[0204] Furthermore, the VVC standard allows encoded images (i.e., the current image) to include slices of different NAL unit types. Whether the current image includes slices of different NAL unit types can be indicated based on the syntax element `mixed_nalu_types_in_pic_flag`. For example, when the current image includes slices of different NAL unit types, the value of the syntax element `mixed_nalu_types_in_pic_flag` can be represented as 1. In this case, the current image must reference a PPS that includes `mixed_nalu_types_in_pic_flag` with a value of 1. The semantics of the flag (`mixed_nalu_types_in_pic_flag`) are as follows:
[0205] When the value of the syntax element mixed_nalu_types_in_pic_flag is 1, it can indicate that each picture in the reference PPS has one or more VCL NAL units, the VCL NAL units do not have the same NAL unit type (nal_unit_type), and the picture is not an IRAP picture.
[0206] When the value of the syntax element mixed_nalu_types_in_pic_flag is 0, it can indicate that each picture in the reference PPS has one or more VCL NAL units, and that the VCL NAL units of each picture in the reference PPS have the same value of NAL unit type (nal_unit_type).
[0207] When the value of no_mixed_nalu_types_in_pic_constraint_flag is 1, the value of mixed_nalu_types_in_pic_flag must be 0. The no_mixed_nalu_types_in_pic_constraint_flag syntax element indicates a constraint regarding whether the value of mixed_nalu_types_in_pic_flag must be 0 for a picture. For example, whether the value of mixed_nalu_types_in_pic_constraint_flag must be 0 can be determined based on no_mixed_nalu_types_in_pic_constraint_flag information signaled from a higher-level syntax (e.g., PPS) or a syntax that includes information about constraints (e.g., GCI; general constraint information).
[0208] In a picture picA that also includes one or more slices of NAL unit type with different values (i.e., when the value of mixed_nalu_types_in_pic_flag of picture picA is 1), the following can be applied for each slice with NAL unit type value nalUnitTypeA in the range from IDR_W_RADL to CRA_NUT (e.g., the value of NAL unit type is 7 to 9 in Table 1 or Table 2).
[0209] - The slice must belong to a subpicA whose corresponding subpic_treated_as_pic_flag value is 1. Here, subpic_treated_as_pic_flag is information about whether the i-th subpic of each picture encoded in CLVS is treated as a picture in the decoding process other than the loop filtering operation. For example, when the value of subpic_treated_as_pic_flag is 1, it can indicate that the i-th subpic is treated as a picture in the decoding process other than the loop filtering operation. Alternatively, when the value of subpic_treated_as_pic_flag is 0, it can indicate that the i-th subpic is not treated as a picture in the decoding process other than the loop filtering operation.
[0210] - A slice must not be a subpicture of a picA that includes a VCL NAL unit with a NAL unit type (nal_unit_type) that is not equal to nalUnitTypeA.
[0211] - For all PUs following CLVS in decoding order, the RefPicList or RefPicList of slices in subpicA must not include images in the active entry that precede picA in decoding order.
[0212] To operate the concepts described above, the following can be specified. For example, the following can be applied to the VCL NAL unit of a specific image.
[0213] - When the value of mixed_nalu_types_in_pic_flag is 0, the value of NAL unit type (nal_unit_type) must be the same for all slice NAL units encoded in the picture. The picture or PU can be regarded as having the same NAL unit type as the slice NAL units encoded in the picture or PU.
[0214] Otherwise (when the value of mixed_nalu_types_in_pic_flag is 1), one or more VCL NAL units must have a NAL unit type with a specific value in the range from IDR_W_RADL to CRA_NUT (e.g., the value of the NAL unit type in Table 1 or Table 2 is 7 to 9), and all other VCL NAL units must have the same NAL unit type as GRA_NUT or a NAL unit type with a specific value in the range from TRAIL_NUT to RSV_VCL_6 (e.g., the value of the NAL unit type in Table 1 or Table 2 is 0 to 6).
[0215] In the current VVC standard, at least the following issues may exist when dealing with images that have mixed NAL unit types.
[0216] 1. When an image includes IDR and non-IRAP NAL units, and when signaling for the Reference Picture List (RPL) is present in the slice header, the signaling must also be present in the IDR slice header. RPL signaling is present in the IDR slice header when the value of `sps_idr_rpl_present_flag` is 1. Currently, this flag (`sps_idr_rpl_present_flag`) can be 0 even when there are one or more images with mixed NAL unit types. Here, the `sps_idr_rpl_present_flag` syntax element can indicate whether an RPL syntax element can exist in the slice header of a slice with NAL unit types such as IDR_N_LP or IDR_W_RADL. For example, when the value of `sps_idr_rpl_present_flag` is 1, it can indicate that an RPL syntax element can exist in the slice header of a slice with NAL unit types such as IDR_N_LP or IDR_W_RADL. Alternatively, when the value of sps_idr_rpl_present_flag is 0, it can indicate that the RPL syntax element does not exist in the slice header of a slice with a NAL unit type such as IDR_N_LP or IDR_W_RADL.
[0217] 2. When the current image references a PPS with a value of 1 for mixed_nalu_types_in_pic_flag, one or more of the VCL NAL cells in the current image must have a NAL cell type with a specific value in the range from IDR_W_RADL to CRA_NUT (e.g., NAL cell types with values of 7 to 9 in Table 1 or Table 2 above), and all other VCL NAL cells must have the same NAL cell type as GRA_NUT or a NAL cell type with a specific value in the range from TRAIL_NUT to RSV_VCL_6 (e.g., NAL cell types with values of 0 to 6 in Table 1 or Table 2 above). This constraint applies only to current images that include a mixture of IRAP and non-IRAP NAL cell types. However, it is not yet correctly applied to images that include a mixture of RASL / RADL and non-IRAP NAL cell types.
[0218] This document provides a solution to the problems mentioned above. Specifically, as described above, an image comprising two or more sub-images (i.e., the current image) can have a mixed NAL unit type. In the current VVC standard, an image with a mixed NAL unit type can have a hybrid form of IRAP NAL unit type and non-IRAP NAL unit type. However, a preceding image associated with a CRA NAL unit type can also have a hybrid form with non-IRAP NAL unit types, and images with this hybrid NAL unit type are not supported under the current standard. Therefore, a solution is needed for images with a hybrid form of CRA NAL unit type and non-IRAP NAL unit type.
[0219] Therefore, this document provides a method for including, in a mixed form, a leading picture NAL unit type (e.g., RASL_NUT, RADL_NUT) and another non-IRAP NAL unit type (e.g., TRAIL_NUT, STSA, NUT). Additionally, this document defines constraints that allow the existence or signaling of a reference picture list when IDR subpicks and other non-IRAP subpicks are mixed. Thus, pictures with mixed NAL unit types can be configured to include not only IRAP but also CRANAL units, providing greater flexibility.
[0220] For example, it can be applied in the following embodiments, thus solving the above-mentioned problems. The following embodiments can be applied independently or in combination.
[0221] In one implementation, when images are allowed to have mixed NAL cell types (when the value of mixed_nal_types_in_pic_flag is 1), the signaling regarding the list of reference images allows it to exist even for slices with IDR-type NAL cell types (e.g., IDR_W_RADL or IDR_N_LP). This constraint can be represented as follows.
[0222] - The value of sps_idr_rpl_present_flag must be 1 when there is at least one PPS that references an SPS whose mixed_nal_types_in_pic_flag value is 1. This constraint can be a bitstream consistency requirement.
[0223] Alternatively, in one implementation, for an image with mixed NAL unit types, it is permissible to include slices of a specific NAL unit type (e.g., RADL or RASL) with a leading image and a specific NAL unit type (non-IRAP) without a leading image. This can be represented as follows.
[0224] The following can be applied to the VCL NAL unit of a specific image.
[0225] - When the value of mixed_nalu_types_in_pic_flag is 0, the value of NAL unit type (nal_unit_type) must be the same for all slice NAL units encoded in the picture. The picture or PU can be regarded as having the same NAL unit type as the slice NAL units encoded in the picture or PU.
[0226] - Otherwise (when the value of mixed_nalu_types_in_pic_flag is 1), one of the following must be satisfied (i.e., one of the following can have a value that is true).
[0227] 1) One or more VCL NAL units must have a NAL unit type (nal_unit_type) with a specific value in the range from IDR_W_RADL to CRA_NUT (e.g., the NAL unit type values are 7 to 9 in Table 1 or Table 2 above), and all other VCL NAL units must have the same NAL unit type as GRA_NUT or a NAL unit type with a specific value in the range from TRAIL_NUT to RSV_VCL_6 (e.g., the NAL unit type values are 0 to 6 in Table 1 or Table 2 above).
[0228] 2) One or more VCL NAL units must all have a specific value of the same NAL unit type as RADL_NUT (e.g., a value of 2 for the NAL unit type in Table 1 or Table 2 above) or RASL_NUT (e.g., a value of 3 for the NAL unit type in Table 1 or Table 2 above), and all other VCL NAL units must have a specific value of the same NAL unit type as TRAIL_NUT (e.g., a value of 0 for the NAL unit type in Table 1 or Table 2 above), STSA_NUT (e.g., a value of 1 for the NAL unit type in Table 1 or Table 2 above), RSV_VCL_4 (e.g., a value of 4 for the NAL unit type in Table 1 or Table 2 above), RSV_VCL_5 (e.g., a value of 5 for the NAL unit type in Table 1 or Table 2 above), RSV_VCL_6 (e.g., a value of 6 for the NAL unit type in Table 1 or Table 2 above), or GRA_NUT.
[0229] Furthermore, this document proposes a method for providing images with the aforementioned hybrid NAL unit type even for single-layer bitstreams. As an implementation, the following constraints can be applied in the case of single-layer bitstreams.
[0230] - In the bitstream, each picture except the first picture in decoding order is considered to be associated with the previous IRAP picture in decoding order.
[0231] - If the image is a leading image of an IRAP image, then it must be a RADL or RASL image.
[0232] - If the image is a trailing image of an IRAP image, then it must be neither a RADL nor a RASL image.
[0233] -RASL images must not exist in the bitstream associated with IDR images.
[0234] - RADL images must not exist in the bitstream associated with an IDR image whose NAL unit type (nal_unit_type) is IDR_N_LP.
[0235] When referenced, and when each parameter set is available, random access can be performed at the location of the IRAP PU by discarding all PUs preceding the IRAP PU (and the IRAP picture and all subsequent non-RASL pictures can be decoded correctly in the decoding order).
[0236] - Images that precede IRAP images in decoding order must precede IRAP images in output order, and must precede the RADL images associated with the IRAP images in output order.
[0237] - The RASL image associated with the CRA image must precede the RADL image associated with the CRA image in the output order.
[0238] - The RASL images associated with the CRA images must be in output order and after the IRAP images in decoding order, which precede the CRA images.
[0239] - If the value of field_seq_flag is 0 and the current image is a preceding image associated with an IRAP image, then it must precede all non-preceding images associated with the same IRAP image in decoding order. Otherwise, when images picA and picB are the first and last preceding images associated with an IRAP image in decoding order, at most one non-preceding image can exist before picA in decoding order, and no non-preceding image can exist between picA and picB in decoding order.
[0240] The following figures are provided to illustrate specific examples of this document. Since the specific terms or names or names of specific devices (e.g., names of grammars / grammatical elements, etc.) described in the figures are presented as examples, the technical features of this document are not limited to the specific names used in the figures below.
[0241] Figure 15 Examples of video / image coding methods applicable to the embodiments described in this document are illustrated. (This can be derived from...) Figure 2 The publicly disclosed encoding device 200 performs Figure 15 The method disclosed in the document.
[0242] Reference Figure 15 The encoding device can determine the NAL unit type of the slice in the image (S1500).
[0243] For example, the encoding device can determine the NAL unit type based on the nature and type of the image or sub-image as described in Tables 1 and 2 above, and based on the NAL unit type of the image or sub-image, it can determine the NAL unit type of each slice.
[0244] For example, when the value of mixed_nalu_types_in_pic_flag is 0, slices in the image associated with PPS can be determined to be of the same NAL unit type. That is, when the value of mixed_nalu_types_in_pic_flag is 0, the NAL unit type defined in the header of the first NAL unit, which includes information about the first slice of the image, is the same as the NAL unit type defined in the header of the second NAL unit, which includes information about the second slice of the same image. Alternatively, when the value of mixed_nalu_types_in_pic_flag is 1, slices in the image associated with PPS can be determined to be of different NAL unit types. Here, the NAL unit type of the slice in the image can be determined based on the method proposed in the above embodiments.
[0245] The encoding device can generate NAL unit type related information (S1510). The NAL unit type related information may include information / syntax elements related to the NAL unit types described in the above embodiments and / or Tables 1 and 2 above. For example, the information related to the NAL unit type may include the mixed_nalu_types_in_pic_flag syntax element included in the PPS. Alternatively, the information related to the NAL unit type may include the nal_unit_type syntax element in the NAL unit header of the NAL unit, which includes information about the encoded slice.
[0246] The encoding device can generate a bitstream (S1520). The bitstream may include at least one NAL unit containing image information about the encoded slice. Additionally, the bitstream may include PPS.
[0247] Figure 16This illustration shows an example of a video / image decoding method applicable to the embodiments described in this document. It can be derived from... Figure 3 The publicly disclosed decoding device 300 performs Figure 16 The method disclosed in the document.
[0248] Reference Figure 16 The decoding device can receive a bitstream (S1600). Here, the bitstream may include at least one NAL unit that includes image information about the encoded slice. Additionally, the bitstream may include PPS.
[0249] The decoding device can obtain NAL unit type-related information (S1610). NAL unit type-related information may include information / syntax elements related to the NAL unit types described in the above embodiments and / or Tables 1 and 2 above. For example, information related to the NAL unit type may include the mixed_nalu_types_in_pic_flag syntax element included in the PPS. Alternatively, information related to the NAL unit type may include the nal_unit_type syntax element in the NAL unit header of the NAL unit, which includes information about the encoded slice.
[0250] The decoding device can determine the NAL unit type of the slice in the image (S1620).
[0251] For example, when the value of mixed_nalu_types_in_pic_flag is 0, the slices in the image associated with PPS use the same NAL unit type. That is, when the value of mixed_nalu_types_in_pic_flag is 0, the NAL unit type defined in the first NAL unit header of the first NAL unit that includes information about the first slice of the image is the same as the NAL unit type defined in the second NAL unit header of the second NAL unit that includes information about the second slice of the same image. Alternatively, when the value of mixed_nalu_types_in_pic_flag is 1, the slices in the image associated with PPS use different NAL unit types. Here, the NAL unit type of the slice in the image can be determined based on the method proposed in the above embodiments.
[0252] The decoding device can decode / reconstruct samples / blocks / slices based on the NAL unit type of the slice (S1630). Samples / blocks in a slice can be decoded / reconstructed based on the NAL unit type of the slice.
[0253] For example, when a first NAL unit type is set for the first slice of the current image and a second NAL unit type (different from the first NAL unit type) is set for the second slice of the current image, samples / blocks in the first slice or the first slice itself can be decoded / reconstructed based on the first NAL unit type, and samples / blocks in the second slice or the second slice itself can be decoded / reconstructed based on the second NAL unit type.
[0254] Figure 17 and Figure 18 Examples of video / image encoding methods and associated components according to embodiments of this document are illustrated.
[0255] Figure 17 The method disclosed in the article can be derived from Figure 2 or Figure 18 The publicly disclosed encoding device 200 is executed. Here, Figure 18 The publicly disclosed encoding device 200 is Figure 2 A simplified representation of the encoding device 200 disclosed herein. Specifically, Figure 17 Steps S1700 to S1730 can be performed by Figure 2 The entropy encoder 240 disclosed herein is executed; furthermore, according to the implementation, each step may be performed by... Figure 2 The image segmenter 210, predictor 220, residual processor 230, adder 340, etc., disclosed herein are executed. Furthermore, implementations including those described in this document can be executed. Figure 17 The method disclosed in [the document]. Therefore, in [the document] Figure 17 In this document, detailed descriptions of content that corresponds to repetitions of the above-described embodiments will be omitted or simplified.
[0256] Reference Figure 17 The encoding device can determine the NAL unit type of the slice in the current image (S1700).
[0257] The current image may include multiple slices, and a slice may include a slice header and slice data. Furthermore, NAL cells can be generated by adding a NAL cell header to the slice (slice header and slice data). The NAL cell header may include NAL cell type information specified based on the slice data included in the corresponding NAL cell.
[0258] As an implementation, the encoding device can generate a first NAL unit for a first slice in the current image and a second NAL unit for a second slice in the current image. Furthermore, the encoding device can determine the type of the first NAL unit for the first slice and the type of the second NAL unit for the second slice based on the types of the first and second slices.
[0259] For example, based on the type of slice data included in the NAL units shown in Table 1 or Table 2 above, NAL unit types can include TRAIL_NUT, STSA_NUT, RADL_NUT, RASL_NUT, IDR_W_RADL, IDR_N_LP, CRA_NUT, etc. Furthermore, the NAL unit type can be signaled based on the nal_unit_type syntax element in the NAL unit header. The nal_unit_type syntax element is syntax information used to specify the NAL unit type and, as shown in Table 1 or Table 2 above, can be represented as a specific value corresponding to a particular NAL unit type.
[0260] The encoding device can determine whether the current image has a mixed NAL unit type based on the NAL unit type (S1710).
[0261] For example, if all NAL unit types in a slice of the current image are the same, the encoding device can determine that the current image does not have a mixed NAL unit type. Alternatively, if all NAL unit types in a slice of the current image are different, the encoding device can determine that the current image has a mixed NAL unit type.
[0262] The encoding device can generate NAL unit type information based on whether the current image has a mixed NAL unit type (S1720).
[0263] NAL unit type related information may include information / syntax elements related to the NAL unit types described in the above embodiments and / or Tables 1 and 2. For example, NAL unit type related information may be information about whether the current image has mixed NAL unit types, and may be represented by the mixed_nalu_types_in_pic_flag syntax element included in PPS. For example, when the value of the mixed_nalu_types_in_pic_flag syntax element is 0, it may indicate that the NAL units in the current image have the same NAL unit type. Alternatively, when the value of the mixed_nalu_types_in_pic_flag syntax element is 1, it may indicate that the NAL units in the current image have different NAL unit types.
[0264] In one implementation, when all NAL unit types of slices in the current image are the same, the encoding device can determine that the current image does not have a mixed NAL unit type and can generate NAL unit type related information (e.g., mixed_nalu_type_in_pic_flag). In this case, the value of the NAL unit type related information (e.g., mixed_nalu_type_in_pic_flag) can be determined to be 0. Alternatively, when the NAL unit types of slices in the current image are not the same, the encoding device can determine that the current image has a mixed NAL unit type and can generate NAL unit type related information (e.g., mixed_nalu_type_in_pic_flag). In this case, the value of the NAL unit type related information (e.g., mixed_nalu_type_in_pic_flag) can be determined to be 1.
[0265] That is, based on NAL unit type information related to the current image with mixed NAL unit types (e.g., the value of mixed_nalu_types_in_pic_flag is 1), the first NAL unit of the first slice of the current image and the second NAL unit of the second slice of the current image can have different NAL unit types. Alternatively, based on NAL unit type information related to the current image without mixed NAL unit types (e.g., the value of mixed_nalu_types_in_pic_flag is 0), the first NAL unit of the first slice of the current image and the second NAL unit of the second slice of the current image can have the same NAL unit type.
[0266] As an example, based on information about the NAL unit types of the current image with mixed NAL unit types (e.g., the value of mixed_nalu_types_in_pic_flag is 1), the first NAL unit of the first slice can have a preceding image NAL unit type, and the second NAL unit of the second slice can have a non-IRAP NAL unit type or a non-preceding image NAL unit type. Here, the preceding image NAL unit type can include a RADL NAL unit type or a RASL NAL unit type, and the non-IRAP NAL unit type or the non-preceding image NAL unit type can include a trail NAL unit type or a STSA NAL unit type.
[0267] Alternatively, as an example, based on information related to the NAL unit types of the current image with mixed NAL unit types (e.g., the value of mixed_nalu_types_in_pic_flag is 1), the first NAL unit of the first slice may have an IRAP NAL unit type, and the second NAL unit of the second slice may have a non-IRAP NAL unit type or a non-leading image NAL unit type. Here, the IRAP NAL unit type may include an IDR NAL unit type (i.e., IDR_N_LPNAL or IDR_W_RADL NAL unit type) or a CRA NAL unit type, and the non-IRAP NAL unit type or non-leading image NAL unit type may include a tracking NAL unit type or an STSA NAL unit type. Furthermore, according to an implementation, the non-IRAP NAL unit type or non-leading image NAL unit type may refer only to the tracking NAL unit type.
[0268] According to the implementation, based on the premise that the current image is allowed to have mixed NAL unit types, for slices in the current image with IDR NAL unit types (e.g., IDR_W_RADL or IDR_N_LP), information related to the signaling reference image list must exist. The information related to the signaling reference image list can indicate whether the syntax element for signaling the reference image list exists in the slice header of the slice. That is, based on a value of 1 for the information related to the signaling reference image list, the syntax element for signaling the reference image list can exist in the slice header of the slice with the IDR NAL unit type. Alternatively, based on a value of 0 for the information related to the signaling reference image list, the syntax element for signaling the reference image list may not exist in the slice header of the slice with the IDR NAL unit type.
[0269] For example, information related to signaling a list of reference images could be the aforementioned `sps_idr_rpl_present_flag` syntax element. When `sps_idr_rpl_present_flag` is 1, it indicates that the syntax element used to signal the list of reference images can exist in the slice header of a slice with a NAL unit type such as `IDR_N_LP` or `IDR_W_RADL`. Alternatively, when `sps_idr_rpl_present_flag` is 0, it indicates that the syntax element used to signal the list of reference images may not exist in the slice header of a slice with a NAL unit type such as `IDR_N_LP` or `IDR_W_RADL`.
[0270] The encoding device can encode image / video information including information related to NAL unit type (S1730).
[0271] For example, when the first NAL unit of the first slice in the current image and the second NAL unit of the second slice in the current image have different NAL unit types, the encoding device can encode image / video information including NAL unit type information with a value of 1 (e.g., mixed_nalu_types_in_pic_flag). Alternatively, when the first NAL unit of the first slice in the current image and the second NAL unit of the second slice in the current image have the same NAL unit type, the encoding device can encode image / video information including NAL unit type information with a value of 0 (e.g., mixed_nalu_types_in_pic_flag).
[0272] Additionally, for example, the encoding device can encode image / video information that includes nal_unit_type information indicating each NAL unit type of slice in the current image.
[0273] Additionally, for example, the encoding device can encode image / video information that includes information related to a list of reference images for signaling (e.g., sps_idr_rpl_present_flag).
[0274] Additionally, for example, the encoding device can encode image / video information that includes NAL units of slices in the current image.
[0275] Image / video information, including the various types of information described above, can be encoded and output as a bitstream. The bitstream can be sent to a decoding device via a network or (digital) storage medium. Here, the network can include broadcast networks, communication networks, and / or the like, and the digital storage medium can include various storage media such as Universal Serial Bus (USB), Secure Digital (SD), Optical Disc (CD), Digital Video Disc (DVD), Blu-ray, Hard Disk Drive (HDD), Solid State Drive (SSD), etc.
[0276] Figure 19 and Figure 20 Examples of video / image decoding methods and associated components according to embodiments of this document are illustrated.
[0277] Figure 19 The method disclosed in the article can be derived from Figure 3 or Figure 20 The publicly disclosed decoding device 300 is executed. Here, Figure 20 The publicly disclosed decoding device 300 is Figure 3A simplified representation of the decoding device 300 disclosed in the document. Specifically, Figure 19 Steps S1900 to S1930 can be performed by Figure 3 The entropy decoder 310 disclosed herein is executed; furthermore, according to the implementation method, each step can be performed by... Figure 3 The residual processor 320, predictor 330, adder 340, etc., disclosed herein are executed. Furthermore, embodiments including those described above in this document can be executed. Figure 19 The method disclosed in [the document]. Therefore, in [the document] Figure 19 In this document, detailed descriptions of content that corresponds to repetitions in the above embodiments will be omitted or simplified.
[0278] Reference Figure 19 The decoding device can obtain image / video information, including information related to the NAL unit type, from the bitstream (S1900).
[0279] For example, a decoding device can parse the bitstream and deduce the information required for image reconstruction (or picture reconstruction) (e.g., video / image information). In this case, the image information may include the aforementioned NAL unit type information (e.g., mixed_nalu_types_in_pic_flag), nal_unit_type information indicating the type of each NAL unit in the current picture, information related to signaling the reference picture list (e.g., sps_idr_rpl_present_flag), the NAL units of the slices within the current picture, etc. That is, the image information may include various information required in the decoding process and may be decoded based on encoding methods such as exponential Golomb coding, CAVLC, or CABAC.
[0280] As described above, NAL unit type information may include information / syntax elements related to the NAL unit types described in the above embodiments and / or Tables 1 and 2. For example, NAL unit type information may be information about whether the current image has mixed NAL unit types, and may be represented by the mixed_nalu_types_in_pic_flag syntax element included in the PPS. For example, when the value of the mixed_nalu_types_in_pic_flag syntax element is 0, it may indicate that the NAL units in the current image have the same NAL unit type. Alternatively, when the value of the mixed_nalu_types_in_pic_flag syntax element is 1, it may indicate that the NAL units in the current image have different NAL unit types.
[0281] The decoding device can determine whether the current image has a mixed NAL unit type based on NAL unit type related information (S1910).
[0282] For example, a decoding device can determine that the current image has a mixed NAL unit type based on a value of 1 for NAL unit type related information (e.g., mixed_nalu_type_in_pic_flag). Alternatively, a decoding device can determine that the current image does not have a mixed NAL unit type based on a value of 0 for NAL unit type related information (e.g., mixed_nalu_type_in_pic_flag).
[0283] The decoding device can determine the NAL unit type of a slice in the current image based on information related to whether the current image has a mixed NAL unit type (S1920).
[0284] The current image can include multiple slices, and a slice can include a slice header and slice data. Additionally, NAL cells can be generated by adding a NAL cell header to a slice (slice header and slice data). The NAL cell header can include NAL cell type information specified based on the slice data included in the corresponding NAL cell.
[0285] For example, based on the type of slice data included in the NAL unit as shown in Table 1 or Table 2 above, the NAL unit type can include TRAIL_NUT, STSA_NUT, RADL_NUT, RASL_NUT, IDR_W_RADL, IDR_N_LP, CRA_NUT, etc. Additionally, the NAL unit type can be signaled based on the nal_unit_type syntax element in the NAL unit header. The nal_unit_type syntax element is syntax information used to specify the NAL unit type and, as shown in Table 1 or Table 2 above, can be represented as a specific value corresponding to a particular NAL unit type.
[0286] In one implementation, the decoding device can determine that the first NAL unit of the first slice of the current image and the second NAL unit of the second slice of the current image can have different NAL unit types based on NAL unit type information about the current image with mixed NAL unit types (e.g., the value of mixed_nalu_types_in_pic_flag is 1). Alternatively, the decoding device can determine that the first NAL unit of the first slice of the current image and the second NAL unit of the second slice of the current image can have the same NAL unit type based on NAL unit type information about the current image without mixed NAL unit types (e.g., the value of mixed_nalu_types_in_pic_flag is 0).
[0287] As an example, based on information about the NAL unit types of the current image with mixed NAL unit types (e.g., the value of mixed_nalu_types_in_pic_flag is 1), the first NAL unit of the first slice can have a preceding image NAL unit type, and the second NAL unit of the second slice can have a non-IRAP NAL unit type or a non-preceding image NAL unit type. Here, the preceding image NAL unit type can include a RADL NAL unit type or a RASL NAL unit type, and the non-IRAP NAL unit type or non-preceding image NAL unit type can include a tracking NAL unit type or a STSA NAL unit type.
[0288] Alternatively, as an example, based on information related to the NAL unit types of the current image with mixed NAL unit types (e.g., the value of mixed_nalu_types_in_pic_flag is 1), the first NAL unit of the first slice may have an IRAP NAL unit type, and the second NAL unit of the second slice may have a non-IRAP NAL unit type or a non-leading image NAL unit type. Here, the IRAP NAL unit type may include an IDR NAL unit type (i.e., IDR_N_LPNAL or IDR_W_RADL NAL unit type) or a CRA NAL unit type, and the non-IRAP NAL unit type or non-leading image NAL unit type may include a tracking NAL unit type or an STSA NAL unit type. Furthermore, according to an implementation, the non-IRAP NAL unit type or non-leading image NAL unit type may refer only to the tracking NAL unit type.
[0289] According to the implementation, based on the premise that the current image is allowed to have mixed NAL unit types, for slices in the current image with IDR NAL unit types (e.g., IDR_W_RADL or IDR_N_LP), information related to the signaling reference image list must exist. The information related to the signaling reference image list can indicate whether the syntax element for signaling the reference image list exists in the slice header of the slice. That is, based on a value of 1 for the information related to the signaling reference image list, the syntax element for signaling the reference image list can exist in the slice header of the slice with the IDR NAL unit type. Alternatively, based on a value of 0 for the information related to the signaling reference image list, the syntax element for signaling the reference image list may not exist in the slice header of the slice with the IDR NAL unit type.
[0290] For example, information related to signaling a list of reference images could be the aforementioned `sps_idr_rpl_present_flag` syntax element. When `sps_idr_rpl_present_flag` is 1, it indicates that the syntax element used to signal the list of reference images can exist in the slice header of a slice with a NAL unit type such as `IDR_N_LP` or `IDR_W_RADL`. Alternatively, when `sps_idr_rpl_present_flag` is 0, it indicates that the syntax element used to signal the list of reference images may not exist in the slice header of a slice with a NAL unit type such as `IDR_N_LP` or `IDR_W_RADL`.
[0291] The decoding device can decode / reconstruct the current image based on the NAL unit type (S1930).
[0292] For example, given a first slice in a current image determined to be of a first NAL unit type and a second slice in a current image determined to be of a second NAL unit type, the decoding device can decode / restore the first slice based on the first NAL unit type and decode / reconstruct the second slice based on the second NAL unit type. Additionally, the decoding device can decode / reconstruct samples / blocks within the first slice based on the first NAL unit type and decode / reconstruct samples / blocks within the second slice based on the second NAL unit type.
[0293] Although the methods have been described in the above embodiments based on flowcharts listing the steps or blocks in sequence, the steps in this document are not limited to a specific order, and specific steps may be performed in different steps or in different orders or simultaneously relative to the steps described above. Furthermore, those skilled in the art will understand that the steps in the flowcharts are not exclusive, and without affecting the scope of this disclosure, one or more steps may be included, or one or more steps may be omitted from the flowcharts.
[0294] The methods mentioned above according to this disclosure can be in the form of software, and the encoding and / or decoding devices according to this disclosure can be included in an apparatus for performing image processing (e.g., TV, computer, smartphone, set-top box, display device, etc.).
[0295] When the embodiments of this disclosure are implemented in software, the methods described above can be implemented using modules (processes or functions) that perform the functions mentioned above. Modules can be stored in memory and executed by a processor. Memory can be installed internally or externally to the processor and can be connected to the processor via various known means. The processor may include application-specific integrated circuits (ASICs), other chipsets, logic circuits, and / or data processing devices. Memory may include read-only memory (ROM), random access memory (RAM), flash memory, memory cards, storage media, and / or other storage devices. In other words, embodiments of this disclosure can be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units illustrated in the corresponding figures can be implemented and executed on a computer, processor, microprocessor, controller, or chip. In this case, information about the implementation (e.g., information about instructions) or algorithms can be stored in a digital storage medium.
[0296] Furthermore, the decoding and encoding devices using the embodiments described in this document can be included in multimedia broadcast transceivers, mobile communication terminals, home theater video devices, digital cinema video devices, surveillance cameras, video chat devices, real-time communication devices such as video communication, mobile streaming devices, storage media, portable cameras, video-on-demand (VoD) service providers, over-the-top (OTT) video devices, internet streaming service providers, 3D video devices, virtual reality (VR) devices, augmented reality (AR) devices, video telephony devices, vehicle-mounted terminals (e.g., vehicle (including autonomous vehicle) terminals, aircraft terminals, or ship terminals), and medical video devices; and can be used to process image signals or data. For example, OTT video devices can include game consoles, Blu-ray players, networked TVs, home theater systems, smartphones, tablet PCs, and digital video recorders (DVRs).
[0297] Furthermore, the processing methods applying the embodiments of this document can be generated in the form of a computer-executable program and can be stored in a computer-readable recording medium. Multimedia data with data structures according to the embodiments of this document can also be stored in a computer-readable recording medium. Computer-readable recording media include all types of storage devices and distributed storage devices storing computer-readable data. Computer-readable recording media can include, for example, Blu-ray discs (BD), Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disks, and optical data storage devices. Computer-readable recording media also include media implemented in the form of carrier waves (e.g., transmission over the Internet). Additionally, bitstreams generated by encoding methods can be stored in computer-readable recording media or transmitted via wired or wireless communication networks.
[0298] Furthermore, the embodiments described in this document can be implemented as a computer program product based on program code, and the program code can be executed on a computer according to the embodiments described in this document. The program code can be stored on a computer-readable medium.
[0299] Figure 21 Examples of content streaming systems to which the implementation methods of this document can be applied are provided.
[0300] Reference Figure 21 The content streaming system using the embodiments described in this document may typically include an encoding server, a streaming server, a network server, a media storage device, a user device, and a multimedia input device.
[0301] An encoding server is used to compress content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data, generate a bitstream, and then transmit it to a streaming server. As another example, if the multimedia input device, such as a smartphone, camera, or camcorder, directly generates the bitstream, the encoding server can be omitted.
[0302] Bitstreams can be generated using the encoding methods or bitstream generation methods applied in the embodiments described in this document. Furthermore, the streaming server can temporarily store the bitstream during the sending or receiving of the bitstream.
[0303] A streaming server transmits multimedia data to a user's device via a web server based on a user's request. The web server acts as a tool to notify the user of available services. When a user requests a desired service, the web server forwards the request to the streaming server, which then delivers the multimedia data to the user. In this respect, the content streaming system may include a separate control server, which in this case controls the commands / responses between the various devices within the content streaming system.
[0304] A streaming server can receive content from media storage devices and / or encoding servers. For example, if content is received from an encoding server, it can be received in real time. In this case, the streaming server can store the bitstream for a predetermined period of time to provide a smooth streaming service.
[0305] For example, user equipment may include mobile phones, smartphones, laptops, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation systems, board PCs, tablet PCs, ultrabooks, wearable devices (e.g., watch-type terminals (smartwatches), glasses-type terminals (smart glasses), head-mounted displays (HMDs)), digital TVs, desktop computers, digital signage, etc.
[0306] Each server in the content streaming system can be operated as a distributed server, and in this case, the data received by each server can be processed in a distributed manner.
[0307] The claims in this specification can be combined in various ways. For example, the technical features in the method claims can be combined to be implemented or performed in a device, and the technical features in the device claims can be combined to be implemented or performed in a method. Furthermore, the technical features in the method claims and the device claims can be combined to be implemented or performed in a device.
Claims
1. An apparatus for decoding image information, the apparatus comprising a memory and at least one processor, the at least one processor being coupled to the memory, the at least one processor being configured to: Obtain the image information, including information related to the network abstraction layer (NAL) unit type, from the bitstream; Based on the NAL unit type information, determine whether the current image has a mixed NAL unit type; The NAL unit type for a slice in the current image is determined based on information related to the NAL unit type for the current image having the hybrid NAL unit type. and The current image is decoded based on the NAL unit type. Specifically, based on the value of the NAL unit type related information being equal to 1, it is determined that the current image has the hybrid NAL unit type. Wherein, based on the NAL unit type of the first slice in the current image being the instantaneous decode refresh IDR NAL unit type, for the first slice in the current image, the value of the information related to signaling the reference image list is equal to 1, and Wherein, the value based on the NAL unit type related information is equal to 1, and the NAL unit type for the first slice in the current image is different from the NAL unit type for the second slice in the current image.
2. An apparatus for encoding image information, the apparatus comprising a memory and at least one processor, the at least one processor being coupled to the memory, the at least one processor being configured to: Determine the NAL unit type for the slice in the current image; Determine whether the current image has a mixed NAL unit type based on the NAL unit type; Based on whether the current image has the hybrid NAL unit type, generate NAL unit type related information; and The image information, including information related to the NAL unit type, is encoded. in, Based on the fact that the current image has the hybrid NAL unit type, the value of the NAL unit type related information is determined to be 1. Wherein, based on the NAL unit type of the first slice in the current image being the instantaneous decode refresh IDR NAL unit type, for the first slice in the current image, the value of the information related to signaling the reference image list is equal to 1, and Wherein, the value based on the NAL unit type related information is equal to 1, and the NAL unit type for the first slice in the current image is different from the NAL unit type for the second slice in the current image.
3. An apparatus for including data containing image information, the apparatus comprising: At least one processor, configured to obtain a bitstream of the image information, wherein the bitstream is generated based on the following operations: determining the NAL unit type for a slice in a current image; determining whether the current image has a mixed NAL unit type based on the NAL unit type; generating NAL unit type-related information based on whether the current image has the mixed NAL unit type; and encoding image information including the NAL unit type-related information; and A transmitter configured to transmit the data comprising the bit stream. Specifically, based on the fact that the current image has the hybrid NAL unit type, the value of the NAL unit type related information is determined to be 1. Wherein, based on the NAL unit type of the first slice in the current image being the instantaneous decode refresh IDR NAL unit type, for the first slice in the current image, the value of the information related to signaling the reference image list is equal to 1, and Wherein, the value based on the NAL unit type related information is equal to 1, and the NAL unit type for the first slice in the current image is different from the NAL unit type for the second slice in the current image.