Image encoding / decoding method and apparatus based on mixed nal unit type, and recording medium storing bitstream
By using an image encoding/decoding method and device based on hybrid NAL unit types, images are segmented into sub-pictures and their NAL unit types are determined, solving the problem of low efficiency in the transmission and storage of high-resolution and high-quality images, and achieving efficient image encoding/decoding and random access.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-23
- Publication Date
- 2026-03-24
AI Technical Summary
Existing technologies are inefficient in the transmission and storage of high-resolution and high-quality images, leading to increased costs and necessitating improved image encoding/decoding methods and devices.
An image encoding/decoding method and device based on hybrid NAL unit types are adopted. By segmenting the image into sub-pictures and determining the NAL unit type of each sub-picture, the decoding of RASL and RADL pictures is realized, generating a randomly accessible bitstream.
It improves the efficiency of image encoding/decoding, supports random access based on decoding order and reference screen decoding, and reduces transmission and storage costs.
Smart Images

Figure CN115668943B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to an image encoding / decoding method and apparatus, and more particularly, to an image encoding and decoding method and apparatus based on a hybrid NAL unit type and a recording medium for storing a bitstream generated by the image encoding method / apparatus of the present disclosure. BACKGROUND
[0002] Recently, the demand for high-resolution and high-quality images such as high definition (HD) images and ultra-high definition (UHD) images is increasing in various fields. As the resolution and quality of image data increase, the amount of information or bits to be transmitted relatively increases compared to existing image data. The increase in the amount of transmitted information or bits results in an increase in transmission and storage costs.
[0003] Therefore, an efficient image compression technique is required to effectively transmit, store, and reproduce information about high-resolution and high-quality images. SUMMARY
[0004] TECHNICAL PROBLEM
[0005] An object of the present disclosure is to provide an image encoding / decoding method and apparatus having improved encoding / decoding efficiency.
[0006] Another object of the present disclosure is to provide an image encoding / decoding method and apparatus based on a hybrid NAL unit type.
[0007] Another object of the present disclosure is to provide an image encoding / decoding method and apparatus based on a sub-picture merge operation.
[0008] Another object of the present disclosure is to provide a method of transmitting a bitstream generated by an image encoding method or apparatus according to the present disclosure.
[0009] Another object of the present disclosure is to provide a recording medium storing a bitstream generated by an image encoding method or apparatus according to the present disclosure.
[0010] Another object of the present disclosure is to provide a recording medium storing a bitstream generated by an image encoding method or apparatus according to the present disclosure.
[0011] The technical problems solved by the present disclosure are not limited to the above-mentioned technical problems, and other technical problems not described herein will be clearly understood by a person skilled in the art from the following description.
[0012] TECHNICAL SOLUTION
[0013] According to an aspect of the disclosure, an image decoding method includes the steps of obtaining, from a bitstream, network abstraction layer (NAL) unit type information of at least one NAL unit including encoded image data of a current picture, determining at least one NAL unit type of one or more slices in the current picture based on the obtained NAL unit type information, and decoding the current picture based on the determined NAL unit type. The current picture can be determined as a random access skipped leading (RASL) picture based on the determined NAL unit type including RASL_NUT. The RASL picture can be decoded based on the RASL picture including one or more slices having RADL_NUT when an intra random access point (IRAP) picture is a first picture in a decoding order.
[0014] According to another aspect of the disclosure, an image decoding apparatus includes a memory and at least one processor. The at least one processor can perform the steps of obtaining, from a bitstream, network abstraction layer (NAL) unit type information of at least one NAL unit including encoded image data of a current picture, determining at least one NAL unit type of one or more slices in the current picture based on the obtained NAL unit type information, and decoding the current picture based on the determined NAL unit type. The current picture can be determined as a random access skipped leading (RASL) picture based on the determined NAL unit type including RASL_NUT. The RASL picture can be decoded based on the RASL picture including one or more slices having RADL_NUT when an intra random access point (IRAP) picture is a first picture in a decoding order.
[0015] According to another aspect of the disclosure, an image encoding method includes the steps of partitioning a current picture into a plurality of sub-pictures, determining a network abstraction layer (NAL) unit type of each of the plurality of sub-pictures, and encoding the plurality of sub-pictures based on the determined NAL unit type. The current picture can be determined as a random access skipped leading (RASL) picture based on the determined NAL unit type including RASL_NUT, and the RASL picture can include at least one sub-picture having RADL_NUT.
[0016] In addition, a computer-readable recording medium according to another aspect of the disclosure can store a bitstream generated by the image encoding apparatus or the image encoding method of the disclosure.
[0017] In a transmitting method according to another aspect of the disclosure, a bitstream generated by the image encoding method or the image encoding apparatus of the disclosure can be transmitted.
[0018] The features of the above brief summary of the present disclosure are merely exemplary aspects of the following detailed description of the present disclosure, and do not limit the scope of the present disclosure.
[0019] Advantageous Effects
[0020] According to the present disclosure, it is possible to provide an image encoding / decoding method and apparatus having improved encoding / decoding efficiency.
[0021] Further, according to the present disclosure, it is possible to provide an image encoding / decoding method and apparatus based on a hybrid NAL unit type.
[0022] Further, according to the present disclosure, it is possible to provide an image encoding / decoding method and apparatus based on a RASL picture that is associated with a first picture in decoding order and is decodable.
[0023] Further, according to the present disclosure, it is possible to provide an image encoding / decoding method and apparatus based on a RASL picture that can be used as a reference picture.
[0024] Further, according to the present disclosure, it is possible to provide a method of transmitting a bitstream generated by an image encoding method or apparatus according to the present disclosure.
[0025] Further, according to the present disclosure, it is possible to provide a recording medium storing a bitstream generated by an image encoding method or apparatus according to the present disclosure.
[0026] Further, according to the present disclosure, it is possible to provide a recording medium storing a bitstream received, decoded, and used for reconstructing an image by an image decoding apparatus according to the present disclosure.
[0027] Those skilled in the art will understand that the effects realized by the present disclosure are not limited to what has been particularly described hereinabove and other advantages of the present disclosure will be more clearly understood from the detailed description. BRIEF DESCRIPTION OF DRAWINGS
[0028] Figure 1 is a view schematically illustrating a video encoding system to which embodiments of the present disclosure are applicable.
[0029] Figure 2 is a view schematically illustrating an image encoding apparatus to which embodiments of the present disclosure are applicable.
[0030] Figure 3 is a view schematically illustrating an image decoding apparatus to which embodiments of the present disclosure are applicable.
[0031] Figure 4 is a schematic flowchart of an image decoding process to which embodiments of the present disclosure are applicable.
[0032] Figure 5is a schematic flowchart of an image encoding process to which embodiments of the present disclosure are applicable.
[0033] Figure 6 is a view illustrating an example of a layer structure of an encoded image / video.
[0034] Figure 7 is a view illustrating a picture parameter set (PPS) according to an embodiment of the present disclosure.
[0035] Figure 8 is a view illustrating a slice header according to an embodiment of the present disclosure.
[0036] Figure 9 is a view illustrating an example of a sub-picture.
[0037] Figure 10 is a view illustrating an example of a picture having a mixed NAL unit type.
[0038] Figure 11 is a view illustrating an example of a picture parameter set (PPS) according to an embodiment of the present disclosure.
[0039] Figure 12 is a view illustrating an example of a sequence parameter set (SPS) according to an embodiment of the present disclosure.
[0040] Figure 13 is a view illustrating a process of constructing a mixed NAL unit type through sub-picture bitstream merging.
[0041] Figure 14 is a view illustrating a decoding order and an output order for each picture type.
[0042] Figure 15a and Figure 15b is a view illustrating a type of a RASL picture according to an embodiment of the present disclosure.
[0043] Figure 16a and Figure 16b is a view illustrating a process of a RASL picture during random access.
[0044] Figure 17 is a flowchart illustrating a decoding process and an output process of a RASL picture according to an embodiment of the present disclosure.
[0045] Figure 18a and Figure 18b is a view illustrating a reference condition of a RASL picture according to an embodiment of the present disclosure.
[0046] Figure 19 is a flowchart illustrating an image encoding method according to an embodiment of the present disclosure.
[0047] Figure 20 FIG. 1 is a flowchart illustrating an image decoding method according to an embodiment of the disclosure.
[0048] Figure 21 FIG. 2 is a view illustrating a content streaming system to which an embodiment of the disclosure is applicable. DETAILED DESCRIPTION
[0049] Hereinafter, embodiments of the disclosure will be described in detail with reference to the accompanying drawings so as to be easily implemented by those skilled in the art. The disclosure may, however, be implemented in various different forms and is not limited to the embodiments described herein.
[0050] In describing the disclosure, if it is determined that a detailed description of related known functions or configurations makes the scope of the disclosure unnecessarily obscure, a detailed description thereof will be omitted. In the drawings, portions unrelated to the description of the disclosure are omitted, and like reference numerals are assigned to like parts.
[0051] In the disclosure, when one component is "connected", "coupled", or "linked" to another component, it can include not only a direct connection relationship but also an indirect connection relationship in which a middle component exists. In addition, when one component "includes" or "has" another component, it means that it can further include the other component unless otherwise specified, rather than excluding the other component.
[0052] In the disclosure, the terms first, second, and the like are used only for the purpose of distinguishing one component from other components, and do not limit the order or importance of the components unless otherwise specified. Accordingly, within the scope of the disclosure, a first component in one embodiment can be referred to as a second component in another embodiment, and similarly, a second component in one embodiment can be referred to as a first component in another embodiment.
[0053] In the disclosure, components distinguished from each other are intended to clearly describe each feature, and do not mean that the components must be separated. That is, a plurality of components can be integrated in one hardware or software unit, or one component can be distributed and implemented in a plurality of hardware or software units. Therefore, even if not specifically described, embodiments in which these components are integrated or distributed are included in the scope of the disclosure.
[0054] In the disclosure, components described in each embodiment are not necessarily essential components, and some components can be optional components. Therefore, embodiments consisting of a subset of the components described in the embodiments are also included in the scope of the disclosure. In addition, embodiments including other components in addition to the components described in the various embodiments are included in the scope of the disclosure.
[0055] The disclosure relates to encoding and decoding of an image, and unless redefined in the disclosure, the terms used in the disclosure can have the general meanings commonly used in the technical field to which the disclosure belongs.
[0056] In the disclosure, a "picture" generally refers to a unit representing one image for a specific time period, and a slice / tile is a coding unit constituting a part of a picture, and one picture can be composed of one or more slices / tiles. In addition, a slice / tile can include one or more coding tree units (CTU).
[0057] In the disclosure, "pixel" or "pel" can mean the smallest unit constituting one picture (or image). In addition, "sample" can be used as a term corresponding to a pixel. One sample can generally represent a pixel or a value of a pixel, and can represent only a pixel / pixel value of a luminance component or only a pixel / pixel value of a chrominance component.
[0058] In the disclosure, "unit" can mean a basic unit of image processing. The unit can include at least one of a specific area of a picture and information related to the area. In some cases, the unit can be used interchangeably with terms such as "sample array", "block", or "region". In general, an MxN block can include a set (or array) of M columns and N rows of samples (or sample array) or transform coefficients.
[0059] In the disclosure, "current block" can mean one of "current coding block", "current coding unit", "coding target block", "decoding target block", or "processing target block". When performing prediction, "current block" can mean "current prediction block" or "prediction target block". When performing transform (inverse transform) / quantization (dequantization), "current block" can mean "current transform block" or "transform target block". When performing filtering, "current block" can mean "filtering target block".
[0060] In addition, in the disclosure, unless explicitly stated as a chrominance block, "current block" can mean a block including both a luminance component block and a chrominance component block or a "luminance block of the current block". The luminance component block of the current block can be represented by including an explicit description of the luminance component block such as "luminance block" or "current luminance block". In addition, "chrominance component block of the current block" can be represented by including an explicit description of the chrominance component block such as "chrominance block" or "current chrominance block".
[0061] In the disclosure, the term " / " or "," can be interpreted to mean "and / or". For example, "A / B" and "A,B" can mean "A and / or B". In addition, "A / B / C" and "A / B / C" can mean "at least one of A, B, and / or C".
[0062] In the disclosure, the term "or" should be interpreted to mean "and / or". For example, the expression "A or B" can include 1) only "A", 2) only "B", or 3) both "A and B". In other words, in the disclosure, "or" should be interpreted to mean "additionally or alternatively".
[0063] Overview of a video coding system
[0064] Figure 1 is a view illustrating a video coding system according to the disclosure.
[0065] The video coding system according to the embodiments can include an encoding apparatus 10 and a decoding apparatus 20. The encoding apparatus 10 can deliver encoded video and / or image information or data in the form of a file or a stream to the decoding apparatus 20 via a digital storage medium or a network.
[0066] The encoding apparatus 10 according to the embodiments can include a video source generator 11, an encoding unit 12, and a transmitter 13. The decoding apparatus 20 according to the embodiments can include a receiver 21, a decoding unit 22, and a renderer 23. The encoding unit 12 can be referred to as a video / image encoding unit, and the decoding unit 22 can be referred to as a video / image decoding unit. The transmitter 13 can be included in the encoding unit 12. The receiver 21 can be included in the decoding unit 22. The renderer 23 can include a display and the display can be configured as a separate device or an external component.
[0067] The video source generator 11 can acquire a video / image through a process of capturing, synthesizing, or generating a video / image. The video source generator 11 can include a video / image capturing device and / or a video / image generating device. The video / image capturing device can include, for example, one or more cameras, a video / image archive including previously captured videos / images, etc. The video / image generating device can include, for example, a computer, a tablet, and a smart phone, and can generate a video / image (electronically). For example, a virtual video / image can be generated through a computer, etc. In this case, the video / image capturing process can be replaced by a process of generating related data.
[0068] The encoding unit 12 can encode an input video / image. For compression and coding efficiency, the encoding unit 12 can perform a series of processes, such as prediction, transformation, and quantization. The encoding unit 12 can output encoded data (encoded video / image information) in the form of a bitstream.
[0069] The transmitter 13 can transmit the encoded video / image information or data, which is output in the form of a bitstream, to the receiver 21 of the decoding apparatus 20 in the form of a file or a stream through a digital storage medium or a network. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter 13 can include an element for generating a media file through a predetermined file format and can include an element for transmission through a broadcasting / communication network. The receiver 21 can extract / receive a bitstream from a storage medium or a network and transmit the bitstream to the decoding unit 22.
[0070] The decoding unit 22 can decode a video / image by performing a series of processes corresponding to the operations of the encoding unit 12, such as dequantization, inverse transform, and prediction.
[0071] The renderer 23 can render the decoded video / image. The rendered video / image can be displayed through a display.
[0072] Overview of an image encoding device
[0073] Figure 2 is a view schematically showing an image encoding apparatus to which embodiments of the present disclosure are applicable.
[0074] As shown in Figure 2 , the image encoding apparatus 100 can include an image partitioner 110, a subtractor 115, a transformer 120, a quantizer 130, a dequantizer 140, an inverse transformer 150, an adder 155, a filter 160, a memory 170, an inter-predictor 180, an intra-predictor 185, and an entropy encoder 190. The inter-predictor 180 and the intra-predictor 185 can be collectively referred to as a "predictor". The transformer 120, the quantizer 130, the dequantizer 140, and the inverse transformer 150 can be included in a residue processor. The residue processor can further include the subtractor 115.
[0075] In some embodiments, all or at least some of the plurality of components configuring the image encoding apparatus 100 can be configured by one hardware component (e.g., an encoder or a processor). Also, the memory 170 can include a decoded picture buffer (DPB) and can be configured by a digital storage medium.
[0076] The image partitioner 110 can partition an input image (or picture or frame) input to the image encoding apparatus 100 into one or more processing units. For example, the processing units can be referred to as coding units (CUs). The coding units can be obtained by recursively partitioning a coding tree unit (CTU) or a largest coding unit (LCU) according to a quadtree binary tree ternary (QT / BT / TT) structure. For example, one coding unit can be partitioned into multiple coding units of a deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. For the partitioning of coding units, a quadtree structure can be applied first, and then a binary tree structure and / or a ternary tree structure can be applied. An encoding process according to the disclosure can be performed based on a final coding unit that is no longer partitioned. A largest coding unit can be used as the final coding unit, and a coding unit of a deeper depth obtained by partitioning the largest coding unit can also be used as the final coding unit. Here, the encoding process can include a process of prediction, transform, and reconstruction that will be described later. As another example, the processing unit of the encoding process can be a prediction unit (PU) or a transform unit (TU). The prediction unit and the transform unit can be divided or partitioned from the final coding unit. The prediction unit can be a sample prediction unit, and the transform unit can be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.
[0077] The predictor (inter prediction 180 or intra prediction 185) can perform prediction on a block (current block) to be processed and generate a prediction block including predicted samples of the current block. The predictor can determine whether to apply intra prediction or inter prediction on a basis of the current block or CU. The predictor can generate various information related to prediction of the current block and transmit the generated information to the entropy encoder 190. The information about prediction can be encoded in the entropy encoder 190 and output in the form of a bitstream.
[0078] The intra predictor 185 can predict the current block by referring to samples in the current picture. The reference samples can be located in the neighbors of the current block or can be placed separately according to the intra prediction mode and / or the intra prediction technique. The intra prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes can include, for example, a DC mode and a planar mode. The directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes according to the level of detail of the prediction direction. However, this is merely an example, and more or less directional prediction modes can be used according to settings. The intra predictor 185 can determine a prediction mode applied to the current block by using a prediction mode applied to a neighboring block.
[0079] The inter predictor 180 can derive a prediction block of the current block based on a reference block (a reference sample array) specified by a motion vector on a reference picture. In this case, to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of a block, a sub-block, or a sample based on the correlation of the motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, bi-prediction, etc.) information. In the case of inter prediction, the neighboring blocks can include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in a reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block can be the same or different. The temporal neighboring block can be referred to as a collocated reference block, a collocated CU (colCU), etc. The reference picture including the temporal neighboring block can be referred to as a collocated picture (colPic). For example, the inter predictor 180 can configure a motion information candidate list based on the neighboring blocks and generate information indicating which candidate to use to derive the motion vector and / or the reference picture index of the current block. The inter prediction can be performed based on various prediction modes. For example, in the case of a skip mode and a merge mode, the inter predictor 180 can use the motion information of the neighboring blocks as the motion information of the current block. In the case of the skip mode, unlike the merge mode, a residual signal can not be transmitted. In the case of a motion vector prediction (MVP) mode, the motion vector of the neighboring block can be used as a motion vector predictor, and the motion vector of the current block can be signaled by encoding a motion vector difference and an indicator of the motion vector predictor. The motion vector difference can mean a difference between the motion vector of the current block and the motion vector predictor.
[0080] The predictor can generate a prediction signal based on various prediction methods and prediction techniques described below. For example, the predictor can not only apply intra prediction or inter prediction, but also simultaneously apply intra prediction and inter prediction to predict the current block. The prediction method of simultaneously applying both intra prediction and inter prediction to predict the current block can be referred to as combined inter and intra prediction (CIIP). In addition, the predictor can perform intra block copy (IBC) to predict the current block. Intra block copy can be used for content image / video encoding of games, etc., for example, screen content coding (SCC). IBC is a method of predicting a current picture using a reference block previously reconstructed in the current picture at a position separated by a predetermined distance. When IBC is applied, the position of the reference block in the current picture can be encoded as a vector (block vector) corresponding to the predetermined distance. IBC basically performs prediction in the current picture, but can be performed similarly to inter prediction in that the reference block is derived within the current picture. That is, IBC can use at least one of the inter prediction techniques described in the disclosure.
[0081] The prediction signal generated by the predictor can be used to generate a reconstructed signal or to generate a residual signal. The subtractor 115 can generate a residual signal (a residual block or a residual sample array) by subtracting the prediction signal (a prediction block or a prediction sample array) output from the predictor from the input image signal (an original block or an original sample array). The generated residual signal can be transmitted to the transformer 120.
[0082] The transformer 120 can generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique can include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a karhunen-loève transform (KLT), a graph-based transform (GBT), or a conditional non-linear transform (CNT). Here, the GBT refers to a transform obtained from a graph when relationship information between pixels is represented by a graph. The CNT refers to a transform obtained based on a prediction signal generated using all previously reconstructed pixels. In addition, the transform process can be applied to a square pixel block having the same size or can be applied to a block having a variable size other than a square.
[0083] The quantizer 130 can quantize the transform coefficients and transmit them to the entropy encoder 190. The entropy encoder 190 can encode the quantized signal (information about the quantized transform coefficients) and output a bitstream. The information about the quantized transform coefficients can be referred to as residual information. The quantizer 130 can rearrange the quantized transform coefficients of the block type into a one-dimensional vector form based on a coefficient scan order, and generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form.
[0084] The entropy encoder 190 can perform various encoding methods, such as exponential golomb, context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), and the like. The entropy encoder 190 can encode information (e.g., values of syntax elements, etc.) required for video / image reconstruction other than the quantized transform coefficients together or individually. The encoded information (e.g., encoded video / image information) can be transmitted or stored in the form of a bitstream in units of a network abstraction layer (NAL). The video / image information can further include information about various parameter sets, such as an adaptive parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information can further include general constraint information. The information signaled, transmitted, and / or syntax elements described in the disclosure can be encoded through the above-described encoding process and included in the bitstream.
[0085] The bitstream can be transmitted through a network or can be stored in a digital storage medium. The network can include a broadcasting network and / or a communication network, and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) that transmits a signal output from the entropy encoder 190 and / or a storage unit (not shown) that stores the signal can be included as an internal / external element of the image encoding apparatus 100. Alternatively, the transmitter can be provided as a component of the entropy encoder 190.
[0086] The quantized transform coefficients output from the quantizer 130 can be used to generate a residual signal. For example, the residual signal (residual block or residual sample) can be reconstructed by applying dequantization and inverse transform to the quantized transform coefficients through the dequantizer 140 and the inverse transformer 150.
[0087] The adder 155 adds the reconstructed residual signal to a prediction signal output from the inter-predictor 180 or the intra-predictor 185 to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). If there is no residual for a block to be processed, for example, in the case of applying a skip mode, a prediction block can be used as a reconstructed block. The adder 155 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-prediction of a next block to be processed in the current picture, and can be used for inter-prediction of a next picture by filtering as described below.
[0088] The filter 160 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 160 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and store the modified reconstructed picture in the memory 170, specifically, in the DPB of the memory 170. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filtering, bilateral filtering, etc. The filter 160 can generate various information related to filtering and transmit the generated information to the entropy encoder 190, as described later in the description of each filtering method. The information related to filtering can be encoded by the entropy encoder 190 and output in the form of a bitstream.
[0089] The modified reconstructed picture transmitted to the memory 170 can be used as a reference picture in the inter-predictor 180. When inter-prediction is applied by the image encoding apparatus 100, prediction mismatch between the image encoding apparatus 100 and an image decoding apparatus can be avoided and coding efficiency can be improved.
[0090] The DPB of the memory 170 can store the modified reconstructed picture to be used as a reference picture in the inter-predictor 180. The memory 170 can store motion information of a block from which motion information in the current picture is derived (or encoded) and / or motion information of a block in the picture that has been reconstructed. The stored motion information can be transferred to the inter-predictor 180 and used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory 170 can store reconstructed samples of a reconstructed block in the current picture and can transfer the reconstructed samples to the intra-predictor 185.
[0091] Overview of an image decoding device
[0092] Figure 3 FIG. 1 is a view schematically illustrating an image encoding apparatus to which embodiments of the present disclosure can be applied.
[0093] As Figure 3 shown, the image decoding apparatus 200 can include an entropy decoder 210, a dequantizer 220, an inverse transformer 230, an adder 235, a filter 240, a memory 250, an inter-predictor 260, and an intra-predictor 265. The inter-predictor 260 and the intra-predictor 265 can be collectively referred to as a "predictor". The dequantizer 220 and the inverse transformer 230 can be included in a residual processor.
[0094] According to embodiments, all or at least some of the plurality of components configuring the image decoding apparatus 200 can be configured by hardware components (e.g., decoders or processors). Also, the memory 250 can include a decoded picture buffer (DPB) or can be configured by a digital storage medium.
[0095] The image decoding apparatus 200 that has received a bitstream including video / image information can reconstruct an image by performing a process corresponding to a process performed by the image encoding apparatus 100. Figure 2 The image decoding apparatus 200 can perform decoding using a processing unit applied in the image encoding apparatus. Accordingly, the processing unit for decoding can be, for example, a coding unit. The coding unit can be acquired by partitioning a coding tree unit or a largest coding unit. The reconstructed image signal decoded and output by the image decoding apparatus 200 can be reproduced through a reproduction apparatus (not shown).
[0096] The image decoding apparatus 200 can receive a bitstream in the form of a stream from the image encoding apparatus 100. Figure 2The received signal can be decoded by the entropy decoder 210. For example, the entropy decoder 210 can parse a bitstream to derive information (e.g., video / image information) required for image reconstruction (or picture reconstruction). The video / image information can further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information can further include general constraint information. The image decoding apparatus can also decode a picture based on the information on the parameter sets and / or the general constraint information. The information and / or the syntax elements described in the disclosure to be signaled / received can be decoded through a decoding process and obtained from the bitstream. For example, the entropy decoder 210 decodes information in the bitstream based on an encoding method such as exponential Golomb coding, CAVLC, or CABAC, and outputs values of syntax elements required for image reconstruction and quantized values of transform coefficients of a residual. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element in the bitstream, determine a context model using information of a decoded target syntax element, decoding information of a neighboring block and a decoded target block, or a symbol / bin decoded in a previous stage, perform arithmetic decoding on the bins by predicting a probability of occurrence of the bins according to the determined context model, and generate a symbol corresponding to a value of each syntax element. In this case, the CABAC entropy decoding method can update the context model by using information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model. Information related to prediction among the information decoded by the entropy decoder 210 can be provided to the predictors (inter-predictor 260 and intra-predictor 265), and residual values, i.e., quantized transform coefficients and related parameter information, on which entropy decoding is performed in the entropy decoder 210 can be input to the dequantizer 220. In addition, information on filtering among the information decoded by the entropy decoder 210 can be provided to the filter 240. Further, a receiver (not shown) for receiving a signal output from the image encoding apparatus can be further configured as an internal / external element of the image decoding apparatus 200, or the receiver can be a component of the entropy decoder 210.
[0097] Further, the image decoding apparatus according to the disclosure can be referred to as a video / image / picture decoding apparatus. The image decoding apparatus can be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder can include the entropy decoder 210. The sample decoder can include at least one of the dequantizer 220, the inverse transformer 230, the adder 235, the filter 240, the memory 250, the inter-predictor 260, or the intra-predictor 265.
[0098] The dequantizer 220 can dequantize the quantized transform coefficients and output the transform coefficients. The dequantizer 220 can rearrange the quantized transform coefficients in the form of a two-dimensional block. In this case, the rearrangement can be performed based on a coefficient scan order performed in the image encoding apparatus. The dequantizer 220 can perform dequantization on the quantized transform coefficients by using a quantization parameter (e.g., quantization step length information) and obtain the transform coefficients.
[0099] The inverse transformer 230 can inverse-transform the transform coefficients to obtain a residual signal (a residual block, a residual sample array).
[0100] The predictor can perform prediction on the current block and generate a prediction block including prediction samples of the current block. The predictor can determine whether to apply intra prediction or inter prediction to the current block based on information about prediction output from the entropy decoder 210, and can determine a specific intra / inter prediction mode (prediction technique).
[0101] The same as described in the predictor of the image encoding apparatus 100, the predictor can generate a prediction signal based on various prediction methods (techniques) described later.
[0102] The intra predictor 265 can predict the current block by referring to samples in the current picture. The description of the intra predictor 185 is equally applicable to the intra predictor 265.
[0103] The inter predictor 260 can derive a prediction block of the current block based on a reference block (a reference sample array) specified by a motion vector on a reference picture. In this case, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of a block, a sub-block, or a sample based on the correlation of the motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, bi-prediction, etc.) information. In the case of inter prediction, the neighboring blocks can include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. For example, the inter predictor 260 can configure a motion information candidate list based on the neighboring blocks, and derive a motion vector and / or a reference picture index of the current block based on received candidate selection information. The inter prediction can be performed based on various prediction modes, and the information about prediction can include information indicating an inter prediction mode of the current block.
[0104] The adder 235 can generate a reconstructed signal (a reconstructed picture, a reconstructed block, a reconstructed sample array) by adding the obtained residual signal to a prediction signal (a prediction block, a prediction sample array) output from the predictor (including the inter-predictor 260 and / or the intra-predictor 265). If the block to be processed has no residual (e.g., in the case where the skip mode is applied), the prediction block can be used as the reconstructed block. The description of the adder 155 is equally applicable to the adder 235. The adder 235 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-prediction of a next block to be processed in the current picture, and can be used for inter-prediction of a next picture by filtering as described below.
[0105] The filter 240 can improve subjective / objective picture quality by applying filtering to the reconstructed signal. For example, the filter 240 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and store the modified reconstructed picture in the memory 250, specifically, in the DPB of the memory 250. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filtering, bilateral filtering, etc.
[0106] The (modified) reconstructed picture stored in the DPB of the memory 250 can be used as a reference picture in the inter-predictor 260. The memory 250 can store motion information of a block from which motion information of the current picture is derived (or decoded) and / or motion information of a block of the picture that has been reconstructed. The stored motion information can be transferred to the inter-predictor 260 to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory 250 can store reconstructed samples of a reconstructed block in the current picture and transfer the reconstructed samples to the intra-predictor 265.
[0107] In the present disclosure, the embodiments described in the filter 160, the inter-predictor 180, and the intra-predictor 185 of the image encoding apparatus 100 can be equally or correspondingly applied to the filter 240, the inter-predictor 260, and the intra-predictor 265 of the image decoding apparatus 200.
[0108] General image / video encoding process
[0109] In image / video encoding, pictures configuring an image / video can be encoded / decoded according to a decoding order. A picture order corresponding to an output order of decoded pictures can be set differently from the decoding order, and based thereon, not only forward prediction but also backward prediction can be performed during inter-prediction.
[0110] Figure 4 An example of an illustrative picture decoding process to which embodiments of the present disclosure are applicable is shown.
[0111] Figure 4The illustrated processes can be performed by an image decoding device of Figure 3 For example, in Figure 4 , step S410 can be performed by the entropy decoder 210, step S420 can be performed by the predictors including the intra predictor 265 and the inter predictor 260, step S430 can be performed by the residual processor including the dequantizer 220 and the inverse transformer 230, step S440 can be performed by the adder 235, and step S450 can be performed by the filter 240. Step S410 can include the information decoding processes described in the present disclosure, step S420 can include the inter / intra prediction processes described in the present disclosure, step S430 can include the residual processing processes described in the present disclosure, step S440 can include the block / picture reconstruction processes described in the present disclosure, and step S450 can include the in-loop filtering processes described in the present disclosure.
[0112] Referring to Figure 4 , the picture decoding processes can illustratively include a process for obtaining image / video information from a bitstream (by decoding) (S410), picture reconstruction processes (S420 to S440), and in-loop filtering processes for the reconstructed pictures (S450). The picture reconstruction processes can be performed based on the prediction samples and the residual samples obtained by the inter / intra prediction (S420) and the residual processing (S430) (dequantization and inverse transformation of quantized transform coefficients) described in the present disclosure. The modified reconstructed pictures can be generated by the in-loop filtering processes for the reconstructed pictures generated by the picture reconstruction processes. The modified reconstructed pictures can be output as decoded pictures, stored in a decoded picture buffer or memory 250 of the decoding device, and used as reference pictures in the inter prediction processes when decoding pictures later. In some cases, the in-loop filtering processes can be omitted. In such cases, the reconstructed pictures can be output as decoded pictures, stored in a decoded picture buffer or memory 250 of the decoding device, and used as reference pictures in the inter prediction processes when decoding pictures later. As described above, the in-loop filtering processes (S450) can include a deblocking filtering process, a sample adaptive offset (SAO) process, an adaptive loop filter (ALF) process, and / or a bilateral filter process, some or all of which can be omitted. Additionally, one or some of the deblocking filtering process, the sample adaptive offset (SAO) process, the adaptive loop filter (ALF) process, and / or the bilateral filter process can be applied in sequence, or all of them can be applied in sequence. For example, the SAO process can be performed after the deblocking filtering process is applied to the reconstructed picture. Alternatively, for example, the ALF process can be performed after the deblocking filtering process is applied to the reconstructed picture. This can be similarly performed in the encoding device as well.
[0113] Figure 5An example of a schematic picture encoding process to which embodiments of the present disclosure are applicable is shown.
[0114] Figure 5 Each of the processes shown can be performed by Figure 2 the image encoding apparatus. For example, step S510 can be performed by a predictor including the intra predictor 185 and the inter predictor 180, step S520 can be performed by the residual processors 115, 120, and 130, and step S530 can be performed by the entropy encoder 190. Step S510 can include the inter / intra prediction processes described in the present disclosure, step S520 can include the residual processing processes described in the present disclosure, and step S530 can include the information encoding processes described in the present disclosure.
[0115] Referring to Figure 5 , the picture encoding process can schematically include not only a process of encoding information (e.g., prediction information, residual information, partition information, etc.) used for picture reconstruction and outputting in the form of a bitstream, but also a process of generating a reconstructed picture for a current picture and a process of applying in-loop filtering to the reconstructed picture (optionally). The encoding apparatus can derive (modified) residual samples from quantized transform coefficients through the dequantizer 140 and the inverse transformer 150, and generate a reconstructed picture based on prediction samples output as a result of step S510 and the (modified) residual samples. The reconstructed picture thus generated can be equal to a reconstructed picture generated in a decoding apparatus. A modified reconstructed picture can be generated through an in-loop filtering process on the reconstructed picture. In this case, the modified reconstructed picture can be stored in a decoded picture buffer of the memory 170, and can be used as a reference picture in an inter prediction process when a picture is encoded later, similarly to a decoding apparatus. As described above, in some cases, some or all of the in-loop filtering processes can be omitted. When the in-loop filtering processes are performed, (in-loop) filtering-related information (parameters) can be encoded in the entropy encoder 190 and output in the form of a bitstream, and a decoding apparatus can perform the in-loop filtering processes using the same method as the encoding apparatus based on the filtering-related information.
[0116] Through such in-loop filtering processes, noise (e.g., block artifacts and ringing artifacts) occurring during image / video encoding can be reduced, and subjective / objective visual quality can be improved. In addition, by performing the in-loop filtering processes in both the encoding apparatus and the decoding apparatus, the same prediction results can be derived by the encoding apparatus and the decoding apparatus, picture encoding reliability can be increased, and the amount of data to be transmitted for picture encoding can be reduced.
[0117] As described above, the picture reconstruction process can be performed not only in the image decoding apparatus but also in the image encoding apparatus. The reconstructed blocks can be generated in units of blocks based on the intra prediction / inter prediction, and a reconstructed picture including the reconstructed blocks can be generated. When the current picture / slice / tile group is an I picture / slice / tile group, the blocks included in the current picture / slice / tile group can be reconstructed based on only the intra prediction. On the other hand, when the current picture / slice / tile group is a P or B picture / slice / tile group, the blocks included in the current picture / slice / tile group can be reconstructed based on the intra prediction or the inter prediction. In this case, the inter prediction can be applied to some blocks in the current picture / slice / tile group, and the intra prediction can be applied to the remaining blocks. The color components of a picture can include a luma component and a chroma component, and the methods and embodiments of the disclosure are applicable to both the luma component and the chroma component unless explicitly limited in the disclosure.
[0118] Example of encoding a layer structure
[0119] The encoded video / image according to the disclosure can be processed, for example, according to the encoding layers and structures to be described below.
[0120] Figure 6 is a view illustrating an example of a layer structure of an encoded image / video.
[0121] The encoded image / video is classified into a video coding layer (VCL) for image / video decoding processing and processing of its own, a lower layer system for transmitting and storing encoded information, and a network abstraction layer (NAL) existing between the VCL and the lower layer system and responsible for a network adaptation function.
[0122] In the VCL, VCL data including compressed image data (slice data) can be generated, or supplemental enhancement information (SEI) messages additionally required for decoding processing of an image or parameter sets including information such as a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS) can be generated.
[0123] In the NAL, header information (NAL unit header) can be added to the raw byte sequence payload (RBSP) generated in the VCL to generate a NAL unit. In this case, the RBSP refers to the slice data, the parameter sets, the SEI messages generated in the VCL. The NAL unit header can include NAL unit type information designated according to the RBSP data included in the corresponding NAL unit.
[0124] As Figure 6As illustrated, the NAL units can be classified into VCL NAL units and non-VCL NAL units according to the type of RBSP generated in the VCL. The VCL NAL unit can mean a NAL unit including information (slice data) about an image, and the non-VCL NAL unit can mean a NAL unit including information (parameter set or SEI message) required to decode an image.
[0125] The VCL NAL unit and the non-VCL NAL unit can be attached with header information according to a data standard of a lower layer system and transmitted through a network. For example, the NAL unit can be modified to have a data format of a predetermined standard (e.g., H.266 / VVC file format, RTP (Real-time Transport Protocol), or TS (Transport Stream)) and transmitted through various networks.
[0126] As described above, in the NAL unit, the NAL unit type can be designated according to the RBSP data structure included in the corresponding NAL unit, and information about the NAL unit type can be stored in the NAL unit header and signaled. For example, this can be roughly classified into a VCL NAL unit type and a non-VCL NAL unit type according to whether the NAL unit includes information (slice data) about an image. The VCL NAL unit type can be classified according to the properties and types of pictures included in the VCL NAL unit, and the non-VCL NAL unit type can be classified according to the types of parameter sets.
[0127] Examples of the NAL unit type designated according to the type of the parameter set / information included in the non-VCL NAL unit type will be listed below.
[0128] - DCI (Decoding Capability Information) NAL unit type (NUT): NAL unit type including DCI
[0129] - VPS (Video Parameter Set) NUT: NAL unit type including VPS
[0130] - SPS (Sequence Parameter Set) NUT: NAL unit type including SPS
[0131] - PPS (Picture Parameter Set) NUT: NAL unit type including PPS
[0132] - APS (Adaptation Parameter Set) NUT: NAL unit type including APS
[0133] - PH (Picture Header) NUT: NAL unit type including picture header
[0134] The above-described NAL unit type can have syntax information for the NAL unit type, and the syntax information can be stored in the NAL unit header and signaled. For example, the syntax information can be a nal_unit_type, and the NAL unit type can be designated using a nal_unit_type value.
[0135] In addition, one picture can include a plurality of slices, and one slice can include a slice header and slice data. In this case, one picture header can be further added to the plurality of slices (slice header and slice data sets) in one picture. The picture header (picture header syntax) can include information / parameters commonly applicable to a picture. The slice header (slice header syntax) can include information / parameters commonly applicable to a slice. The APS (APS syntax) or the PPS (PPS syntax) can include information / parameters commonly applicable to one or more slices or pictures. The SPS (SPS syntax) can include information / parameters commonly applicable to one or more sequences. The VPS (VPS syntax) can be information / parameters commonly applicable to a plurality of layers. The DCI (DCI syntax) can include information / parameters related to decoding capability.
[0136] In the disclosure, the high-level syntax (HLS) can include at least one of the APS syntax, the PPS syntax, the SPS syntax, the VPS syntax, the DCI syntax, the picture header syntax, or the slice header syntax. In addition, in the disclosure, the low-level syntax (LLS) can include, for example, a slice data syntax, a CTU syntax, a coding unit syntax, a transform unit syntax, and the like.
[0137] In the disclosure, the image / video information encoded in the encoding apparatus and signaled to the decoding apparatus in the form of a bitstream can include not only intra partitioning related information, intra / inter prediction information, residual information, in-loop filtering information, but also information on a slice header, information on a picture header, information on an APS, information on a PPS, information on an SPS, information on a VPS, and / or information on a DCI. In addition, the image / video information can further include general constraint information and / or information on a NAL unit header.
[0138] Overview of signaling of entry points
[0139] As described above, the VCL NAL unit can include slice data as an RBSP (raw byte sequence payload). The slice data can be byte-aligned in the VCL NAL unit, and can include one or more subsets. At least one entry point for random access (RA) can be defined for the subset, and parallel processing can be performed based on the entry point.
[0140] The VVC standard supports wavefront parallel processing (WPP) as one of various parallel processing techniques. Multiple slices in a picture can be parallel coded / decoded based on WPP.
[0141] To enable parallel processing capability, entry point information can be signaled. An image decoding apparatus can directly access a starting point of a data segment included in a NAL unit based on the entry point information. Here, the starting point of the data segment can mean a starting point of a tile in a slice or a starting point of a CTU row in a slice.
[0142] The entry point information can be signaled in a high-level syntax (e.g., a picture parameter set (PPS)) and / or a slice header.
[0143] Figure 7 is a view illustrating a picture parameter set (PPS) according to an embodiment of the disclosure, Figure 8 is a view illustrating a slice header according to an embodiment of the disclosure.
[0144] First, referring to Figure 7 , a picture parameter set (PPS) can include entry_point_offsets_present_flag as a syntax element indicating whether entry point information is signaled.
[0145] The entry_point_offsets_present_flag can indicate whether signaling of entry point information is present in a slice header referring to the picture parameter set (PPS). For example, the entry_point_offsets_present_flag having a first value (e.g., 0) can indicate that signaling of entry point information of a tile or a specific CTU row in a tile is not present in the slice header. In contrast, the entry_point_offsets_present_flag having a second value (e.g., 1) can indicate that signaling of entry point information of a tile or a specific CTU row in a tile is present in the slice header.
[0146] Further, although Figure 7 shows a case in which the entry_point_offsets_present_flag is included in the picture parameter set (PPS), this is an example, and thus embodiments of the disclosure are not limited thereto. For example, the entry_point_offsets_present_flag can be included in a sequence parameter set (SPS).
[0147] Next, referring to Figure 8 , a slice header can include offset_len_minus1 and entry_point_offset_minus1[i] as syntax elements identifying an entry point.
[0148] offset_len_minus1 can indicate a value obtained by subtracting 1 from a bit length of entry_point_offset_minus1[i]. A value of offset_len_minus1 can have a range from 0 to 31. In an example, offset_len_minus1 can be signaled based on a variable NumEntryPoints indicating a total number of entry points. For example, offset_len_minus1 can be signaled only when NumEntryPoints is greater than 0. In addition, offset_len_minus1 can be signaled based on entry_point_offsets_present_flag described above with reference to Figure 7
[0149] entry_point_offset_minus1[i] can indicate an i-th entry point offset in units of bytes, and can be represented by adding 1 bit to offset_len_minus1. Slice data in a NAL unit can include a number of subsets identical to a value obtained by adding 1 to NumEntryPoints, and an index value indicating each subset can have a range from 0 to NumEntryPoints. A first byte of slice data in a NAL unit can be represented by 0.
[0150] When entry_point_offset_minus1[i] is signaled, a dummy prevention byte included in slice data in a NAL unit can be counted as a slice data portion for subset identification. Subset 0 as a first subset of slice data can have a configuration from byte 0 to entry_point_offset_minus1[0]. Similarly, subset k as a k-th subset of slice data can have a configuration from firstByte[k] to lastByte[k]. Here, firstByte[k] can be derived as shown in Equation 1 below, and lastByte[k] can be derived as shown in Equation 2 below.
[0151] [Equation 1]
[0152]
[0153] [Equation 2]
[0154] lastByte[k] = firstByte[k] + sh_entry_point_offset_minus1[k]
[0155] In Equation 1 and Equation 2, k can have a range from 1 to a value obtained by subtracting 1 from NumEntryPoints.
[0156] The last subset of slice data (i.e., the NumEntryPoints-th subset) can consist of the remaining bytes of the slice data.
[0157] In addition, if the predetermined synchronization process on the context variables is not performed before decoding the CTU including the first CTB of the CTB row in each tile (e.g., sps_entropy_coding_sync_enabled_flag == 0) and the slice includes one or more complete tiles, each subset of the slice header can consist of all coding bits of all CTUs in the same tile. In this case, the total number of subsets of the slice data can be equal to the total number of tiles in the slice.
[0158] In contrast, when the predetermined synchronization process is not performed and the slice includes one subset of CTB rows in a single tile, NumEntryPoints can be 0. In this case, one subset of the slice data can consist of all coding bits of all CTUs in the slice.
[0159] In contrast, when the predetermined synchronization process is performed (e.g., sps_entropy_coding_sync_enabled_flag == 1), each subset can consist of all coding bits of all CTUs of one CTB row in one tile. In this case, the total number of subsets of the slice data can be equal to the total number of CTB rows of each tile in the slice.
[0160] Overview of mixed NAL unit types
[0161] In general, one NAL unit type can be set for one picture. As described above, syntax information indicating the NAL unit type can be stored in the NAL unit header of the NAL unit and signaled. For example, the syntax information can be nal_unit_type, and the NAL unit type can be designated using the nal_unit_type value. Examples of NAL unit types to which embodiments of the present disclosure are applicable are shown in Table 1 below.
[0162] [Table 1]
[0163]
[0164]
[0165] Referring to Table 1, the VCL NAL unit types can be classified into NAL unit types 0 to 12 according to the properties and types of pictures. In addition, the non-VCL NAL unit types can be classified into NAL unit types 13 to 31 according to the types of parameter sets.
[0166] Detailed examples of the VCL NAL unit types are as follows.
[0167] - IRAP (intra random access point) NAL unit type (NUT): NAL unit type of an IRAP picture, which is set to the range from IDR_W_RADL to CRA_NUT.
[0168] - IDR (instantaneous decoding refresh) NUT: NAL unit type of an IDR picture, which is set to IDR_W_RADL or IDR_N_LP.
[0169] - CRA (clean random access) NUT: NAL unit type of a CRA picture, which is set to CRA_NUT.
[0170] - RADL (random access decodable leading) NUT: NAL unit type of a RADL picture, which is set to RADL_NUT.
[0171] - RASL (random access skipped leading) NUT: NAL unit type of a RASL picture, which is set to RASL_NUT.
[0172] - End NUT: NAL unit type of an end picture, which is set to TRAIL_NUT.
[0173] - GDR (gradual decoding refresh) NUT: NAL unit type of a GDR picture, which is GDR_NUT.
[0174] - STSA (step-wise temporal sub-layer access) NUT: NAL unit type of an STSA picture, which is set to STSA_NUT.
[0175] In addition, the VVC standard allows one picture to include multiple slices having different NAL unit types. For example, one picture can include at least one first slice having a first NAL unit type and at least two second slices having a second NAL unit type different from the first NAL unit type. In this case, the NAL unit type of the picture can be referred to as a mixed NAL unit type. When the VVC standard supports the mixed NAL unit type, it can be easier to reconstruct / synthesize multiple pictures in content synthesis processing, encoding / decoding processing, etc.
[0176] However, according to the existing scheme for the mixed NAL unit type, there is a problem that the type of a picture having the mixed NAL unit type based on RASL_NUT and RADL_NUT can not be clear. According to the type of the picture, the constraint on the picture order and more specifically on the reference picture list needs to be modified to ensure the correct picture order in the bitstream.
[0177] To solve such a problem, according to an embodiment of the disclosure, a picture having the mixed NAL unit type based on RASL_NUT and RADL_NUT can be regarded as a RASL picture. When an IRAP picture associated with the RASL picture is the first picture in the decoding order or the first IRAP picture, the output process of the RASL picture can be skipped. However, the decoding process of the RASL picture can be performed, and the RASL picture can be used as a reference picture for inter prediction of a RADL picture under a predetermined reference condition.
[0178] Hereinafter, embodiments of the disclosure will be described in detail with reference to the accompanying drawings.
[0179] When one picture has a mixed NAL unit type, the picture can include a plurality of sub-pictures having different NAL unit types. For example, the picture can include at least one first sub-picture having a first NAL unit type and at least one second sub-picture having a second NAL unit type.
[0180] A sub-picture can include one or more slices and constitute a rectangular region in a picture. The size of a sub-picture in a picture can be variously set. On the other hand, for all pictures belonging to one sequence, the size and position of a certain individual sub-picture can be set to be equal to each other.
[0181] Figure 9 is a view illustrating an example of a sub-picture.
[0182] Referring to Figure 9 , one picture can be divided into 18 tiles. 12 tiles can be arranged at the left side of the picture, and each tile can include one slice constituted by 4x4 CTUs. In addition, 6 tiles can be arranged at the right side of the picture, and each tile can include two slices, each slice including 2x2 CTUs and being stacked in the vertical direction. As a result, the picture can include 24 sub-pictures and 24 slices, and each sub-picture can include one slice.
[0183] In an embodiment, each sub-picture in a picture can be treated as a picture to support mixed NAL unit types. When a sub-picture is treated as a picture, the sub-picture can be independently coded / decoded regardless of the result of coding / decoding another sub-picture. Here, independent coding / decoding can mean that a block partitioning structure (e.g., a single tree structure, a dual tree structure, etc.), a prediction mode type (e.g., intra prediction, inter prediction, etc.), a decoding order, etc. of the sub-picture are different from another sub-picture. For example, when a first sub-picture is coded / decoded based on an intra prediction mode, a second sub-picture adjacent to the first sub-picture and treated as a picture can be coded / decoded based on an inter prediction mode.
[0184] Information about whether a sub-picture is treated as a picture can be signaled using a syntax element in a high-level syntax. For example, sps_subpic_treated_as_pic_flag[i] indicating whether a current sub-picture is treated as a picture can be signaled through a sequence parameter set (SPS). When the current sub-picture is treated as a picture, sps_subpic_treated_as_pic_flag[i] can have a second value (e.g., 1).
[0185] Figure 10 is a view illustrating an example of a picture having mixed NAL unit types.
[0186] Referring to Figure 10 One picture 1000 can include first to third sub-pictures 1010 to 1030. Each of the first and third sub-pictures 1010 and 1030 can include two slices. In contrast, the second sub-picture 1020 can include four slices.
[0187] When each of the first to third sub-pictures 1010 to 1030 is treated as a picture, the first to third sub-pictures 1010 to 1030 can be independently coded to constitute different bitstreams. For example, coded slice data of the first sub-picture 1010 can be encapsulated into one or more NAL units having the same NAL unit type as RASL_NUT to constitute a first bitstream (bitstream 1). In addition, coded slice data of the second sub-picture 1020 can be encapsulated into one or more NAL units having the same NAL unit type as RADL_NUT to constitute a second bitstream (bitstream 2). In addition, coded slice data of the third sub-picture 1030 can be encapsulated into one or more NAL units having the same NAL unit type as RASL_NUT to constitute a third bitstream (bitstream 3). As a result, one picture 1000 can have a mixed NAL unit type of RASL_NUT and RADL_NUT mixed.
[0188] In an embodiment, all slices included in each sub-picture in a picture can be restricted to have the same NAL unit type. For example, both slices included in the first sub-picture 1010 can have the same NAL unit type as RASL NUT. In addition, all four slices included in the second sub-picture 1020 can have the same NAL unit type as RADL NUT. In addition, both slices included in the third sub-picture 1030 can have the same NAL unit type as RASL NUT.
[0189] The information about mixed NAL unit types and the information about picture partitioning can be signaled in, for example, high-level syntax of a picture parameter set (PPS) and / or a sequence parameter set (SPS).
[0190] Figure 11 is a view illustrating an example of a picture parameter set (PPS) according to an embodiment of the disclosure, and Figure 12 is a view illustrating an example of a sequence parameter set (SPS) according to an embodiment of the disclosure.
[0191] First, referring to Figure 11 , a picture parameter set (PPS) can include pps_mixed_nalu_types_in_pic_flag as a syntax element for mixed NAL unit types.
[0192] The pps_mixed_nalu_types_in_pic_flag can indicate whether the current picture has mixed NAL unit types. For example, the pps_mixed_nalu_types_in_pic_flag having a first value (e.g., 0) can indicate that the current picture does not have mixed NAL unit types. In this case, the current picture can have the same NAL unit type, e.g., the same NAL unit type as the coded slice NAL unit, for all VCL NAL units. In contrast, the pps_mixed_nalu_types_in_pic_flag having a second value (e.g., 1) can indicate that the current picture has mixed NAL unit types. In this case, the VCL NAL units of the current picture can be restricted to not have the same NAL unit type as GDR NUT. In addition, when any of the VCL NAL units of the current picture has the same NAL unit type as IDR W RADL, IDR N LP, or CRA NUT, all different VCL NAL units of the current picture can be restricted to have the same NAL unit type as IDR W RADL, IDR N LP, CRA NUT, or TRAIL NUT.
[0193] When the current picture has mixed NAL unit types (e.g., pps_mixed_nalu_types_in_pic_flag = = 1), each sub-picture in the current picture can have any of the VCL NAL unit types described above with reference to Table 1. For example, when a sub-picture in the current picture is an IDR sub-picture, the sub-picture can have the same NAL unit type as IDR W RADL or IDR N LP. Alternatively, when a sub-picture in the current picture is an end sub-picture, the sub-picture can have the same NAL unit type as TRAIL NUT.
[0194] pps_mixed_nalu_types_in_pic_flag having the second value (e.g., 1) can indicate that a picture referring to a picture parameter set (PPS) can include slices having different NAL unit types. Here, the picture can result from a sub-picture bitstream merge operation in which the encoder ensures matching of bitstream structures and alignment between parameters of the original bitstream. As an example of alignment, when a reference picture list (RPL) syntax element for a slice having the same NAL unit as IDR W RADL or IDR N LP is not present in the slice header (e.g., sps_idr_rpl_present_flag = = 0) and the current picture including the slice does not have mixed NAL unit types (e.g., pps_mixed_nalu_type_in_pic_flag = = 1), the current picture can be restricted to not include slices having the same NAL unit type as IDR W RADL or IDR N LP.
[0195] Further, pps_mixed_nalu_types_in_pic_flag can have the first value (e.g., 0) when there is a restriction that mixed NAL unit types are not applied for all pictures in an output layer set (OLS) (e.g., gci_no_mixed_nalu_types_in_pic_constraint_flag = = 1).
[0196] In addition, a picture parameter set (PPS) can include pps_no_pic_partition_flag as a syntax element for picture partitioning.
[0197] The pps_no_pic_partition_flag can indicate whether picture partitioning is applicable to the current picture. For example, the pps_no_pic_partition_flag having a first value (e.g., 0) can indicate that the current picture cannot be partitioned. In contrast, the pps_no_pic_partition_flag having a second value (e.g., 1) can indicate that the current picture can be partitioned into two or more tiles or slices. When the current picture has mixed NAL unit types (e.g., pps_mixed_nalu_types_in_pic_flag == 1), the current picture can be restricted to be partitioned into two or more tiles or slices (e.g., pps_no_pic_partition_flag = 1).
[0198] In addition, the picture parameter set (PPS) can include a pps_num_subpics_minus1 as a syntax element indicating a number of sub-pictures.
[0199] The pps_num_subpics_minus1 can indicate a value obtained by subtracting 1 from the number of sub-pictures included in the current picture. The pps_num_subpics_minus1 can be signaled only when picture partitioning is applicable to the current picture (e.g., pps_no_pic_partition_flag == 1). When the pps_num_subpics_minus1 is not signaled, a value of the pps_num_subpics_minus1 can be inferred to be 0. In addition, a syntax element indicating the number of sub-pictures can be signaled in a high-level syntax (e.g., a sequence parameter set (SPS)) other than the picture parameter set (PPS).
[0200] In an embodiment, when the current picture includes only one sub-picture (e.g., pps_num_subpics_minus1 == 0), the current picture can be restricted to not have mixed NAL unit types (e.g., pps_mixed_nalu_types_in_pic_flag == 0). That is, when the current picture has mixed NAL unit types (e.g., pps_mixed_nalu_types_in_pic_flag == 1), the current picture can be restricted to include two or more sub-pictures (e.g., pps_num_subpics_minus1 > 0).
[0201] Next, referring to FIG. 2, Figure 12A sequence parameter set (SPS) can include sps_subpic_treated_as_pic_flag[i] as a syntax element related to the treatment of sub-pictures during encoding / decoding.
[0202] sps_subpic_treated_as_pic_flag[i] can indicate whether each sub-picture in the current picture is treated as a picture. For example, sps_subpic_treated_as_pic_flag[i] having a first value (e.g., 0) can indicate that the ith sub-picture in the current picture is not treated as a picture. In contrast, sps_subpic_treated_as_pic_flag[i] having a second value (e.g., 1) can indicate that the ith sub-picture in the current picture is treated as a picture in encoding / decoding processes other than the in-loop filtering operation. When sps_subpic_treated_as_pic_flag[i] is not signaled, sps_subpic_treated_as_pic_flag[i] can be inferred to have the second value (e.g., 1).
[0203] In an embodiment, when the current picture includes two or more sub-pictures (e.g., pps_num_subpics_minus1 > 0) and at least one of the sub-pictures is not treated as a picture (e.g., sps_subpic_treated_as_pic_flag[i] == 0), the current picture can be restricted to not have mixed NAL unit types (e.g., pps_mixed_nalu_types_in_pic_flag = 0). That is, when the current picture has mixed NAL unit types (e.g., pps_mixed_nalu_types_in_pic_flag == 1), all of the sub-pictures in the current picture can be restricted to be treated as a picture.
[0204] Hereinafter, the NAL unit types according to the embodiments of the disclosure will be described in detail for each picture type.
[0205] (1) IRAP (Intra Random Access Point) picture
[0206] The IRAP picture is a picture that is randomly accessible, and can have the same NAL unit types described above with reference to Table 1, e.g., IDR_W_RADL, IDR_N_LP, or CRA_NUT. The IRAP picture can not refer to pictures other than the IRAP picture used for inter prediction in the decoding process. The IRAP picture can include an IDR (Instantaneous Decoding Refresh) picture and a CRA (Clean Random Access) picture.
[0207] The first picture in the bitstream in decoder order can be constrained to be an IRAP picture or a GDR (gradual decoding refresh) picture. For a single-layer bitstream, when the underlying parameter set to be referred to is available, the IRAP picture and all non-RASL pictures following the IRAP picture in decoding order can be correctly decoded, although pictures preceding the IRAP picture in decoding order are not decoded at all.
[0208] In an embodiment, the IRAP picture can not have mixed NAL unit types. That is, for the IRAP picture, pps_mixed_nalu_types_in_pic_flag can have a first value (e.g., 0), and all slices in the IRAP picture can have the same NAL unit type in the range from IDR W RADL to CRA NUT. As a result, when a first slice received from a predetermined picture has a NAL unit type in the range from IDR W RADL to CRA NUT, the picture can be determined to be an IRAP picture.
[0209] (2) CRA (clean random access) picture
[0210] The CRA picture is one of the IRAP pictures, and can have the same NAL unit type as CRA NUT, as described above with reference to Table 1. The CRA picture can not refer to pictures other than the CRA picture used for inter prediction in the decoding process.
[0211] The CRA picture can be the first picture in the bitstream in decoding order or a picture following the first picture. The CRA picture can be associated with a RADL or RASL picture.
[0212] When NoIncorrectPicOutputFlag has a second value (e.g., 1) for a CRA picture, the reference pictures of the RASL pictures associated with the CRA picture that are not present in the bitstream can not be decoded, and as a result, can not be output by the image decoding device. Here, NoIncorrectPicOutputFlag can indicate whether the pictures that precede the recovery point picture in decoding order are output before the recovery point picture. For example, NoIncorrectPicOutputFlag having a first value (e.g., 0) can indicate that the pictures that precede the recovery point picture in decoding order are output before the recovery point picture. In this case, the CRA picture can not be the first picture in the bitstream or the first picture in decoding order after an end of sequence (EOS) NAL unit, which can mean that random access does not occur. In contrast, NoIncorrectPicOutputFlag having a second value (e.g., 1) can indicate that the pictures that precede the recovery point picture in decoding order cannot be output before the recovery point picture. In this case, the CRA picture can be the first picture in the bitstream or the first picture in decoding order after an end of sequence (EOS) NAL unit, which can mean that random access occurs. Also, in some embodiments, NoIncorrectPicOutputFlag can be referred to as NoOutputBeforeRecoveryFlag.
[0213] For all picture units (PUs) within a CLVS (coded layer video sequence) that follow the current picture in decoding order, the reference picture lists 0 (e.g., RefPicList[0]) and 1 (e.g., RefPicList[1]) of one slice included in the CRA sub-picture belonging to the picture unit (PU) can be restricted to not include any picture within the active entry that precedes the picture including the CRA sub-picture in decoding order. Here, picture unit (PU) can mean a set of NAL units that includes multiple NAL units that are associated with each other according to a predetermined classification rule and are consecutive in decoding order for one coded picture.
[0214] (3) IDR (Instantaneous Decoding Refresh) picture
[0215] An IDR picture is one of IRAP pictures, and can have the same NAL unit type as IDR W RADL or IDR N LP, as described above with reference to Table 1. An IDR picture can not refer to pictures other than IDR pictures used for inter prediction in a decoding process.
[0216] An IDR picture can be the first picture in a bitstream in decoding order, or can be a picture following the first picture in decoding order. Each IDR picture can be the first picture in a CVS (coded video sequence) in decoding order.
[0217] When an IDR picture has the same NAL unit type as IDR W RADL for each NAL unit, the IDR picture can have an associated RADL picture. In contrast, when an IDR picture has the same NAL unit type as IDR N LP for each NAL unit, the IDR picture can not have an associated leading picture. Furthermore, an IDR picture can not be associated with a RASL picture.
[0218] For all picture units (PUs) following the current picture in decoding order within a CLVS (coded layer video sequence), the reference picture list 0 (e.g., RefPicList[0]) and the reference picture list 1 (e.g., RefPicList[1]) belonging to one slice included in the IDR sub-picture of the picture unit (PU) can be restricted to not include any picture within the active entry that precedes the picture including the IDR sub-picture in decoding order.
[0219] (4) RADL (random access decodable leading) picture
[0220] A RADL picture is one of the leading pictures, and can have the same NAL unit type as RADL NUT, as described above with reference to Table 1.
[0221] In decoding processing of a trailing picture having the same associated IRAP picture, a RADL picture can not be used as a reference picture. When field_seq_flag has a first value (e.g., 0) for the RADL picture, the RADL picture can precede all non-leading pictures having the same associated IRAP picture in decoding order. Here, field_seq_flag can indicate whether a CLVS (coded layer video sequence) conveys pictures indicating fields or pictures indicating frames. For example, field_seq_flag having a first value (e.g., 0) can indicate that the CLVS conveys pictures indicating frames. In contrast, field_seq_flag having a second value (e.g., 1) can indicate that the CLVS conveys pictures indicating fields.
[0222] (5) RASL (random access skipped leading) picture
[0223] A RASL picture is one of the leading pictures, and can have the same NAL unit type as RASL NUT, as described above with reference to Table 1.
[0224] In an example, all RASL pictures can be leading pictures of an associated CRA picture. When NoIncorrectPicOutputFlag has the second value (e.g., 1) for the CRA picture, the RASL picture refers to a picture that is not present in the bitstream, and thus can not be decoded, and as a result, can not be output by the picture decoding device.
[0225] In the decoding process of non-RASL pictures, a RASL picture can not be used as a reference picture. However, when there is a RADL picture that belongs to the same layer as the RASL picture and is associated with the same CRA picture, the RASL picture can be used as a collocated reference picture for inter prediction of a RADL sub-picture included in the RADL picture.
[0226] When field_seq_flag has the first value (e.g., 0) for the RASL picture, the RASL picture can precede, in decoding order, all non-leading pictures of the CRA picture associated with the RASL picture.
[0227] (6) End picture
[0228] An end picture is a non-IRAP picture that follows, in output order, an associated IRAP picture or GDR picture, and can not be a STSA picture. In addition, an end picture can follow, in decoding order, the associated IRAP picture. That is, an end picture that follows, in output order, the associated IRAP picture but precedes, in decoding order, the associated IRAP picture can not be allowed.
[0229] (7) GDR (Gradual Decoding Refresh) picture
[0230] A GDR picture is a picture that can be randomly accessible, and can have the same NAL unit type as GDR NUT, as described above with reference to Table 1.
[0231] (8) STSA (Stepwise Temporal Sub-Layer Access) picture
[0232] A STSA picture is a picture that can be randomly accessible, and can have the same NAL unit type as STSA NUT, as described above with reference to Table 1.
[0233] A STSA picture can not refer to a picture that has the same Temporalld as STSA for inter prediction. Here, Temporalld can be an identifier that indicates a temporal layer (e.g., a temporal sub-layer in scalable video coding). In an embodiment, a STSA picture can be restricted to have a Temporalld greater than 0.
[0234] For inter prediction, a picture that has the same TemporalID as the STSA picture and that follows the STSA picture in decoding order can not refer to pictures that have the same TemporalID as the STSA picture and that precede the STSA picture in decoding order. The STSA picture can enable an up-switch from an immediately lower sublayer to the current sublayer of the sublayer to which the STSA picture belongs.
[0235] Figure 13 is a view illustrating a process of constructing a mixed NAL unit type through sub-picture bitstream merging.
[0236] Referring to Figure 13 In the image encoding process, different bitstreams Bitstream 1 to Bitstream 3 can be generated for a plurality of sub-pictures included in one picture. For example, when one picture includes first to third sub-pictures, a first bitstream Bitstream 1 can be generated for the first sub-picture, a second bitstream Bitstream 2 can be generated for the second sub-picture, and a third bitstream Bitstream 3 can be generated for the third sub-picture. In this case, each bitstream independently generated for each sub-picture can be referred to as a sub-picture bitstream.
[0237] In an embodiment, all coded slices can be restricted to have the same NAL unit type for each sub-picture. Accordingly, all VCL NAL units included in one sub-picture bitstream can have the same NAL unit type. For example, all VCL NAL units included in the first bitstream Bitstream 1 have the same NAL unit type as RASL_NUT, all VCL NAL units included in the second bitstream Bitstream 2 have the same NAL unit type as RADL_NUT, and all VCL NAL units included in the third bitstream Bitstream 3 have the same NAL unit type as RASL_NUT.
[0238] In the image encoding process, a plurality of sub-pictures included in a current picture can be merged. That is, a plurality of sub-picture bitstreams corresponding to the plurality of sub-pictures can be merged into a single bitstream, and the current picture can be decoded based on the single bitstream. In this case, it can be determined whether the current picture has a mixed NAL unit type based on the plurality of merged sub-picture bitstreams having different NAL unit types. For example, when the first bitstream Bitstream 1 to the third bitstream Bitstream 3 are merged to constitute a single bitstream, the current picture can have a mixed NAL unit type based on RASL_NUT and RADL_NUT. In contrast, when the first bitstream Bitstream 1 to the third bitstream Bitstream 3 are merged to constitute a single bitstream, the current picture can have a non-mixed NAL unit type based on RASL_NUT.
[0239] Figure 14 is a view illustrating a decoding order and an output order of each picture type.
[0240] A plurality of pictures can be classified into an I picture, a P picture, or a B picture according to a prediction method. The I picture can refer to a picture to which only intra prediction is applied, and can be decoded without reference to another picture. The I picture can be referred to as an intra picture and can include the IRAP picture described above. The P picture can refer to a picture to which intra prediction and single-direction inter prediction are applied, and can be decoded by referring to another picture. The B picture can refer to a picture to which intra prediction and bi-directional / single-direction inter prediction are applied, and can be decoded using one or two other pictures. The P picture and the B picture can be referred to as inter pictures, and can include the RADL picture, the RASL picture, and the end picture described above.
[0241] The inter picture can be further classified into a leading picture (LP) or a non-leading picture (NLP) according to a decoding order and an output order. The leading picture can refer to a picture that is after an IRAP picture in the decoding order and is before the IRAP picture in the output order, and can include the RADL picture and the RASL picture described above. The non-leading picture can refer to a picture that is after the IRAP picture in the decoding order and the output order, and can include the end picture described above.
[0242] In Figure 14 , each picture name can indicate the picture type described above. For example, the I5 picture can be an I picture, the B0, B2, B3, B4, and B6 pictures can be B pictures, and the P1 and P7 pictures can be P pictures. In addition, in Figure 14 , each arrow can indicate a reference direction between pictures. For example, the B0 picture can be decoded by referring to the P1 picture.
[0243] Referring to Figure 14The I5 picture can be an IRAP picture (e.g., a CRA picture). When random access to the I5 picture is performed, the I5 picture can be the first picture in decoding order.
[0244] The B0 and P1 pictures can precede the I5 picture in decoding order and form a separate video sequence from the I5 picture. The B2, B3, B4, B6, and P7 pictures can follow the I5 picture in decoding order and form a video sequence with the I5 picture.
[0245] The B2, B3, and B4 pictures follow the I5 picture in decoding order and precede the I5 picture in output order, and thus can be classified as leading pictures. The B2 picture can be decoded by referring to the P1 picture that precedes the I5 picture in decoding order. Thus, when random access to the I5 picture is performed, the B2 picture cannot be correctly decoded by referring to the P1 picture that is not present in the bitstream. The picture type such as the B2 picture can be referred to as a RASL picture. In contrast, the B3 picture can be decoded by referring to the I5 picture that precedes the B3 picture in decoding order. Thus, when random access to the I5 picture is performed, the B3 picture can be correctly decoded by referring to the I5 picture that is decoded in advance. In addition, the B4 can be decoded by referring to the I5 picture and the B3 picture that precede the B4 picture in decoding order. Thus, when random access to the I5 picture is performed, the B4 picture can be correctly decoded by referring to the I5 picture and the B3 picture that are decoded in advance. The picture type such as the B3 and B4 pictures can be referred to as a RADL picture.
[0246] In addition, the B6 and P7 pictures follow the I5 picture in decoding order and output order, and thus can be classified as non-leading pictures. The B6 picture can be decoded by referring to the I5 picture and the P7 picture that precede the B6 picture in decoding order. Thus, when random access to the I5 picture is performed, the B6 picture can be correctly decoded by referring to the I5 picture and the P7 picture that are decoded in advance.
[0247] In one video sequence, the decoding process and the output process can be performed in different orders based on the picture type. For example, when a video sequence includes an IRAP picture, leading pictures, and non-leading pictures, the decoding process can be performed in the order of the IRAP picture, the leading pictures, and the normal pictures, and the output process can be performed in the order of the leading pictures, the IRAP picture, and the normal pictures.
[0248] Hereinafter, the type of the RASL picture and the process during random access will be described in detail.
[0249] Figure 15a and Figure 15b are views illustrating the type of the RASL picture according to an embodiment of the disclosure.
[0250] First, referring to Figure 15a , a RASL picture can have a single NAL unit type of RASL NUT. That is, when all VCL NAL units in a bitstream have the same NAL unit type as RASL NUT, a picture corresponding to the bitstream can be determined as a RASL picture. In the disclosure, a RASL picture having a single NAL unit type of RASL NUT can be referred to as a pure RASL picture.
[0251] In addition, in an embodiment, a RASL picture can have a mixed NAL unit type of RASL NUT and RADL NUT. Referring to Figure 15b , even when all VCL NAL units in a bitstream have the same NAL unit type as RASL NUT or RADL NUT, a picture corresponding to the bitstream can be considered as a RASL picture. In the disclosure, a RASL picture having a mixed NAL unit type of RASL NUT and RADL NUT can be referred to as a mixed RASL picture.
[0252] A RASL picture can include a pure RASL picture having a single NAL unit type of RASL NUT (see Figure 15a ) and a mixed RASL picture having a mixed NAL unit type of RASL NUT and RADL NUT (see Figure 15b ).
[0253] Figure 16a and Figure 16b are views illustrating processing of a RASL picture during random access. Specifically, Figure 16a shows a case of a pure RASL picture, and Figure 16b shows a case of a mixed RASL picture. In Figure 16a and Figure 16b , the basic configuration of multiple pictures (e.g., picture types and picture numbers) has been described above with reference to Figure 14 , and thus repetitive description thereof will be omitted.
[0254] When an IRAP picture associated with a pure RASL picture is a starting point of decoding processing or random access or is a first IRAP picture in decoding order, both decoding processing and output processing of the pure RASL picture can be skipped. Thus, the pure RASL picture can not be used as a reference picture for inter prediction.
[0255] For example, referring to Figure 16aIn contrast, in embodiments, when the IRAP picture associated with the mixed RASL picture is a starting point of decoding process or random access or is the first IRAP picture in decoding order, only the output process of the mixed RASL picture can be skipped. That is, the decoding process of the mixed RASL picture can be performed, and thus the mixed RASL picture can be used as a reference picture for inter prediction.
[0256] In contrast, in embodiments, when the IRAP picture associated with the mixed RASL picture is a starting point of decoding process or random access or is the first IRAP picture in decoding order, only the output process of the mixed RASL picture can be skipped. That is, the decoding process of the mixed RASL picture can be performed, and thus the mixed RASL picture can be used as a reference picture for inter prediction.
[0257] For example, referring to Figure 16b , the B2 picture can be a mixed RASL picture having a mixed NAL unit type of RASL NUT and RADL NUT. The B2 picture can be associated with an I5 picture that is an IRAP picture, and a first VCL NAL unit having the RADL NUT in the B2 picture can be decoded without reference to a P1 picture that precedes the I5 picture in decoding order. Thus, when the I5 is a starting point of random access, the decoding process of the B2 picture can be performed, and the B2 picture can be referenced by another picture (e.g., a B4 picture).
[0258] However, since a second VCL NAL unit having the RASL NUT in the B2 picture can still be decoded by reference to the P1 picture, the decoding process of the B2 picture can not ensure a correct decoding result of the second VCL NAL unit. Thus, when the I5 picture is a starting point of random access, the output process of the B2 picture can be skipped.
[0259] When the IRAP picture is a start point of decoding process or random access or is the first IRAP picture in decoding order, the output process of the RASL picture associated with the IRAP picture can be skipped regardless of whether the RASL picture is a pure RASL picture or a mixed RASL picture. In embodiments, whether the IRAP picture is a start point of decoding process or random access or the first IRAP picture in decoding order can be determined based on NoIncorrecPicOutputFlag (or NoOutputBeforeRecoveryFlag) described above. For example, when NoIncorrecPicOutputFlag has a first value (e.g., 0), the IRAP picture can be determined as neither a start point nor a first IRAP picture. In this case, the output process of the RASL picture associated with the IRAP picture can not be skipped. In contrast, when NoIncorrecPicOutputFlag has a second value (e.g., 1), the IRAP picture can be determined as a start point or a first IRAP picture. In this case, the output process of the RASL picture associated with the IRAP picture can be skipped.
[0260] In addition, the decoding process of the RASL picture associated with the IRAP picture can be skipped only when the RASL picture is a pure RASL picture. In embodiments, when an alternative timing for bitstream conformance test is used, the pure RASL picture associated with the first IRAP picture in decoding order can be excluded from the decoding process.
[0261] Figure 17 The decoding process and the output process according to the picture type of the RASL picture are shown in FIG. 10.
[0262] Figure 17 is a flowchart illustrating the decoding process and the output process of the RASL picture according to embodiments of the disclosure.
[0263] Figure 17 The decoding process and the output process of FIG. 10 can be performed by Figure 3 the image decoding apparatus of FIG. 1. For example, the decoding process can be performed by at least one of the image decoding apparatus, and the output process can be performed by the output interface of the image decoding apparatus under the control of the processor.
[0264] Referring to Figure 17 , the image decoding apparatus can determine whether the IRAP picture is the first picture in decoding order (e.g., a start point of decoding process or random access) in a current video sequence (S1710).
[0265] When the IRAP picture is not the first picture in the current video sequence (NO of S1710), the image decoding device can perform general decoding processing and output processing with respect to the current picture (S1720). For example, when the current picture is a leading picture, the current picture can be decoded after the IRAP picture but can be output before the IRAP picture. In contrast, when the current picture is a non-leading picture, the current picture can be decoded and output after the IRAP picture.
[0266] When the IRAP picture is the first picture in the current video sequence (YES of S1710), the image decoding device can determine whether the current picture is a RASL picture (S1730).
[0267] When the current picture is not a RASL picture (NO of S1730), the image decoding device can perform the general decoding processing and output processing described above with respect to the current picture (S1720).
[0268] When the current picture is a RASL picture (YES of S1730), the image decoding device can determine whether the current picture has a mixed NAL unit type (S1740).
[0269] When the current picture does not have a mixed NAL unit type (that is, when the current picture is a pure RASL picture) (NO of S1740), the image decoding device can skip both decoding processing and output processing for the current picture (S1750). As described above, a pure RASL picture can be decoded by referring to another picture that precedes the IRAP picture in decoding order. Thus, when the IRAP picture is the first picture in decoding order, the pure RASL picture can not be correctly decoded. Therefore, in this case, both decoding processing and output processing for the pure RASL picture associated with the IRAP picture can be skipped.
[0270] In contrast, when the current picture has the mixed NAL unit type (that is, when the current picture is a mixed RASL picture) (YES in S1740), the image decoding apparatus can skip only the output process for the current picture (S1760). As described above, the mixed RASL picture can include at least one first sub-picture having the same NAL unit type as the RADL_NUT, and the first sub-picture can be treated as one picture (e.g., subpic_treated_as_pic[i] = 1). Unlike the clean RASL picture, the first sub-picture can be decoded without referring to another picture that precedes the IRAP picture in the decoding order. As a result, the first sub-picture can be correctly decoded even when the IRAP picture is the first picture in the decoding order. Therefore, in this case, the decoding process for the mixed RASL picture associated with the IRAP picture can be performed. Further, the mixed RASL picture can also include at least one second sub-picture having the same NAL unit type as the RASL_NUT. Similar to the clean RASL picture, the second sub-picture can be decoded by referring to another picture that precedes the IRAP picture in the decoding order. As a result, the second sub-picture can not be correctly decoded when the IRAP picture is the first picture in the decoding order. Therefore, in this case, the output process for the mixed RASL picture associated with the IRAP picture can be skipped.
[0271] As described above, when the IRAP picture is a starting point of the decoding process or random access or is the first IRAP picture in the decoding order, the output process for the RASL picture associated with the IRAP picture can be skipped regardless of whether the RASL picture is a clean RASL picture or a mixed RASL picture. In contrast, the decoding process for the RASL picture associated with the IRAP picture can be skipped only when the RASL picture is a clean RASL picture.
[0272] Further, although the step S1710 is shown as being performed before the step S1720 in Figure 17 , this can be variously modified according to the embodiment. For example, the step S1710 can be performed simultaneously with the step S1720 or after the step S1710.
[0273] Figure 18a and Figure 18b are views illustrating the reference condition of the RASL picture according to the embodiment of the present disclosure. Specifically, Figure 18a illustrates the case of the clean RASL picture, and Figure 17 b illustrates the case of the mixed RASL picture.
[0274] Referring to Figure 18a and Figure 18bA RASL picture can be restrictively referred to by a RADL picture under a predetermined reference condition. The reference condition of a RASL picture according to an embodiment of the disclosure is as follows.
[0275] (1) Reference condition 1
[0276] A RASL picture referred to by a RADL picture can be restricted to a mixed RASL picture. For example, when a RASL picture includes a slice having RASL_NUT, the RASL picture can not be referred to by a RADL picture (see Figure 18a ). In contrast, when a RASL picture includes both a slice having the same NAL unit type as RASL_NUT and a slice having the same NAL unit type as RADL_NUT, the RASL picture can be referred to by a RADL picture (see Figure 18b ).
[0277] (2) Reference condition 2
[0278] a) A RADL picture referring to a RASL picture can be restricted to contain at least two sub-pictures.
[0279] b) In addition, each sub-picture can be restricted to be treated as one independent picture (e.g., subpic_treated_as_pic_flag[i] = 1).
[0280] c) In addition, a RASL picture referred to by a RADL picture can include at least two sub-pictures, and at least one of the sub-pictures can be restricted to have the same NAL unit type as RADL_NUT.
[0281] d) In addition, a sub-picture in a RASL picture having the same NAL unit type as RADL_NUT can be restricted to be treated as one independent picture (e.g., subpic_treated_as_pic_flag[i] = 1).
[0282] e) In addition, when a RASL picture is referred to by a specific sub-picture (subpicA) in a RADL picture, a collocated sub-picture of the specific sub-picture subpicA in the RASL picture can be restricted to a sub-picture having the same NAL unit type as RADL_NUT.
[0283] For example, as Figure 18bAs shown in the middle, a mixed RASL picture (at least one of which (Subpic_2) has RADL_NUT) including only two or more sub-pictures (Subpic_1, Subpic_2, and Subpic_3) can be referred to by a RADL picture. In this case, in the mixed RASL picture, the second sub-picture (Subpic2) with RADL_NUT can be considered as one independent picture. In addition, only the second sub-picture (Subpic2) can be the collocated sub-picture of the particular sub-picture in the RADL picture.
[0284] Furthermore, when the current picture has a mixed NAL unit type like a mixed RASL picture, multiple sub-picture bitstreams can be generated for multiple sub-pictures in the current picture. The multiple sub-picture bitstreams can be merged in the decoding process, thereby constructing a single bitstream. The single bitstream can be referred to as a single-layer bitstream, and the following constraints apply in order to meet bitstream conformance.
[0285] - (Constraint 1) Each picture within the bitstream, except the first picture in decoding order, is considered to be associated with the previous IRAP picture in decoding order.
[0286] - (Constraint 2) When one picture is a leading picture of an IRAP picture, the picture shall be a RADL picture or a RASL picture.
[0287] - (Constraint 3) When one picture is an end-of-IRAP picture, the picture shall not be a RADL picture or a RASL picture.
[0288] - (Constraint 4) Any RASL picture associated with an IDR picture shall not be present in the bitstream.
[0289] - (Constraint 5) Any RADL picture associated with an IDR picture having the same NAL unit type as IDR N LP shall not be present in the bitstream. In this case, when the necessary parameter sets to be referred to are available (within the bitstream or by external means), random access (and correct decoding of the IRAP picture and all non-RASL pictures consecutive to it) can be possible at the location of the IRAP picture unit (PU) by discarding all picture units (PUs) preceding the IRAP picture unit (PU).
[0290] - (Constraint 6) All pictures preceding the IRAP picture in decoding order shall precede the IRAP picture in output order and precede all RADL pictures associated with the IRAP picture in output order.
[0291] - (Constraint 7) All RASL pictures associated with a CRA picture shall precede all RADL pictures associated with the CRA picture in output order.
[0292] - (Constraint 8) All RASL pictures associated with a CRA picture shall follow, in output order, all IRAP pictures that precede the CRA picture in decoding order.
[0293] - (Constraint 9) When field_seq_flag has a first value (e.g., 0) and the current picture is a leading picture associated with an IRAP picture, the picture shall precede, in decoding order, all non-leading pictures associated with the IRAP picture. Alternatively, for a first leading picture picA and a last leading picture picB associated with an IRAP picture, one non-leading picture shall exist, in decoding order, before picA and no non-leading picture shall exist, in decoding order, between picA and picB.
[0294] Hereinafter, an image encoding / decoding method according to an embodiment of the disclosure will be described with reference to Figure 19 and Figure 20 The image encoding / decoding method according to an embodiment of the disclosure will be described in detail.
[0295] Figure 19 is a flowchart illustrating an image encoding method according to an embodiment of the disclosure.
[0296] Figure 19 The image encoding method of Figure 2 may be performed by an image encoding apparatus of For example, step S1910 can be performed by the image partitioner 110, and steps S1920 and S1930 can be performed by the entropy encoder 190.
[0297] Referring to Figure 19 , the image encoding apparatus can partition a current picture into two or more sub-pictures (S1910). Partitioning information of the current picture can be signaled using one or more syntax elements in a high-level syntax. For example, by a picture parameter set (PPS) described above with reference to Figure 7 , no_pic_partition_flag indicating whether the current picture is partitioned and pps_num_subpics_minus1 indicating a number of sub-pictures included in the current picture can be signaled. When the current picture is partitioned into two or more sub-pictures, no_pic_partition_flag can have a first value (e.g., 0) and pps_num_subpics_minus can have a value greater than 0.
[0298] Each subpicture in the current picture can be treated as a picture. When a subpicture is treated as a picture, the subpicture can be independently coded / decoded regardless of the result of coding / decoding another subpicture. During coding / decoding, information related to the processing of a subpicture can be signaled using syntax elements in high-level syntax. For example, by referring to the above description Figure 12 A sequence parameter set (SPS) described above, sps_subpic_treated_as_pic_flag[i] indicating whether each subpicture in the current picture is treated as a picture, can be signaled. When the i-th subpicture in the current picture is treated as a picture in the coding / decoding process excluding the in-loop filtering operation, sps_subpic_treated_as_pic_flag[i] can have a second value (e.g., 1). In addition, when sps_subpic_treated_as_pic_flag[i] is not signaled, sps_subpic_treated_as_pic_flag[i] can be inferred to have the second value (e.g., 1).
[0299] At least some of the subpictures in the current picture can have different NAL unit types. For example, when the current picture includes a first subpicture and a second subpicture, the first subpicture can have a first NAL unit type, and the second subpicture can have a second NAL unit type different from the first NAL unit type. Examples of the NAL unit type of a subpicture are described above with reference to Table 1.
[0300] All slices included in each subpicture in the current picture can have the same NAL unit type. In the above example, all slices included in the first subpicture can have the first NAL unit type, and all slices included in the second subpicture can have the second NAL unit type.
[0301] The picture coding device can determine a NAL unit type of each subpicture in the current picture (S1920).
[0302] In an embodiment, the NAL unit type of a subpicture can be determined based on a subpicture type. For example, when the subpicture is an IRAP subpicture, the NAL unit type of the subpicture can be determined as IDR W RADL, IDR N LP, or CRA NUT. Alternatively, when the subpicture is a RASL subpicture, the NAL unit type of the subpicture can be determined as RASL NUT.
[0303] Further, the combination of the NAL unit types of the sub-pictures in the current picture can be determined based on predetermined mixed constraints. For example, when one or more slices in the current picture have the same NAL unit type as IDR_W_RADL, IDR_N_LP, or CRA_NUT, all other slices in the current picture can have the same NAL unit type as IDR_W_RADL, IDR_N_LP, CRA_NUT, or TRAIL_NUT. In addition, when one or more slices in the current picture have the same NAL unit type such as RASL_NUT, all other slices in the current picture can have the same NAL unit type as RADL_NUT.
[0304] In an embodiment, the current picture having the mixed NAL unit types based on RASL_NUT and RADL_NUT can be regarded as a RASL picture associated with a predetermined intra random access point (IRAP) picture, and can be referred to as a mixed RASL picture. Here, the IRAP picture can include a clean random access (CRA) picture.
[0305] In an embodiment, the mixed RASL picture can be used as a reference picture for a RADL picture that refers to the same IRAP picture as the mixed RASL picture. In this case, the predetermined reference condition described above with reference to Figure 18a and Figure 18b is applicable.
[0306] The image encoding apparatus can encode the sub-pictures in the current picture based on the determined NAL unit types (S1930). In this case, as described above, the encoding process for each slice can be performed in a coding unit (CU) based on a predetermined prediction mode. Further, each sub-picture can be independently encoded to constitute different (sub-picture) bitstreams. For example, a first (sub-picture) bitstream including the encoding information of the first sub-picture and a second (sub-picture) bitstream including the encoding information of the second sub-picture can be constituted.
[0307] According to an embodiment of the disclosure, a picture having a mixed NAL unit type based on RASL_NUT and RADL_NUT can be regarded as a RASL picture. Unlike a RASL picture having a single NAL unit type of RASL_NUT, the RASL picture can be referred to by a RADL picture under a predetermined reference condition.
[0308] Figure 20 is a flowchart illustrating an image decoding method according to an embodiment of the disclosure.
[0309] Figure 20 The image decoding method of Figure 3The image decoding apparatus can perform the steps S2010 and S2020, and the step S2030. For example, the steps S2010 and S2020 can be performed by the entropy decoder 210, and the step S2030 can be performed by the dequantizer 220 to the intra predictor 265.
[0310] Referring to Figure 20 The image decoding apparatus can acquire, from the bitstream, NAL unit type information of at least one NAL unit including encoded picture data (S2010). Here, the encoded picture data can include slice data, and the NAL unit including the encoded picture data can mean a VCL NAL unit. In addition, the NAL unit type information can include a syntax element related to a NAL unit type. For example, the NAL unit type information can include a syntax element nal_unit_type acquired from a NAL unit header. In addition, the NAL unit type information can include a syntax element pps_mixed_nalu_types_in_pic_flag acquired from a picture parameter set (PPS).
[0311] The image decoding apparatus can determine at least one NAL unit type of one or more slices in a current picture based on the acquired NAL unit type information (S2020).
[0312] When the pps_mixed_nalu_types_in_pic_flag has a first value (e.g., 0), the current picture can not have mixed NAL unit types. In this case, all of the slices in the current picture can have the same NAL unit type determined based on a value of the nal_unit_type.
[0313] In contrast, when the pps_mixed_nalu_types_in_pic_flag has a second value (e.g., 1), the current picture can have mixed NAL unit types. In this case, the current picture can include two or more sub-pictures each regarded as one picture. Some of the sub-pictures can have different NAL unit types. For example, when the current picture includes first to third sub-pictures, the first sub-picture can have a first NAL unit type, the second sub-picture can have a second NAL unit type, and the third sub-picture can have a third NAL unit type. In this case, the first to third NAL unit types can be different from each other, or the first NAL unit type can be equal to the second NAL type but can be different from the third NAL unit type. All of the slices included in the sub-pictures can have the same NAL unit type.
[0314] Further, the combination of the NAL unit types of the slices in the current picture can be determined based on predetermined mixed constraints. For example, when one or more slices in the picture have the same NAL unit type as IDR W RADL, IDR N LP, or CRA NUT, all other slices in the picture can have the same NAL unit type as IDR W RADL, IDR N LP, CRA NUT, or TRAIL NUT. In addition, when one or more slices in the current picture have the same NAL unit type as RASL NUT, all other slices in the current picture can have the same NAL unit type as RADL NUT.
[0315] Further, when the current picture has mixed NAL unit types, the type of the current picture can be determined based on the NUT values of the slices. For example, when the slices in the current picture have NAL unit types of CRA NUT and IDR RADL, the type of the current picture can be determined as an IRAP picture. In addition, when the slices in the current picture have NAL unit types of RASL NUT and RADL NUT, the type of the current picture can be determined as a RASL picture. That is, the type of the current picture can be determined as a RASL picture based on at least one slice in the current picture having RASL NUT.
[0316] In an embodiment, when the current picture has mixed NAL unit types, the current picture can be restricted to have two or more sub-pictures. In addition, each sub-picture in the current picture can be restricted to be considered as one picture. For example, when the current picture is a RASL picture having mixed NAL unit types (i.e., a mixed RASL picture), the current picture can include two or more sub-pictures each of which is considered as one picture. In this case, at least one of the sub-pictures can include one or more slices having the same NAL unit type as RADL NUT.
[0317] The image decoding apparatus can decode each slice in the current picture based on the determined NAL unit types (S2030). In this case, the decoding process of each slice can be performed in a coding unit (CU) based on the predetermined prediction mode as described above.
[0318] Further, when the current picture is a RASL picture and the IRAP picture associated with the current picture is the first picture in decoding order (e.g., the start point of decoding process or random access), the current picture can be selectively decoded based on whether one or more slices with RADL NUT are included. For example, when the current picture is a pure RASL picture with a single NAL unit type of RASL NUT and random access is performed with respect to the IRAP picture associated with the current picture, the decoding process of the current picture can be skipped. In contrast, when the current picture is a mixed RASL picture with a mixed NAL unit type based on RASL NUT and RADL NUT and random access is performed with respect to the IRAP picture associated with the current picture, the decoding process of the current picture can be performed. Here, the IRAP picture can include a CRA picture.
[0319] In an embodiment, the mixed RASL picture can include a plurality of sub-pictures. In this case, each of the sub-pictures can be regarded as an independent picture. Information indicating whether the sub-pictures are regarded as an independent picture can be signaled using a syntax element in a high-level syntax. For example, sps_subpic_treated_as_pic_flag[i] indicating whether the current sub-picture is regarded as a picture can be signaled through a sequence parameter set (SPS). When the current sub-picture is regarded as a picture, sps_subpic_treated_as_pic_flag[i] can have a second value (e.g., 1). In addition, all slices included in the respective sub-pictures can have the same NAL unit type.
[0320] In an embodiment, at least one sub-picture having RADL NUT among the plurality of sub-pictures included in the mixed RASL picture can be used as a reference picture of a RADL picture associated with the same IRAP picture as the mixed RASL picture.
[0321] In an embodiment, a first sub-picture having RASL NUT among the plurality of sub-pictures included in the mixed RASL picture can precede a second sub-picture having RADL NUT among the plurality of sub-pictures in output order. Further, when the IRAP picture associated with the mixed RASL picture is the first picture in decoding order (e.g., the start point of decoding process or random access), the output process of the mixed RASL picture can be skipped.
[0322] According to embodiments of the present disclosure, a picture having a mixed NAL unit type based on RASL_NUT and RADL_NUT can be regarded as a RASL picture. The RASL picture can be referred to as a mixed RASL picture, and a RASL picture having a single NAL unit type of RASL_NUT can be referred to as a pure RASL picture. When an IRAP picture associated with the mixed RASL picture is a first picture in decoding order, only output processing of the mixed RASL picture can be skipped. Accordingly, the mixed RASL picture can be used as a reference picture for a RADL picture under a predetermined reference condition. In this regard, the mixed RASL picture can be different from the pure RASL picture whose both decoding and output processing are skipped.
[0323] The names of the syntax elements described in the present disclosure can include information about a position in which a corresponding syntax element is signaled. For example, a syntax element starting with "sps_" can mean that the corresponding syntax element is signaled in a sequence parameter set (SPS). In addition, syntax elements starting with "pps_," "ph_," and "sh_" can mean that the corresponding syntax elements are signaled in a picture parameter set (PPS), a picture header, and a slice header, respectively.
[0324] Although the above-described exemplary methods of the present disclosure are represented as a series of operations for clarity of description, the order of performing the steps is not intended to be limited, and the steps can be performed simultaneously or in a different order if necessary. To implement the methods according to the present disclosure, the described steps can further include other steps, can include the remaining steps except for some steps, or can include other additional steps except for some steps.
[0325] In the present disclosure, an image encoding apparatus or an image decoding apparatus that performs a predetermined operation (step) can perform an operation (step) that confirms an execution condition or a situation of the corresponding operation (step). For example, if it is described that a predetermined operation is performed when a predetermined condition is satisfied, the image encoding apparatus or the image decoding apparatus can perform the predetermined operation after determining whether the predetermined condition is satisfied.
[0326] The various embodiments of the present disclosure are not a list of all possible combinations and are intended to describe representative aspects of the present disclosure, and matters described in the various embodiments can be applied independently or in combination of two or more.
[0327] Various embodiments of the present disclosure can be implemented in hardware, firmware, software, or combinations thereof. In the case of hardware implementation, the present disclosure can be implemented by application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), general purpose processors, controllers, micro-controllers, microprocessors, and the like.
[0328] In addition, the image decoding apparatus and the image encoding apparatus to which embodiments of the present disclosure are applied can be included in a multimedia broadcast transmitting and receiving device, a mobile communication terminal, a home theater video device, a digital theater video device, a surveillance camera, a video chat device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camcorder, a video on demand (VoD) service providing device, an over the top video (OTT video) device, an Internet streaming service providing device, a three-dimensional (3D) video device, a video phone video device, a medical video device, and the like, and can be used to process a video signal or a data signal. For example, the OTT video device can include a game console, a Blu-ray player, an Internet access television, a home theater system, a smart phone, a tablet PC, a digital video recorder (DVR), and the like.
[0329] Figure 21 is a view illustrating a content streaming system to which embodiments of the present disclosure can be applied.
[0330] As Figure 21 As shown in
[0331] The encoding server compresses content input from a multimedia input device such as a smart phone, a camera, a camcorder, and the like, into digital data to generate a bitstream and transmits the bitstream to the streaming server. As another example, when a multimedia input device such as a smart phone, a camera, a camcorder, and the like, directly generates a bitstream, the encoding server can be omitted.
[0332] The bitstream can be generated by an image encoding method or an image encoding apparatus to which embodiments of the present disclosure are applied, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.
[0333] The streaming server transmits multimedia data to the user device based on a request of the user through the web server, and the web server serves as a medium to inform the user of a service. When the user requests a desired service to the web server, the web server can deliver it to the streaming server, and the streaming server can transmit multimedia data to the user. In this case, the content streaming system can include a separate control server. In this case, the control server is used to control commands / responses between devices in the content streaming system.
[0334] The streaming server can receive content from the media storage device and / or the encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store a bitstream for a predetermined time.
[0335] Examples of the user device can include a mobile phone, a smart phone, a laptop computer, a digital broadcast terminal, a personal digital assistant (PDA), a portable multimedia player (PMP), a navigation device, a slate PC, a tablet PC, an ultrabook, a wearable device (e.g., a smart watch, smart glasses, a head-mounted display), a digital TV, a desktop computer, a digital signage, etc.
[0336] The respective servers in the content streaming system can operate as distributed servers, in which case data received from the respective servers can be distributed.
[0337] The scope of the present disclosure includes software or machine-executable commands (e.g., operating systems, applications, firmware, programs, etc.) for enabling operations of the methods according to various embodiments to be performed on a device or computer, a non-transitory computer-readable medium having such software or commands stored thereon and executable on a device or computer.
[0338] Industrial applicability
[0339] Embodiments of the present disclosure can be used to encode or decode an image.
Claims
1. An image decoding method performed by an image decoding device, the image decoding method comprising the following steps: Obtain NAL unit type information from at least one Network Abstraction Layer (NAL) unit, including encoded image data of the current frame, from the bitstream; The NAL unit type of each of the multiple slices in the current frame is determined based on the obtained NAL unit type information; as well as The multiple slices are decoded based on the determined NAL unit type. Specifically, if at least one of the multiple slices has RASL_NUT and all other slices have either RASL_NUT or RADL_NUT, the current frame is considered a random access skipping of the preceding RASL frame. Wherein, based on at least one of the plurality of slices having the RADL_NUT instead of the RASL_NUT, the current frame of the RASL frame is considered a randomly accessable decodeable preceding RADL frame reference, and The current frame, which is considered the RASL frame, includes two or more sub-frames, and at least one of the sub-frames includes one or more slices having the RADL_NUT.
2. The image decoding method according to claim 1, in, Since the intra-frame random access point (IRAP) frame associated with the RASL frame is the first frame in the decoding order, the output processing for the RASL frame is skipped.
3. The image decoding method according to claim 2, in, Whether to skip the output processing for the RASL screen is determined based on the flag information for the IRAP screen.
4. The image decoding method according to claim 1, in, At least one sub-frame of the RASL frame is considered as a frame.
5. The image decoding method according to claim 1, in, The RASL screen is referenced based on at least one first sub-screen in the RADL screen, wherein the first sub-screen in the RASL screen comprises one or more slices having the RADL_NUT.
6. The image decoding method according to claim 1, in, The RADL screen comprises two or more sub-screens, each of which is considered a screen.
7. The image decoding method according to claim 1, in, Based on the fact that the intra-frame random access point (IRAP) frame associated with the RASL frame is the first frame in decoding order and all of the multiple slices have the RASL_NUT, the RASL frame is omitted from the decoding process.
8. An image encoding method performed by an image encoding device, the image encoding method comprising the following steps: Divide the current screen into multiple sub-screens; Determine the Network Abstraction Layer (NAL) unit type for each of the plurality of sub-pictures; as well as The multiple sub-pictures are encoded based on the determined NAL unit type. Specifically, based on the fact that at least one slice among the multiple slices included in the plurality of sub-pictures has RASL_NUT and all other slices among the plurality of slices have RASL_NUT or RADL_NUT, the current picture is considered as random access skipping the preceding RASL picture. Wherein, based on at least one of the plurality of slices having the RADL_NUT instead of the RASL_NUT, the current frame of the RASL frame is considered a randomly accessable decodeable preceding RADL frame reference, and The current frame, which is considered the RASL frame, includes two or more sub-frames, and at least one of the sub-frames includes one or more slices having the RADL_NUT.
9. The image encoding method according to claim 8, in, Since the intra-frame random access point (IRAP) frame associated with the RASL frame is the first frame in the decoding order, the output processing for the RASL frame is skipped.
10. A non-transitory computer-readable recording medium storing a program that, when executed by a processor, implements the image encoding method according to claim 8.
11. A method for transmitting a bit stream, the method comprising the following steps: Divide the current screen into multiple sub-screens; Determine the Network Abstraction Layer (NAL) unit type for each of the plurality of sub-pictures; The multiple sub-pictures are encoded based on the determined NAL unit type to generate the bitstream; as well as Send data including the bit stream. Specifically, based on the fact that at least one slice among the multiple slices included in the plurality of sub-pictures has RASL_NUT and all other slices among the plurality of slices have RASL_NUT or RADL_NUT, the current picture is considered as random access skipping the preceding RASL picture. Wherein, based on at least one of the plurality of slices having the RADL_NUT instead of the RASL_NUT, the current frame of the RASL frame is considered a randomly accessable decodeable preceding RADL frame reference, and The current frame, which is considered the RASL frame, includes two or more sub-frames, and at least one of the sub-frames includes one or more slices having the RADL_NUT.
Citation Information
Patent Citations
Image encoding / decoding method and apparatus based on hybrid NAL unit type, and method for transmitting bit stream
CN115244936A
Compact network abstraction layer (NAL) unit header
US20220303558A1