Video coding method and apparatus
By employing a video decoding and encoding method that utilizes varying NAL unit types for subpictures, the challenge of high data volume in high-resolution video is addressed, achieving efficient data compression and reduced transmission costs.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SAMSUNG ELECTRONICS CO LTD
- Filing Date
- 2026-01-14
- Publication Date
- 2026-04-10
AI Technical Summary
The increasing demand for high-resolution video leads to higher data transmission and storage costs due to the larger amount of encoded data, necessitating efficient encoding and decoding methods to reduce data volume.
A video decoding method that involves obtaining NAL unit type information to determine the type of the current NAL unit, allowing for decoding based on whether a mixed NAL unit type is applied, and a video encoding method that determines subpicture types for encoding, generating NAL units with varying NAL unit types for different subpictures within a single picture.
This approach enables efficient data compression by allowing different NAL unit types for subpictures, facilitating easier composition of images from multiple sequences without requiring equal NUT values, thus reducing data volume and transmission costs.
Smart Images

Figure 2026063102000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a subpicture segmentation method for synthesis with other sequences and a slice segmentation method for bitstream packing.
Background Art
[0002] The demand from users for high-resolution, high-quality video is increasing. Since the encoded data of high-resolution video has a larger amount of information than the encoded data of low-resolution or medium-resolution video, the cost for transmitting or storing this data increases.
[0003] Research on encoding and decoding methods for effectively reducing the amount of encoded data of high-resolution video to solve such problems has continued.
Summary of the Invention
Problems to be Solved by the Invention
[0004] This specification presents a subpicture segmentation method for synthesis with other sequences and a slice segmentation method for bitstream packing.
Means for Solving the Problems
[0005] To solve the above-mentioned problems, a video decoding method performed by a video decoding device according to one embodiment of the present invention includes the steps of: obtaining NAL unit type information indicating the type of the current NAL (network abstraction layer) unit from a bitstream; and decoding the video slice based on whether a mixed NAL unit type is applied to the current picture, if the NAL unit type information indicates that the NAL unit type of the current NAL unit is encoded data for a video slice. Here, the step of decoding the video slice may be performed by determining whether the NAL unit type of the current NAL unit indicates an attribute of a subpicture for the current video slice, based on whether the mixed NAL unit type is applied.
[0006] Furthermore, an image decoding device according to one embodiment of the present invention for solving the above-mentioned problems is an image decoding device including a memory and at least one processor, wherein the at least one processor obtains NAL unit type information indicating the type of the current NAL unit from a bitstream, and if the NAL unit type information indicates that the NAL unit type of the current NAL unit is encoded data for an image slice, the image slice may be decoded based on whether or not a hybrid NAL unit type is applied to the current picture. In this case, the decoding of the image slice may be performed by determining whether or not the NAL unit type of the current NAL unit indicates an attribute of a subpicture for the current image slice, based on whether or not the hybrid NAL unit type is applied.
[0007] Furthermore, a video encoding method performed by a video encoding device according to one embodiment of the present invention for solving the above-mentioned problems may include, when the current picture is encoded based on a hybrid NAL unit type, the steps of determining the type of subpictures to divide the picture, and, based on the type of subpictures, encoding at least one current video slice constituting the subpictures to generate a current NAL unit. Here, the step of encoding the video slice may be performed by encoding the NAL unit type of the current NAL unit to indicate the attributes of the subpicture to the current video slice, when the current picture is encoded based on the hybrid NAL unit type.
[0008] Furthermore, a transmission method according to one embodiment of the present invention for solving the above-mentioned problems may transmit a bitstream generated by the video encoding device or video encoding method of this disclosure.
[0009] Furthermore, a computer-readable recording medium according to one embodiment of the present invention for solving the above-mentioned problems may store a bitstream generated by the video encoding method or video encoding apparatus of this disclosure. [Effects of the Invention]
[0010] This invention presents a method for generating a single picture by combining it with many other sequences. A picture within a sequence is divided into numerous sub-pictures, and these divided sub-pictures from other pictures are combined to generate a new picture.
[0011] By applying the present invention, the NAL unit type values for two or more sub-pictures constituting a single picture may be different from each other. This has the advantage that when composing different contents, it is not necessary to make the NUTs of the numerous sub-pictures constituting a single image equal, thus making it easier to compose / compose an image. [Brief explanation of the drawing]
[0012] [Figure 1] This diagram schematically shows the configuration of a video encoding device to which the present invention is applied. [Figure 2] This figure shows an example of a video encoding method performed by a video encoding device. [Figure 3] This diagram schematically shows the configuration of a video decoding device to which the present invention is applied. [Figure 4] This figure shows an example of a video decoding method performed by a decoding device. [Figure 5] This figure shows an example of a NAL packet for slicing. [Figure 6] This figure shows an example of a hierarchical GOP structure. [Figure 7] This figure shows an example of the display output order and decoding order. [Figure 8] This figure shows examples of reading pictures and normal pictures. [Figure 9] This figure shows examples of RASL and RADL pictures. [Figure 10] This figure shows the syntax for slice segment headers. [Figure 11] This figure shows an example of the content synthesis process. [Figure 12] This figure shows an example of a subpicture ID and slice address. [Figure 13] This figure shows an example of a sub-picture / slice-specific NUT. [Figure 14] This figure shows one embodiment of the syntax for Picture Parameter Set (PPS). [Figure 15] This figure shows one embodiment of the syntax for a slice header. [Figure 16] This diagram shows the syntax of a picture header structure. [Figure 17] This diagram shows the syntax for obtaining a list of referenced pictures. [Figure 18] This is a diagram showing an example of content synthesis. [Figure 19] This is a sequence diagram for explaining a decoding method and an encoding method according to an embodiment of the present invention. [Figure 20] This is a sequence diagram for explaining a decoding method and an encoding method according to an embodiment of the present invention.
Embodiments for Carrying Out the Invention
[0013] The present invention may be subjected to various modifications and may have various embodiments. Therefore, specific embodiments are illustrated in the drawings and described in detail. However, this is not intended to limit the present invention to specific embodiments. The terms used in this specification are merely used to describe specific embodiments and are not intended to limit the technical idea of the present invention. Singular expressions include plural expressions unless the context clearly indicates a different meaning. Terms such as "including" or "having" in this specification are intended to specify the existence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and it should not be understood as precluding the possibility of the existence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0014] On the other hand, each configuration on the drawings described in the present invention is shown independently for the convenience of explaining separate characteristic functions, and it does not mean that each configuration is realized by separate hardware or separate software. For example, two or more of the configurations may be combined to form one configuration, or one configuration may be divided into multiple configurations. Embodiments in which each configuration is integrated and / or separated are also included in the scope of the present invention as long as they do not depart from the essence of the present invention.
[0015] Preferred embodiments of the present invention will be described in further detail below with reference to the attached drawings. Hereafter, the same reference numerals will be used for identical components in the drawings, and redundant descriptions of identical components will be omitted.
[0016] On the other hand, the present invention relates to video / image coding. For example, the methods / embodiments disclosed in the present invention may be applied to methods disclosed in the VVC (versatile video coding) standard, the EVC (Essential Video Coding) standard, the AV1 (AOMedia Video 1) standard, the AVS2 (2nd generation of audio video coding standard), or next-generation video / image coding standards (e.g., H.267, H.268, etc.).
[0017] In this specification, an Access unit (AU) refers to a unit representing multiple sets of pictures belonging to different layers that are output from a Decoded Picture Buffer (DPB) at the same time. A picture generally refers to a unit representing a single image at a specific time, and a slice is a unit that constitutes a part of a picture in coding. A single picture may consist of multiple slices, and pictures and slices may be mixed together as needed.
[0018] A pixel or pel may refer to the smallest unit that makes up a picture (or image). Alternatively, the term "sample" may be used as a counterpart to pixel. A sample may generally represent a pixel or a pixel value, or only the pixel / pixel value of the lumen component, or only the pixel / pixel value of the chroma component.
[0019] A unit represents a basic unit of image processing. A unit may include at least one of the following: a specific area of a picture and information about that area. A unit may be used interchangeably with terms such as block or area, as appropriate. In general, an M×N block may represent a set of samples or transform coefficients consisting of M columns and N rows.
[0020] Figure 1 is a schematic diagram showing the configuration of a video encoding device to which the present invention is applied.
[0021] Referring to Figure 1, the video encoding device 100 may include a picture splitting unit 105, a prediction unit 110, a residual processing unit 120, an entropy encoding unit 130, an addition unit 140, a filter unit 150, and a memory 160. The residual processing unit 120 may include a subtraction unit 121, a conversion unit 122, a quantization unit 123, a realignment unit 124, an inverse quantization unit 125, and an inverse transformation unit 126.
[0022] The picture splitting unit 105 splits the input picture into at least one processing unit.
[0023] For example, a processing unit may be called a coding unit (CU). In this case, a coding unit may be recursively partitioned from a coding tree unit using a quad-tree binary-tree (QTBT) structure. For example, a coding tree unit may be partitioned into multiple nodes of deeper depth based on a quad-tree structure and / or a binary tree structure. In this case, for example, the quad-tree structure may be applied first and the binary tree structure later, or the binary tree structure may be applied first. Decoding may be performed on nodes that cannot be further partitioned, and in this way, a coding unit may be determined for nodes that cannot be further partitioned. Since a coding tree unit is a unit for partitioning coding units, a coding tree unit may be named a coding unit. In this case, since a coding unit is determined by partitioning a coding tree unit, a coding tree unit may be named the largest coding unit (LCU).
[0024] Thus, the coding procedure according to the present invention may be performed based on a final coding unit that cannot be further divided. In this case, based on coding efficiency due to video characteristics, the coding tree unit may be immediately used as the final coding unit, or, if necessary, the coding unit may be recursively divided into coding units of lower depth, and the coding unit of the optimal size may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration described later.
[0025] As another example, the processing unit may include a coding unit (CU), a prediction unit (PU), or a transform unit (TU). The coding unit may be split from the coding tree unit into lower-depth coding units using a quad-tree structure. In this case, based on coding efficiency due to video characteristics, the coding tree unit may be used immediately as the final coding unit, or, if necessary, the coding unit may be recursively split into further lower-depth coding units so that the optimally sized coding unit is used as the final coding unit. If a minimum coding unit (min coding unit, min CU) is set, the coding unit is not split into coding units smaller than the minimum coding unit. Here, the final coding unit means the underlying coding unit that is partitioned or split into prediction units or transform units. The prediction unit is a unit partitioned from the coding unit and may be a sample prediction unit. In this case, the prediction unit may be divided into subblocks. The transformation unit may be divided from the coding unit by a quad-tree structure and may be a unit that derives transformation coefficients and / or a unit that derives residual signals from transformation coefficients. Hereinafter, the coding unit may be called a coding block (CB), the prediction unit a prediction block (PB), and the transformation unit a transform block (TB). A prediction block or prediction unit means a specific block-shaped region within a picture and may include an array of prediction samples. Similarly, a transformation block or transformation unit means a specific block-shaped region within a picture and may include an array of transformation coefficients or residual samples.
[0026] The prediction unit 110 makes predictions for the block to be processed (hereinafter referred to as the current block) and generates a predicted block that includes prediction samples for the current block. The unit of prediction performed by the prediction unit 110 may be a coding block, a conversion block, or a prediction block.
[0027] The prediction unit 110 determines whether intra-prediction or inter-prediction is applied to the current block. For example, the prediction unit 110 determines whether intra-prediction or inter-prediction is applied on a CU basis.
[0028] In the case of intra-prediction, the prediction unit 110 can derive prediction samples for the current block based on reference samples outside the current block within the picture to which the current block belongs (hereinafter referred to as the current picture). At this time, the prediction unit 110 may (i) derive prediction samples based on the average or interpolation of neighboring reference samples of the current block, or (ii) derive prediction samples based on reference samples among the neighboring reference samples of the current block that are located in a specific (prediction) direction relative to the prediction sample. Case (i) is called a non-directional mode or non-angular mode, and case (ii) is called a directional mode or angular mode. The prediction modes in intra-prediction may, for example, have 33 directional prediction modes and at least 2 non-directional modes. The non-directional modes may include DC prediction modes and planar modes. The prediction unit 110 may use the prediction modes applied to neighboring blocks to determine the prediction mode to be applied to the current block.
[0029] In interpretation, the prediction unit 110 can guide predicted samples for the current block based on samples identified by motion vectors on the reference picture. The prediction unit 110 can guide predicted samples for the current block by applying one of the following modes: skip mode, merge mode, and MVP (motion vector prediction). In skip mode and merge mode, the prediction unit 110 may use motion information from adjacent blocks as motion information for the current block. In skip mode, unlike merge mode, the difference (residual) between the predicted sample and the original sample is not transmitted. In MVP mode, the motion vector of adjacent blocks can be used as a motion vector predictor and used as a motion vector predictor for the current block to guide the motion vector of the current block.
[0030] In interpretation, adjacent blocks may include spatially adjacent blocks currently present in the picture and temporally adjacent blocks present in the reference picture. The reference picture containing the temporally adjacent blocks may be called a collocated picture (colPic). Motion information may include motion vectors and reference picture indices. Prediction mode information and motion information, etc., may be (entropy) encoded and output in the form of a bitstream.
[0031] When motion information of temporally adjacent blocks is used in skip mode and merge mode, the top-level picture on the reference picture list may be used as the reference picture. Reference pictures included in the reference picture list may be sorted based on the difference in POC (Picture Order Count) between the current picture and the reference picture. POC corresponds to the display order of the pictures and may be distinct from the coding order.
[0032] The subtraction unit 121 generates a residual sample, which is the difference between the original sample and the predicted sample. If skip mode is applied, it is not necessary to generate a residual sample, as described above.
[0033] The transformation unit 122 transforms residual samples in units of transformation blocks to generate transformation coefficients. The transformation unit 122 can perform the transformation according to the size of the transformation block and the prediction mode applied to the coding block or prediction block that spatially overlaps with the transformation block. For example, if intraprediction is applied to the coding block or prediction block that overlaps with the transformation block, and the transformation block is a 4x4 residual array, the residual samples are transformed using a DST (Discrete Sine Transform) transformation kernel. Otherwise, the residual samples are transformed using a DCT (Discrete Cosine Transform) transformation kernel.
[0034] The quantization unit 123 quantizes the conversion coefficients and generates quantized conversion coefficients.
[0035] The realignment unit 124 realigns the quantized transformation coefficients. The realignment unit 124 can realign the block-shaped quantized transformation coefficients into a one-dimensional vector form using a coefficient scanning method. Here, although the realignment unit 124 is described in a separate configuration, it may also be part of the quantization unit 123.
[0036] The entropy encoding unit 130 performs entropy encoding on the quantized conversion coefficients. Entropy encoding may include encoding methods such as exponential Golomb, CAVLC (context-adaptive variable length coding), and CABAC (context-adaptive binary arithmetic coding). In addition to the quantized conversion coefficients, the entropy encoding unit 130 may encode information necessary for video restoration (e.g., the values of syntax elements) together with or separately from the quantized conversion coefficients. The entropy-encoded information may be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer) units.
[0037] The inverse quantization unit 125 inversely quantizes the values (quantized conversion coefficients) quantized by the quantization unit 123, and the inverse transformation unit 126 inversely transforms the values inversely quantized by the inverse quantization unit 125 to generate residual samples.
[0038] The addition unit 140 reconstructs the picture by combining the residual sample and the predicted sample. The residual sample and the predicted sample are added in block units to generate a reconstructed block. Here, although the addition unit 140 has been described in a separate configuration, it may also be part of the prediction unit 110. On the other hand, the addition unit 140 is also called the reconstruction unit or the reconstructed block generation unit.
[0039] The filter unit 150 can apply a deblocking filter and / or a sample adaptive offset to the reconstructed picture. Deblocking filtering and / or a sample adaptive offset can correct block boundary artifacts and distortions from the quantization process within the reconstructed picture. The sample adaptive offset may be applied on a sample-by-sample basis or after the deblocking filtering process is complete. The filter unit 150 can also apply an Adaptive Loop Filter (ALF) to the reconstructed picture. The ALF may be applied to the reconstructed picture after the deblocking filter and / or a sample adaptive offset has been applied.
[0040] Memory 160 stores a restored picture (decoded picture) or information necessary for encoding / decoding. Here, the restored picture may be a restored picture that has undergone filtering by the filter unit 150. The stored restored picture may be used as a reference picture for (inter)prediction of other pictures. For example, memory 160 can store a (reference) picture used for interprediction. In this case, the picture used for interprediction may be specified by a reference picture set or a reference picture list.
[0041] Figure 2 shows an example of a video encoding method performed by a video encoding device. Referring to Figure 2, the video encoding method may include block partitioning, intra / inter prediction, transform, quantization, and entropy encoding processes. For example, the current picture may be divided into multiple blocks, a predicted block of the current block may be generated by intra / inter prediction, and a residual block of the current block may be generated by subtracting the input block of the current block from the predicted block. Subsequently, a coefficient block, i.e., a transformation coefficient of the current block, may be generated by transforming the residual block. The transformation coefficient may be quantized and entropy encoded and stored in a bitstream.
[0042] Figure 3 is a schematic diagram illustrating the configuration of a video decoding device to which the present invention is applied.
[0043] Referring to Figure 3, the video decoding device 300 may include an entropy decoding unit 310, a residual processing unit 320, a prediction unit 330, an addition unit 340, a filter unit 350, and a memory 360. Here, the residual processing unit 320 may include a realignment unit 321, an inverse quantization unit 322, and an inverse transformation unit 323.
[0044] When a bitstream containing video information is input, the video decoding device 300 can restore the video in accordance with the process by which the video information was processed by the video encoding device.
[0045] For example, the video decoding device 300 can perform video decoding using the processing units applied in the video encoding device. Therefore, the processing unit block for video decoding may be a coding unit as one example, or it may be a coding unit, a prediction unit, or a conversion unit as another example. The coding unit may be divided from the coding tree unit by a quad-tree structure and / or a binary tree structure.
[0046] Prediction units and transformation units may be used as appropriate, in which case the prediction block may be a block derived or partitioned from the coding unit and may be a unit for sample prediction. In this case, the prediction unit may be divided into subblocks. The transformation unit may be divided from the coding unit by a quad-tree structure and may be a unit for deriving transformation coefficients or a unit for deriving residual signals from transformation coefficients.
[0047] The entropy decoding unit 310 parses the bitstream and outputs the information necessary for video or picture restoration. For example, the entropy decoding unit 310 can decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the values of the syntax elements necessary for video restoration and the quantized values of the conversion coefficients for the residuals.
[0048] More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element in a bitstream, determines a context model using the syntax element information to be decoded and the decoding information of adjacent and decoded blocks or symbol / bin information decoded in a previous step, and uses the determined context model to predict the probability of bin occurrence and perform arithmetic decoding of the bins to generate symbols corresponding to the values of each syntax element. At this point, after determining the context model, the CABAC entropy decoding method can update the context model using the decoded symbol / bin information for the context model of the next symbol / bin.
[0049] The information decoded by the entropy decoding unit 310 that relates to predictions is provided to the prediction unit 330, and the residual values from the entropy decoding performed by the entropy decoding unit 310, i.e., the quantized transformation coefficients, are input to the re-alignment unit 421.
[0050] The realignment unit 421 realigns the quantized transformation coefficients into a two-dimensional block form. The realignment unit 421 can perform realignment in response to coefficient scanning performed by the encoding device. Here, although the realignment unit 321 has been described in a separate configuration, it may also be part of the inverse quantization unit 322.
[0051] The inverse quantization unit 322 can inverse quantize the quantized conversion coefficients based on (inverse) quantization parameters and output the conversion coefficients. In this case, information for inducing the quantization parameters may be signaled from the encoding device.
[0052] The inverse transformation unit 323 can inversely transform the transformation coefficients to derive the residual sample.
[0053] The prediction unit 330 can perform predictions on the current block and generate a predicted block that includes prediction samples for the current block. The unit of prediction performed by the prediction unit 330 may be a coding block, a conversion block, or a prediction block.
[0054] The prediction unit 330 determines whether to apply intra-prediction or inter-prediction based on the prediction information. In this case, the unit for determining which of intra-prediction and inter-prediction to apply is different from the unit for generating prediction samples. In addition, the unit for generating prediction samples is also different for inter-prediction and intra-prediction. For example, the decision on which of inter-prediction and intra-prediction to apply can be made at the CU (Unit of Reference) level. Alternatively, for example, in inter-prediction, the prediction mode may be determined at the PU (Unit of Reference) level and prediction samples may be generated, and in intra-prediction, the prediction mode may be determined at the PU level and prediction samples may be generated at the TU (Unit of Reference) level.
[0055] In the case of intra-prediction, the prediction unit 330 can guide prediction samples for the current block based on adjacent reference samples in the current picture. The prediction unit 330 can guide prediction samples for the current block by applying a directional mode or a non-directional mode based on adjacent reference samples in the current block. In this case, the prediction mode to be applied to the current block may be determined using the intra-prediction mode of the adjacent block.
[0056] In the case of interpretation, the prediction unit 330 can guide prediction samples for the current block based on the samples identified on the reference picture by motion vectors on the reference picture. The prediction unit 330 can guide prediction samples for the current block by applying one of the following modes: skip mode, merge mode, and MVP mode. At this time, motion information necessary for interpretation of the current block provided by the video encoding device, such as motion vectors and reference picture indexes, may be acquired or guided based on the prediction information.
[0057] In skip mode and merge mode, the movement information of adjacent blocks may be used as the movement information of the current block. In this case, adjacent blocks may include spatially adjacent blocks and temporally adjacent blocks.
[0058] The prediction unit 330 constructs a merge candidate list using motion information of available adjacent blocks, and may use the information indicated by the merge index on the merge candidate list as the motion vector of the current block. The merge index may be signaled from the encoding device. The motion information may include motion vectors and reference pictures. In skip mode and merge mode, when motion information of temporally adjacent blocks is used, the top-level picture on the reference picture list may be used as the reference picture.
[0059] In skip mode, unlike merge mode, the difference (residual) between the predicted sample and the original sample is not transmitted.
[0060] In MVP mode, the motion vector of the current block may be induced by using the motion vector of an adjacent block as a motion vector predictor. In this case, adjacent blocks may include both spatially adjacent blocks and temporally adjacent blocks.
[0061] As an example, when merge mode is applied, a merge candidate list may be generated using the motion vectors of the restored spatially adjacent blocks and / or the motion vectors corresponding to the temporally adjacent block, Col. In merge mode, the motion vectors of the candidate blocks selected from the merge candidate list are used as the motion vectors of the current block. The prediction information may include a merge index that indicates the candidate block having the optimal motion vector selected from among the candidate blocks included in the merge candidate list. In this case, the prediction unit 330 may derive the motion vector of the current block using the merge index.
[0062] As another example, when the MVP (Motion Vector Prediction) mode is applied, a list of motion vector predictor candidates may be generated using the motion vectors of the restored spatially adjacent blocks and / or the motion vectors corresponding to the temporally adjacent block, Col. That is, the motion vectors of the restored spatially adjacent blocks and / or the motion vectors corresponding to the temporally adjacent block, Col, may be used as motion vector candidates. The prediction information may include a predicted motion vector index that indicates the optimal motion vector selected from among the motion vector candidates included in the list. In this case, the prediction unit 330 can use the motion vector index to select the predicted motion vector for the current block from among the motion vector candidates included in the motion vector candidate list. The prediction unit of the encoding device calculates the motion vector difference (MVD) between the motion vector of the current block and the motion vector predictor, encodes it, and outputs it in the form of a bitstream. That is, the MVD is obtained by subtracting the motion vector predictor from the motion vector of the current block. At this time, the prediction unit 330 can obtain the motion vector difference included in the prediction information and derive the motion vector of the current block by adding the motion vector difference and the motion vector predictor. The prediction unit can also obtain or derive a reference picture index that indicates a reference picture from the prediction information.
[0063] The adder 340 restores the current block or current picture by adding the residual sample and the predicted sample. The adder 440 may restore the current picture by adding the residual sample and the predicted sample in block units. If skip mode is applied, the residual is not transmitted, so the predicted sample becomes the restored sample. Here, the adder 340 is described as a separate configuration, but it may be part of the prediction unit 330. On the other hand, the adder 340 is also called the restoration unit or the restored block generation unit.
[0064] The filter unit 350 may apply deblocking filtering, sample-adaptive offset and / or ALF to the restored picture. In this case, the sample-adaptive offset may be applied on a sample-by-sample basis or after deblocking filtering. ALF may also be applied after deblocking filtering and / or sample-adaptive offset.
[0065] Memory 360 stores the restored picture (decoded picture) or information necessary for decoding. Here, the restored picture may be a restored picture that has undergone filtering by the filter unit 350. For example, memory 360 may store a picture used for inter prediction. In this case, the picture used for inter prediction may be specified by a reference picture set or a reference picture list. The restored picture may be used as a reference picture for other pictures. Memory 360 may also output the restored pictures in output order.
[0066] Figure 4 shows an example of a video decoding method performed by a decoding device. Referring to Figure 4, the video decoding method may include entropy decoding, inverse quantization, inverse transform, and intra / inter prediction processes. For example, the decoding device may perform the reverse process of the encoding method. Specifically, quantized transformation coefficients may be obtained by entropy decoding of the bitstream, and the coefficient block of the current block, i.e., the transformation coefficients, may be obtained by an inverse quantization process on the quantized transformation coefficients. The residual block of the current block may be derived by an inverse transform on the transformation coefficients, and the reconstructed block of the current block may be derived by adding the predicted block of the current block derived by intra / inter prediction and the residual block.
[0067] On the one hand, the operator in the embodiments described below may be defined as follows in the table below.
[0068]
Table 1
[0069] Referring to Table 1, Floor(x) indicates the largest integer value less than or equal to x, Log2(u) indicates the logarithm value of u with base 2, and Ceil(x) indicates the smallest integer value greater than or equal to x. For example, in the case of Floor(5.93), the largest integer value less than or equal to 5.93 is 5, so it indicates 5.
[0070] Also, referring to Table 1, x>>y indicates an operator that right-shifts x by y bits, and x<<y indicates an operator that left-shifts x by y bits.
[0071] <遂行><Introduction> The HEVC standard proposes two types of screen partitioning methods.
[0072] 1) Slice: It provides a function of dividing a single image into coding tree unit (CTU) units in raster scan order for encoding / decoding, and slice header information exists.
[0073] 2) Tile: It provides a function of partitioning a single image into multiple columns and rows in CTU units for encoding / decoding. The partitioning method can be either equal division or individual division. The header for the tile does not exist separately.
[0074] A slice is a bitstream packing unit. That is, a single slice may be generated from a single NAL (network abstraction layer) bitstream. As shown in Figure 5, a NAL packet for a slice consists of a NAL header, a slice header, and slice data in that order. In this case, the NAL header information includes a NAL unit type (NUT).
[0075] Table 2 shows the NUTs for slices proposed in one embodiment of the HEVC standard. In Table 2, the NUTs for inter-slices where inter-prediction is performed are 0 through 9, and the NUTs for intra-slices where intra-prediction is performed are 16 through 21. Here, inter-slice means encoded using the inter-frame prediction method, and intra-slice means encoded using the intra-frame prediction method. Each slice is configured to have one NUT, and multiple slices within a single picture may all be configured to have the same NUT value. For example, if a single picture is divided into four slices and encoded using the intra-prediction method, the NUT values for all four slices within that picture may all be set to the same value, 19:IDR_W_RADL.
[0076] [Table 2]
[0077] In Table 2 above, abbreviations are defined as follows: -TSA(Temporal sub-layer Switching Access) -STSA(Step-wise Temporal sub-layer Switching Access) -RADL(Random Access Decodable Leading) -RASL(Random Access Skipped Leading) -BLA (Broken Link Access) -IDR(Instantaneous Decoding Refresh) - CRA (Clean Random Access) -LP (Leading Picture) -_N(No reference) -_R(Reference) -_W_LP / RADL(With LP / RADL) -_N_LP (No LP, without LP)
[0078] The BLA, IDR, and CRA, which are NUTs for intra-slice, are referred to as IRAP (Intra Random Access Point). IRAP is an intermediate position in the bitstream and represents a picture that can be accessed randomly. In other words, it is a picture that allows for abrupt changes in playback position during video playback. Intra-slice exists only in I-slice type.
[0079] An interslice is divided into P-slice or B-slice based on either unidirectional prediction (P: predictive) or bidirectional prediction (B: bi-predictive). The prediction and coding processes are performed in GOP (group of picture) units, but the HEVC standard uses a hierarchical GOP structure to perform the coding / decoding process, including prediction. Figure 6 shows an example of a hierarchical GOP structure, where each picture is divided into I, P, or B-picture (slice) depending on the prediction method.
[0080] Due to the bidirectional prediction capabilities of the B-slice and / or hierarchical GOP structure, the decoding and display order of pictures within the sequence differs (see Figure 7). In Figure 7, IRAP represents the intra-slice, and B and P represent the inter-slice, confirming that the playback and restoration order have been completely altered.
[0081] Among interslices, a picture whose restoration order is later than IRAP and whose playback order is earlier than IRAP is called an LP (leading picture) (see Figure 8). LPs are divided into RADL and RASL depending on the situation. RADL is defined as an LP that can be decoded when random access occurs, and RASL is defined as an LP that cannot be decoded during random access and whose restoration process must be skipped. In Figure 8, pictures of the same color are defined as one GOP.
[0082] The distinction between RADL and RASL is determined by the position of the reference picture during inter-slice prediction (see Figure 9). Specifically, RASL refers to an interpicture that uses a restored picture in another GOP as a reference picture, or a picture restored using a restored picture in another GOP as a reference picture. In this case, since a restored picture in another GOP is used (directly or indirectly) as a reference picture, it is called an open GOP. RASL and RADL are set in the NUT information for the interslice.
[0083] The NUT for an intraslice is divided into other intraslice NUTs based on the NUTs of preceding and / or succeeding interslices in the regeneration and / or restoration order of that intraslice. When examining the IDR NUT, IDRs are divided into IDR_W_RADL, which has RADL, and IDR_N_LP, which does not have LP. That is, IDRs are either of the type that do not have LP, or of the type that has only RADL among LP, and IDRs cannot have RASL. On the other hand, CRAs are of the type that have all of RADL and / or RASL among LP. That is, CRAs are of the type that can support open GOP.
[0084] Generally, intra-slices do not require reference picture information for that intra-slice because they only perform in-screen prediction. Here, the reference picture is used during inter-screen prediction. However, due to its feature supporting the open GOP structure, a CRA NUT slice inserts reference picture information into the NAL bitstream of the CRA, even though it is an intra-slice. This reference picture information is not for use in the CRA slice itself, but rather information for reference pictures that are planned to be used in subsequent inter-slices (in terms of the restoration order). This is because the reference picture is not removed in the DPB (decoded picture buffer). For example, if the NUT of the intra-slice is IDR, the DPB is reset. That is, all restoration pictures present in the DPB at that time are removed. Figure 10 shows the syntax for the slice segment header. As shown in Figure 10, if the NUT of the slice is not IDR, reference picture information can be written to the bitstream. That is, if the NUT of the slice is CRA, reference picture information can be written.
[0085] This invention presents a subpicture splitting method for compositing with other sequences and a slice splitting method for bitstream packing.
[0086] In this invention, a slice refers to an encoding / decoding region and is a data packing unit that generates a single NAL bitstream. For example, a single picture is divided into multiple slices, and each slice is generated as a single NAL packet through an encoding process.
[0087] In this invention, a subpicture is a region division for compositing with other content. Figure 11 shows an example of compositing with other content. There are three types of content: white, gray, and black. One image (AU: access unit) of each type of content is divided into four slice regions and packetized. As shown in the image on the right in Figure 11, the upper left portion may be white content, the lower left portion gray content, and the right portion black content, and these may be combined to generate a new image. Here, the white and gray regions consist of one subpicture per slice, and the black region consists of one subpicture per two slices. That is, one subpicture may contain at least one slice. To create a new image (to combine content), the BEAMe (Bit-stream Extractor And Merger) extracts regions from different content on a subpicture-by-subpicture basis and combines them. The combined image in Figure 11 may be divided into four slices and consist of three subpictures.
[0088] A sub-picture refers to an area that has the same sub-picture ID and / or sub-picture index value. In other words, a sub-picture area can be defined as the smallest single slice that has the same sub-picture ID and / or sub-picture index value. Here, the slice header information includes the sub-picture ID and / or sub-picture index value. The sub-picture index value may be set in the raster scan order. Figure 12 shows an example where one picture consists of six (rectangular) slices and four (color-coded) sub-picture areas. Here, A, B, C, and D are examples of sub-picture IDs, and 0 and 1 represent the slice addresses within the sub-picture. That is, the slice address value is the slice index value in the raster scan order within the sub-picture. For example, B-0 refers to the 0th slice within sub-picture B, and B-1 refers to the 1st slice within sub-picture B.
[0089] In this invention, the NUT values for two or more subpictures that constitute a single image may be different. For example, in Figure 12, a white subpicture (slice) in one image may be an intra-slice, and the gray subpicture (slice) and black subpicture (slice) may be inter-slices.
[0090] This has the advantage of making it easy to compose / combine images because, when combining different contents, it is not necessary to make the NUTs of the numerous sub-pictures that make up a single image equal. This function is called a mixed NAL Unit Type in a picture, and may be simply called a mixed NUT. The mixed_nalu_type_in_pic_flag can be used to enable / disable this function. This flag can be set in one or more of the following locations: SPS (sequence parameter set), PPS (picture parameter set), PH (picture header), and SH (slice header). For example, if the flag is set in the PPS, it is named pps_mixed_nalu_types_in_pic_flag.
[0091] If the aforementioned flag value is disabled (egmixed_nalu_type_in_pic_flag==0), then all subpictures and / or slices within the picture may have the same NUT value. For example, the NUT for all VCL (video coding layer) NAL units for a single picture may be set to have the same value. Also, a picture or picture unit (PU) may be referred to as having the same NUT as the encoded slice NAL unit for it. Here, VCL means the NAL type for a slice containing slice data values.
[0092] On the other hand, if the flag value is enabled (egmixed_nalu_type_in_pic_flag==1), the picture may consist of two or more subpictures. The picture may also have other NUT values. Furthermore, if the flag value is enabled, the VCL NAL units of the picture may be restricted from having a NUT of type GDR_NUT. Also, if the NUT (e.g., first NUT) of any one of the VCL NAL units (e.g., first NAL unit) of the picture is one of IDR_W_RADL, IDR_N_LP, or CRA_NUT, the NUT (e.g., second NUT) of the other VCL NAL units (e.g., second NAL unit) of the picture may be restricted to being set to one of IDR_W_RADL, IDR_N_LP, CRA_NUT, or TRAIL_NUT. For example, the second NUT may be restricted to being set to one of the values of the first NUT or TRAIL_NUT.
[0093] Referring to Figures 12 and 13, an example will be described in which the VCL NAL unit of the picture has at least two different NUT values. In one embodiment, two or more subpictures may have two or more different NUT values. In this case, the NUT values for all slices contained in a single subpicture may be equally limited. For example, as shown in Figure 13, the NUT values for the two slices in subpicture B in Figure 12 may be set to be equal to CRA, and the NUT values for the two slices in subpicture C may also be set to be equal to TRAIL, and subpictures A, B, C, and D may be set to have at least two different NUT values. Thus, as shown in Figure 13, the NUT values for the slices in subpictures A, C, and D may be set to be TRAIL, which is different from CRA, the NUT of subpicture B.
[0094] In the present invention, the NUTs for intra-slice and inter-slice are as shown in Table 3. As in the embodiment of Table 3, the definitions and functions for RADL, RASL, IDR, CRA, etc. may be set to be the same as the HEVC standard (Table 1). In the case of Table 3, a mixed NUT type is added. In Table 3, the disabled value of mixed_nalu_type_in_pic_flag (eg0) indicates the NUT for slices in a picture (similar to HEVC), and the enabled value of mixed_nalu_type_in_pic_flag (eg1) indicates the NUT for slices in subpictures. For example, if the value of mixed_nalu_type_in_pic_flag is 0 and the NUT of the VCL NAL unit is TRAIL_NUT, the NUT of the current picture is identified as TRAIL_NUT, and the NUTs of other subpictures belonging to the current picture may also be directed to be TRAIL_NUT. Furthermore, if the value of mixed_nalu_type_in_pic_flag is 1 and the NUT of the VCL NAL unit is TRAIL_NUT, then the NUT of the current subpicture is identified as TRAIL_NUT, and it is predicted that at least one of the other subpictures belonging to the current picture will not have a NUT that is TRAIL_NUT.
[0095] [Table 3]
[0096] As described above, if the value of mixed_nalu_type_in_pic_flag indicates enable (e.g., 1), then if any one VCL NAL unit (e.g., 1st NAL unit) belonging to a picture has one of the following NUT values (e.g., 1st NUT): IDR_W_RADL, IDR_N_LP, or CRA_NUT, then at least one of the other VCL NAL units (e.g., 2nd NAL unit) of the same picture may have one of the following NUT values (e.g., 2nd NUT): IDR_W_RADL, IDR_N_LP, CRA_NUT, or TRAIL_NUT, other than the 1st NUT.
[0097] Thus, if the VCL NAL unit (e.g., first NAL unit) for a first sub-picture belonging to a single picture has one of the following NUT values (e.g., first NUT): IDR_W_RADL, IDR_N_LP, or CRA_NUT, then the VCL NAL unit (e.g., second NAL unit) for a second sub-picture of the same picture may have one of the following NUT values (e.g., second NUT): IDR_W_RADL, IDR_N_LP, CRA_NUT, or TRAIL_NUT, other than the first NUT.
[0098] For example, if the value of mixed_nalu_type_in_pic_flag indicates activation (eg1), the NUT values of the VCL NAL units for two or more subpictures may be configured as follows. The following description is illustrative only and not limiting.
[0099] Combination 1) IRAP + non-IRAP (inter) Combination 2) non_IRAP(inter) + non-IRAP(inter) Combination 3) IRAP + IRAP = IDR + CRA (limited by embodiment)
[0100] Combination 1) is an embodiment in which at least one sub-picture in a picture has an IRAP (IDR or CRA) NUT value and at least one other sub-picture has a non-IRAP (inter-slice) NUT value. Here, the inter-slice NUT value may be a value excluding LP (RASL and RADL). For example, LP (RASL or RADL) may not be allowed as an inter-slice NUT value. In this way, the bitstream associated with the IDR or CRA sub-picture may be restricted so that the RASL and RADL sub-pictures are not encoded.
[0101] In another embodiment, only the TRAIL value may be allowed as the interslice NUT value. Alternatively, in another embodiment, all interslice VCL NUTs may be allowed as the interslice NUT value.
[0102] Combination 2) is an embodiment in which at least one sub-picture in the picture has a non-IRAP (inter-slice) NUT value, and at least one other sub-picture has a different non-IRAP (inter-slice) NUT value. For example, at least one sub-picture may have a RASL NUT value, and at least one other sub-picture may have a RADL NUT value. In the embodiment according to Combination 2), the following limitations may apply depending on the embodiment.
[0103] -In one embodiment, LP (RASL and RADL) and non-LP (TRAIL) are not used together. For example, the NUT of at least one subpicture cannot be RASL (or RADL) while the NUT of another at least one subpicture is TRAIL. If the NUT of at least one subpicture is RASL (or RADL), then RASL or RADL must not be used as the NUT of another at least one subpicture. For example, the leading subpicture of an IRAP subpicture may be forced to be a RADL or RASL subpicture.
[0104] -In another embodiment, LP (RASL and RADL) and non-LP (TRAIL) may be used together. For example, at least one subpicture may be RASL (or RADL) while at least one other subpicture is TRAIL.
[0105] -In another embodiment, exceptionally under condition 2), all subpictures may have the same interslice NUT value. For example, all subpictures within a picture may have a TRAIL NUT value. As another example, all subpictures within a picture may have a RASL (or RADL) NUT value.
[0106] Combination 3) represents an embodiment in which all subpictures or slices within a picture are composed of IRAP. For example, if the NUT value for a slice in a first subpicture is IDR_W_RADL, IDR_N_LP, or CRA_NUT, the NUT value for a slice in a second subpicture may be composed of IDR_W_RADL, IDR_N_LP, and CRA_NUT, but not the NUT value of the first subpicture. For example, the NUT value for a slice in at least one subpicture may be IDR, while the NUT value for a slice in at least one other subpicture may be CRA.
[0107] On the other hand, depending on the embodiment, the application of embodiments such as combination 3) may be restricted. In one embodiment, all pictures belonging to an IRAP or GDR access unit may be restricted to have the same NUT. That is, if the current access unit is an IRAP access unit composed only of IRAP pictures, or if the current access unit is a GDR access unit composed only of GDR pictures, all pictures belonging to it may be restricted to have the same NUT. For example, the NUT value for a slice in at least one subpicture may be IDR, while the NUT value for a slice in another at least one subpicture may not be CRA. In this way, when combination 3) is restricted and combinations 1) and 2) described above are applied, the at least one subpicture in the picture may be restricted to have a NUT value for non-IRAP (interslice). For example, during the encoding and decoding process, all subpictures in the picture may be restricted to not have a NUT value for IDR. Alternatively, some subpictures in the picture may have a NUT value for IDR, while other subpictures may not have a CRA NUT value.
[0108] The following describes the relevant syntax and semantics for signaling encoded information when a mixed NUT (mixed NAL unit type) is applied within a picture. The decoding process using these syntax and semantics is also described. As mentioned above, when mixed_nalu_type_in_pic_flag=1, the picture described by the NUT may also refer to a subpicture (see Table 3).
[0109] On the other hand, as described above, if the value of mixed_nalu_type_in_pic_flag indicates that a mixed NUT is applied, a single picture may be split into at least two subpictures. In this case, information about the subpictures for that picture may be signaled by a bitstream. In this respect, mixed_nalu_type_in_pic_flag can indicate whether or not a picture is currently being split. For example, if the value of mixed_nalu_type_in_pic_flag indicates that a mixed NUT is applied, it can indicate that a picture is currently being split.
[0110] The following explanation will be given with reference to the syntax in Figure 14. Figure 14 shows one embodiment of the Picture Parameter Set (PPS) syntax. For example, a bitstream may signal a flag (egpps_no_pic_partition_flag) in the Picture Parameter Set (PPS) indicating whether or not the picture is currently partitioned. A value (eg1) indicating enable for pps_no_pic_partition_flag can indicate that no picture partitioning is applied to pictures that currently reference the PPS. A value (eg0) indicating disable for pps_no_pic_partition_flag can indicate that picture partitioning using slices or tiles is applied to pictures that currently reference the PPS. In such an embodiment, if the value of mixed_nalu_type_in_pic_flag indicates that a mixed NUT is applied, the value of pps_no_pic_partition_flag may be forced to a value (eg0) indicating disable.
[0111] If pps_no_pic_partition_flag indicates that the picture is currently being partitioned, information about the number of subpictures (egpps_num_subpics_minus1) may be obtained from the bitstream. pps_num_subpics_minus1 can represent the number of subpictures currently contained in the picture minus 1. If pps_no_pic_partition_flag indicates that the picture is not currently being partitioned, the value of pps_num_subpics_minus1 may not be obtained from the bitstream and may be guided to 0. Based on the subpicture count information thus determined, encoding information for each subpicture may be signaled, corresponding to the number of subpictures contained in a single picture. For example, a subpicture identifier (egpps_subpic_id) to identify each subpicture and / or a flag value (subpic_treated_as_pic_flag[i]) indicating whether the encoding / decoding process for each subpicture is independent may be specified and signaled.
[0112] A hybrid NUT may be applied when a single picture consists of two or more subpictures. In this case, a flag value (subpic_treated_as_pic_flag[i]) indicating whether the encoding / decoding process of each subpicture is independent may be specified and signaled, equal to the number of subpictures (i) contained in a single picture. Decoding a subpicture independently means that it was treated as a separate picture during decoding. That is, if the flag value is on (egsubpic_treated_as_pic_flag=1), the subpicture may be decoded independently of other subpictures in all decoding processes except the in-loop filter process. Conversely, if the flag value is off (egsubpic_treated_as_pic_flag=0), the subpicture may reference other subpictures within the picture during the interprediction process. Here, a separate flag can be used to control whether the in-loop filter process is independent or referenced. The flag (subpic_treated_as_pic_flag) may be justified in one or more of the following locations: SPS, PPS, and PH. For example, if the flag is justified in SPS, it may be named sps_subpic_treated_as_pic_flag.
[0113] Furthermore, in the present invention, if other NUTs exist within a single picture (egmixed_nalu_type_in_pic_flag=1), each subpicture within the picture must be encoded / decoded independently due to the characteristic that different types of NUTs must be used between subpictures within a single picture. For example, in the case of a picture where mixed_nalu_type_in_pic_flag=1, if the picture contains one or more inter(P or B) slices, the subpic_treated_as_pic_flag value of all subpictures within the picture may be forced to be set to 1 or to be induced to a value of 1. Alternatively, if mixed_nalu_type_in_pic_flag=1, subpic_treated_as_pic_flag may be forced not to have a value of 0. For example, in the case of a picture where mixed_nalu_type_in_pic_flag=1, if one or more inter-slices are included within that picture, the subpic_treated_as_pic_flag value may be reset to 1 for all subpictures of that picture, regardless of the parsed value. Conversely, in the case of a picture where mixed_nalu_type_in_pic_flag=1 but subpic_treated_as_pic_flag=0, no inter-slices are included within that picture. That is, in the case of a picture where mixed_nalu_type_in_pic_flag=1 but subpic_treated_as_pic_flag=0, the slice type within that picture must be intra.
[0114] In another embodiment, if mixed_nalu_type_in_pic_flag=1, and the current picture's NUT is RASL, then the subpic_treated_as_pic_flag for the current picture may be forced to be set to 1. As another example, if mixed_nalu_type_in_pic_flag=1, and the current picture's NUT is RADL, but the referenced picture's NUT is RASL, then the subpic_treated_as_pic_flag for the current picture may be forced to be 1.
[0115] The mixed NUT function may restrict all subpictures (or slices) within a single picture from being composed of IRAP. In this case, the flag value (gdr_or_irap_pic_flag) indicating that all slices within a single picture are composed of IRAP or that the picture is a GDR (Gradual Decoding Refresh) picture may be forced to 0. That is, in the present invention, if other NUTs exist within a single picture (mixed_nalu_type_in_pic_flag=1), the flag value (gdr_or_irap_pic_flag) may be set to 0 or guided to a value of 0. Alternatively, if mixed_nalu_type_in_pic_flag=1, gdr_or_irap_pic_flag may be forced to have a value of 1. The flag (gdr_or_irap_pic_flag) may be enforced at one or more of the SPS, PPS, and PH locations.
[0116] Furthermore, with the application of the hybrid NUT function, at least one sub-picture within a single picture may have an IRAP (IDR or CRA) NUT value, while at least one other sub-picture may have a non-IRAP (inter-slice) NUT value. In other words, intra-slice and inter-slice may coexist within a single picture. In the case of the existing HEVC standard, if the NUT of the intra-slice was IDR, the DPB was reset. This removed all restored pictures present in the DPB at that time.
[0117] However, in the present invention, if mixed_nalu_type_in_pic_flag=1, intra-slice and inter-slice can coexist simultaneously within a single picture, so even if a single picture is an IDR NUT, there are cases where the DPB cannot be reset. As a result, in one embodiment, if the slice is an IDR NUT, reference picture list (RPL) may be inserted into the NAL bitstream as the slice header information of the IDR, as in CRA. For this reason, even though it is an IDR NUT, the value of the flag (idr_rpl_present_flag) indicating the presence of RPL information may be set to 1. When the value of the flag (idr_rpl_present_flag) is 1, the RPL exists as the slice header information of the IDR. Conversely, when the value of the flag (idr_rpl_present_flag) is 0, the RPL does not exist as the slice header information of the IDR.
[0118] On the other hand, in the present invention, if other NUTs exist within a single picture (mixed_nalu_type_in_pic_flag=1) but RPL information for an IDR picture is not permitted (idr_rpl_present_flag=0), the NUT for that picture must not have an IDR_W_RADL or IDR_N_LP value.
[0119] The aforementioned flag (idr_rpl_present_flag) may be justified at one or more of the following locations: SPS, PPS, and PH. For example, if the flag is justified at SPS, the flag may be named sps_idr_rpl_present_flag. For example, even if the NUT of a slice is currently IDR_W_RADL or IDR_N_RADL, the slice header information may be signaled using the slice header syntax of Figure 15 to signal the RPL based on the value of sps_idr_rpl_present_flag. Here, the first value of sps_idr_rpl_present_flag (eg0) indicates that the slice header of a slice whose NUT is IDR_N_LP or IDR_W_RADL does not provide an RPL syntax element. The second value of sps_idr_rpl_present_flag (eg1) indicates that the RPL syntax elements are provided by the slice header of a slice where NUT is IDR_N_LP or IDR_W_RADL.
[0120] On the other hand, in another embodiment, if mixed_nalu_type_in_pic_flag=1, the RPL may be signaled in the picture header information. For example, in the syntax application in Figure 14, if the value of mixed_nalu_type_in_pic_flag indicates that a mixed NUT is applied, the value of pps_no_pic_partition_flag may be forced to a value indicating disable (eg0). In addition, the value of a flag (pps_rpl_info_in_ph_flag) indicating whether or not RPL information is provided from the picture header may be obtained from the bitstream. If pps_rpl_info_in_ph_flag indicates enable (eg1), the RPL information may be obtained from the picture header as shown in Figures 16 and 17. In this way, RPL information is obtained based on the value of mixed_nalu_type_in_pic_flag, regardless of the type of picture. On the other hand, if pps_rpl_info_in_ph_flag indicates disabled (eg0), RPL information cannot be obtained from the picture header. For example, if the value of pps_rpl_info_in_ph_flag is 0, but the slice NUT is IDR_N_LP or IDR_W_RADL, and the value of sps_idr_rpl_present_flag is 0, then the RPL information for that slice will not be obtained. In other words, since there is no RPL information for that slice, the RPL information may be redirected to an initialized and empty one.
[0121] As described above, a single picture may be signaled by different types of NAL units. Since NAL units with different NUTs are used to signal a single picture, a method is required to determine the type of picture based on the type of NAL unit. This may determine whether the picture can be successfully restored and output during random access (RA).
[0122] In the decoding process according to one embodiment, if each VCL NAL unit corresponding to a picture is a CRA_NUT type NAL unit, the picture may be determined to be a CRA picture. Also, if each VCL NAL unit corresponding to a picture is an IDR_W_RADL or IDR_N_LP type NAL unit, the picture may be determined to be an IDR picture. Furthermore, if each VCL NAL unit corresponding to a picture is an IDR_W_RADL, IDR_N_LP, or CRA_NUT type NAL unit, the picture may be determined to be an IRAP picture.
[0123] Furthermore, if each VCL NAL unit corresponding to a single picture is a NAL unit of type RADL_NUT, the picture may be determined to be a RADL (Random Access decodable leading) picture. Also, if each VCL NAL unit corresponding to a single picture is a NAL unit of type TRAIL_NUT, the picture may be determined to be a trailing picture. Furthermore, if at least one of the VCL NAL units corresponding to a single picture is of type RASL_NUT, and all other VCL NAL units are of type RASL_NUT or RADL_NUT, the picture may be determined to be a RASL (random access skipped leading) picture.
[0124] On the other hand, in the decoding process according to another embodiment, if the minimum number of subpictures within a picture is RASL and the minimum number of other subpictures is RADL, the picture may be determined to be a RASL picture. For example, if the minimum number of subpictures within a picture is RASL and the minimum number of other subpictures is RADL, the picture may be set to a RASL picture during the decoding process. Here, if the type of the VCL NAL unit corresponding to the subpicture is RASL_NUT, the subpicture may be determined to be RASL. As a result, during RA, all RASL subpictures and RADL subpictures may be treated as RASL pictures, and therefore, the picture may not be output.
[0125] On the other hand, in the decoding process according to another embodiment, if the minimum number of sub-pictures within a picture is RASL, the picture may be set as a RASL picture. For example, if the minimum number of sub-pictures within a picture is RASL, and the other minimum number of sub-pictures is TRAIL, the picture may be set as a RASL picture during the decoding process. As a result, during RA, the picture may be treated as a RASL picture, and the picture may not be output.
[0126] Here, the occurrence of RA may be determined by the NoOutputBeforeRecoveryFlag value of the IRAP picture linked to the interslice (RADL, RASL, or TRAIL). If the flag value is 1 (NoOutputBeforeRecoveryFlag=1), it means that RA has occurred, and if the flag value is 0 (NoOutputBeforeRecoveryFlag=0), it means that normal playback has occurred. The flag value may be set for IRAP as follows:
[0127] - The process of setting the NoOutputBeforeRecoveryFlag value when the current picture is IRAP. 1. If the picture is the first picture in the bitstream, set NoOutputBeforeRecoveryFlag to 1. 2. If the picture is IDR, set NoOutputBeforeRecoveryFlag to 1. 3. If the picture is a CRA (Critical Recovery Arrest) and an RA (Recovery Arrest) is reported from an external source, set NoOutputBeforeRecoveryFlag to 1. 4. If the picture is a CRA (Critical Recovery Assurance) but no external notification of RA is received, set NoOutputBeforeRecoveryFlag to 0.
[0128] In one embodiment, the decoding device may signal the occurrence of random access from an external terminal. For example, the external terminal may signal the occurrence of random access by setting the value of the random access occurrence information to 1. The decoding device may set the value of HandleCraAsClvsStartFlag, a flag indicating whether it has received the occurrence of random access from the external terminal, to 1 based on the random access occurrence information received from the external terminal. The decoding device may set the value of NoOutputBeforeRecoveryFlag to the same value as the value of HandleCraAsClvsStartFlag. As a result, if the picture is currently a CRA picture and the value of HandleCraAsClvsStartFlag is 1, the decoding device may determine that random access has occurred for the CRA picture, or perform decoding as if the CRA was located at the beginning of the bitstream.
[0129] During RA (Router Adapter) processing, the process for setting the PictureOutputFlag, which determines whether or not the current picture is output, is as follows. For example, the PictureOutputFlag for the current picture may be set in the following order. Here, the first value of PictureOutputFlag (eg0) indicates that the current picture is not output. The second value of PictureOutputFlag (eg1) indicates that the current picture is output.
[0130] (1) If the picture is currently RASL and the NoOutputBeforeRecoveryFlag of the associated IRAP picture is 1, set PictureOutputFlag to 0. (2) If the current picture is a GDR picture with a NoOutputBeforeRecoveryFlag value of 1, or a restored picture thereof, set PictureOutputFlag to 0. (3) In addition, set the value of PictureOutputFlag to the same value as the pic_output_flag value in the bitstream. Here, pic_output_flag may be obtained at one or more positions in PH and SH.
[0131] Figure 18 illustrates an example of the synthesis of three different contents as presented in the present invention. Figure 18(a) shows a sequence for three different contents, where, for convenience, one picture is shown as one packet, but one picture may be divided into multiple slices, resulting in multiple palettes. Figures 18(b) and 18(c) show the synthesized image results for the picture indicated by the dotted line in Figure 18(a). In Figure 18, the same color represents the same picture / subpicture / slice. Also, the P slice and B slice may have one of the interNUT values.
[0132] As described above, with the present invention, when compositing multiple pieces of content, it is not always necessary to make the positions of the intra-slices (pictures) equal. By simply aligning the hierarchical GOP structure, content can be composed quickly and easily without delay.
[0133] Encoding and Decoding Embodiments The following describes how the video decoding device decodes video using the method described above. Figures 19 and 20 are sequential diagrams illustrating the decoding method and encoding method according to one embodiment of the present invention.
[0134] A video decoding device according to one embodiment may include memory and at least one processor, and the following decoding method can be performed by the operation of the processor. First, the decoding device obtains NAL unit type information indicating the current type of NAL (network abstraction layer) unit from the bitstream (S1910).
[0135] Next, the decoding device decodes the video slice based on whether a mixed NAL unit type is applied to the current picture, if the NAL unit type information indicates that the current NAL unit's NAL unit type is encoded data for the video slice (S1920).
[0136] Here, the decoding device can decode the video slice by determining whether the NAL unit type of the current NAL unit indicates the attributes of the subpicture for the current video slice, based on whether a hybrid NAL unit type is applicable.
[0137] Whether a mixed NAL unit type is applied may be identified based on a first flag (egpps_mixed_nalu_types_in_pic_flag) obtained from the picture parameter set. If a mixed NAL unit type is applied, the current picture to which the current video slice belongs may be split into at least two subpictures.
[0138] Furthermore, depending on whether a hybrid NAL unit type is applied, decoding information for subpictures may be included in the bitstream. In one embodiment, a second flag (egpps_no_pic_partition_flag) is obtained from the bitstream to indicate whether the picture is currently not partitioned. If the second flag indicates that the picture is currently partitionable (egpps_no_pic_partition_flag==0), a third flag (egpps_rpl_info_in_ph_flag) is obtained from the bitstream to indicate whether reference picture list information is provided from the picture header.
[0139] In such an example, when the hybrid NAL unit type is applied, the second flag (pps_no_pic_partition_flag) is forced to be 0 because the picture is now forced to be split into at least two subpictures, and the third flag (egpps_rpl_info_in_ph_flag), which indicates whether or not reference picture list information is provided from the picture header, is obtained from the bitstream regardless of the value of the second flag (pps_no_pic_partition_flag) actually obtained from the bitstream. As a result, if the third flag indicates that reference picture list information is provided from the picture header (egpps_rpl_info_in_ph_flag==1), the reference picture list information is obtained from the bitstream relating to the picture header.
[0140] Furthermore, when a hybrid NAL unit type is applied, the current picture may be decoded based on a first subpicture and a second subpicture having different NAL unit types. Here, if the NAL unit type of the first subpicture has one of the values IDR_W_RADL (Instantaneous Decoding Refresh_With_Random Access Decodable Leading), IDR_N_LP (Instantaneous Decoding Refresh_No reference_Leading Picture), and CRA_NUT (Clean Random Access_NAL Unit Type), the available NAL unit types that can be selected as the second subpicture NUT may include NAL unit types from IDR_W_RADL, IDR_N_LP, and CRA_NUT that were not selected from the first subpicture.
[0141] Alternatively, if the NAL unit type of the first sub-picture has one of the following values: IDR_W_RADL (Instantaneous Decoding Refresh_With_Random Access Decodable Leading), IDR_N_LP (Instantaneous Decoding Refresh_No reference_Leading Picture), and CRA_NUT (Clean Random Access_NAL Unit Type), the available NAL unit types of the second sub-picture may include TRAIL_NUT (Trail_NAL Unit Type).
[0142] On the other hand, when a hybrid NAL unit type is applied, the first and second subpictures that currently constitute the picture may be decoded independently. For example, the first and second subpictures containing B or P slices may be treated and decoded as a single picture. For example, the first subpicture may be decoded without using the second subpicture as a reference picture.
[0143] More specifically, a fourth flag (egsps_subpic_treated_as_pic_flag) may be obtained from the bitstream to indicate whether the first subpicture is treated as a picture during the decoding process. If the fourth flag indicates that the first subpicture is treated as a picture during the decoding process (egsps_subpic_treated_as_pic_flag==1), then the first subpicture may be treated as a picture and decoded during the decoding process. In such a process, if a hybrid NAL unit type is applied to the current picture and the current picture containing the first subpicture contains at least one P slice or B slice, the fourth flag may be forced to have a value indicating that the first subpicture is treated as a picture during the decoding process. On the other hand, if the hybrid NAL unit type is currently applied to the picture and the fourth flag indicates that the first subpicture will not be treated as a picture during the decoding process (egsps_subpic_treated_as_pic_flag==0), then the slice type currently belonging to the picture must be intra.
[0144] If the fourth flag indicates that the first subpicture will be treated as a picture during the decoding process, it can be determined that the decoding process of the first subpicture is independent of the other subpictures. For example, if the fourth flag indicates that the first subpicture will be decoded independently of the other subpictures during the decoding process, the first subpicture may be decoded without using the other subpictures as reference pictures.
[0145] Furthermore, if the first subpicture is a RASL (Random Access Skipped Leading) subpicture, the current picture may be determined to be a RASL picture based on whether or not the second subpicture is a RADL (Random Access Decodable Leading) subpicture. Here, if the type of NAL unit corresponding to the first subpicture is RASL_NUT (Random Access Skipped Leading_NAL Unit Type), the first subpicture may be determined to be a RASL subpicture.
[0146] Furthermore, if the third flag (egpps_rpl_info_in_ph_flag) indicates that the reference picture list information is not obtained from the picture header but from the slice header (egpps_rpl_info_in_ph_flag==0), and the NAL unit type of the first subpicture has one of the values IDR_W_RADL (Instantaneous Decoding Refresh_With_Random Access Decodable Leading) or IDR_N_LP (Instantaneous Decoding Refresh_No reference_Leading Picture), the reference picture list information may be obtained from the bitstream relating to the slice header based on the fifth flag (egsps_idr_rpl_present_flag), which indicates whether or not the reference picture list information for the IDR picture exists in the slice header. Here, the fifth flag may be obtained from the bitstream relating to the sequence parameter set.
[0147] On the other hand, if random access is performed to an IRAP (Intra Random Access Point) picture associated with the current picture, and the current picture is a RASL (Random Access Skipped Leading) subpicture, the current picture does not need to be output (displayed).
[0148] An image encoding device according to one embodiment may include a memory and at least one processor, and the processor's operation can perform an encoding method corresponding to the decoding method described above. For example, if the current picture is encoded based on a hybrid NAL unit type, the encoding device determines the type of sub-pictures to divide the picture (S2010). The encoding device also encodes at least one current image slice constituting the sub-picture based on the sub-picture type to generate a current NAL unit (S2020). At this time, if the current picture is encoded based on a hybrid NAL unit type, the encoding device encodes the image slice by encoding the NAL unit type of the current NAL unit to indicate the attribute of the sub-picture to the current image slice.
[0149] Furthermore, the present invention can be implemented as a code readable by a computer (including all devices having information processing functions) on a computer-readable recording medium. A computer-readable recording medium includes all types of recording devices on which data read by a computer system is stored. Examples of computer-readable recording devices include ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage devices, etc.
[0150] Although the present invention has been described with reference to the embodiments shown in the drawings, these are merely illustrative, and it will be understood by those with ordinary skill in the art that various modifications and equivalent other embodiments are possible therefrom. Accordingly, the true scope of technical protection of the present invention shall be determined by the technical idea of the appended claims.
Claims
1. A video decoding method performed by a video decoding device, The steps include obtaining a first flag indicating whether the current picture contains subpictures having different NAL (Network Abstraction Layer) unit types, The step includes obtaining a second flag indicating whether the subpicture was treated as a separate picture during the decoding process, If, based on the first flag, it is determined that the picture includes subpictures having different NAL unit types, the second flag is forced to have a first value for subpictures that include at least one P-slice or B-slice from among the subpictures. The VCL (Video Coding Layer) NAL unit of the current picture is forced to have a different NAL unit type than GDR_NUT. The first value indicates that the subpicture is treated as a separate picture during the decoding process, in a video decoding method.
2. A video decoding device, Memory to store one or more instructions, Includes at least one processor, The at least one processor performs one or more instructions, Currently, a first flag is obtained indicating whether the picture contains subpictures that have different NAL (Network Abstraction Layer) unit types. A second flag is obtained indicating whether the aforementioned subpicture is treated as a separate picture during the decoding process. If, based on the first flag, it is determined that the picture includes subpictures having different NAL unit types, the second flag is forced to have a first value for subpictures that include at least one P-slice or B-slice from among the subpictures. The VCL (Video Coding Layer) NAL unit of the current picture is forced to have a different NAL unit type than GDR_NUT. The first value indicates that the subpicture is treated as a separate picture during the decoding process, in this video decoding device.
3. A video encoding method performed by a video encoding device, Currently, the process involves determining whether the picture contains subpictures that have different NAL (Network Abstraction Layer) unit types. A step of determining whether the aforementioned subpicture is treated as a separate picture during the decoding process, The step includes generating a first flag indicating whether the current picture includes subpictures having different NAL unit types and a second flag indicating whether the subpictures are treated as separate pictures during the decoding process, If, based on the first flag, it is determined that the picture includes subpictures having different NAL unit types, the second flag is forced to have a first value for subpictures that include at least one P-slice or B-slice from among the subpictures. The VCL (Video Coding Layer) NAL unit of the current picture is forced to have a different NAL unit type than GDR_NUT. The first value indicates that the subpicture is treated as a separate picture during the decoding process, a video encoding method.
4. A method for transmitting an encoded bitstream, Currently, the process involves determining whether the picture contains subpictures that have different NAL (Network Abstraction Layer) unit types. A step of determining whether the aforementioned subpicture is treated as a separate picture during the decoding process, A step of generating a bitstream including a first flag indicating whether the current picture includes subpictures having different NAL unit types and a second flag indicating whether the subpictures are treated as separate pictures during the decoding process, The step includes transmitting the bitstream from the video encoding device to the video decoding device, If, based on the first flag, it is determined that the picture includes subpictures having different NAL unit types, the second flag is forced to have a first value for subpictures that include at least one P-slice or B-slice from among the subpictures. The VCL (Video Coding Layer) NAL unit of the current picture is forced to have a different NAL unit type than GDR_NUT. The first value is a method that indicates that the subpicture is treated as a separate picture during the decoding process.