Video decoding method and device
The video decoding and encoding methods enable efficient data reduction and synthesis of high-resolution video by allowing mixed NAL unit types within a picture, addressing the challenge of high data volume in high-quality video transmission and storage.
Patent Information
- Application Number
- JP2025079371
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-11-17
- Filing Date
- 2025-05-12
- Publication Date
- 2025-07-17
- Estimated Expiration
- 2040-12-16
AI Technical Summary
The increasing demand for high-resolution and high-quality video leads to a significant increase in the amount of encoded data, resulting in higher costs for transmission and storage, necessitating more efficient encoding and decoding methods to reduce data volume.
A video decoding method that involves obtaining NAL unit type information to determine the type of current NAL units, allowing for mixed NAL unit types within a picture, and a video encoding method that determines sub-picture types for encoding, enabling efficient bit-stream packing and synthesis of sub-pictures with different NAL unit types.
This approach allows for easier configuration and synthesis of images by allowing different NAL unit types within a picture, reducing the need to equalize NAL unit types across sub-pictures, thereby optimizing data efficiency and simplifying the composition process.
Smart Images

Figure 2025107407000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a sub-picture splitting method for composition with other sequences and a slice splitting method for bit-stream packing.
Background Art
[0002] The demand of users for high-resolution and high-quality video is increasing. Since the encoded data of high-resolution video has a larger amount of information than the encoded data of low-resolution or medium-resolution video, the cost for transmitting or storing this increases.
[0003] Research on encoding and decoding methods for effectively reducing the amount of encoded data of high-resolution video to solve such problems has been continuing.
Summary of the Invention
Problems to be Solved by the Invention
[0004] This specification presents a sub-picture splitting method for composition with other sequences and a slice splitting method for bit-stream packing.
Means for Solving the Problems
[0005] A video decoding method performed by a video decoding apparatus according to an embodiment of the present invention for solving the above-described problem includes: obtaining NAL unit type information indicating the type of a current NAL (network abstraction layer) unit from a bitstream; and when the NAL unit type information indicates that the NAL unit type of the current NAL unit is encoded data for a video slice, decoding the video slice based on whether a mixed NAL unit type is applied to the current picture. Here, the step of decoding the video slice may be performed by determining whether the NAL unit type of the current NAL unit indicates the attributes of a sub-picture for the current video slice based on whether the mixed NAL unit type is applied.
[0006] Further, a video decoding apparatus according to an embodiment of the present invention for solving the above-described problem is a video decoding apparatus including a memory and at least one processor, where the at least one processor obtains NAL unit type information indicating the type of a current NAL unit from a bitstream, and when the NAL unit type information indicates that the NAL unit type of the current NAL unit is encoded data for a video slice, may decode the video slice based on whether a mixed NAL unit type is applied to the current picture. At this time, the decoding of the video slice may be performed by determining whether the NAL unit type of the current NAL unit indicates the attributes of a sub-picture for the current video slice based on whether the mixed NAL unit type is applied.
[0007] Also, a video encoding method performed by a video encoding apparatus according to an embodiment of the present invention for solving the above-described problems may include: when a current picture is encoded based on a hybrid NAL unit type, determining a type of a sub-picture for dividing the picture; and encoding at least one current video slice constituting the sub-picture based on the type of the sub-picture to generate a current NAL unit. Here, the step of encoding the video slice may be performed by encoding such that the NAL unit type of the current NAL unit indicates an attribute of the sub-picture for the current video slice when the current picture is encoded based on the hybrid NAL unit type.
[0008] Also, a transmission method according to an embodiment of the present invention for solving the above-described problems may transmit a bitstream generated by the video encoding apparatus or video encoding method of this disclosure.
[0009] Also, a computer-readable recording medium according to an embodiment of the present invention for solving the above-described problems may store a bitstream generated by the video encoding method or video encoding apparatus of this disclosure.
Advantages of the Invention
[0010] The present invention presents a method for generating one picture by synthesizing with many other sequences. The pictures in the sequence are divided into a number of sub-pictures, and the divided sub-pictures of other pictures are synthesized to generate a new picture.
[0011] By applying the present invention, the NAL unit type values for two or more sub-pictures constituting one picture may be different from each other. This has the advantage that when synthesizing different contents, it is not necessary to make the NUTs of a number of sub-pictures constituting one image equal, so that the image can be easily configured / synthesized.
Brief Description of the Drawings
[0012]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Mode for Carrying Out the Invention
[0013] The present invention may be subject to various modifications and may have various embodiments. Therefore, specific embodiments are illustrated in the drawings and described in detail. However, this is not intended to limit the present invention to specific embodiments. The terms used in this specification are merely used to describe specific embodiments and are not intended to limit the technical idea of the present invention. Singular expressions include plural expressions unless the context clearly indicates otherwise. Terms such as "including" or "having" in this specification are intended to specify the existence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and should not be understood as precluding the possibility of the existence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0014] On the other hand, each configuration on the drawings described in the present invention is shown independently for the convenience of explaining separate characteristic functions, and does not mean that each configuration is realized by separate hardware or separate software. For example, two or more of the configurations may be combined to form one configuration, or one configuration may be divided into a plurality of configurations. Embodiments in which each configuration is integrated and / or separated are also included in the scope of the present invention as long as they do not depart from the essence of the present invention.
[0015] Hereinafter, with reference to the accompanying drawings, preferred embodiments of the present invention will be described in more detail. Hereinafter, the same reference numerals will be used for the same components on the drawings, and overlapping descriptions for the same components will be omitted.
[0016] On the other hand, the present invention relates to video / video coding. For example, the methods / embodiments disclosed in the present invention can be applied to the methods disclosed in the VVC (versatile video coding) standard, EVC (Essential Video Coding) standard, AV1 (AOMedia Video 1) standard, AVS2 (2nd generation of audio video coding standard), or the next-generation video / image coding standard (for example, H.267, H.268, etc.).
[0017] In this specification, an access unit (AU) means a unit indicating a set of a plurality of pictures belonging to different layers output from the DPB (Decoded picture buffer) at the same time. A picture generally means a unit indicating one video in a specific time period, and a slice is a unit constituting a part of a picture in coding. One picture may be composed of a plurality of slices, and if necessary, pictures and slices may be used interchangeably.
[0018] A pixel or pel may mean the smallest unit constituting one picture (or video). Also, the term "sample" may be used as a term corresponding to a pixel. A sample may generally indicate a pixel or a pixel value, or only a pixel / pixel value of a luminance component, or only a pixel / pixel value of a chroma component.
[0019] A "unit" represents the basic unit of video processing. A unit may include at least one of a specific area of a picture and information regarding the area. Optionally, the term "unit" may be used interchangeably with terms such as "block" or "area". In a general case, an M×N block may represent a set of samples or transform coefficients consisting of M columns and N rows.
[0020] FIG. 1 is a diagram schematically showing the configuration of a video encoding apparatus to which the present invention is applied.
[0021] Referring to FIG. 1, the video encoding apparatus 100 may include a picture splitting unit 105, a prediction unit 110, a residual processing unit 120, an entropy encoding unit 130, an addition unit 140, a filter unit 150, and a memory 160. The residual processing unit 120 may include a subtraction unit 121, a transformation unit 122, a quantization unit 123, a reordering unit 124, an inverse quantization unit 125, and an inverse transformation unit 126.
[0022] The picture splitting unit 105 splits the input picture into at least one processing unit.
[0023] As an example, the processing unit may be referred to as a coding unit (CU). In this case, the coding unit may be recursively divided from a Coding Tree Unit by a QTBT (Quad-tree binary-tree) structure. For example, one coding tree unit may be divided into a plurality of nodes at a deeper depth based on a quad-tree structure and / or a binary-tree structure. In this case, for example, the quad-tree structure may be applied first, and the binary-tree structure may be applied later. Alternatively, the binary-tree structure may be applied first. Decoding may be performed on nodes that are no longer divisible, and thus, a coding unit may be determined for nodes that are no longer divisible. Since the coding tree unit is a unit for dividing the coding unit, the coding tree unit may be named the coding unit. In this case, since the coding unit is determined by dividing the coding tree unit, the coding tree unit may be named the largest coding unit (LCU).
[0024] Thus, based on the final coding unit that is no longer divisible, the coding procedure according to the present invention may be performed. In this case, based on the coding efficiency and the like according to the video characteristics, the coding tree unit may be immediately used as the final coding unit, or, if necessary, the coding unit may be recursively divided into coding units at a deeper depth so that the coding unit of an optimal size is used as the final coding unit. Here, the coding procedure may include procedures such as prediction, conversion, and restoration, which will be described later.
[0025] As another example, the processing unit may include a coding unit (CU), a prediction unit (PU), or a transform unit (TU). The coding unit may be split from a coding tree unit into coding units of a lower depth according to a quad tree structure. In this case, based on coding efficiency according to video characteristics or the like, the coding tree unit may be immediately used as the final coding unit, or, if necessary, the coding unit may be recursively split into coding units of a lower depth, and a coding unit of an optimal size may be used as the final coding unit. When a minimum coding unit (min CU) is set, the coding unit is not split into coding units smaller than the minimum coding unit. Here, the final coding unit means a coding unit that serves as a basis for partitioning or splitting into a prediction unit or a transform unit. The prediction unit is a unit partitioned from the coding unit and may be a unit for sample prediction. At this time, the prediction unit may be divided into subblocks. The transform unit may be split from the coding unit according to a quad tree structure and may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients. Hereinafter, the coding unit may be referred to as a coding block (CB), the prediction unit may be referred to as a prediction block (PB), and the transform unit may be referred to as a transform block (TB). The prediction block or prediction unit means a specific area in block form within a picture and may include an array of prediction samples. Also, the transform block or transform unit means a specific area in block form within a picture and may include an array of transform coefficients or residual samples.
[0026] The prediction unit 110 makes a prediction for a block to be processed (hereinafter referred to as the current block), and generates a predicted block including a prediction sample for the current block. The unit of prediction performed by the prediction unit 110 may be a coding block, a transform block, or a prediction block.
[0027] The prediction unit 110 determines whether intra prediction or inter prediction is applied to the current block. As an example, the prediction unit 110 determines whether intra prediction or inter prediction is applied in the CU unit.
[0028] In the case of intra prediction, the prediction unit 110 can derive a prediction sample for the current block based on reference samples outside the current block within the picture (hereinafter referred to as the current picture) to which the current block belongs. At this time, the prediction unit 110 may (i) derive a prediction sample based on the average or interpolation of neighboring reference samples of the current block, or (ii) derive the prediction sample based on reference samples existing in a specific (prediction) direction with respect to the prediction sample among the neighboring reference samples of the current block. In the case of (i), it is called a non-directional mode or a non-angle mode, and in the case of (ii), it is called a directional mode or an angular mode. The prediction mode in intra prediction may have, for example, 33 directional prediction modes and at least 2 non-directional modes. The non-directional modes may include a DC prediction mode and a Planar mode. The prediction unit 110 may determine the prediction mode applied to the current block using the prediction mode applied to the adjacent block.
[0029] In the case of inter prediction, the prediction unit 110 can derive a prediction sample for the current block based on samples specified by motion vectors on the reference picture. The prediction unit 110 can apply any one of a skip mode, a merge mode, and an MVP (motion vector prediction) mode to derive a prediction sample for the current block. In the case of the skip mode and the merge mode, the prediction unit 110 may use the motion information of an adjacent block as the motion information of the current block. In the case of the skip mode, unlike the merge mode, the difference (residual) between the prediction sample and the original sample is not transmitted. In the case of the MVP mode, the motion vector of an adjacent block can be used as a motion vector predictor, and the motion vector of the current block can be derived by using it as the motion vector predictor of the current block.
[0030] In the case of inter prediction, the adjacent blocks may include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in a reference picture. The reference picture including the temporal neighboring blocks may be referred to as a collocated picture (colPic). The motion information may include a motion vector and a reference picture index. Information such as prediction mode information and motion information may be (entropy) encoded and output in the form of a bitstream.
[0031] When motion information of temporally adjacent blocks is used in skip mode and merge mode, the top picture on the reference picture list may be used as a reference picture. The reference pictures included in the reference picture list may be sorted based on the difference in POC (Picture Order Count) between the current picture and the reference picture. The POC corresponds to the display order of pictures and can be distinguished from the coding order.
[0032] The subtraction unit 121 generates a residual sample that is the difference between the original sample and the predicted sample. When the skip mode is applied, it may not be necessary to generate the residual sample as described above.
[0033] The conversion unit 122 converts the residual samples in units of conversion blocks to generate transform coefficients. The conversion unit 122 can perform conversion according to the size of the conversion block and the prediction mode applied to the coding block or prediction block that spatially overlaps with the conversion block. For example, if intra prediction is applied to the coding block or the prediction block that overlaps with the conversion block, and the conversion block is a 4×4 residual array, the residual samples are converted using a DST (Discrete Sine Transform) conversion kernel, and in other cases, the residual samples are converted using a DCT (Discrete Cosine Transform) conversion kernel.
[0034] The quantization unit 123 quantizes the transform coefficients to generate quantized transform coefficients.
[0035] The reordering unit 124 reorders the quantized transform coefficients. The reordering unit 124 can reorder the quantized transform coefficients in block form into a one-dimensional vector form by a coefficient scanning method. Here, although the reordering unit 124 has been described separately, it may also be a part of the quantization unit 123.
[0036] The entropy encoding unit 130 performs entropy encoding on the quantized transform coefficients. The entropy encoding may include encoding methods such as, for example, exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. The entropy encoding unit 130 may encode, together or separately, information necessary for video restoration (e.g., values of syntax elements, etc.) in addition to the quantized transform coefficients. The entropy-encoded information may be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer) units.
[0037] The inverse quantization unit 125 inverse-quantizes the value (quantized transform coefficient) quantized by the quantization unit 123, and the inverse transform unit 126 inverse-transforms the value inverse-quantized by the inverse quantization unit 125 to generate residual samples.
[0038] The addition unit 140 combines the residual samples and the prediction samples to restore the picture. The residual samples and the prediction samples are added in block units to generate a restored block. Here, although the addition unit 140 has been described with a separate configuration, it may be a part of the prediction unit 110. On the other hand, the addition unit 140 is also called a restoration unit or a restored block generation unit.
[0039] For the reconstructed picture, the filter unit 150 can apply a deblocking filter and / or a sample adaptive offset. Through deblocking filtering and / or sample adaptive offset, artifacts at block boundaries in the reconstructed picture and distortions in the quantization process can be corrected. The sample adaptive offset may be applied on a sample-by-sample basis or after the deblocking filtering process is completed. The filter unit 150 can also apply an ALF (Adaptive Loop Filter) to the reconstructed picture. The ALF may be applied to the reconstructed picture after the deblocking filter and / or the sample adaptive offset have been applied.
[0040] The memory 160 stores the reconstructed picture (decoded picture) or information necessary for encoding / decoding. Here, the reconstructed picture may be the one on which the filtering procedure by the filter unit 150 has been completed. The stored reconstructed picture may be utilized as a reference picture for (inter) prediction of other pictures. For example, the memory 160 can store the (reference) picture used for inter prediction. At this time, the picture used for inter prediction may be specified by a reference picture set or a reference picture list.
[0041] Figure 2 shows an example of a video encoding method performed by a video encoding apparatus. Referring to Figure 2, the video encoding method may include a block partitioning, intra / inter prediction, transform, quantization, and entropy encoding process. For example, a current picture may be divided into a plurality of blocks, and a predicted block of the current block may be generated by intra / inter prediction. A residual block of the current block may be generated by subtracting the input block of the current block from the predicted block. Thereafter, a coefficient block, i.e., a transform coefficient of the current block, may be generated by a transform on the residual block. The transform coefficient may be quantized and entropy encoded and stored in a bitstream.
[0042] Figure 3 is a diagram schematically explaining the configuration of a video decoding apparatus to which the present invention is applied.
[0043] Referring to Figure 3, the video decoding apparatus 300 may include an entropy decoding unit 310, a residual processing unit 320, a prediction unit 330, an addition unit 340, a filter unit 350, and a memory 360. Here, the residual processing unit 320 may include a reordering unit 321, an inverse quantization unit 322, and an inverse transform unit 323.
[0044] When a bitstream including video information is input, the video decoding apparatus 300 can restore the video corresponding to the process in which the video information is processed by the video encoding apparatus.
[0045] For example, the video decoding device 300 can perform video decoding using the processing unit applied in the video encoding device. Therefore, the processing unit block for video decoding may be, as an example, a coding unit, or may be, as another example, a coding unit, a prediction unit, or a conversion unit. The coding unit may be divided from the coding tree unit by a quad tree structure and / or a binary tree structure.
[0046] The prediction unit and the conversion unit may be further used as appropriate. In this case, the prediction block is a block derived or partitioned from the coding unit and may be a unit for sample prediction. At this time, the prediction unit may be divided into sub-blocks. The conversion unit may be divided from the coding unit by a quad tree structure and may be a unit for deriving conversion coefficients or a unit for deriving a residual signal from the conversion coefficients.
[0047] The entropy decoding unit 310 parses the bitstream and outputs information necessary for video restoration or picture restoration. For example, the entropy decoding unit 310 can decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the value of the syntax element necessary for video restoration and the quantized value of the conversion coefficient for the residual.
[0048] More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element in a bitstream, determines a context model using the syntax element information to be decoded, the decoding information of adjacent and decoding target blocks, or the information of symbols / bins decoded in the previous step, predicts the occurrence probability of the bin based on the determined context model, and performs arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element. At this time, after determining the context model, the CABAC entropy decoding method can update the context model using the information of the decoded symbols / bins for the context model of the next symbol / bin.
[0049] Among the information decoded by the entropy decoding unit 310, the information related to prediction is provided to the prediction unit 330, and the residual value obtained by performing entropy decoding in the entropy decoding unit 310, that is, the quantized transform coefficient, is input to the reordering unit 421.
[0050] The reordering unit 421 reorders the quantized transform coefficients in a two-dimensional block form. The reordering unit 421 can perform reordering corresponding to the coefficient scanning performed in the encoding device. Here, although the reordering unit 321 is described separately, it may be a part of the inverse quantization unit 322.
[0051] The inverse quantization unit 322 can perform inverse quantization on the quantized transform coefficients based on (inverse) quantization parameters and output the transform coefficients. At this time, the information for deriving the quantization parameters may be signaled from the encoding device.
[0052] The inverse transform unit 323 can perform inverse transform on the transform coefficients to derive residual samples.
[0053] The prediction unit 330 can perform a prediction on the current block and generate a predicted block including a prediction sample for the current block. The unit of prediction performed by the prediction unit 330 may be a coding block, a transform block, or a prediction block.
[0054] Based on the information regarding the prediction, the prediction unit 330 determines whether to apply intra prediction or inter prediction. At this time, the unit for determining which of intra prediction and inter prediction to apply is different from the unit for generating a prediction sample. Also, in inter prediction and intra prediction, the unit for generating a prediction sample is different. For example, which of inter prediction and intra prediction to apply can be determined in CU units. Also, for example, in inter prediction, the prediction mode may be determined in PU units and a prediction sample may be generated, and in intra prediction, the prediction mode may be determined in PU units and a prediction sample may be generated in TU units.
[0055] In the case of intra prediction, the prediction unit 330 can derive a prediction sample for the current block based on adjacent reference samples within the current picture. The prediction unit 330 can derive a prediction sample for the current block by applying a directional mode or a non - directional mode based on the adjacent reference samples of the current block. At this time, the intra - prediction mode of an adjacent block may be used to determine the prediction mode to be applied to the current block.
[0056] In the case of inter prediction, the prediction unit 330 can derive a prediction sample for the current block based on samples specified on the reference picture by the motion vector on the reference picture. The prediction unit 330 can apply any one of the skip mode, merge mode, and MVP mode to derive a prediction sample for the current block. At this time, motion information necessary for inter prediction of the current block provided from the video encoding device, for example, information regarding a motion vector, a reference picture index, etc., may be obtained or derived based on the information regarding the prediction.
[0057] In the case of the skip mode and the merge mode, the motion information of adjacent blocks may be used as the motion information of the current block. At this time, the adjacent blocks may include spatially adjacent blocks and temporally adjacent blocks.
[0058] The prediction unit 330 may configure a merge candidate list with the motion information of available adjacent blocks, and use the information indicated by the merge index on the merge candidate list as the motion vector of the current block. The merge index may be signaled from the encoding device. The motion information may include a motion vector and a reference picture. In the skip mode and the merge mode, when the motion information of temporally adjacent blocks is used, the top picture on the reference picture list may be used as the reference picture.
[0059] In the case of the skip mode, different from the merge mode, the difference (residual) between the prediction sample and the original sample is not transmitted.
[0060] In the case of the MVP mode, the motion vector of the current block may be derived by using the motion vector of an adjacent block as a motion vector predictor. At this time, the adjacent blocks may include spatially adjacent blocks and temporally adjacent blocks.
[0061] As an example, when the merge mode is applied, a merge candidate list may be generated using the motion vectors of the restored spatially adjacent blocks and / or the motion vectors corresponding to the Col blocks which are temporally adjacent blocks. In the merge mode, the motion vector of the candidate block selected from the merge candidate list is used as the motion vector of the current block. The information regarding the prediction may include a merge index indicating a candidate block having an optimal motion vector selected from among the candidate blocks included in the merge candidate list. At this time, the prediction unit 330 may derive the motion vector of the current block using the merge index.
[0062] As another example, when the MVP (Motion Vector Prediction) mode is applied, a motion vector predictor candidate list may be generated using the motion vectors of the restored spatially adjacent blocks and / or the motion vectors corresponding to the Col blocks which are temporally adjacent blocks. That is, the motion vectors of the restored spatially adjacent blocks and / or the motion vectors corresponding to the Col blocks which are temporally adjacent blocks may be used as motion vector candidates. The information regarding the prediction may include a predicted motion vector index indicating the optimal motion vector selected from among the motion vector candidates included in the list. At this time, the prediction unit 330 can select the predicted motion vector of the current block from among the motion vector candidates included in the motion vector candidate list using the motion vector index. The prediction unit of the encoding device obtains the motion vector difference (MVD) between the motion vector of the current block and the motion vector predictor, encodes this, and outputs it in the form of a bit stream. That is, the MVD is obtained as the value obtained by subtracting the motion vector predictor from the motion vector of the current block. At this time, the prediction unit 330 can obtain the motion vector difference included in the information regarding the prediction and derive the motion vector of the current block by adding the motion vector difference and the motion vector predictor. Also, the prediction unit can obtain or derive a reference picture index indicating a reference picture, etc. from the information regarding the prediction.
[0063] The addition unit 340 adds the residual samples and the predicted samples to restore the current block or the current picture. The addition unit 440 may add the residual samples and the predicted samples in block units to restore the current picture. When the skip mode is applied, since no residual is transmitted, the predicted samples become the restored samples. Here, the addition unit 340 has been described as a separate configuration, but it may also be a part of the prediction unit 330. On the other hand, the addition unit 340 is also called a restoration unit or a restored block generation unit.
[0064] The filter unit 350 may apply a deblocking filter filtering sample adaptive offset and / or ALF or the like to the restored picture. At this time, the sample adaptive offset may be applied in units of samples, or may be applied after deblocking filter filtering. ALF may be applied after deblocking filter filtering and / or the sample adaptive offset.
[0065] The memory 360 stores the restored picture (decoded picture) or information necessary for decoding. Here, the restored picture may be a restored picture for which the filtering procedure has been completed by the filter unit 350. For example, the memory 360 may store a picture used for inter prediction. At this time, the picture used for inter prediction may be specified by a reference picture set or a reference picture list. The restored picture may be used as a reference picture for other pictures. Also, the memory 360 may output the restored pictures in output order.
[0066] FIG. 4 is a diagram showing an example of a video decoding method performed by a decoding apparatus. Referring to FIG. 4, the video decoding method may include an entropy decoding, an inverse quantization, an inverse transform, and an intra / inter prediction process. For example, in the decoding apparatus, an inverse process of the encoding method may be performed. Specifically, the quantization conversion coefficients for the bit stream may be obtained by entropy decoding, and the coefficient block of the current block, that is, the conversion coefficients, may be obtained by an inverse quantization process for the quantization conversion coefficients. The residual block of the current block may be derived by inverse conversion of the conversion coefficients, and the restored block of the current block may be derived by adding the prediction block of the current block derived by intra / inter prediction and the residual block.
[0067] On the other hand, the operator in the embodiments described below may be defined as follows in the following table.
[0068]
Table 1
[0069] Referring to Table 1, Floor(x) represents the largest integer value less than or equal to x, Log2(u) represents the logarithm value of u with base 2, and Ceil(x) represents the smallest integer value greater than or equal to x. For example, in the case of Floor(5.93), since the largest integer value less than or equal to 5.93 is 5, it represents 5.
[0070] Also, referring to Table 1, x>>y represents the operator that right-shifts x by y bits, and x<<y represents the operator that left-shifts x by y bits.
[0071] <Introduction> The HEVC standard proposes two types of screen segmentation methods.
[0072] 1) Slice: It provides the function of dividing a single image in raster scan order into coding tree unit (CTU) units for encoding / decoding, and slice header information exists.
[0073] 2) Tile: It provides the function of dividing a single image into multiple columns and rows in CTU units for encoding / decoding. The partitioning method can be either equal partitioning or individual partitioning. There is no separate header for tiles.
[0074] A slice serves as a bit-stream packing unit. That is, one slice may be generated from one NAL (network abstraction layer) bit-stream. As shown in Figure 5, the NAL packet for a slice is composed in the order of NAL header, slice header, and slice data. At this time, the NAL header information includes the NAL unit type (NUT).
[0075] In the HEVC standard proposed according to one embodiment, the NUT for a slice is as shown in Table 2. In Table 2, the NUTs for inter slices where inter prediction is performed are from 0 to 9, and the NUTs for intra slices where intra prediction is performed are from 16 to 21. Here, an inter slice means that it is encoded by an inter-picture prediction method, and an intra slice means that it is encoded by an intra-picture prediction method. One slice is defined to have one NUT, and multiple slices within one picture may be set to all have the same NUT value. For example, if one picture is divided into 4 slices and encoded by the intra prediction method, the NUT values for the 4 slices within the picture may all be set to the same value of 19:IDR_W_RADL.
[0076]
Table 2
[0077] In the above Table 2, the abbreviations are defined as follows. -TSA (Temporal sub-layer Switching Access) -STSA (Step-wise Temporal sub-layer Switching Access) -RADL (Random Access Decodable Leading) -RASL (Random Access Skipped Leading) -BLA (Broken Link Access) -IDR (Instantaneous Decoding Refresh) -CRA (Clean Random Access) -LP (Leading Picture) -_N (No reference) -_R (Reference) -_W_LP / RADL (With LP / RADL) -_N_LP (No LP, without LP)
[0078] BLA, IDR, and CRA, which are NUTs for intra slices, are referred to as IRAP (Intra Random Access Point). IRAP is a middle position in the bitstream and means a picture that allows random access. That is, it refers to a picture that enables a sudden change in the playback position during video playback. Intra slices exist only in the I slice type.
[0079] Inter slices are divided into P slices or B slices by unidirectional prediction (P: predictive) or bidirectional prediction (B: bi-predictive). The prediction and encoding processes are performed in units of GOP (group of picture), but the HEVC standard uses a hierarchical GOP structure to perform the encoding / decoding process including prediction. Figure 6 shows an example of a hierarchical GOP structure, and each picture is divided into I, P, or B pictures (slices) according to the prediction method.
[0080] Due to the B slice that performs two-way prediction and / or the hierarchical GOP structure characteristics, the decoding order and display order of pictures in the sequence are different (see Figure 7). In Figure 7, IRAP means intra-slice, B and P mean inter-slice, and it is confirmed that the playback order and the restoration order are completely different.
[0081] Among the inter-slices, a picture whose restoration order is later than IRAP and whose playback order is earlier than IRAP is called an LP (leading picture) (see Figure 8). The LP is divided into RADL and RASL according to the situation. When random access occurs, an LP that can be decoded is defined as RADL, and an LP that cannot be decoded during random access and whose restoration process of the picture must be skipped is defined as RASL. In Figure 8, pictures of the same color are defined as one GOP.
[0082] The classification of RADL and RASL is determined by the position of the reference picture during inter-picture prediction (see Figure 9). That is, RASL means an inter-picture that uses a restored picture in another GOP as a reference picture in addition to the current GOP, or uses a picture restored using a restored picture in another GOP as a reference picture as a reference picture. In this case, since a restored picture in another GOP is used as a reference picture (directly or indirectly), it is called an open GOP. RASL and RADL are set in the NUT information for the inter-slice.
[0083] The NUT for an intra-slice is divided into other intra-slice NUTs by the NUTs of the preceding and / or succeeding inter-slices in the playback order and / or restoration order of the intra-slice. When examining the IDR NUT, the IDR is divided into IDR_W_RADL having RADL and IDR_N_LP not having LP. That is, the IDR is of a type that does not have LP or a type that has only RADL among LPs, and the IDR cannot have RASL. On the other hand, the CRA is of a type that has all of RADL and / or RASL among LPs. That is, the CRA is a type that can support an open GOP.
[0084] Generally, an intra-slice only performs in-picture prediction, so reference picture information for the intra-slice is not required. Here, the reference picture is used during inter-picture prediction. However, due to the feature of the CRA NUT slice to support an open GOP structure, the CRA slice inserts reference picture information into the NAL bitstream of the CRA even though it is an intra-slice. The reference picture information is not for use in the CRA slice itself, but is information for the reference picture that is scheduled to be used in the inter-slices after the CRA (in the restoration order). This is because the reference picture is not removed in the DPB (decoded picture buffer). For example, when the NUT of the intra-slice is an IDR, the DPB is reset. That is, all the restored pictures existing in the DPB at that time are removed. Figure 10 is a diagram showing the syntax for the slice segment header. As shown in Figure 10, if the NUT of the slice is not an IDR, the reference picture information can be described in the bitstream. That is, if the NUT of the slice is a CRA, the reference picture information can be described.
[0085] The present invention presents a sub-picture segmentation method for composition with other sequences and a slice segmentation method for bitstream packing.
[0086] In the present invention, a slice means an area for encoding / decoding and is a data packing unit that generates one NAL bitstream. For example, one picture is divided into a number of slices, and each slice is generated as one NAL packet through the encoding process.
[0087] In the present invention, a sub-picture is an area division for composition with other contents. In FIG. 11, an example of composition with other contents is shown. There are three contents of white, gray, and black, and one image (AU: access unit) of each content is divided into four slice areas for packet generation. As shown in the image on the right side of FIG. 11, the upper left part may be composed of white content, the lower left part of gray content, and the right side of black content to generate a new image. Here, the white area and the gray area are each composed of one slice as one sub-picture, and the black area is composed of two slices as one sub-picture. That is, one sub-picture may include at least one slice. To create a new image (to compose contents), BEAMer (Bit-stream Extractor And Merger) extracts areas from different contents in units of sub-pictures and composes them. The composed image in FIG. 11 may be divided into four slices and composed of three sub-pictures.
[0088] One sub-picture means an area having the same sub-picture ID (sub picture ID) and / or sub-picture index value. In other words, at least one slice having the same sub-picture ID and / or sub-picture index value can be said to be one sub-picture area. Here, among the slice header information, the sub-picture ID and / or sub-picture index value are included. The sub-picture index value may be set in raster scan order. FIG. 12 shows an example in which one picture is composed of six (quadrilateral) slices and four (color-separated) sub-picture areas. Here, A, B, C, D show an example for the sub-picture ID, and 0, 1 show the slice address within the sub-picture. That is, the slice address value is the slice index value in the raster scan order within the sub-picture. For example, B-0 means the 0th slice within the B sub-picture, and B-1 means the 1st slice within the B sub-picture.
[0089] In the present invention, the NUT values for two or more sub-pictures constituting one image may be different. For example, in FIG. 12, the white sub-picture (slice) within one image may be an intra-slice, and the gray sub-picture (slice) and the black sub-picture (slice) may be inter-slices.
[0090] This has the advantage that when synthesizing different contents, it is not necessary to equalize the NUTs of multiple sub-pictures that make up one image, so the image can be easily composed / synthesized. This function can be referred to as a mixed NAL Unit Type in a picture, and can be simply named mixed NUT for short. By setting the mixed_nalu_type_in_pic_flag, the enabling / disabling of this function can be set. This flag can be defined at one or more positions among SPS (sequence parameter set), PPS (picture parameter set), PH (picture header), and SH (slice header). For example, when the flag is defined in PPS, the flag can be named pps_mixed_nalu_types_in_pic_flag.
[0091] When the flag value is disabled (e.g., mixed_nalu_type_in_pic_flag == 0), the NUTs for all sub-pictures and / or slices within the picture may have the same value. For example, the NUTs for all VCL (video coding layer) NAL units for one picture may be set to have the same value. Also, a picture or picture unit (PU, picture unit) may be referred to as having the same NUT as the encoded slice NAL unit for it. Here, VCL means the NAL type for a slice that includes slice data values.
[0092] On the one hand, when the flag value is enabled (e.g., mixed_nalu_type_in_pic_flag == 1), the picture may be composed of two or more sub - pictures. Also, the picture may have other NUT values. Further, when the flag value is enabled, the VCL NAL unit of the picture may be restricted so as not to have a GDR_NUT - type NUT. Also, when the NUT (e.g., the first NUT) of any one VCL NAL unit (e.g., the first NAL unit) of the picture is any one of IDR_W_RADL, IDR_N_LP, or CRA_NUT, the NUT (e.g., the second NUT) of the other VCL NAL units (e.g., the second NAL unit) of the picture may be restricted to be set to any one of IDR_W_RADL, IDR_N_LP, CRA_NUT, or TRAIL_NUT. For example, the second NUT may be restricted to be set to one of the values of the first NUT or TRAIL_NUT.
[0093] Referring to FIGS. 12 and 13, an example in which the VCL NAL unit of the picture has at least two different NUT values will be described. In one embodiment, two or more sub - pictures may have two or more different NUT values. At this time, the NUT values for all slices included in one sub - picture may be equally restricted. For example, as shown in FIG. 13, the NUT values for two slices in the B sub - picture of FIG. 12 may be equally set to CRA, and the NUT values for two slices in the C sub - picture may also be equally set to TRAIL. The A, B, C, and D sub - pictures may be set to have at least two different NUT values. Thereby, as shown in FIG. 13, the NUT values for slices in the A, C, and D sub - pictures may be TRAIL, and may be set to have NUT values different from CRA, which is the NUT of the B sub - picture.
[0094] In the present invention, the NUTs for intra-slice and inter-slice are as shown in Table 3. Similar to the embodiments in Table 3, the syntax and functions for RADL, RASL, IDR, CRA, etc. may be set to be the same as those in the HEVC standard (Table 1). In the case of Table 3, a hybrid NUT type is added. In Table 3, the disable value (e.g., 0) of mixed_nalu_type_in_pic_flag indicates the NUT for a slice within a picture (similar to HEVC), and the enable value (e.g., 1) of mixed_nalu_type_in_pic_flag indicates the NUT for a slice within a sub-picture. For example, when the value of mixed_nalu_type_in_pic_flag is 0 and the NUT of the VCL NAL unit is TRAIL_NUT, the NUT of the current picture is identified as TRAIL_NUT, and the NUTs of other sub-pictures belonging to the current picture may also be induced to be those with TRAIL_NUT. Also, when the value of mixed_nalu_type_in_pic_flag is 1 and the NUT of the VCL NAL unit is TRAIL_NUT, the NUT of the current sub-picture is identified as TRAIL_NUT, and it is predicted that at least one of the NUTs of other sub-pictures belonging to the current picture is not TRAIL_NUT.
[0095] [Table 3]
[0096] As described above, when the value of mixed_nalu_type_in_pic_flag indicates enabled (e.g., 1), if any one VCL NAL unit (e.g., the first NAL unit) belonging to a picture has a value of any one of IDR_W_RADL, IDR_N_LP, or CRA_NUT as the NUT (e.g., the first NUT), then at least one VCL NAL unit (e.g., the second NAL unit) among the other VCL NAL units of the said picture may have a value of any one NUT value other than the first NUT among IDR_W_RADL, IDR_N_LP, CRA_NUT, or TRAIL_NUT as the NUT (e.g., the second NUT).
[0097] Thus, if the VCL NAL unit (e.g., the first NAL unit) for the first sub-picture belonging to a picture has a value of any one of IDR_W_RADL, IDR_N_LP, or CRA_NUT as the NUT (e.g., the first NUT), then the VCL NAL unit (e.g., the second NAL unit) for the second sub-picture of the said picture may have a value of any one NUT value other than the first NUT among IDR_W_RADL, IDR_N_LP, CRA_NUT, or TRAIL_NUT as the NUT (e.g., the second NUT).
[0098] For example, when the value of mixed_nalu_type_in_pic_flag indicates activation (e.g., 1), the NUT values of the VCL NAL units for two or more sub-pictures may be configured as follows. The following description is only illustrative and not limited thereto.
[0099] Combination 1) IRAP + non - IRAP (inter) Combination 2) non_IRAP (inter) + non - IRAP (inter) Combination 3) IRAP + IRAP = IDR + CRA (restricted by the embodiment)
[0100] Combination 1) is an embodiment in which at least one sub-picture in the picture has an IRAP (IDR or CRA) NUT value, and at least one other sub-picture has a non-IRAP (inter-slice) NUT value. Here, as the inter-slice NUT value, a value excluding LP (RASL and RADL) may be allowed. For example, as the inter-slice NUT value, LP (RASL or RADL) may not be allowed. In this way, in the bitstream related to the IDR or CRA sub-picture, the RASL and RADL sub-pictures may be restricted from being encoded.
[0101] In another embodiment, as the inter-slice NUT value, only the TRAIL value may be allowed. Alternatively, in another embodiment, as the inter-slice NUT value, all inter-slice VCL NUTs may be allowed.
[0102] Combination 2) is an embodiment in which at least one sub-picture in the picture has a non-IRAP (inter-slice) NUT value, and at least one other sub-picture has another non-IRAP (inter-slice) NUT value. For example, at least one sub-picture may have a RASL NUT value, and at least one other sub-picture may have a RADL NUT value. In the case of the embodiment according to Combination 2), the following restrictions may be applied according to the embodiment.
[0103] - In one embodiment, LP (RASL and RADL) and non-LP (TRAIL) are not used together. For example, while the NUT of at least one sub-picture is RASL (or RADL), the NUT of at least one other sub-picture must not be TRAIL. When the NUT of at least one sub-picture is RASL (or RADL), RASL or RADL must not be used as the NUT of at least one other sub-picture. For example, the leading sub-picture of the IRAP sub-picture may be forced by the RADL or RASL sub-picture.
[0104] - In another embodiment, LP (RASL and RADL) and non-LP (TRAIL) may be used together. For example, at least one sub-picture may be RASL (or RADL), while at least one other sub-picture may be TRAIL.
[0105] - In another embodiment, except for the case of condition 2), all sub-pictures may have the same inter-slice NUT value. For example, all sub-pictures within a picture may have a TRAIL NUT value. As another example, all sub-pictures within a picture may have an RASL (or RADL) NUT value.
[0106] Combination 3) shows an embodiment in which all sub-pictures or slices within a picture are composed of IRAP. For example, if the NUT value for the slice in the first sub-picture is IDR_W_RADL, IDR_N_LP, or CRA_NUT, the NUT value for the slice in the second sub-picture may be composed of a value that is not the NUT of the first sub-picture among IDR_W_RADL, IDR_N_LP, and CRA_NUT. For example, while the NUT value for the slice in at least one sub-picture is IDR, the NUT value for the slice in at least one other sub-picture may be composed of CRA.
[0107] On the one hand, according to an embodiment, the application of an embodiment such as combination 3) may be restricted. In one embodiment, the pictures belonging to the IRAP or GDR access unit may be restricted to all have the same NUT. That is, if the current access unit is an IRAP access unit composed only of IRAP pictures, or if the current access unit is a GDR access unit composed only of GDR pictures, the pictures belonging to it may be restricted to all have the same NUT. For example, it may be restricted such that the NUT value for the slices within at least one sub-picture is IDR, while the NUT value for the slices within at least one other sub-picture is not composed of CRA. In this way, when combination 3) is restricted and combinations 1) and 2) described above are applied, at least one sub-picture within the picture may be restricted to have a NUT value for non-IRAP (inter-slice). For example, in the encoding and decoding processes, it may be restricted such that all sub-pictures within the picture do not have a NUT value for IDR. Alternatively, it may be restricted such that some sub-pictures within the picture have a NUT value for IDR and other sub-pictures do not have a CRA NUT value.
[0108] Hereinafter, the related syntax and semantics for signaling the encoding information when a mixed NAL unit type within a picture is applied will be described. Also, the decoding process using this will be described. As described above, when mixed_nalu_type_in_pic_flag = 1, the picture described by the NUT may mean a sub-picture (see Table 3).
[0109] On the one hand, as described above, when the value of mixed_nalu_type_in_pic_flag indicates that the hybrid NUT is applicable, one picture may be divided into at least two sub-pictures. Thereby, the information of the sub-pictures for the picture may be signaled by the bitstream. In this regard, mixed_nalu_type_in_pic_flag can indicate whether the current picture is divided. For example, when the value of mixed_nalu_type_in_pic_flag indicates that the hybrid NUT is applicable, it can indicate that the current picture is divided.
[0110] Hereinafter, it will be described with reference to the syntax of FIG. 14. FIG. 14 is a diagram showing an embodiment of the syntax of a picture parameter set (PPS). For example, a flag (e.g., pps_no_pic_partition_flag) indicating whether the current picture is not divided may be signaled by the picture parameter set (PPS) by the bitstream. A value (e.g., 1) indicating the enable of pps_no_pic_partition_flag can indicate that picture division is not applied to the picture referring to the current PPS. A value (e.g., 0) indicating the disable of pps_no_pic_partition_flag can indicate that picture division using slices or tiles is applied to the picture referring to the current PPS. In such an embodiment, when the value of mixed_nalu_type_in_pic_flag indicates that the hybrid NUT is applicable, the value of pps_no_pic_partition_flag may be forced to be a value indicating disable (e.g., 0).
[0111] If the pps_no_pic_partition_flag indicates that the current picture is partitioned, the number information of sub-pictures (e.g., pps_num_subpics_minus1) may be obtained from the bitstream. pps_num_subpics_minus1 can indicate a value obtained by subtracting 1 from the number of sub-pictures included in the current picture. If the pps_no_pic_partition_flag indicates that the current picture is not partitioned, the value of pps_num_subpics_minus1 may not be obtained from the bitstream and may be derived to be 0. Based on the number information of sub-pictures determined in this way, only the number of sub-pictures included in one picture may be signaled for the encoding information for each sub-picture. For example, a sub-picture identifier (e.g., pps_subpic_id) for identifying each sub-picture and / or a flag value (subpic_treated_as_pic_flag[i]) indicating whether the encoding / decoding process of each sub-picture is independent may be specified and signaled.
[0112] The hybrid NUT may be applied when one picture is composed of two or more sub - pictures. At this time, the value of the flag (subpic_treated_as_pic_flag[i]) indicating whether the encoding / decoding process of each sub - picture can be independent may be specified and signaled only by the number (i) of sub - pictures included in one picture. That one sub - picture is independently decoded means that the sub - picture is decoded by treating it as a separate picture. That is, when the flag value is on (e.g., subpic_treated_as_pic_flag = 1), the sub - picture may be independently decoded from other sub - pictures in all other decoding processes except the in - loop filter process. Conversely, when the flag value is off (e.g., subpic_treated_as_pic_flag = 0), the sub - picture may refer to other sub - pictures within the picture in the inter - prediction process. Here, for the in - loop filter process, a separate flag can be set to control whether it can be independent or refer. The flag (subpic_treated_as_pic_flag) may be defined at one or more positions among the SPS, PPS, and PH. For example, when the flag is defined in the SPS, the flag may be named sps_subpic_treated_as_pic_flag.
[0113] Also, in the present invention, when other NUTs exist within one picture (e.g., mixed_nalu_type_in_pic_flag = 1), due to the characteristic that different types of NUTs must be used between sub-pictures within one picture, each sub-picture within the said picture must be encoded / decoded independently. For example, in the case of a picture where mixed_nalu_type_in_pic_flag = 1, if the picture contains one or more inter (P or B) slices, the subpic_treated_as_pic_flag values of all sub-pictures within the said picture may be forced to be set to 1 or induced to the 1 value. Alternatively, when mixed_nalu_type_in_pic_flag = 1, subpic_treated_as_pic_flag may be forced not to have a 0 value. For example, in the case of a picture where mixed_nalu_type_in_pic_flag = 1, if the picture contains one or more inter slices, regardless of the parsed value, the subpic_treated_as_pic_flag value may be reset to 1 for all sub-pictures of the said picture. Conversely, in the case of a picture where mixed_nalu_type_in_pic_flag = 1 but subpic_treated_as_pic_flag = 0, the picture must not contain inter slices. That is, in the case of a picture where mixed_nalu_type_in_pic_flag = 1 but subpic_treated_as_pic_flag = 0, the slice type within the said picture must be intra.
[0114] In another embodiment, when mixed_nalu_type_in_pic_flag = 1, if the NUT of the current picture is RASL, the subpic_treated_as_pic_flag for the current picture may be forced to be set to 1. As another example, when mixed_nalu_type_in_pic_flag = 1, if the NUT of the current picture is RADL while the NUT of the reference picture is RASL, the subpic_treated_as_pic_flag for the current picture may be forced to be 1.
[0115] The hybrid NUT function may limit that all subpictures (or slices) within a picture are composed of IRAP. At this time, if all slices within a picture are composed of IRAP, or the flag (gdr_or_irap_pic_flag) value indicating that the picture is a GDR (Gradual Decoding Refresh) picture may be forced to be 0. That is, in the present invention, when there are other NUTs within a picture (mixed_nalu_type_in_pic_flag = 1), the flag (gdr_or_irap_pic_flag) value may be set to 0 or induced to a 0 value. Alternatively, when mixed_nalu_type_in_pic_flag = 1, gdr_or_irap_pic_flag may be forced to have a value of 1. The flag (gdr_or_irap_pic_flag) may be defined at one or more positions among SPS, PPS, and PH.
[0116] In addition, when the hybrid NUT function is applied, at least one sub-picture in one picture may have an IRAP (IDR or CRA) NUT value, and at least one other sub-picture may have a non-IRAP (inter slice) NUT value. That is, in one picture, an intra-slice and an inter-slice may coexist. In the case of the existing HEVC standard, when the NUT of the intra-slice is IDR, the DPB is reset. As a result, all the reconstructed pictures existing in the DPB at that time are removed.
[0117] However, in the case of the present invention, when mixed_nalu_type_in_pic_flag = 1, since an intra-slice and an inter-slice can coexist in one picture, there may be a case where the DPB cannot be reset even if one picture is an IDR NUT. Thus, in one embodiment, when the slice is an IDR NUT, reference picture information (RPL: reference picture list) may be inserted into the NAL bitstream as the slice header information of the IDR, such as CRA. Therefore, a flag (idr_rpl_present_flag) value indicating the existence of RPL information may be set to 1 even though it is an IDR NUT. When the flag (idr_rpl_present_flag) value is 1, RPL exists as the slice header information of IDR. Conversely, when the flag (idr_rpl_present_flag) value is 0, RPL does not exist as the slice header information of IDR.
[0118] On the other hand, in the present invention, when there is no other NUT in one picture (mixed_nalu_type_in_pic_flag = 1) and the RPL information of the IDR picture is not allowed (idr_rpl_present_flag = 0), the NUT for the picture shall not have an IDR_W_RADL or IDR_N_LP value.
[0119] The flag (idr_rpl_present_flag) may be justified at one or more positions among SPS, PPS, and PH. For example, when the flag is justified at SPS, the flag may be named sps_idr_rpl_present_flag. For example, even if the NUT of the current slice is IDR_W_RADL or IDR_N_RADL, slice header information may be signaled using the syntax of the slice header in FIG. 15 to signal RPL according to the value of sps_idr_rpl_present_flag. Here, the first value of sps_idr_rpl_present_flag (e.g., 0) indicates that the RPL syntax element is not provided by the slice header of a slice where the NUT is IDR_N_LP or IDR_W_RADL. The second value of sps_idr_rpl_present_flag (e.g., 1) indicates that the RPL syntax element is provided by the slice header of a slice where the NUT is IDR_N_LP or IDR_W_RADL.
[0120] On the other hand, in another embodiment, when mixed_nalu_type_in_pic_flag = 1, the RPL may be signaled with the picture header information. For example, in the application of the syntax in FIG. 14, when the value of mixed_nalu_type_in_pic_flag indicates that the hybrid NUT is applied, the value of pps_no_pic_partition_flag may be forced to a value indicating disable (e.g., 0). Also, thereby, the value of a flag (pps_rpl_info_in_ph_flag) indicating whether the RPL information is provided from the picture header may be obtained from the bitstream. When pps_rpl_info_in_ph_flag indicates enable (e.g., 1), the RPL information may be obtained from the picture header as shown in FIGS. 16 and 17. Thus, based on the value of mixed_nalu_type_in_pic_flag, the RPL information is obtained regardless of the type of the picture. On the other hand, when pps_rpl_info_in_ph_flag indicates disable (e.g., 0), the RPL information cannot be obtained from the picture header. For example, while the pps_rpl_info_in_ph_flag value is 0, if the slice NUT is IDR_N_LP or IDR_W_RADL and the value of sps_idr_rpl_present_flag is 0, the RPL information of the slice is not obtained. That is, since there is no RPL information for the slice, the RPL information may be induced to be initialized and empty.
[0121] As described above, one picture may be signaled with different types of NAL units. Thus, since NAL units having different NUTs from each other are used to signal one picture, a method for determining the type of the picture according to the type of the NAL unit is required. Thereby, at the time of random access (RA: Random Access), it may be determined whether the picture can be normally restored and output.
[0122] In the decoding process according to one embodiment, when each VCL NAL unit corresponding to a picture is a NAL unit of CRA_NUT type, the picture may be determined as a CRA picture. Also, when each VCL NAL unit corresponding to a picture is IDR_W_RADL or a NAL unit of IDR_N_LP type, the picture may be determined as an IDR picture. Also, when each VCL NAL unit corresponding to a picture is a NAL unit of IDR_W_RADL, IDR_N_LP, or CRA_NUT type, the picture may be determined as an IRAP picture.
[0123] Also, when each VCL NAL unit corresponding to a picture is a NAL unit of RADL_NUT type, the picture may be determined as a RADL (Random Access decodable leading) picture. Also, when each VCL NAL unit corresponding to a picture is a NAL unit of TRAIL_NUT type, the picture may be determined as a trailing picture. Also, when the type of at least one VCL NAL unit among the VCL NAL units corresponding to a picture is of RASL_NUT type and the types of all other VCL NAL units are of RASL_NUT type or RADL_NUT type, the said picture may be determined as a RASL (random access skipped leading) picture.
[0124] On the other hand, in the decoding process according to another embodiment, if at least one sub-picture in a picture is RASL while at least one other sub-picture is RADL, the picture may be determined as a RASL picture. For example, if at least one sub-picture in a picture is RASL while at least one other sub-picture is RADL, during the decoding process, the picture may be set as a RASL picture. Here, if the type of the VCL NAL unit corresponding to the sub-picture is RASL_NUT, the sub-picture may be determined as RASL. Thereby, at the time of RA, all RASL sub-pictures and RADL sub-pictures may be treated as RASL pictures, and thus, the picture may not be output.
[0125] On the other hand, in the decoding process according to another embodiment, if at least one sub-picture in a picture is RASL, the picture may be set as a RASL picture. For example, if at least one sub-picture in a picture is RASL while at least one other sub-picture is TRAIL, during the decoding process, the picture may be set as a RASL picture. Thereby, at the time of RA, the picture may be treated as a RASL picture, and the picture may not be output.
[0126] Here, the occurrence of RA may be determined by the NoOutputBeforeRecoveryFlag value of the IRAP picture concatenated (related) with the inter-slice (RADL, RASL, or TRAIL). When the flag value is 1 (NoOutputBeforeRecoveryFlag = 1), it means the occurrence of RA, and when the flag value is 0 (NoOutputBeforeRecoveryFlag = 0), it means normal playback. The flag value may be set for IRAP as follows.
[0127] - When the current picture is IRAP, the setting process of the NoOutputBeforeRecoveryFlag value 1. If the picture is the first picture in the bitstream, set NoOutputBeforeRecoveryFlag to 1. 2. If the picture is an IDR picture, set NoOutputBeforeRecoveryFlag to 1. 3. If the picture is a CRA picture and an RA is signaled externally, set NoOutputBeforeRecoveryFlag to 1. 4. If the picture is a CRA picture and no RA is signaled externally, set NoOutputBeforeRecoveryFlag to 0.
[0128] In one embodiment, the decoding device may be signaled by an external terminal about the occurrence of a random access. For example, the external terminal may signal the occurrence of a random access to the decoding device by setting the value of the occurrence information of the random access to 1 and signaling it to the decoding device. The decoding device may set the value of HandleCraAsClvsStartFlag, which is a flag indicating whether it has received the occurrence of a random access from the external terminal, to 1 according to the occurrence information of the random access received from the external terminal. The decoding device may set the value of NoOutputBeforeRecoveryFlag to the same value as the value of HandleCraAsClvsStartFlag. Thereby, when the current picture is a CRA picture and the value of HandleCraAsClvsStartFlag is 1, the decoding device may determine that a random access has occurred for the CRA picture, or may perform decoding assuming that the CRA is located at the beginning of the bitstream.
[0129] In the case of RA, the process of setting a flag (PictureOutputFlag) for determining whether the current picture is to be output is as follows. For example, the PictureOutputFlag for the current picture may be set in the following order. Here, the first value of PictureOutputFlag (e.g., 0) indicates that the current picture is not to be output. The second value of PictureOutputFlag (e.g., 1) indicates that the current picture is to be output.
[0130] (1) If the current picture is a RASL and the NoOutputBeforeRecoveryFlag of the related IRAP picture is 1, set PictureOutputFlag to 0. (2) If the current picture is a GDR picture with the NoOutputBeforeRecoveryFlag value of 1 or its recovered picture, set PictureOutputFlag to 0. (3) Otherwise, set the value of PictureOutputFlag to the same value as the pic_output_flag value in the bitstream. Here, pic_output_flag may be obtained at one or more positions of PH and SH.
[0131] FIG. 18 shows an illustration of the composition of three different contents presented in the present invention. FIG. 18(a) shows the sequences for three different contents. For convenience, one picture is shown as one packet, but one picture may be divided into multiple slices and there may be multiple pallets. FIGS. 18(b) and 18(c) show the synthesized image results for the pictures indicated by the dotted lines in FIG. 18(a). In FIG. 18, the same color means the same picture / sub-picture / slice. Also, the P slice and the B slice may have one value among the INTER NUTs.
[0132] As described above, according to the present invention, when synthesizing a large number of contents, it is not always necessary to make the positions of the intra slices (pictures) equal. By simply matching the hierarchical GOP structure, the contents can be synthesized quickly and easily without delay.
[0133] Embodiments of Symbolization and Decryption Hereinafter, a method for a video decoding apparatus to decode video by the above-described method will be described. FIGS. 19 and 20 are sequence diagrams for explaining a decoding method and an encoding method according to an embodiment of the present invention.
[0134] A video decoding apparatus according to an embodiment may include a memory and at least one processor, and by the operation of the processor, the following decoding method can be performed. First, the decoding apparatus acquires NAL unit type information indicating the type of the current NAL (network abstraction layer) unit from the bitstream (S1910).
[0135] Next, when the NAL unit type information indicates that the NAL unit type of the current NAL unit is encoded data for a video slice, the decoding apparatus decodes the video slice based on whether a mixed NAL unit type is applied to the current picture (S1920).
[0136] Here, the decoding apparatus can decode the video slice by determining whether the NAL unit type of the current NAL unit indicates the attribute of the sub-picture for the current video slice based on whether the mixed NAL unit type is applied.
[0137] Whether the hybrid NAL unit type is applied may be identified based on a first flag (e.g., pps_mixed_nalu_types_in_pic_flag) obtained from the picture parameter set. When the hybrid NAL unit type is applied, the current picture to which the current video slice belongs may be divided into at least two sub-pictures.
[0138] Furthermore, decoding information for the sub-picture may be included in the bitstream based on whether the hybrid NAL unit type is applied. In one embodiment, a second flag (e.g., pps_no_pic_partition_flag) indicating whether the current picture is divided from the bitstream is obtained. Also, when the second flag indicates that the current picture is dividable (e.g., pps_no_pic_partition_flag == 0), a third flag (e.g., pps_rpl_info_in_ph_flag) indicating whether reference picture list information is provided from the picture header is obtained from the bitstream.
[0139] In such an example, when the hybrid NAL unit type is applied, the division of the current picture into at least two sub-pictures is forced, so that the value of the second flag (pps_no_pic_partition_flag) is forced to 0, and a third flag (e.g., pps_rpl_info_in_ph_flag) indicating whether reference picture list information is provided from the picture header is obtained from the bitstream regardless of the value of the second flag (pps_no_pic_partition_flag) actually obtained from the bitstream. Thereby, when the third flag indicates that reference picture list information is provided from the picture header (e.g., pps_rpl_info_in_ph_flag == 1), the reference picture list information is obtained from the bitstream regarding the picture header.
[0140] Also, when the hybrid NAL unit type is applied, the current picture may be decoded based on a first sub-picture and a second sub-picture having different NAL unit types from each other. Here, when the NAL unit type of the first sub-picture has any one value among IDR_W_RADL (Instantaneous Decoding Refresh_With_Random Access Decodable Leading), IDR_N_LP (Instantaneous Decoding Refresh_No reference_Leading Picture), and CRA_NUT (Clean Random Access_NAL Unit Type), the available NAL unit types selectable as the second sub-picture NUT may include the NAL unit type not selected from the first sub-picture among IDR_W_RADL, IDR_N_LP, and CRA_NUT.
[0141] Alternatively, when the NAL unit type of the first sub-picture has any one value among IDR_W_RADL (Instantaneous Decoding Refresh_With_Random Access Decodable Leading), IDR_N_LP (Instantaneous Decoding Refresh_No reference_Leading Picture), and CRA_NUT (Clean Random Access_NAL Unit Type), the available NAL unit type of the second sub-picture may include TRAIL_NUT (Trail_NAL Unit Type).
[0142] On the other hand, when the hybrid NAL unit type is applied, the first sub-picture and the second sub-picture constituting the current picture may be independently decoded. For example, the first sub-picture and the second sub-picture including B or P slices may be treated as one picture and decoded. For example, the first sub-picture may be decoded without using the second sub-picture as a reference picture.
[0143] More specifically, in the decoding process, a fourth flag (e.g., sps_subpic_treated_as_pic_flag) indicating whether the first sub-picture is treated as a picture may be obtained from the bitstream. If the fourth flag indicates that the first sub-picture is treated as a picture in the decoding process (e.g., sps_subpic_treated_as_pic_flag == 1), the first sub-picture may be treated as a picture and decoded in the decoding process. In such a process, if the hybrid NAL unit type is applied to the current picture and the current picture including the first sub-picture includes at least one P slice or B slice, the fourth flag may be forced to have a value indicating that the first sub-picture is treated as a picture in the decoding process. On the other hand, if the fourth flag indicates that the first sub-picture is not treated as a picture in the decoding process (e.g., sps_subpic_treated_as_pic_flag == 0) when the hybrid NAL unit type is applied to the current picture, the slice type belonging to the current picture must be intra.
[0144] If the fourth flag indicates that the first sub-picture is treated as a picture in the decoding process, it can be determined that the decoding process of the first sub-picture is independent of other sub-pictures. For example, if the fourth flag indicates that the first sub-picture is decoded independently of other sub-pictures in the decoding process, the first sub-picture may be decoded without using other sub-pictures as reference pictures.
[0145] Also, when the first sub-picture is a RASL (Random Access Skipped Leading) sub-picture, the current picture may be determined to be a RASL picture based on whether the second sub-picture is a RADL (Random Access Decodable Leading) sub-picture. Here, when the type of the NAL unit corresponding to the first sub-picture is RASL_NUT (Random Access Skipped Leading_NAL Unit Type), the first sub-picture may be determined to be a RASL sub-picture.
[0146] Also, when the third flag (e.g., pps_rpl_info_in_ph_flag) indicates that the reference picture list information is not obtained from the picture header but from the slice header (e.g., pps_rpl_info_in_ph_flag == 0), and when the NAL unit type of the first sub-picture has a value of either IDR_W_RADL (Instantaneous Decoding Refresh_With_Random Access Decodable Leading) or IDR_N_LP (Instantaneous Decoding Refresh_No reference_Leading Picture), the reference picture list information may be obtained from the bitstream related to the slice header based on the fifth flag (e.g., sps_idr_rpl_present_flag) indicating whether the reference picture list information for the IDR picture exists in the slice header. Here, the fifth flag may be obtained from the bitstream related to the sequence parameter set.
[0147] On the other hand, when random access is performed on an IRAP (Intra Random Access Point) picture related to the current picture, if the current picture is a RASL (Random Access Skipped Leading) sub-picture, the current picture may not be output (displayed).
[0148] A video encoding device according to an embodiment may include a memory and at least one processor, and by the operation of the processor, an encoding method corresponding to the above-described decoding method can be performed. For example, when the current picture is encoded based on the hybrid NAL unit type, the encoding device determines the type of sub-picture for dividing the picture (S2010). Further, the encoding device encodes at least one current video slice constituting the sub-picture based on the type of sub-picture to generate a current NAL unit (S2020). At this time, when the current picture is encoded based on the hybrid NAL unit type, the encoding device encodes the video slice by encoding the NAL unit type of the current NAL unit so as to indicate the attribute of the sub-picture for the current video slice.
[0149] Further, the present invention can be realized as a computer-readable code on a computer-readable recording medium (including all devices having an information processing function). The computer-readable recording medium includes all types of recording devices in which data read by a computer system is stored. Examples of computer-readable recording devices include ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage device, and the like.
[0150] The present invention has been described with reference to the embodiments shown in the drawings, which are merely exemplary, and it will be understood by those having ordinary knowledge in this technical field that various modifications and equivalent other embodiments will be possible hereinafter. Therefore, the true technical protection scope of the present invention should be determined by the technical idea of the appended claims.
Claims
1. A video decoding method performed by a video decoding apparatus, comprising: obtaining a first flag indicating whether a current picture includes sub-pictures having different NAL (Network Abstraction Layer) unit types; obtaining a second flag indicating whether the sub-pictures have been treated as separate pictures during the decoding process, wherein when it is determined based on the first flag that the picture includes sub-pictures having different NAL unit types, the second flag is forced to have a first value for at least one sub-picture including at least one of a P slice or a B slice among the sub-pictures, wherein the first value indicates that the sub-picture has been treated as a separate picture during the decoding process, wherein the first flag is obtained from a picture parameter set and the second flag is obtained from a sequence parameter set. A video decoding method.
2. A video decoding apparatus, comprising: a memory storing one or more instructions; at least one processor, wherein by executing the one or more instructions by the at least one processor, a first flag indicating whether a current picture includes sub-pictures having different NAL (Network Abstraction Layer) unit types is obtained, a second flag indicating whether the sub-pictures are treated as separate pictures during the decoding process is obtained, wherein when it is determined based on the first flag that the picture includes sub-pictures having different NAL unit types, the second flag is forced to have a first value for at least one sub-picture including at least one of a P slice or a B slice among the sub-pictures, wherein the first value indicates that the sub-picture is treated as a separate picture during the decoding process, wherein the first flag is obtained from a picture parameter set and the second flag is obtained from a sequence parameter set. A video decoding apparatus.
3. A video encoding method performed by a video encoding apparatus, comprising: determining whether a current picture includes sub-pictures having different NAL (Network Abstraction Layer) unit types, a step of determining whether the sub-picture is treated as a separate picture during the decoding process; a step of generating a first flag indicating whether the current picture includes sub-pictures having different NAL unit types from each other and a second flag indicating whether the sub-picture is treated as a separate picture during the decoding process; when it is determined based on the first flag that the picture includes sub-pictures having different NAL unit types from each other, the second flag is forced to have a first value for a sub-picture including at least one of a P slice or a B slice among the sub-pictures; the first value indicates that the sub-picture is treated as a separate picture during the decoding process; the first flag is obtained from a picture parameter set, and the second flag is obtained from a sequence parameter set, a video encoding method. [
4. ] A method for transmitting an encoded bitstream, comprising: a step of determining whether the current picture includes sub-pictures having different NAL (Network Abstraction Layer) unit types from each other; a step of determining whether the sub-picture is treated as a separate picture during the decoding process; a step of generating a bitstream including a first flag indicating whether the current picture includes sub-pictures having different NAL unit types from each other and a second flag indicating whether the sub-picture is treated as a separate picture during the decoding process; a step of transmitting the bitstream from a video encoding device to a video decoding device; when it is determined based on the first flag that the picture includes sub-pictures having different NAL unit types from each other, the second flag is forced to have a first value for a sub-picture including at least one of a P slice or a B slice among the sub-pictures; the first value indicates that the sub-picture is treated as a separate picture during the decoding process; the first flag is obtained from a picture parameter set, and the second flag is obtained from a sequence parameter set, a method.
Citation Information
Patent Citations
Pictures with mixed NAL unit types
WO2020185922A1