Image decoding method and device
By acquiring NAL unit type information and decoding the image slices based on the mixed NAL unit type, the problem of large amount of data encoding for high-resolution images is solved, and the data volume and transmission and storage cost are reduced.
Patent Information
- Application Number
- CN202510256762.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2020-11-17
- Filing Date
- 2020-12-16
- Publication Date
- 2025-05-06
AI Technical Summary
The large amount of encoded data of high-resolution images leads to an increase in transmission and storage costs, and it is difficult for the prior art to effectively reduce the amount of encoded data of high-resolution images.
An image decoding method is adopted to obtain NAL unit type information, determine whether the current NAL unit is the encoded data of the image slice, and decode the image slice based on whether the mixed NAL unit type is applicable.
Through the slice segmentation method of sub-graph segmentation and bitstream packaging, the amount of image encoding data is reduced, efficient image transmission and storage is realized, and the cost is reduced.
Smart Images

Figure CN119946283A_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application with application date of December 16, 2020, application number 202080005988.4, and invention name “Image decoding method and device”. Technical Field
[0002] The present invention relates to a sub-image segmentation method for synthesis with other sequences and a slice segmentation method for bit stream packaging. Background Art
[0003] The user demand for high-resolution and high-quality images is increasing. Since the coded data of high-resolution images has more information than the coded data of low-resolution or medium-resolution images, the transmission or storage cost increases.
[0004] In order to solve the above-mentioned problems, research on encoding and decoding methods that effectively reduce the amount of encoded data of high-resolution images is continuously being conducted. Summary of the invention
[0005] Technical issues The present specification discloses a sub-picture partitioning method for synthesis with other sequences and a slice partitioning method for bitstream packing.
[0006] Technical Solution An image decoding method performed by an image decoding device of an embodiment of the present invention for solving the above-mentioned problem comprises the following steps: obtaining NAL unit type information showing the type of a current NAL (network abstraction layer) unit from a bitstream; and decoding the image slice based on whether a mixed NAL unit type (mixed NAL unit type) is applicable to the current image, in the case where the NAL unit type information shows that the NAL unit type of the current NAL unit is the coded data of an image slice. Here, the step of decoding the image slice is performed by determining whether the NAL unit type of the current NAL unit shows the attributes of a sub-image of the current image slice based on whether the mixed NAL unit type is applicable.
[0007] Furthermore, an image decoding device according to an embodiment of the present invention for solving the above-mentioned problem is an image decoding device including a memory and at least one processor, wherein the at least one processor obtains NAL unit type information indicating the type of a current NAL (network abstraction layer) unit from a bitstream, and decodes the image slice based on whether a mixed NAL unit type is applicable to the current image when the NAL unit type information indicates that the NAL unit type of the current NAL unit is coded data of an image slice. At this time, the decoding of the image slice is performed based on whether the mixed NAL unit type is applicable, by determining whether the NAL unit type of the current NAL unit indicates an attribute of a sub-image of the current image slice.
[0008] Furthermore, an image encoding method performed by an image encoding device of an embodiment of the present invention for solving the above problem includes the following steps: in the case where the current image is encoded based on a mixed NAL unit type, determining the type of a sub-image that divides the image; and based on the type of the sub-image, encoding at least one current image slice constituting the sub-image to generate a current NAL unit. Here, in the step of encoding the image slice, in the case where the current image is encoded based on the mixed NAL unit type, the NAL unit type of the current NAL unit is encoded in such a way that it displays the attributes of the sub-image of the current image slice.
[0009] Furthermore, a transmission method according to an embodiment of the present invention for solving the above-mentioned problem transmits a bit stream generated by the image encoding device or the image encoding method of the present application.
[0010] Furthermore, a computer-readable recording medium according to an embodiment of the present invention for solving the above-mentioned problem stores a bit stream generated by the video encoding method or the video encoding device of the present application.
[0011] Beneficial Effects The present invention discloses a method for generating an image by synthesizing with various different sequences. An image in a sequence is divided into a plurality of sub-images, and a new image is generated by synthesizing the divided sub-images of other images.
[0012] According to the application of the present invention, the NAL (network abstraction layer) unit type values of two or more sub-images constituting an image are different from each other. When different contents are synthesized, it is not necessary to also match the NUTs of the multiple sub-images constituting an image, which has the advantage of being easy to construct / synthesize images. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1A diagram schematically showing a structure of a video encoding device to which the present invention is applied.
[0014] Figure 2 The figure is a diagram showing an example of a video encoding method performed by a video encoding device.
[0015] Figure 3 A diagram schematically showing a structure of a video decoding device to which the present invention is applied.
[0016] Figure 4 A diagram showing an example of a video decoding method performed by a decoding device.
[0017] Figure 5 The figure is an illustration showing an example of a NAL data packet for a slice.
[0018] Figure 6 A diagram showing an example of a hierarchical GOP structure.
[0019] Figure 7 This is a diagram showing an example of the display output order and the decoding order.
[0020] Figure 8 This is a diagram showing an example of a read image and a standard image.
[0021] Fig. 9 The following is a diagram showing an example of a RASL image and a RADL image.
[0022] Fig.10 An illustration showing the syntax of a slice segment title.
[0023] Fig.11 The figure is a diagram showing an example of the content synthesis process.
[0024] Fig.12 This is a diagram showing an example of a sub-image ID and a slice address.
[0025] Fig.13 A diagram showing an example of NUT displayed in sub-graphs / slices.
[0026] Fig.14 FIG. 1 is a diagram showing an embodiment of the syntax of a picture parameter set (PPS).
[0027] Fig.15 A diagram showing an embodiment of the syntax of a slice header.
[0028] Fig.16 Figure 2 is a diagram showing the syntax of the image title structure.
[0029] Fig.17 The figure is a diagram showing the syntax for obtaining a reference image list.
[0030] Fig.18 The following is a diagram showing an example of content synthesis.
[0031] Fig.19 and Fig. 20 A sequence diagram for illustrating a decoding method and an encoding method according to an embodiment of the present invention. DETAILED DESCRIPTION
[0032] The present invention is subject to various changes and has various embodiments, and specific embodiments are illustrated in the drawings and described in detail. However, the present invention is not limited to specific embodiments. The terms used in this specification are used only to illustrate specific embodiments and are not intended to limit the technical ideas of the present invention. Expressions in the singular that are not clearly defined differently in the context include plural expressions. In this specification, the terms "including" or "having" specify the existence of features, numbers, steps, actions, constituent elements, parts or features, numbers, steps, actions, constituent elements, and combinations of parts recorded in the specification, and should be understood as not excluding in advance the existence or additional possibility of one or more other features, numbers, steps, actions, constituent elements, parts or features, numbers, steps, actions, constituent elements, and combinations of parts.
[0033] In addition, the various structures on the drawings described in the present invention are independently displayed for the convenience of explaining the functions of different features, and the various structures are not implemented by separate hardware or other software. For example, two or more structures in each structure can be combined to form a structure, and a structure can also be divided into multiple structures. The embodiments of each structure in combination and / or separation do not deviate from the essence of the present invention, but are included in the scope of the claims of the present invention.
[0034] The preferred embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. The same reference numerals are used for the same components in the drawings, and repeated description of the same components is omitted.
[0035] In addition, the present invention relates to video / image coding. For example, the method / embodiment disclosed in the present invention is applicable to methods disclosed in VVC (universal video coding) standard, EVC (Essential Video Coding) standard, AV1 (video coding format, AOMedia Video 1) standard, AVS2 (2nd generation of audio video coding standard) or new generation video / image coding standards (e.g., H.267, H.268, etc.).
[0036] In this specification, an access unit (AU) refers to a unit that displays multiple sets of pictures belonging to different layers output from the DPB (Decoded Picture Buffer) at the same time. A picture generally refers to a unit that displays an image in a specific time period, and a slice refers to a unit that constitutes a part of a picture during encoding. A picture is composed of multiple slices, and pictures and slices can be mixed and used as needed.
[0037] A pixel or a picture element (pel) is the smallest unit that constitutes an image (or video). In addition, a "sample" is used as a term corresponding to a pixel. A sample generally displays a pixel or a pixel value, but it can also display only the pixel / pixel value of the brightness (luma) component, or it can also display only the pixel / pixel value of the chromaticity (chroma) component.
[0038] Unit represents the basic unit of image processing. Unit includes at least one of a specific area of the image and information related to the corresponding area. Unit is used interchangeably with terms such as block or area depending on the situation. In general, an MxN block represents a set of samples or transform coefficients consisting of M columns and N rows.
[0039] Figure 1 A diagram briefly illustrating the structure of a video encoding device to which the present invention is applied.
[0040] Reference Figure 1 The video encoding device 100 includes: an image segmentation unit 105, a prediction unit 110, a residual processing unit 120, an entropy coding unit 130, an addition unit 140, a filter unit 150 and a memory 160. The residual processing unit 120 includes: a subtraction unit 121, a conversion unit 122, a quantization unit 123, a rearrangement unit 124, an inverse quantization unit 125 and an inverse conversion unit 126.
[0041] The image segmentation unit 105 segments the input image into at least one processing unit.
[0042] For example, a processing unit is called a coding unit (CU). In this case, the coding unit is recursively split from a coding tree unit (Coding Tree Unit) according to a QTBT (quad-tree binary-tree) structure. For example, a coding tree unit is split into multiple nodes at a deeper level based on a quadtree structure and / or a binary tree structure. In this case, for example, the quadtree structure is first applied, and the binary tree structure is finally applied. Or the binary tree structure can also be applied first. For nodes that cannot be split any further, decoding is performed, thereby determining a coding unit for nodes that cannot be split any further. The coding tree unit can also be named as a coding unit in terms of being a unit for splitting a coding unit. In this case, in terms of determining a coding unit by splitting a coding tree unit, the coding tree unit can also be named as the largest coding unit (LCU).
[0043] Thus, the encoding procedure of the present invention is executed based on the final coding unit that cannot be further divided. In this case, based on the coding efficiency of the image characteristics, the coding tree unit is directly used as the final coding unit, or the coding unit is recursively divided into coding units of lower depths as needed, and the coding unit of the most appropriate size is used as the final coding unit. Here, the so-called encoding procedure includes the following prediction, conversion and recovery procedures.
[0044] In another example, the processing unit also includes: a coding unit (CU), a prediction unit (PU) or a transform unit (TU). The coding unit is split into deeper coding units from the coding tree unit according to a quadtree structure. In this case, based on the coding efficiency according to the image characteristics, the coding tree unit is directly used as the final coding unit, or the coding unit is recursively split into deeper coding units as needed, and the coding unit of the optimal size is finally used as the coding unit. In the case of setting the minimum coding unit (min coding unit, minCU), the coding unit is split into coding units smaller than the minimum coding unit. Here, the so-called final coding unit refers to a coding unit based on partitioning or partitioning into prediction units or transform units. The prediction unit is a unit of sample prediction as a unit partitioning from the coding unit. At this time, the prediction unit can also be divided into sub-blocks. The transform unit is split from the coding unit according to the quadtree structure, and is a unit that guides the residual signal from the unit that guides the transform parameter and / or the transform parameter. Hereinafter, the coding unit is referred to as a coding block (CB), the prediction unit is referred to as a prediction block (PB), and the transform unit is referred to as a transform block (TB). A prediction block or a prediction unit refers to a block-shaped specific area within an image, including an array of prediction samples. Furthermore, a transform block or a transform unit refers to a block-shaped specific area within an image, including an array of transform parameters or residual samples.
[0045] The prediction unit 110 performs prediction on a processing target block (hereinafter referred to as a current block) to generate a predicted block (predicted block) including prediction samples of the current block. The unit of prediction performed in the prediction unit 110 is a coding block, which may also be a conversion block or a prediction block.
[0046] The prediction unit 110 determines whether to apply intra prediction or inter prediction in the current block. For example, the prediction unit 110 determines whether to apply intra prediction or inter prediction in units of CUs.
[0047] In the case of intra-frame prediction, the prediction unit 110 guides the prediction samples of the current block based on the reference samples outside the current block in the image to which the current block belongs (hereinafter referred to as the current image). At this time, the prediction unit 110 (i) guides the prediction samples based on the average or interpolation of the neighboring reference samples of the current block, and (ii) can also guide the prediction samples in the neighboring reference samples of the current block based on the reference samples existing in a specific (prediction) direction. The case of (i) is called a non-directional mode or a non-angular mode, and the case of (ii) is called a directional mode or an angular mode. In intra-frame prediction, the prediction mode has, for example, 33 directional prediction modes and at least two non-directional modes. The non-directional mode includes a DC prediction mode and a planar mode. The prediction unit 110 can also determine the prediction mode applicable to the current block by using the prediction mode applicable to the neighboring blocks.
[0048] In the case of inter-frame prediction, the prediction unit 110 can guide the prediction sample of the current block based on the sample specified by the motion vector on the reference image. The prediction unit 110 applies any one of the skip mode, merge mode and MVP (motion vector prediction) mode to guide the prediction sample of the current block. In the case of the skip mode and the merge mode, the prediction unit 110 uses the motion information of the surrounding blocks as the motion information of the current block. In the case of the skip mode, unlike the merge mode, the residual between the prediction sample and the original sample is not transmitted. In the case of the MVP mode, the motion vector of the surrounding block is used as a motion vector predictor (Motion Vector Predictor), and the motion vector of the current block is guided using the motion vector predictor of the current block.
[0049] In the case of inter-frame prediction, the neighboring blocks include spatial neighboring blocks in the current image and temporal neighboring blocks in the reference picture. The reference picture containing the temporal neighboring blocks is also called a collocated picture (colPic). Motion information includes motion vectors and reference picture indexes. Information such as prediction mode information and motion information is (entropy) encoded and output in the form of a bitstream.
[0050] In skip mode and merge mode, the top picture on the reference picture list can also be used as a reference picture when using the temporal motion information of the surrounding blocks. The reference pictures included in the reference picture list are arranged based on the POC (Picture Order Count) difference between the current picture and the corresponding reference picture. POC corresponds to the display order of the pictures and is distinguished from the encoding order.
[0051] The subtraction unit 121 generates residual samples which are differences between original samples and predicted samples. When the skip mode is applied, the residual samples are not generated as described above.
[0052] The conversion unit 122 converts the residual samples in units of conversion blocks and generates a conversion parameter (transform coefficient). The conversion unit 122 performs conversion according to the size of the corresponding conversion block, the coding block spatially overlapping with the corresponding conversion block, or the prediction mode applicable to the prediction block. For example, in the case where intra prediction is applied in the coding block or the prediction block overlapping with the conversion block, and the conversion block is a 4×4 residual array, the residual samples are converted using a DST (Discrete Sine Transform) conversion kernel. For other cases, the residual samples are converted using a DCT (Discrete Cosine Transform) conversion kernel.
[0053] The quantization unit 123 quantizes the conversion parameter to generate a quantized conversion parameter.
[0054] The rearrangement unit 124 rearranges the quantization conversion parameters. The rearrangement unit 124 rearranges the block-shaped quantization conversion parameters in a one-dimensional vector form by a scanning parameter method. Here, the rearrangement unit 124 is described with another structure, but the rearrangement unit 124 is a part of the quantization unit 123.
[0055] The entropy coding unit 130 performs entropy coding on the quantized conversion parameters. Entropy coding includes, for example, coding methods such as exponential Golomb, CAVLC (context-adaptive variable length coding), and CABAC (context-adaptive binary arithmetic coding). In addition to the quantized conversion parameters, the entropy coding unit 130 can also encode information required for video recovery (for example, the value of a syntax element) together or separately. The entropy coded information is in the form of a bit stream and is transmitted or stored in units of NAL (network abstraction layer).
[0056] The inverse quantization unit 125 inversely quantizes the value quantized by the quantization unit 123 (quantized conversion parameter), and the inverse conversion unit 126 inversely converts the value inversely quantized by the inverse quantization unit 125 to generate residual samples.
[0057] The adding unit 140 combines the residual samples and the prediction samples to restore the image. The residual samples and the prediction samples are added in units of blocks to generate a restored block. Here, the adding unit 140 is described as a separate structure, but the adding unit 140 is a part of the prediction unit 110. In addition, the adding unit 140 can also be called a restoration unit or a restoration block generation unit.
[0058] For the reconstructed picture, the filter unit 150 applies deblocking filtering and / or sample adaptive offset. The artifacts of block boundaries or distortions of the quantization process in the reconstructed picture are corrected by deblocking filtering and / or sample adaptive offset. Sample adaptive offset is applied in units of samples and is applied after the deblocking filtering process is completed. The filter unit 150 can also apply to the reconstructed picture after the ALF (Adaptive Loop Filter) is applied. ALF is applied to the reconstructed picture after the deblocking filter and / or sample adaptive offset are applied.
[0059] The memory 160 stores a restored image (decoded image) or information required for encoding / decoding. Here, the restored image is a restored image that has been filtered through the filter unit 150. The stored restored image is used as a reference image (inter-frame) for predicting other images. For example, the memory 160 stores a (reference) image used in inter-frame prediction. At this time, the image used in the inter-frame prediction is specified by a reference picture set or a reference picture list.
[0060] Figure 2 An example of a video encoding method performed by a video encoding device is shown. Figure 2 , the image encoding method includes: block partitioning, intra / inter prediction, transform, quantization and entropy encoding processes. For example, the current image is divided into multiple blocks, and a prediction block of the current block is generated by intra / inter prediction, and a residual block of the current block is generated by subtracting the input block of the current block and the prediction block. After that, a parameter block, that is, a transformation parameter of the current block, is generated by transforming the residual block. The transformation parameters are quantized and entropy encoded and stored in the bit stream.
[0061] Figure 3 A diagram schematically illustrating the structure of a video decoding device to which the present invention is applied.
[0062] Reference Figure 3 The video decoding device 300 includes an entropy decoding unit 310, a residual processing unit 320, a prediction unit 330, an addition unit 340, a filter unit 350, and a memory 360. Here, the residual processing unit 320 includes a rearrangement unit 321, an inverse quantization unit 322, and an inverse conversion unit 323.
[0063] When a bit stream including video information is input, the video decoding device 300 restores the video in accordance with a program for processing the video information in the video encoding device.
[0064] For example, the video decoding device 300 performs video decoding using a processing unit applicable to a video encoding device. Therefore, the processing unit block of the video decoding is, for example, a coding unit, another example is a coding unit, a prediction unit, or a conversion unit. The coding unit is divided from the coding tree unit according to a quadtree structure and / or a binary tree structure.
[0065] The prediction unit and the conversion unit are further applicable according to this case, for which the prediction block is a block derived or partitioned from the coding unit, and is a unit of sample prediction. In this case, the prediction unit can also be divided into sub-blocks. The conversion unit is a unit that is split from the coding unit according to the quadtree structure and leads to a unit of a residual signal from a unit that guides a conversion parameter or a conversion parameter.
[0066] The entropy decoding unit 310 parses the bitstream and outputs information required for video restoration or image restoration. For example, the entropy decoding unit 310 decodes the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and outputs the values of syntax elements required for video restoration and the quantized values of conversion parameters related to the residual.
[0067] More specifically, the CABAC entropy decoding method receives a binary corresponding to each sentence element from a bit stream, determines a context model using information about the decoding object sentence element and decoding information of surrounding and decoding object blocks or information about symbols / bins decoded in the previous step, and predicts the probability of occurrence of a binary according to the determined context model, performs arithmetic decoding of the binary, and generates a symbol corresponding to the value of each sentence element. At this time, after determining the context model, the CABAC entropy decoding method updates the context model using information about symbols / bins decoded using the context model for symbols / bins.
[0068] Information related to prediction among the information decoded in the entropy decoding unit 310 is provided to the prediction unit 330 , and residual values obtained by entropy decoding in the entropy decoding unit 310 , that is, quantized conversion parameters, are input to the rearrangement unit 321 .
[0069] The rearrangement unit 321 rearranges the quantized conversion parameters in a two-dimensional block. The rearrangement unit 321 performs rearrangement in accordance with the parameter scan performed in the encoding device. Here, the rearrangement unit 321 is described as a separate structure, but the rearrangement unit 321 is a part of the inverse quantization unit 322.
[0070] The inverse quantization unit 322 inversely quantizes the quantized conversion parameter based on the (inverse) quantization parameter and outputs the conversion parameter. At this time, information for guiding the quantization parameter is signaled by the encoding device.
[0071] The inverse transformation unit 323 performs inverse transformation on the transformation parameters to induce residual samples.
[0072] The prediction unit 330 performs prediction on the current block and generates a predicted block including prediction samples of the current block. The unit of the prediction performed in the prediction unit 330 can be a coding block, a transform block, or a prediction block.
[0073] The prediction unit 330 determines whether to apply intra-frame prediction or inter-frame prediction based on the predicted information. At this time, the unit for determining whether intra-frame prediction or inter-frame prediction is applicable is different from the unit for generating prediction samples. Moreover, the unit for generating prediction samples is also different for inter-frame prediction and intra-frame prediction. For example, it is determined by the CU unit to which either inter-frame prediction or intra-frame prediction is applied. Moreover, for example, for inter-frame prediction, the prediction mode is determined by the PU unit and the prediction sample is generated, and for intra-frame prediction, the prediction mode is determined by the PU unit, and the prediction sample can also be generated by the TU unit.
[0074] In the case of intra prediction, the prediction unit 330 guides the prediction samples of the current block based on the surrounding reference samples in the current image. The prediction unit 330 applies a directional mode or a non-directional mode based on the surrounding reference samples of the current block, thereby guiding the prediction samples for the current block. At this time, the prediction mode applicable to the current block can also be determined by using the intra prediction mode of the surrounding blocks.
[0075] In the case of inter-frame prediction, the prediction unit 330 guides the prediction sample of the current block based on the sample specified on the reference image by the motion vector on the reference image. The prediction unit 330 applies any one of the skip mode, merge mode and MVP mode to guide the prediction sample of the current block. At this time, the motion information required for the inter-frame prediction of the current block provided by the video encoding device, such as the motion vector, the reference image index and other related information, is obtained or guided based on the prediction-related information.
[0076] In the case of skip mode and merge mode, the motion information of the neighboring blocks is used as the motion information of the current block. At this time, the neighboring blocks include spatial neighboring blocks and temporal neighboring blocks.
[0077] The prediction unit 330 forms a merge alternative list from the motion information of the available surrounding blocks, and uses the information indicated by the merge index on the merge alternative list as the motion vector of the current block. The merge index is signaled by the encoding device. The motion information includes a motion vector and a reference image. When the motion information of the temporal surrounding blocks is used in the skip mode and the merge mode, the highest image on the reference image list is used as the reference image.
[0078] In the case of skip mode, unlike merge mode, the difference (residual) between the predicted sample and the original sample is not transmitted.
[0079] In the case of the MVP mode, the motion vector of the surrounding blocks is used as a motion vector predictor to guide the motion vector of the current block. At this time, the surrounding blocks include spatial surrounding blocks and temporal surrounding blocks.
[0080] For example, in the case where the merge mode is applied, a merge alternative list is generated using the motion vectors of the restored spatial neighboring blocks and / or the motion vectors corresponding to the temporal neighboring blocks, i.e., Col blocks. In the merge mode, the motion vector of the substitute block selected from the merge alternative list is used as the motion vector of the current block. The prediction-related information includes a merge index indicating a substitute block having the best motion vector selected from the substitute blocks included in the merge alternative list. At this time, the prediction unit 330 uses the merge index to obtain the motion vector of the current block.
[0081] In another example, in the case where the MVP (motion vector prediction) mode is applied, a motion vector prediction value alternative list is generated using the motion vector of the restored spatial neighboring block and / or the motion vector corresponding to the temporal neighboring block, i.e., Col block. That is, the motion vector corresponding to the restored spatial neighboring block and / or the motion vector corresponding to the temporal neighboring block, i.e., Col block, is used as a motion vector alternative. The information related to the prediction includes a predicted motion vector index indicating the best motion vector selected from the motion vector alternatives contained in the list. At this time, the prediction unit 330 uses the motion vector index to select the predicted motion vector of the current block from the motion vector alternatives contained in the motion vector alternative list. The prediction unit of the encoding device seeks the motion vector difference (MVD) between the motion vector of the current block and the motion vector prediction value, encodes it, and outputs it in the form of a bit stream. That is, the MVD is calculated by subtracting the motion vector prediction value from the motion vector of the current block. At this time, the prediction unit 330 obtains the motion vector difference contained in the information related to the prediction, and derives the motion vector of the current block by adding the motion vector difference and the motion vector prediction value. The prediction unit acquires or guides a reference picture index or the like of the reference picture from the information related to the prediction.
[0082] The adding unit 340 adds the residual sample and the prediction sample to restore the current block or the current image. The adding unit 340 can also restore the current image by adding the residual sample and the prediction sample in units of blocks. When the skip mode is applied, the residual is not transmitted, so the prediction sample is a restored sample. Here, the adding unit 340 is described as a separate structure, but the adding unit 340 can also be a part of the prediction unit 330. In addition, the adding unit 340 can also be called a restoration unit or a restoration block generation unit.
[0083] The filtering unit 350 applies sample adaptive offset for deblocking filtering to the restored image, and / or ALF, etc. At this time, the sample adaptive offset is also applied on a sample-by-sample basis and can also be applied after deblocking filtering. ALF can also be applied after deblocking filtering and / or sample adaptive offset.
[0084] The memory 360 stores the restored image (decoded image) or information required for decoding. Here, the restored image is the restored image that has completed the filtering process by the filtering unit 350. For example, the memory 360 stores the image used for inter prediction. At this time, the image used for inter prediction can also be specified by referring to the reference picture set or reference picture list. The restored image is used as a reference image for other images. Also, the memory 360 can output the restored image according to the output order.
[0085] Figure 4 An example of a video decoding method performed by a decoding device is shown. Refer to Figure 4 , the video decoding method includes: entropy decoding, inverse quantization, inverse transform, and intra / inter prediction processes. For example, in the decoding device, the inverse process of the encoding method is performed. Specifically, the quantized transform parameters are obtained through entropy decoding of the bitstream, and the parameter block of the current block, that is, the transform parameters, are obtained through the inverse quantization process of the quantized transform parameters. The residual block of the current block is derived through the inverse transform of the transform parameters, and the reconstructed block of the current block is derived by adding the prediction block of the current block derived through intra / inter prediction and the residual block.
[0086] In addition, the keywords in the following embodiments are defined as follows.
[0087]
Table 1
[0088] Referring to Table 1, Floor(x) shows the largest integer value less than or equal to x, Log2(u) shows the logarithm of u with base 2, and Ceil(x) shows the smallest integer value greater than or equal to x. For example, for Floor(5.93), since the largest integer value less than 5.93 is 5, it shows 5.
[0089] Also, referring to Table 1, x>>y shows the operator that shifts x to the right by y axes, and x<<y shows the operator that shifts x to the left by y axes.
[0090] <Import> The HEVC standard proposes two types of screen segmentation methods.
[0091] 1) Slice: Provides the function of dividing an image into CTU (coding tree unit) units in raster scan order for encoding / decoding, and there is slice header information.
[0092] 2) Tile: Provides a function to divide an image into multiple columns and rows in CTU units for encoding / decoding. All division methods can be divided equally or individually. There is no header for tiles.
[0093] A slice is a bit-stream packaging unit. That is, a slice is generated according to a NAL (network abstraction layer) bit stream. Figure 5 The NAL packet for slicing is composed of a NAL header, a slice header, and slice data in this order. At this time, the NAL header information contains a NAL unit type (NUT).
[0094] The NUT for slices proposed by the HEVC standard in one embodiment is shown in Table 2. In Table 2, the NUT of an inter slice for performing inter prediction is 0 to 9 times, and the NUT of an intra slice for performing intra prediction is 16 to 21 times. Here, an inter slice refers to a slice that is encoded by an inter-frame prediction method, and an intra slice refers to a slice that is encoded by a prediction method within a picture. A slice is defined in a manner having one NUT, and multiple slices within an image are set to have all the same NUT values. For example, an image is divided into four slices, and when encoded by an intra prediction method, the NUT values of the four slices in the corresponding image are all set to "19:IDR_W_RADL".
[0095]
Table 2
[0096] In Table 2, the abbreviations are defined as follows.
[0097] -TSA (Temporal sub-layer Switching Access) -STSA (Step-wise Temporalsub-layer Switching Access) -RADL (Random Access Decodable Leading) -RASL (Random Access Skipped Leading) -BLA (Broken Link Access) -IDR (Instantaneous Decoding Refresh) -CRA (Clean Random Access) -LP (Leading Picture) -_N (No reference) -_R (Reference) -_W_LP / RADL (With LP / RADL) -_N_LP (no previous image, No LP, without LP) The NUT of intra slices, namely BLA, IDR and CRA, is called IRAP (Intra Random Access Point). IRAP refers to an image that can be randomly accessed in the middle of the bitstream. In other words, it refers to an image that can suddenly change the playback position during video playback. Intra slices only exist as I slice types.
[0098] Inter-frame slices are classified as P slices or B slices according to unidirectional prediction (P: predictive) or bidirectional prediction (B: bi-predictive). The prediction and encoding process is performed in GOP (group of picture) units. The HEVC standard uses a hierarchical GOP structure to perform encoding / decoding processes including prediction. Figure 6 An example of a hierarchically divided GOP structure is shown, where each picture is divided into an I, P or B picture (slice) according to the prediction method.
[0099] Due to the characteristics of bi-directionally predicted B slices and / or hierarchical GOP structures, the decoding order and display order of pictures in a sequence are different (see Figure 7 ). Figure 7IRAP refers to intra-frame slices, B and P refer to inter-frame slices, and the complete conversion of playback order and recovery order is confirmed.
[0100] In the inter-frame slice, the picture that is delayed than IRAP in the recovery order and is played before IRAP is called LP (leading picture) (refer to Figure 8 ). It is divided into LP, RADL and RASL. When random access occurs, the LP that can be decoded is defined as RADL, and the LP that cannot be decoded during random access and should skip the recovery process of the corresponding image is called RASL. Figure 8 An image of the indicated hue is defined as one GOP.
[0101] The distinction between RADL and RASL is determined by the position of the reference image when predicting between pictures (see Fig. 9 ). That is, RASL refers to an inter-frame image that uses a restored image as a reference image in other GOPs other than the corresponding GOP, or uses a restored image as a reference image in other GOPs. In this case, using a restored image (directly / indirectly) as a reference image in other GOPs is called open GOP (open group of pictures). RASL and RADL are set according to the NUT information of the corresponding inter-frame slice.
[0102] The NUT of an intra-frame slice is distinguished into other intra-frame slice NUTs according to the NUT of an inter-frame slice preceding and / or lagging in the playback order and / or recovery order of the corresponding intra-frame slice. When observing the IDR NUT, the IDR is divided into IDR_W_RADL with RADL and IDR_N_LP without LP. That is, the IDR is a type without LP, or as a type with only RADL in LP, the IDR does not have RASL. On the contrary, CRA is a type with RADL and / or RASL in all LPs. That is, CRA refers to a type that supports open GOP.
[0103] In general, intra slices perform only intra-frame prediction without the need for reference image information for the corresponding intra slices. Here, the reference image is used for prediction between frames. However, since CRA NUT slices support the feature of the open GOP structure, although CRA slices are intra slices, reference image information can be inserted into the NAL bitstream of the corresponding CRA. The reference image information does not refer to the information used in the corresponding CRA slice, but is the information of the predetermined reference image used in the inter slice after the corresponding CRA (in the recovery order). The reference image is not removed in the DPB (decoded picture buffer). For example, in the case where the NUT of the corresponding intra slice is IDR, the DPB is reset. That is, all the restored images existing in the DPB at the corresponding time point are removed. Fig.10 FIG. 1 is a diagram showing the syntax of a slice segment title. Fig.10 As shown in the figure, when the NUT of the corresponding slice is not IDR, the reference image information is recorded in the bitstream. That is, when the NUT of the corresponding slice is not CRA, the reference image information is recorded. The invention discloses a sub-picture segmentation method for synthesis with other sequences and a slice segmentation method for bit stream packaging.
[0104] In the present invention, a slice refers to a coding / decoding area and a data packet unit for generating a NAL bit stream. For example, a picture is divided into multiple slices, and each slice is generated into a NAL data packet through a coding process.
[0105] In the present invention, a sub-picture refers to a region partitioned for synthesis with other contents. Fig.11 There are three contents: white, gray, and black. One image (AU: access unit) of each content is divided into four slices by area to generate data packets. Fig.11 As shown in the image on the right, the upper left side is synthesized with white content, the lower left side is synthesized with gray content, and the right side is synthesized with black content to generate a new image. Here, the white area and the gray area are composed of one slice to form a sub-image, and the two slices of the black area are composed of one sub-image. That is, a sub-image includes at least one slice. In order to create a new image (in order to synthesize content), BEAmer (Bit-stream Extractor And Merger) extracts areas from different contents in sub-image units and synthesizes them. Fig.11 The synthesized image in is divided into four slices and consists of three sub-images.
[0106] A sub-picture is an area with the same sub-picture ID and / or sub-picture index value. In other words, at least one slice with the same sub-picture ID and / or sub-picture index value becomes a sub-picture area. Here, the sub-picture ID and / or sub-picture index value in the slice header information is included. The sub-picture index value is set in a raster scan order. Fig.12 An example of an image consisting of six (quadrilateral) slices and four (hue) sub-image areas is shown. Here, "A", "B", "C", and "D" show an example of sub-image IDs, and "0" and "1" show the slice addresses in the corresponding sub-images. That is, the slice address value refers to the slice index value in the raster scan order in the corresponding sub-image. For example, "B-0" refers to the zeroth slice in the B sub-image, and "B-1" shows the first slice in the B sub-image.
[0107] In the present invention, the NUT values of two or more sub-images constituting an image are different. For example, Fig.12 In the image, the white sub-image (slice) is an intra-frame slice, and the gray sub-image (slice) and the black sub-image (slice) are inter-frame slices.
[0108] When synthesizing different contents, it is not necessary to identically match the NUTs of multiple sub-images constituting an image, which has the advantage of easy composition / synthesis of images. The corresponding function refers to a mixed NAL Unit Type in a single image, and in short, it can also be named mixed NUT. Set mixed_NALu_type_in_pic_flag to set the enabled / disabled status of the corresponding function. The corresponding flag is defined in one or more positions of SPS (sequence parameter set), PPS (picture parameter set), PH (picture header), and SH (slice header). For example, for the case where the corresponding flag is defined in PPS, the corresponding flag can also be named pps_mixed_NALu_types_in_pic_flag.
[0109] For the case where the flag value is disabled, (e.g., mixed_NALu_type_in_pic_flag==0), the NUT of all sub-pictures and / or slices in the corresponding picture has the same value. For example, the NUT of all VCL (video coding layer) NAL units of a picture is set to have the same value. And, a picture or picture unit (PU) is referenced with the same NUT as its corresponding coded slice NAL unit. Here, VCL refers to the NAL type of the slice containing the slice data value.
[0110] In addition, for the case where the flag value is enabled (e.g., mixed_NALu_type_in_pic_flag==1), the corresponding image is composed of more than two sub-images. And, the corresponding images have different NUT values. And, for the case where the flag value is enabled, the VCL NAL unit of the corresponding image is restricted to not have a NUT of the GDR_NUT type. And, for the case where the NUT (e.g., the first NUT) of any VCL NAL unit (e.g., the first NAL unit) of the corresponding image is any one of IDR_W_RADL, IDR_N_LP, or CRA_NUT, the NUT (e.g., the second NUT) of the other VCL NAL unit (e.g., the second NUT) of the corresponding image is restricted to be set to any one of IDR_W_RADL, IDR_N_LP, CRA_NUT, or TRAIL_NUT. For example, the second NUT is restricted to be set to a value of the first NUT or TRAIL_NUT.
[0111] Reference Fig.12 and Fig.13 , the example of the VCL NAL unit of the corresponding image having at least two different NUT values is used for explanation. In one embodiment, two or more sub-images have two or more different NUT values. In this case, the NUT values of all slices included in one sub-image are restricted to be the same. For example, Fig.13 As shown in the display, Fig.12 The NUT values of the two slices in the B sub-graph are set to CRA, the NUT values of the two slices in the C sub-graph are also set to TRAIL, and the A, B, C, and D sub-graphs are set to have two or more different NUT values. Fig.13 As shown, the NUT value of the slices in sub-images A, C, and D is TRAIL, and is set in a manner to have a different NUT value from the NUT of sub-image B, namely CRA.
[0112] In the present invention, the NUT of intra-frame slices and inter-frame slices is clear as shown in Table 3. As shown in the embodiment of Table 3, the definitions and functions of RADL, RASL, IDR, CRA, etc. can also be set in the same way as the HEVC standard (Table 1). For the case of Table 3, a mixed NUT type is added. In Table 3, an infeasible value (for example, 0) of mixed_NALu_type_in_pic_flag shows (the same as HEVC) the NUT of the slice within the image, and a feasible value (for example, 1) of mixed_NALu_type_in_pic_flag shows the NUT of the slice within the sub-image. For example, when the value of mixed_NALu_type_in_pic_flag is 0 and the NUT of the VCL NAL unit is TRAIL_NUT, the NUT of the current image is identified as TRAIL_NUT, and the NUT of other sub-images belonging to the current image is also TRAIL_NUT. Furthermore, when the value of mixed_NALu_type_in_pic_flag is 1 and the NUT of the VCL NAL unit is TRAIL_NUT, the NUT of the current sub-picture is identified as TRAIL_NUT, and it is predicted that at least one NUT in other sub-pictures belonging to the current picture is not TRAIL_NUT.
[0113]
Table 3
[0114] As described above, for a case where the value of mixed_NALu_type_in_pic_flag is shown as feasible (e.g., 1), when any VCL NAL unit belonging to an image (e.g., the first NAL unit) has any value of IDR_W_RADL, IDR_N_LP, or CRA_NUT as the NUT (e.g., the first NUT), at least one VCL NAL unit among the other VCL NAL units of the corresponding image (e.g., the second NAL unit) has any NUT value other than the first NUT among IDR_W_RADL, IDR_N_LP, CRA_NUT, or TRAIL_NUT as the NUT (e.g., the second NUT).
[0115] Thus, in the case where a VCL NAL unit belonging to a first sub-image of an image (e.g., a first NAL unit) has any one value of IDR_W_RADL, IDR_N_LP, or CRA_NUT as NUT (e.g., a first NUT), a VCL NAL unit of a second sub-image of the corresponding image (e.g., a second NAL unit) has any one NUT value of IDR_W_RADL, IDR_N_LP, CRA_NUT, or TRAIL_NUT other than the first NUT as NUT (e.g., a second NUT).
[0116] For example, when the value of mixed_NALu_type_in_pic_flag indicates activation (eg, 1), the NUT value of the VCL NAL unit of two or more sub-pictures is constructed as follows. The following is for illustrative purposes only and is not intended to be limiting.
[0117] Combination 1) IRAP+non-IRAP (inter) Combination 2) non_IRAP(inter)+non-IRAP(inter) Combination 3) IRAP+IRAP=IDR+CRA (limited by embodiment) Combination 1) is an embodiment in which at least one sub-picture in a picture has an IRAP (IDR or CRA) NUT value, and at least one other sub-picture has a non-IRAP (inter-frame slice) NUT value. Here, values other than LP (RASL and RADL) are allowed to be used as the inter-frame slice NUT value. For example, LP (RASL or RADL) is not used as the inter-frame slice NUT value. Therefore, there is no restriction on encoding RASL and RADL sub-pictures in the bitstream associated with IDR or CRA sub-pictures.
[0118] In another embodiment, only TRAIL values are allowed to be used as inter-slice NUT values. In another embodiment, all inter-slice VCL NUTs are allowed as inter-slice NUT values.
[0119] Combination 2) is an embodiment in which at least one sub-image in the image has a non-IRAP (inter-frame slice) NUT value, and at least one other sub-image has another non-IRAP (inter-frame slice) NUT value. For example, at least one sub-image has a RASL NUT value, and at least one other sub-image has a RADL NUT value. For the embodiment of combination 2, the following restrictions apply according to the embodiment.
[0120] - In one embodiment, LP (RASL and RADL) and non-LP (TRAIL) cannot be used together. For example, the NUT of at least one subgraph is RASL (or RADL), and the NUT of at least one other subgraph is not TRAIL. In the case where the NUT of at least one subgraph is RASL (or, RADL), RASL or RADL is used as the NUT of at least one other subgraph. For example, the leading subgraph of an IRAP subgraph is forced to be a RADL or RASL subgraph.
[0121] - In another embodiment, LP (RASL and RADL) and non-LP (TRAIL) are used together. For example, at least one subgraph is RASL (or RADL), and at least one other subgraph is TRAIL.
[0122] - In another embodiment, except for the case of condition 2), all sub-images have the same inter-slice NUT value. For example, all sub-images in the image have the TRAIL NUT value. In another example, all sub-images in the image have the RASL (or RADL) NUT value.
[0123] Combination 3) is an embodiment in which all sub-images or slices in the display image are composed of IRAP. For example, when the NUT value of the slice in the first sub-image is IDR_W_RADL, IDR_N_LP or CRA_NUT, the NUT value of the slice in the second sub-image is composed of the value of IDR_W_RADL, IDR_N_LP and CRA_NUT that is not the NUT value of the first sub-image. For example, the NUT value of the slice in at least one sub-image is IDR, and the NUT value of the slice in at least another sub-image is composed of CRA.
[0124] In addition, according to an embodiment, the embodiment shown in combination 3 is limited to be applicable. In one embodiment, the images to which the IRAP or GDR access unit belongs are limited to have all the same NUT. That is, for the case where the current access unit is an IRAP access unit consisting only of IRAP images, or the current access unit is a GDR access unit consisting only of GDR images, the images to which they belong are limited to have all the same NUT. For example, the NUT value of the slice in at least one sub-image is limited to IDR, and the NUT value of the slice in at least another sub-image cannot be composed of CRA. Therefore, combination 3) is limited, and for the case where the aforementioned combination 1) and combination 2 are applicable, at least one sub-image in the corresponding image is limited to have a NUT value of non-IRAP (inter-frame slice). For example, during the encoding and decoding process, it is limited so that all sub-images in the corresponding image have a NUT value of IDR. Or a part of the sub-images in the corresponding image are limited to have a NUT value of IDR, and another sub-image does not have a CRA NUT value.
[0125] Next, the syntax and semantics of the signaling of the coded information for the case where the mixed NAL unit type (NUT) in the picture is applied are described. In addition, the decoding process using it is described. As shown above, for the case where mixed_NALu_type_in_pic_flag=1, the picture to which the NUT belongs refers to a sub-picture (see Table 3).
[0126] In addition, as described above, for the case where the value of mixed_NALu_type_in_pic_flag indicates that mixed NUT is applicable, one image is divided into at least two sub-images. Thus, information of the sub-images of the corresponding image is signaled through the bitstream. In this regard, mixed_NALu_type_in_pic_flag indicates whether the current image is divided. For example, for the case where the value of mixed_NALu_type_in_pic_flag indicates that mixed NUT is applicable, the current image is displayed as being divided.
[0127] Below, refer to Fig.14 The syntax of . Fig.14 A diagram showing an embodiment of the syntax of a picture parameter set (PPS). For example, a flag (e.g., pps_no_pic_partition_flag) indicating whether the current picture is not partitioned is signaled through a picture parameter set (PPS) through a bitstream. A possible value (e.g., 1) of pps_no_pic_partition_flag is displayed to indicate that partitioning of the picture is not applicable for pictures referring to the current PPS. An infeasible value (e.g., 0) of pps_no_pic_partition_flag is displayed to indicate that partitioning of the picture using slices or tiles is applicable for pictures referring to the current PPS. In this embodiment, for the case where the value of mixed_NALu_type_in_pic_flag indicates that mixed NUT is applicable, the value of pps_no_pic_partition_flag is forced to display an infeasible value (e.g., 0).
[0128] When pps_no_pic_partition_flag indicates that the current picture is partitioned, information on the number of subpictures (e.g., pps_num_subpics_minus1) is obtained from the bitstream. pps_num_subpics_minus1 indicates the value of the number of subpictures contained in the current picture minus 1. When pps_no_pic_partition_flag indicates that the current picture is not partitioned, the value of pps_num_subpics_minus1 is not obtained from the bitstream and is guided to 0. Based on the information on the number of subpictures determined in this way, coding information of each subpicture is signaled according to the number of subpictures included in one picture. For example, a subpicture identifier (e.g., pps_subpic_id) for identifying each subpicture and / or a flag (subpic_treated_as_pic_flag[i]) value indicating whether the coding / decoding process of each subpicture is independent is specified and signaled.
[0129] Hybrid NUT is applicable when one image is composed of two or more sub-images. In this case, the value of the flag (subpic_treated_as_pic_flag[i]) that indicates whether the encoding / decoding process of each sub-image is independent is specified according to the number of sub-images (i) included in one image, and signaling is sent. In the case where a sub-image is decoded independently, the corresponding sub-image is decoded as another image process. That is, in the case where the corresponding flag value is "on" (for example, subpic_treated_as_pic_flag=1), the corresponding sub-image is decoded independently from other sub-images in all decoding processes except the in-loop filter process. On the contrary, in the process where the corresponding flag value is "off" (for example, subpic_treated_as_pic_flag=0), the corresponding sub-image refers to other sub-images in the image during the inter-frame prediction process. Here, for the in-loop filter process, another flag is set to control whether it is independent or referenced. The corresponding flag (subpic_treated_as_pic_flag) is defined in one or more of the SPS, PPS, and PH. For example, in the case where the corresponding flag is defined in the SPS, the corresponding flag is named sps_subpic_treated_as_pic_flag.
[0130] Furthermore, in the present invention, in the case where other NUTs exist in one picture (for example, mixed_NALu_type_in_pic_flag=1), due to the characteristics of two NUTs to be used between sub-pictures in one picture, each sub-picture in the picture is independently encoded / decoded. For example, in the case of a picture with mixed_NALu_type_in_pic_flag=1, when the corresponding picture contains more than one inter-frame (P or B) slice, the subpic_treated_as_pic_flag value of all sub-pictures in the corresponding filter is forcibly set to "1" or guided by a "1" value. Or in the case of mixed_NALu_type_in_pic_flag=1, subpic_treated_as_pic_flag is forced so as not to have a "0" value. For example, in the case of a picture with mixed_NALu_type_in_pic_flag=1, if one or more inter slices are included in the corresponding picture, the subpic_treated_as_pic_flag value is reset to "1" for all subpictures of the corresponding picture, regardless of the issued value. Conversely, in the case of a picture with mixed_NALu_type_in_pic_flag=1 and subpic_treated_as_pic_flag=0, no inter slices are included in the corresponding picture. That is, in the case of a picture with mixed_NALu_type_in_pic_flag=1 and subpic_treated_as_pic_flag=0, the slice type in the corresponding picture is intra.
[0131] Furthermore, in another embodiment, for the case where mixed_NALu_type_in_pic_flag=1, when the NUT of the current image is RASL, the subpic_treated_as_pic_flag of the current image is forcibly set to "1". By way of another example, for the case where mixed_NALu_type_in_pic_flag=1, when the NUT of the current image is RADL and the NUT of the referenced image is RASL, the subpic_treated_as_pic_flag of the current image is forcibly set to "1".
[0132] Mixed NUT function All sub-images (or slices) in one picture are composed of IRAP. At this time, all slices in one picture are composed of IRAP, or the flag (gdr_or_irap_pic_flag) value of the corresponding picture displayed as a GDR (Gradual Decoding Refresh) picture is forced to be "0". That is, in the present invention, for the case where other NUTs exist in one picture (mixed_NALu_type_in_pic_flag=1), the flag (gdr_or_irap_pic_flag) value is set to "0" or guided by the "0" value. Or for the case where mixed_NALu_type_in_pic_flag=1, gdr_or_irap_pic_flag is forced not to have a "1" value. The flag (gdr_or_irap_pic_flag) can be defined in more than one position in SPS, PPS and PH.
[0133] Furthermore, the hybrid NUT function has at least one sub-image in one picture having an IRAP (IDR or CRA) NUT value, and at least one other sub-image having a non-IRAP (inter slice) NUT value, depending on the application. That is, intra slices and inter slices exist simultaneously in one picture. In the case of the existing HEVC standard, when the NUT of the corresponding intra slice is IDR, a DPB is set. Thus, all restored images existing in the DPB at the corresponding time point are removed.
[0134] However, for the case of the present invention, for the case where mixed_NALu_type_in_pic_flag=1, intra-frame slices and inter-frame slices can exist simultaneously within an image, and even if an image is an IDR NUT, there is a situation where the DPB cannot be reset. Therefore, in one embodiment, for the case where the corresponding slice is an IDR NUT, such as CRA, the reference picture information (RPL: reference picture list) is inserted into the NAL bit stream as the slice header information of the corresponding IDR. For this reason, although it is an IDR NUT, the flag (idr_rpl_present_flag) value that indicates the existence of RPL information is also set to "1". For the case where the flag (idr_rpl_present_flag) value is "1", there is RPL as the slice header information of the IDR. On the contrary, for the case where the flag (idr_rpl_present_flag) value is "0", there is no RPL as the slice header information of the IDR.
[0135] In addition, in the present invention, there are other NUTs in one picture, and (mixed_NALu_type_in_pic_flag=1), for the case where the use of RPL information of the IDR picture is not allowed (idr_rpl_present_flag=0), the NUT of the corresponding picture does not have an IDR_W_RADL or IDR_N_LP value.
[0136] The flag (idr_rpl_present_flag) is defined in one or more of the SPS, PPS, and PH. For example, when the flag is defined in the SPS, the flag is named sps_idr_rpl_present_flag. For example, even if the NUT of the current slice is IDR_W_RADL or IDR_N_RADL, in order to send RPL signaling according to the value of sps_idr_rpl_present_flag, Fig.15 Here, a first value (e.g., 0) of sps_idr_rpl_present_flag indicates that no RPL syntax elements are provided through the slice header of a slice whose NUT is IDR_N_LP or IDR_W_RADL. A second value (e.g., 1) of sps_idr_rpl_present_flag indicates that RPL syntax elements are provided through the slice header of a slice whose NUT is IDR_N_LP or IDR_W_RADL.
[0137] In addition, in another embodiment, for the case where mixed_NALu_type_in_pic_flag=1, RPL sends signaling according to the picture header information. Fig.14 The syntax of the application, for the case where the value of mixed_NALu_type_in_pic_flag indicates that mixed NUT is applicable, the value of pps_no_pic_partition_flag is forced to a value indicating that it is not feasible (e.g., 0). And, thereby, the value of the flag (pps_rpl_info_in_ph_flag) indicating whether RPL information is provided from the picture header is obtained from the bitstream. In the case where pps_rpl_info_in_ph_flag indicates that it is feasible (e.g., 1), the RPL information is as follows Fig.16 and Fig.17As shown in the display, it can be obtained from the image header. Thus, based on the value of mixed_NALu_type_in_pic_flag, RPL information is obtained regardless of the type of the corresponding image. However, when pps_rpl_info_in_ph_flag is displayed as not feasible (for example, 0), RPL information cannot be obtained from the image header. For example, when the value of pps_rpl_info_in_ph_flag is "0", the slice NUT is IDR_N_LP or IDR_W_RADL, and the value of sps_idr_rpl_present_flag is "0", the RPL information of the corresponding slice cannot be obtained. That is, since there is no RPL information for the corresponding slice, the RPL information is initialized and guided to empty.
[0138] As shown above, one image is signaled by dual NAL units. Therefore, in order to signal one image, a method of determining the type of image according to the type of NAL unit is required in which NAL units with different NUTs are used. Therefore, during random access (RA), the corresponding image is restored normally and whether it can be output (output) is determined.
[0139] In the decoding process of an embodiment, when each VCL NAL unit corresponding to a picture is a NAL unit of the CRA_NUT type, the corresponding picture is determined to be a CRA picture. Also, when each VCL NAL unit corresponding to a picture is a NAL unit of the IDR_W_RADL type or the IDR_N_LP type, the corresponding picture is determined to be an IDR picture. Also, when each VCL NAL unit corresponding to a picture is a NAL unit of the IDR_W_RADL type, the IDR_N_LP type, or the CRA_NUT type, the corresponding picture is determined to be an IRAP picture.
[0140] Furthermore, for a case where each VCL NAL unit corresponding to an image is a NAL unit of the RADL_NUT type, the corresponding image is determined as a RADL (Random Access decodable leading) image. Furthermore, for a case where each VCL NAL unit corresponding to an image is a NAL unit of the TRAIL_NUT type, the corresponding image is determined as a trailing image. Furthermore, for a case where, among the VCL NAL units corresponding to an image, at least one VCL NAL unit is of the RASL_NUT type, and all other VCL NAL units are of the RASL_NUT type or the RADL_NUT type, the corresponding image is determined as a RASL (random access skipped leading) image.
[0141] In addition, in the decoding process of another embodiment, for a case where at least one sub-image in an image is RASL and at least one other sub-image is RADL, the corresponding image is determined to be a RASL image. For example, for a case where at least one sub-image in an image is RASL and at least one other sub-image is RADL, the corresponding image in the decoding process is determined to be a RASL image. Here, when the type of the VCL NAL unit corresponding to the sub-image is RASL_NUT, the corresponding sub-image is determined to be RASL. Therefore, when RA is used, the RASL sub-image and the RADL sub-image are all processed as RASL images, and thus, the corresponding image is not output.
[0142] In addition, in the decoding process of another embodiment, when at least one sub-image in an image is RASL, the corresponding image is set as a RASL image. For example, when at least one sub-image in an image is RASL and at least one other sub-image is TRAIL, the corresponding image in the decoding process is set as a RASL image. Thus, during RA, the corresponding image is obtained as a RASL image and is not output.
[0143] Here, RA occurrence is determined by the NoOutputBeforeRecoveryFlag value of the IRAP picture connected (associated) to the corresponding inter-frame slice (RADL, RASL, or TRAIL). When the corresponding flag value is "1", (NoOutputBeforeRecoveryFlag = 1), it means that RA occurs, and when the corresponding flag value is "0", (NoOutputBeforeRecoveryFlag = 0), it means normal playback. The corresponding flag value is set as follows for IRAP.
[0144] -When the current image is IRAP, the NoOutputBeforeRecoveryFlag value setting process ① If the image is the first image in the bitstream, set NoOutputBeforeRecoveryFlag to "1" ② If the image is IDR, set NoOutputBeforeRecoveryFlag to "1" ③ When the image is CRA and RA is notified from the outside, set NoOutputBeforeRecoveryFlag to "1" ④ If the image is CRA and RA is not notified from the outside, set NoOutputBeforeRecoveryFlag to "0" In one embodiment, the decoding device obtains a signaling of the occurrence of random access from the external terminal. For example, the external terminal sets the value of the random access occurrence information to 1 and sends a signaling to the decoding device, thereby sending a signaling of the occurrence of random access to the decoding device. The decoding device sets the value of the flag HandleCraAsClvsStartFlag indicating whether the random access is received from the external terminal to 1 according to the random access occurrence information received from the external terminal. The decoding device sets the value of NoOutputBeforeRecoveryFlag to the same value as the value of HandleCraAsClvsStartFlag. Thus, when the current image is a CRA image and the value of HandleCraAsClvsStartFlag is "1", the decoding device determines that a random access has occurred for the corresponding CRA image, or performs decoding with the corresponding CRA being in the initial processing of the bit stream.
[0145] At RA, the process of setting the flag (PictureOutputFlag) of whether to output the current image is as follows. For example, the PictureOutputFlag of the current image is set according to the following sequence. Here, the first value (for example, 0) of PictureOutputFlag indicates that the current image is not output. The second value (for example, 1) of PictureOutputFlag indicates that the current image is output.
[0146] (1) If the current image is RASL and the NoOutputBeforeRecoveryFlag of the associated IRAP image is "1", set PictureOutputFlag to "0" (2) For the current image, the value of NoOutputBeforeRecoveryFlag is "1", i.e., for a GDR image, or for the recovered image, PictureOutputFlag is set to "0" (3) Otherwise, the value of PictureOutputFlag is set to the same value as the value of pic_output_flag in the bitstream. Here, pic_output_flag can be obtained from one or more positions of PH and SH.
[0147] Fig.18 This is an example showing the synthesis of three different contents disclosed in the present invention. Fig.18 - (a) A sequence of three different contents is displayed. For convenience, one image is displayed as one data packet, or one image is divided into multiple slices and multiple data packets are displayed. Fig.18 -(b) and Fig.18 -(c) Displayed in Fig.18 -The image result of the synthesis of the image shown by the dotted line in (a). Fig.18 The same hue in refers to the same picture / sub-picture / slice. Also, P slices and B slices have one value in the inter-frame NUT.
[0148] As described above, when a plurality of contents are synthesized by the present invention, there is no need to also match the positions of slices (images) within a frame, and the contents can be synthesized quickly and easily without delay by simply matching the hierarchically divided GOP structure.
[0149] Encoding and decoding embodiments Next, a method for decoding an image by the image decoding device will be described based on the above method. Fig.19 and Fig. 20 A sequence diagram illustrating a decoding method and an encoding method according to an embodiment of the present invention is shown.
[0150] An image decoding device of an embodiment includes a memory and at least one processor, and the following decoding method is executed by the operation of the processor: First, the decoding device obtains NAL unit type information indicating the type of the current NAL (network abstraction layer) unit from a bit stream (S1910).
[0151] Next, when the NAL unit type information indicates that the NAL unit type of the current NAL unit is coded data of a video slice, the decoding apparatus decodes the video slice based on whether a mixed NAL unit type is applicable to the current image ( S1920 ).
[0152] Here, the decoding apparatus determines whether the NAL unit type of the current NAL unit indicates the attribute of the sub-picture of the current video slice based on whether the mixed NAL unit type is applicable, and performs decoding of the video slice.
[0153] Whether the mixed NAL unit type is applicable or not is identified based on the first flag (eg, pps_mixed_NALu_types_in_pic_flag) obtained from the picture parameter set. In the case where the mixed NAL unit type is applicable, the current picture to which the current image slice belongs is divided into at least two sub-pictures.
[0154] Furthermore, based on whether the mixed NAL unit type is applicable, the decoding information of the sub-picture is included in the bitstream. In one embodiment, a second flag (e.g., pps_no_pic_partition_flag) indicating whether the current picture is not partitioned is obtained from the bitstream. And, in the case where the second flag indicates that the current picture is partitioned (e.g., pps_no_pic_partition_flag==0), a third flag (e.g., pps_rpl_info_in_ph_flag) indicating whether the reference picture list information is provided from the picture header can be obtained from the bitstream.
[0155] In this example, in the case where the mixed NAL unit type is applied, the current picture is forcibly partitioned into at least two sub-pictures, the value of the second flag (pps_no_pic_partition_flag) is forced to be 0, and the third flag (for example, pps_rpl_info_in_ph_flag) indicating whether the reference picture list information is provided from the picture header can be obtained from the bitstream regardless of the value of the second flag (pps_no_pic_partition_flag) actually obtained from the bitstream. Therefore, in the case where the third flag indicates that the reference picture list information is provided from the picture header (for example, pps_rpl_info_in_ph_flag==1), the reference picture list information can be obtained from the bitstream associated with the picture header.
[0156] Furthermore, in the case where a mixed NAL unit type is applicable, the current picture is decoded based on the first sub-picture and the second sub-picture having different NAL unit types. Here, in the case where the NAL unit type of the first sub-picture has any value of IDR_W_RADL (Instantaneous Decoding Refresh_With_Random Access Decodable Leading), IDR_N_LP (Instantaneous Decoding Refresh_No reference_Leading Picture), and CRA_NUT (Clean Random Access_NAL Unit Type), the available NAL unit type selected by the second sub-picture NUT includes the NAL unit type that is not selected in the first sub-picture from IDR_W_RADL, IDR_N_LP, and CRA_NUT.
[0157] Alternatively, for the case where the NAL unit type of the first sub-image has any one of the values of IDR_W_RADL (Instantaneous Decoding Refresh_With_Random Access DecodableLeading), IDR_N_LP (Instantaneous Decoding Refresh_Noreference_LeadingPicture) and CRA_NUT (Clean Random Access_NAL Unit Type), the available NAL unit types of the second sub-image include TRAIL_NUT (Trail_NAL Unit Type).
[0158] In addition, in the case where a mixed NAL unit type is applicable, the first sub-image and the second sub-image constituting the current image can be decoded independently. For example, the first sub-image and the second sub-image including a B or P slice can be obtained from one image and decoded. For example, the first sub-image does not use the second sub-image as a reference image and is decoded.
[0159] More specifically, the first sub-image is obtained from the bitstream by indicating whether the fourth flag (e.g., sps_subpic_treated_as_pic_flag) is processed as an image during decoding. In the case where the first sub-image is processed as an image during decoding and the fourth flag is indicated (e.g., sps_subpic_treated_as_pic_flag==1), the first sub-image is decoded as an image during decoding. In this process, when the mixed NAL unit type is applied to the current image and the current image including the first sub-image includes at least one P slice or B slice, the fourth flag is forced to have a value indicating that the first sub-image is processed as an image during decoding. When the mixed NAL unit type is applied to the current image and the fourth flag indicates that the first sub-image is not processed as an image during decoding (e.g., sps_subpic_treated_as_pic_flag==0), the slice type to which the current image belongs should be intra.
[0160] When the fourth flag indicates that the first sub-image is processed as an image during the decoding process, it is determined that the decoding process of the first sub-image is independent of other sub-images. For example, when the fourth flag indicates that the first sub-image is decoded independently of other sub-images during the decoding process, the first sub-image is decoded without using other sub-images as reference images.
[0161] Furthermore, in the case where the first sub-image is a RASL (Random Access Skipped Leading) sub-image, the second sub-image is determined to be a RASL (Random Access Decodable Leading) sub-image based on whether it is a RADL (Random Access Decodable Leading) sub-image. Here, in the case where the type of the NAL unit corresponding to the first sub-image is RASL_NUT (Random Access Skipped Leading_NAL Unit Type), the first sub-image is determined to be a RASL sub-image.
[0162] Furthermore, the third flag (e.g., pps_rpl_info_in_ph_flag) indicates that the reference picture list information is not obtained from the picture header, but is obtained from the slice header (e.g., pps_rpl_info_in_ph_flag==0), and in the case where the NAL unit type of the first sub-picture has any value of IDR_W_RADL (Instantaneous Decoding Refresh_With_Random Access Decodable Leading) and IDR_N_LP (Instantaneous Decoding Refresh_Noreference_Leading Picture), the reference picture list information can be obtained from the bitstream of the slice header based on whether the reference picture list information of the IDR picture is present in the fifth flag (egsps_idr_rpl_present_flag) of the slice header. Here, the fifth flag can be obtained from the bitstream associated with the sequence parameter set.
[0163] In addition, in the case where random access is performed on an IRAP (Intra Random Access Point) image associated with the current image, if the current image is a RASL (Random Access Skipped Leading) sub-image, the current image is not output (displayed).
[0164] An image encoding device of one embodiment includes a memory and at least one processor, and executes an encoding method corresponding to the above-mentioned decoding method through the operation of the processor. For example, the encoding device determines the type of the sub-image of the segmented image (S2010) when the current image is encoded based on a mixed NAL unit type. And, the encoding device encodes at least one current image slice constituting the sub-image based on the type of the sub-image, thereby generating a current NAL unit (S2020). At this time, when the current image is encoded based on a mixed NAL unit type, the encoding device encodes the NAL unit type of the current NAL unit to display the attributes of the sub-image of the current image slice, and encodes the image slice.
[0165] Furthermore, the present invention is implemented as a code that can be read by a computer (including a device having an information processing function) through a computer-readable recording medium. Computer-readable recording media include all kinds of recording devices that store data read by a computer system. Examples of computer-readable recording devices include ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage device, etc.
[0166] The present invention is described with reference to the embodiments shown in the accompanying drawings, but they are only for illustration, and those skilled in the art should understand that various modifications and other equivalent embodiments can be made therefrom. Therefore, the true technical protection scope of the present invention should be defined by the technical concept of the claims.
Claims
1. An image decoding method performed by an image decoding device, comprising the following steps: acquiring a first flag indicating whether the current picture includes sub-pictures having mutually different NAL unit types; and acquiring second flags respectively corresponding to the sub-images, wherein each of the second flags indicates whether the corresponding sub-image is processed as an image during decoding, in, When the first flag indicates that the current picture includes sub-pictures having mutually different NAL unit types, the second flag for all sub-pictures including P slices or B slices among the sub-pictures is restricted to have a first value, and the VLC NAL unit of the current picture is restricted to have a NAL unit type different from GDR_NUT, The first value indicates that the corresponding sub-image is processed as an image during the decoding process.
2. An image decoding device, comprising: a memory configured to store one or more instructions; as well as The processor is configured as: Acquire a first flag indicating whether a current picture includes sub-pictures having mutually different NAL unit types; and acquiring second flags respectively corresponding to the sub-images, wherein each of the second flags indicates whether the corresponding sub-image is processed as an image during decoding, wherein, when the first flag indicates that the current picture includes sub-pictures having mutually different NAL unit types, the second flag for all sub-pictures including P slices or B slices in the sub-pictures is restricted to have a first value, and the VCL NAL unit of the current picture is restricted to have a NAL unit type different from that of the GDR NUT, The first value indicates that the corresponding sub-image is processed as an image during the decoding process.
3. A video encoding method performed by a video encoding device, comprising the following steps: determining whether a current picture includes sub-pictures having mutually different NAL unit types; determining whether the sub-image is to be processed as an image during decoding; and generating a first flag indicating whether the current image includes sub-pictures having mutually different NAL unit types and second flags respectively corresponding to the sub-pictures, in, Each of the second flags indicates whether the corresponding sub-picture is processed as an image during decoding, wherein, when the first flag indicates that the current picture includes sub-pictures having mutually different NAL unit types, the second flag for all sub-pictures including at least one of a P slice or a B slice in the sub-picture is restricted to have a first value, and the VCL NAL unit of the current picture is restricted to have a NAL unit type different from a GDR NUT, The first value indicates that the sub-image is processed as an image during the decoding process.
4. A method for transmitting a bit stream, comprising the steps of: determining whether the current picture contains sub-pictures having mutually different NAL unit types; Determine whether the sub-image is processed as an image during decoding; generating a first flag indicating whether the current image includes sub-pictures having mutually different NAL unit types and second flags respectively corresponding to the corresponding sub-pictures; and transmitting a bit stream including the first flag and the second flag, in, Each of the second flags indicates whether the corresponding sub-picture is processed as an image during decoding, wherein, when the first flag indicates that the current picture includes sub-pictures having mutually different NAL unit types, the second flag for all sub-pictures including at least one of a P slice or a B slice in the sub-picture is restricted to have a first value, and the VCL NAL unit of the current picture is restricted to have a NAL unit type different from a GDR NUT, The first value indicates that the sub-image is processed as an image during the decoding process.