Image encoding / decoding method and apparatus for signaling information related to a sub-picture and a picture header, and method for transmitting a bitstream
By efficiently signaling information about sub-pictures and picture heads in the image encoding and decoding methods, the problem of low high-resolution and high-quality image encoding/decoding efficiency in the prior art is solved, and more efficient image transmission and storage are achieved.
Patent Information
- Application Number
- CN202180019325.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-01-14
- Filing Date
- 2021-01-14
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2041-01-14
AI Technical Summary
The prior art is difficult to effectively improve the encoding/decoding efficiency of high resolution and high-quality images, resulting in increased transmission and storage costs.
By efficiently signaling information about the sub-picture and the picture head in the image encoding and decoding methods, encoding/decoding efficiency is improved. The specific implementation includes including the signs about the sub-picture and the picture header information in the slice header in the bitstream, thereby optimizing the encoding and decoding process.
Improves the efficiency of image encoding/decoding, reduces transmission and storage costs, and supports efficient transmission and storage of bitstreams generated by the image encoding method.
Smart Images

Figure CN115244937B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an image encoding / decoding method and apparatus, and more particularly, to an image encoding and decoding method and apparatus for signaling information about a sub-picture and a picture header, and a method for transmitting a bit stream generated by the image encoding method / apparatus of the present disclosure. Background Art
[0002] Recently, the demand for high-resolution and high-quality images, such as high-definition (HD) images and ultra-high-definition (UHD) images, is increasing in various fields. As the resolution and quality of image data are improved, the amount of information or bits transmitted is relatively increased compared to existing image data. The increase in the amount of information or bits transmitted leads to an increase in transmission cost and storage cost.
[0003] Therefore, efficient image compression technology is needed to effectively transmit, store, and reproduce information about high-resolution and high-quality images. Summary of the invention
[0004] Technical issues
[0005] An object of the present disclosure is to provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.
[0006] Another object of the present disclosure is to provide an image encoding / decoding method and apparatus for improving encoding / decoding efficiency by efficiently signaling information about a sub-picture and a picture header.
[0007] Another object of the present disclosure is to provide a method for transmitting a bit stream generated by the image encoding method or apparatus according to the present disclosure.
[0008] Another object of the present disclosure is to provide a recording medium storing a bit stream generated by the image encoding method or apparatus according to the present disclosure.
[0009] Another object of the present disclosure is to provide a recording medium storing a bit stream received and decoded by the image decoding apparatus according to the present disclosure and used to reconstruct an image.
[0010] The technical problems solved by the present disclosure are not limited to the above-mentioned technical problems, and other technical problems not described here will be clear to those skilled in the art through the following description.
[0011] Technical Solution
[0012] An image decoding method performed by an image decoding device according to an aspect of the present disclosure may include: obtaining a first flag specifying whether information about a sub-picture exists in a bitstream; obtaining a second flag specifying whether picture header information exists in a slice header; and decoding the bitstream based on the first flag and the second flag. When the first flag specifies that information about a sub-picture exists in the bitstream, the second flag may have a value specifying that picture header information does not exist in the slice header.
[0013] In the image decoding method according to the present disclosure, when the first flag specifies that information about a sub-picture exists in the bitstream, the slice header may include an identifier of the sub-picture including a slice related to the slice header.
[0014] The image decoding method according to the present disclosure may further include: when the second flag specifies that the picture header information exists in the slice header, acquiring the picture header information from the slice header.
[0015] In the image decoding method according to the present disclosure, the second flag may have the same value with respect to all slices in a coding layer video sequence (CLVS).
[0016] In the image decoding method according to the present disclosure, when the second flag specifies that picture header information exists in the slice header, a network abstraction layer (NAL) unit for transmitting the picture header information may not exist in the coding layer video sequence (CLVS).
[0017] In the image decoding method according to the present disclosure, when the second flag specifies that picture header information does not exist in the slice header, the picture header information may be acquired from a network abstraction layer (NAL) unit whose NAL unit type is equal to PH_NUT.
[0018] In the image decoding method according to the present disclosure, the first flag may be signaled at a higher level of a slice, and the second flag may be included in a slice header and signaled.
[0019] An image decoding device according to another aspect of the present disclosure may include a memory and at least one processor. The at least one processor may be configured to: obtain a first flag specifying whether there is information about a sub-picture in a bitstream; obtain a second flag specifying whether there is picture header information in a slice header; and decode the bitstream based on the first flag and the second flag. When the first flag specifies that there is information about a sub-picture in the bitstream, the second flag may have a value specifying that there is no picture header information in the slice header.
[0020] According to another aspect of the present disclosure, an image encoding method may include: encoding a first flag specifying whether information about a sub-picture exists in a bitstream; encoding a second flag specifying whether picture header information exists in a slice header; and encoding the bitstream based on the first flag and the second flag. When the first flag specifies that information about a sub-picture exists in the bitstream, the second flag may have a value specifying that picture header information does not exist in the slice header.
[0021] In the image encoding method according to the present disclosure, when the first flag specifies that information about a sub-picture exists in the bitstream, the slice header may include an identifier of the sub-picture including a slice related to the slice header.
[0022] The image encoding method according to the present disclosure may further include: when the second flag specifies that the picture header information exists in the slice header, encoding the picture header information in the slice header.
[0023] In the image encoding method according to the present disclosure, the second flag may have the same value with respect to all slices in a coding layer video sequence (CLVS).
[0024] In the image encoding method according to the present disclosure, when the second flag specifies that picture header information does not exist in the slice header, the picture header information may be signaled through a network abstraction layer (NAL) unit whose NAL unit type is equal to PH_NUT.
[0025] In the image encoding method according to the present disclosure, the first flag may be signaled at a higher level of a slice, and the second flag may be included in a slice header and signaled.
[0026] In addition, a transmission method according to another aspect of the present disclosure may transmit a bit stream generated by the image encoding device or the image encoding method of the present disclosure.
[0027] In addition, a computer-readable recording medium according to another aspect of the present disclosure may store a bit stream generated by the image encoding device or the image encoding method of the present disclosure.
[0028] The features described above in brief summary of the present disclosure are merely exemplary aspects of the following detailed description of the present disclosure and do not limit the scope of the present disclosure.
[0029] Beneficial Effects
[0030] According to the present disclosure, an image encoding / decoding method and apparatus with improved encoding / decoding efficiency may be provided.
[0031] In addition, according to the present disclosure, an image encoding / decoding method and apparatus that improves encoding / decoding efficiency by efficiently signaling information about a sub-picture and a picture header may be provided.
[0032] In addition, according to the present disclosure, a method of transmitting a bit stream generated by the image encoding method or apparatus according to the present disclosure may be provided.
[0033] In addition, according to the present disclosure, a recording medium storing a bit stream generated by the image encoding method or apparatus according to the present disclosure may be provided.
[0034] In addition, according to the present disclosure, there may be provided a recording medium storing a bit stream received and decoded by the image decoding apparatus according to the present disclosure and used to reconstruct an image.
[0035] Those skilled in the art will understand that the effects that can be achieved through the present disclosure are not limited to what has been specifically described above, and other advantages of the present disclosure will be more clearly understood from the detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 FIG. 4 is a diagram schematically showing a video encoding system to which an embodiment of the present disclosure is applicable.
[0037] Figure 2 is a view schematically showing an image encoding device to which an embodiment of the present disclosure is applicable.
[0038] Figure 3 is a view schematically showing an image decoding device to which an embodiment of the present disclosure is applicable.
[0039] Figure 4 is a view showing a segmentation structure of an image according to an embodiment.
[0040] Figure 5 is a view showing an embodiment of a partition type of a block according to a multi-type tree structure.
[0041] Figure 6 is a diagram illustrating a signaling mechanism of block partitioning information in a quadtree having a nested multi-type tree structure according to the present disclosure.
[0042] Figure 7 is a diagram showing an example of partitioning a CTU into a plurality of CUs.
[0043] Figure 8 : is a flowchart illustrating a method in which an image encoding apparatus encodes an image using slices / tiles according to an embodiment of the present disclosure.
[0044] Fig. 9 The flowchart illustrates a method for decoding an image using slices / tiles by an image decoding apparatus according to an embodiment of the present disclosure.
[0045] Fig.10 is a diagram of an example of the present disclosure showing signaling and syntax elements in a picture header.
[0046] Fig.11 is a diagram showing a syntax structure of a slice header according to an embodiment of the present disclosure.
[0047] Fig.12 This example illustrates parsing and decoding Fig.11 Flowchart of the method of slicing header.
[0048] Fig.13 is an example of Fig.11 Flowchart of a method for encoding a slice header.
[0049] Fig.14 is a diagram showing a syntax structure of a slice header according to another embodiment of the present disclosure.
[0050] Fig.15 This example illustrates parsing and decoding Fig.14 Flowchart of the method of slicing header.
[0051] Fig.16 is an example of Fig.14 Flowchart of a method for encoding a slice header.
[0052] Fig.17 is a diagram showing a syntax structure of a slice header according to another embodiment of the present disclosure.
[0053] Fig.18 This example illustrates parsing and decoding Fig.17 Flowchart of the method of slicing header.
[0054] Fig.19 is an example of Fig.17 Flowchart of a method for encoding a slice header.
[0055] Fig. 20 is a diagram showing a content streaming system to which an embodiment of the present disclosure is applicable. DETAILED DESCRIPTION
[0056] Hereinafter, the embodiments of the present disclosure will be described in detail with reference to the accompanying drawings to facilitate implementation by those skilled in the art. However, the present disclosure can be implemented in various forms and is not limited to the embodiments described herein.
[0057] When describing the present disclosure, if it is determined that the detailed description of related known functions or configurations makes the scope of the present disclosure unnecessarily ambiguous, the detailed description thereof will be omitted. In the drawings, parts irrelevant to the description of the present disclosure are omitted, and like reference numerals are given to like parts.
[0058] In the present disclosure, when a component is "connected", "coupled" or "linked" to another component, it may include not only a direct connection relationship but also an indirect connection relationship with intermediate components. In addition, when a component "includes" or "has" other components, unless otherwise specified, it means that other components may also be included, rather than excluding other components.
[0059] In the present disclosure, the terms first, second, etc. are used only for the purpose of distinguishing one component from other components, and do not limit the order or importance of the components unless otherwise specified. Accordingly, within the scope of the present disclosure, the first component in one embodiment may be referred to as the second component in another embodiment, and similarly, the second component in one embodiment may be referred to as the first component in another embodiment.
[0060] In the present disclosure, components that are distinguished from each other are intended to clearly describe each feature and do not mean that the components must be separated. That is, multiple components can be integrated and implemented in one hardware or software unit, or one component can be distributed and implemented in multiple hardware or software units. Therefore, even if not specifically stated, implementations in which these components are integrated or distributed are also included in the scope of the present disclosure.
[0061] In the present disclosure, the components described in each embodiment are not necessarily indispensable components, and some components may be optional components. Therefore, the embodiments consisting of a subset of the components described in the embodiments are also included in the scope of the present disclosure. In addition, the embodiments including other components in addition to the components described in the various embodiments are included in the scope of the present disclosure.
[0062] The present disclosure relates to encoding and decoding of images. Unless otherwise defined in the present disclosure, terms used in the present disclosure may have general meanings commonly used in the technical field to which the present disclosure belongs.
[0063] In the present disclosure, a "picture" generally refers to a unit representing an image within a specific time period, and a slice / tile is a coding unit that constitutes a part of a picture. A picture may be composed of one or more slices / tiles. In addition, a slice / tile may include one or more coding tree units (CTUs).
[0064] In the present disclosure, "pixel" or "pel" may mean the smallest single element constituting a picture (or image). In addition, "sample" may be used as a term corresponding to a pixel. A sample may generally represent a pixel or a pixel value, or may represent only a pixel / pixel value of a luminance component or only a pixel / pixel value of a chrominance component.
[0065] In the present disclosure, a "unit" may refer to a basic unit of image processing. The unit may include at least one of a specific area of a picture and information related to the area. In some cases, the unit may be used interchangeably with terms such as "sample array", "block" or "area". In general, an M×N block may include M columns and N rows of samples (or sample arrays) or a set (or array) of transform coefficients.
[0066] In the present disclosure, "current block" may mean one of "current coding block", "current coding unit", "coding target block", "decoding target block" or "processing target block". When prediction is performed, "current block" may mean "current prediction block" or "prediction target block". When transform (inverse transform) / quantization (dequantization) is performed, "current block" may mean "current transform block" or "transform target block". When filtering is performed, "current block" may mean "filtering target block".
[0067] In addition, in the present disclosure, unless explicitly stated as a chroma block, “current block” may mean “luminance block of the current block.” “Chroma block of the current block” may be expressed by including an explicit description of a chroma block such as “chroma block” or “current chroma block.”
[0068] In the present disclosure, the slash " / " or "," may be interpreted as indicating "and / or". For example, "A / B" and "A, B" may mean "A and / or B". In addition, "A / B / C" and "A, B, C" may mean "at least one of A, B and / or C".
[0069] In the present disclosure, the term "or" should be interpreted to indicate "and / or". For example, the expression "A or B" may include 1) only "A", 2) only "B", or 3) both "A and B". In other words, in the present disclosure, "or" should be interpreted to indicate "additionally or alternatively".
[0070] Video Coding System Overview
[0071] Figure 1 is a diagram schematically illustrating a video encoding system according to the present disclosure.
[0072] The video encoding system according to the embodiment may include an encoding device 10 and a decoding device 20. The encoding device 10 may deliver the encoded video and / or image information or data to the decoding device 20 via a digital storage medium or a network in the form of a file or a stream.
[0073] The encoding device 10 according to the embodiment may include a video source generator 11, an encoding unit 12, and a transmitter 13. The decoding device 20 according to the embodiment may include a receiver 21, a decoding unit 22, and a renderer 23. The encoding unit 12 may be called a video / image encoding unit, and the decoding unit 22 may be called a video / image decoding unit. The transmitter 13 may be included in the encoding unit 12. The receiver 21 may be included in the decoding unit 22. The renderer 23 may include a display and the display may be configured as a separate device or an external component.
[0074] The video source generator 11 can obtain the video / image by the process of capturing, synthesizing or generating the video / image. The video source generator 11 may include a video / image capturing device and / or a video / image generating device. The video / image capturing device may include, for example, one or more cameras, a video / image archive including previously captured videos / images, etc. The video / image generating device may include, for example, a computer, a tablet computer, and a smart phone, and may (electronically) generate the video / image. For example, a virtual video / image may be generated by a computer, etc. In this case, the video / image capturing process may be replaced by a process of generating relevant data.
[0075] The encoding unit 12 may encode the input video / image. For compression and encoding efficiency, the encoding unit 12 may perform a series of processes such as prediction, transformation, and quantization. The encoding unit 12 may output encoded data (encoded video / image information) in the form of a bitstream.
[0076] The transmitter 13 may transmit the encoded video / image information or the data output in the form of a bit stream to the receiver 21 of the decoding device 20 in the form of a file or stream through a digital storage medium or a network. The digital storage medium may include various storage media, such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter 13 may include an element for generating a media file in a predetermined file format and may include an element for transmitting through a broadcast / communication network. The receiver 21 may extract / receive a bit stream from a storage medium or a network and transmit the bit stream to the decoding unit 22.
[0077] The decoding unit 22 may decode a video / image by performing a series of processes corresponding to the operations of the encoding unit 12, such as dequantization, inverse transformation, and prediction.
[0078] The renderer 23 may render the decoded video / image. The rendered video / image may be displayed through a display.
[0079] Overview of Image Coding Device
[0080] Figure 2FIG. 1 is a diagram schematically showing an image encoding device to which an embodiment of the present disclosure is applicable.
[0081] like Figure 2 As shown, the image encoding device 100 may include an image segmenter 110, a subtractor 115, a transformer 120, a quantizer 130, a dequantizer 140, an inverse transformer 150, an adder 155, a filter 160, a memory 170, an inter-frame predictor 180, an intra-frame predictor 185, and an entropy encoder 190. The inter-frame predictor 180 and the intra-frame predictor 185 may be collectively referred to as a "predictor". The transformer 120, the quantizer 130, the dequantizer 140, and the inverse transformer 150 may be included in a residual processor. The residual processor may also include a subtractor 115.
[0082] In some embodiments, all or at least some of the components configuring the image encoding apparatus 100 may be configured by one hardware component (eg, an encoder or a processor). In addition, the memory 170 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium.
[0083] The image segmenter 110 may segment the input image (or picture or frame) input to the image encoding device 100 into one or more processing units. For example, the processing unit may be referred to as a coding unit (CU). The coding unit may be obtained by recursively segmenting a coding tree unit (CTU) or a maximum coding unit (LCU) according to a quadtree, binary tree, ternary tree (QT / BT / TT) structure. For example, a coding unit may be segmented into a plurality of coding units of a deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. For the segmentation of the coding unit, the quadtree structure may be applied first, and then the binary tree structure and / or the ternary tree structure may be applied. The encoding process according to the present disclosure may be performed based on the final coding unit that is no longer segmented. The maximum coding unit may be used as the final coding unit, and the coding unit of a deeper depth obtained by segmenting the maximum coding unit may also be used as the final coding unit. Here, the encoding process may include the prediction, transformation, and reconstruction processes described later. As another example, the processing unit of the encoding process may be a prediction unit (PU) or a transformation unit (TU). The prediction unit and the transform unit may be divided or partitioned from the final coding unit. The prediction unit may be a sample prediction unit, and the transform unit may be a unit for deriving a transform coefficient and / or a unit for deriving a residual signal from a transform coefficient.
[0084] The predictor (inter predictor 180 or intra predictor 185) may perform prediction on the block to be processed (current block) and generate a prediction block including prediction samples of the current block. The predictor may determine whether to apply intra prediction or inter prediction based on the current block or CU. The predictor may generate various information related to the prediction of the current block and transmit the generated information to the entropy encoder 190. Information about the prediction may be encoded in the entropy encoder 190 and output in the form of a bitstream.
[0085] The intra-frame predictor 185 can predict the current block by referring to the samples in the current picture. Depending on the intra-frame prediction mode and / or the intra-frame prediction technology, the reference samples can be located in the neighbors of the current block or can be placed separately. The intra-frame prediction mode may include multiple non-directional modes and multiple directional modes. The non-directional mode may include, for example, a DC mode and a plane mode. Depending on the level of detail of the prediction direction, the directional mode may include, for example, 33 directional prediction modes or 65 directional prediction modes. However, this is only an example, and more or fewer directional prediction modes may be used according to the settings. The intra-frame predictor 185 may determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.
[0086] The inter-frame predictor 180 may derive a prediction block of the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, in order to reduce the amount of motion information transmitted in the inter-frame prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may also include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, the neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block may be referred to as a collocated reference block, a collocated CU (colCU), etc. The reference picture including the temporal neighboring block may be referred to as a collocated picture (colPic). For example, the inter-frame predictor 180 may configure a motion information candidate list based on the neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter-frame prediction may be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the inter-frame predictor 180 may use the motion information of the neighboring blocks as the motion information of the current block. In the case of skip mode, unlike the merge mode, the residual signal may not be transmitted. In the case of motion vector prediction (MVP) mode, the motion vector of the neighboring block may be used as a motion vector predictor, and the motion vector of the current block may be signaled by encoding the motion vector difference and an indicator of the motion vector predictor. The motion vector difference may mean the difference between the motion vector of the current block and the motion vector predictor.
[0087] The predictor may generate a prediction signal based on various prediction methods and prediction techniques described below. For example, the predictor may apply not only intra prediction or inter prediction, but also intra prediction and inter prediction simultaneously to predict the current block. The prediction method of applying both intra prediction and inter prediction simultaneously to predict the current block may be referred to as combined inter and intra prediction (CIIP). In addition, the predictor may perform intra block copying (IBC) to predict the current block. Intra block copying may be used for content image / video encoding of games, etc., such as screen content coding (SCC). IBC is a method of predicting the current picture using a previously reconstructed reference block in the current picture at a position separated by a predetermined distance from the current block. When IBC is applied, the position of the reference block in the current picture may be encoded as a vector (block vector) corresponding to a predetermined distance. IBC basically performs prediction in the current picture, but may be performed similarly to inter prediction because the reference block is derived within the current picture. That is, IBC may use at least one of the inter prediction techniques described in the present disclosure. IBC basically performs prediction in the current picture, but may be performed similarly to inter prediction because the reference block is derived within the current picture. That is, IBC may use at least one of the inter-frame prediction techniques described in this disclosure.
[0088] The prediction signal generated by the predictor can be used to generate a reconstruction signal or to generate a residual signal. The subtractor 115 can generate a residual signal (residual block or residual sample array) by subtracting the prediction signal (prediction block or prediction sample array) output from the predictor from the input image signal (original block or original sample array). The generated residual signal can be transmitted to the transformer 120.
[0089] The transformer 120 may generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a karhunen-loève transform (KLT), a graph-based transform (GBT), or a conditional nonlinear transform (CNT). Here, GBT refers to a transform obtained from a graph when relationship information between pixels is represented by a graph. CNT refers to a transform obtained based on a prediction signal generated using all previously reconstructed pixels. In addition, the transform process may be applied to square pixel blocks of the same size or may be applied to blocks of variable size rather than square.
[0090] The quantizer 130 may quantize the transform coefficients and transmit them to the entropy encoder 190. The entropy encoder 190 may encode the quantized signal (information about the quantized transform coefficients) and output a bitstream. The information about the quantized transform coefficients may be referred to as residual information. The quantizer 130 may rearrange the quantized transform coefficients in the block form into a one-dimensional vector form based on the coefficient scanning order, and generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form.
[0091] The entropy encoder 190 may perform various encoding methods, such as exponential Golomb, context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), etc. The entropy encoder 190 may encode information required for video / image reconstruction (e.g., values of syntax elements, etc.) together or separately in addition to quantized transform coefficients. The encoded information (e.g., encoded video / image information) may be transmitted or stored in units of a network abstraction layer (NAL) in the form of a bitstream. The video / image information may also include information about various parameter sets, such as an adaptive parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may also include general constraint information. The signaled information, the transmitted information, and / or the syntax elements described in the present disclosure may be encoded and included in the bitstream through the above-mentioned encoding process.
[0092] The bitstream may be transmitted over a network or may be stored in a digital storage medium. The network may include a broadcast network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) that transmits a signal output from the entropy encoder 190 and / or a storage unit (not shown) that stores the signal may be included as an internal / external element of the image encoding device 100. Alternatively, a transmitter may be provided as a component of the entropy encoder 190.
[0093] The quantized transform coefficients output from the quantizer 130 may be used to generate a residual signal. For example, the residual signal (residual block or residual sample) may be reconstructed by applying dequantization and inverse transformation to the quantized transform coefficients through the dequantizer 140 and the inverse transformer 150.
[0094] The adder 155 adds the reconstructed residual signal to the prediction signal output from the inter-frame predictor 180 or the intra-frame predictor 185 to generate a reconstructed signal (reconstructed image, reconstructed block, reconstructed sample array). If there is no residual in the block to be processed, such as when the skip mode is applied, the prediction block can be used as a reconstructed block. The adder 155 can be called a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current picture, and can be used for inter-frame prediction of the next picture by filtering as described below.
[0095] The filter 160 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 160 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and store the modified reconstructed picture in the memory 170, specifically, in the DPB of the memory 170. Various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filtering, bilateral filtering, etc. The filter 160 can generate various information related to filtering and transmit the generated information to the entropy encoder 190, as described later in the description of each filtering method. The information related to filtering can be encoded by the entropy encoder 190 and output in the form of a bit stream.
[0096] The modified reconstructed picture transferred to the memory 170 may be used as a reference picture in the inter predictor 180. When inter prediction is applied by the image encoding device 100, prediction mismatch between the image encoding device 100 and the image decoding device may be avoided and encoding efficiency may be improved.
[0097] The DPB of the memory 170 may store the modified reconstructed picture for use as a reference picture in the inter-frame predictor 180. The memory 170 may store the motion information of the block from which the motion information in the current picture is derived (or encoded) and / or the motion information of the reconstructed block in the picture. The stored motion information may be transmitted to the inter-frame predictor 180 and used as the motion information of the spatial neighboring block or the motion information of the temporal neighboring block. The memory 170 may store the reconstructed samples of the reconstructed blocks in the current picture and may transmit the reconstructed samples to the intra-frame predictor 185.
[0098] Overview of image decoding device
[0099] Figure 3 FIG. 1 is a diagram schematically showing an image decoding device to which an embodiment of the present disclosure is applicable.
[0100] like Figure 3As shown, the image decoding device 200 may include an entropy decoder 210, a dequantizer 220, an inverse transformer 230, an adder 235, a filter 240, a memory 250, an inter-frame predictor 260, and an intra-frame predictor 265. The inter-frame predictor 260 and the intra-frame predictor 265 may be collectively referred to as a "predictor". The dequantizer 220 and the inverse transformer 230 may be included in a residual processor.
[0101] According to an embodiment, all or at least some of the plurality of components configuring the image decoding apparatus 200 may be configured by hardware components (eg, a decoder or a processor). In addition, the memory 250 may include a decoded picture buffer (DPB) or may be configured by a digital storage medium.
[0102] The image decoding apparatus 200 having received a bit stream including video / image information may perform the same operation as that performed by Figure 2 The image may be reconstructed by processing corresponding to the processing performed by the image encoding device 100. For example, the image decoding device 200 may perform decoding using a processing unit applied in the image encoding device. Therefore, the processing unit of decoding may be, for example, a coding unit. The coding unit may be obtained by splitting a coding tree unit or a maximum coding unit. The reconstructed image signal decoded and output by the image decoding device 200 may be reproduced by a reproduction device (not shown).
[0103] The image decoding apparatus 200 may receive the image in the form of a bit stream from Figure 2The received signal may be decoded by the entropy decoder 210. For example, the entropy decoder 210 may parse the bitstream to derive information (e.g., video / image information) required for image reconstruction (or picture reconstruction). The video / image information may also include information about various parameter sets, such as an adaptive parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may also include general constraint information. The image decoding device may also decode the picture based on the parameter set information and / or the general constraint information. The information and / or syntax elements signaled / received described in the present disclosure may be decoded and obtained from the bitstream through a decoding process. For example, the entropy decoder 210 may decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the values of the syntax elements required for image reconstruction and the quantized values of the transform coefficients of the residual. More specifically, the CABAC entropy decoding method can receive a bin corresponding to each syntax element in the bitstream, use the decoding target syntax element information, the decoding information of the neighboring block and the decoding target block, or the information of the symbol / bin decoded in the previous stage to determine the context model, and perform arithmetic decoding on the bin by predicting the probability of occurrence of the bin according to the determined context model, and generate a symbol corresponding to the value of each syntax element. In this case, the CABAC entropy decoding method can update the context model by using the information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model. The information related to the prediction in the information decoded by the entropy decoder 210 can be provided to the predictor (inter-frame predictor 260 and intra-frame predictor 265), and the residual value for which entropy decoding is performed in the entropy decoder 210, that is, the quantized transform coefficient and related parameter information can be input to the dequantizer 220. In addition, information about filtering among the information decoded by the entropy decoder 210 can be provided to the filter 240. Meanwhile, a receiver (not shown) for receiving a signal output from the image encoding device may be further configured as an internal / external element of the image decoding device 200 , or the receiver may be a component of the entropy decoder 210 .
[0104] Meanwhile, the image decoding device according to the present disclosure may be referred to as a video / image / picture decoding device. The image decoding device may be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include an entropy decoder 210. The sample decoder may include at least one of a dequantizer 220, an inverse transformer 230, an adder 235, a filter 240, a memory 250, an inter-frame predictor 260, or an intra-frame predictor 265.
[0105] The dequantizer 220 may dequantize the quantized transform coefficient and output the transform coefficient. The dequantizer 220 may rearrange the quantized transform coefficient in the form of a two-dimensional block. In this case, the rearrangement may be performed based on the coefficient scanning order performed in the image encoding device. The dequantizer 220 may perform dequantization on the quantized transform coefficient by using a quantization parameter (e.g., quantization step size information) and obtain the transform coefficient.
[0106] The inverse transformer 230 may inversely transform the transform coefficients to obtain a residual signal (residual block, residual sample array).
[0107] The predictor may perform prediction on the current block and generate a prediction block including prediction samples of the current block. The predictor may determine whether to apply intra prediction or inter prediction to the current block based on information about the prediction output from the entropy decoder 210, and may determine a specific intra / inter prediction mode (prediction technique).
[0108] As described in the predictor of the image encoding device 100 , the predictor can generate a prediction signal based on various prediction methods (techniques) described later.
[0109] The intra predictor 265 may predict the current block by referring to samples in the current picture. The description of the intra predictor 185 is also applicable to the intra predictor 265.
[0110] The inter-frame predictor 260 may derive a prediction block of the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, in order to reduce the amount of motion information transmitted in the inter-frame prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may also include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, the neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. For example, the inter-frame predictor 260 may configure a motion information candidate list based on the neighboring blocks, and derive a motion vector and / or a reference picture index of the current block based on the received candidate selection information. Inter-frame prediction may be performed based on various prediction modes, and information about the prediction may include information indicating the inter-frame prediction mode of the current block.
[0111] The adder 235 can generate a reconstructed block by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor (including the inter-frame predictor 260 and / or the intra-predictor 265). If the block to be processed has no residual, such as when the skip mode is applied, the prediction block can be used as the reconstructed block. The description of the adder 155 is also applicable to the adder 235. The adder 235 can be called a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current picture, and can be used for inter-frame prediction of the next picture by filtering as described below.
[0112] The filter 240 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 240 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and store the modified reconstructed picture in the memory 250, specifically, in the DPB of the memory 250. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filtering, bilateral filtering, etc.
[0113] The (modified) reconstructed picture stored in the DPB of the memory 250 can be used as a reference picture in the inter-frame predictor 260. The memory 250 can store the motion information of the block from which the motion information in the current picture is derived (or decoded) and / or the motion information of the reconstructed block in the picture. The stored motion information can be transmitted to the inter-frame predictor 260 to be used as the motion information of the spatial neighboring block or the motion information of the temporal neighboring block. The memory 250 can store the reconstructed samples of the reconstructed block in the current picture and transmit the reconstructed samples to the intra-frame predictor 265.
[0114] In the present disclosure, the embodiments described in the filter 160 , the inter-frame predictor 180 , and the intra-frame predictor 185 of the image encoding device 100 may be equally or correspondingly applied to the filter 240 , the inter-frame predictor 260 , and the intra-frame predictor 265 of the image decoding device 200 .
[0115] Overview of Image Segmentation
[0116] The video / image encoding method according to the present disclosure can be performed based on the image segmentation structure as follows. Specifically, the prediction, residual processing ((inverse) transform, (de)quantization, etc.), syntax element encoding and filtering processes described later can be performed based on the CTU, CU (and / or TU, PU) derived from the image segmentation structure. The image can be segmented in units of blocks and the block segmentation process can be performed in the image segmentor 110 of the encoding device. The segmentation related information can be encoded by the entropy encoder 190 and sent to the decoding device in the form of a bitstream. The entropy decoder 210 of the decoding device can derive the block segmentation structure of the current picture based on the segmentation related information obtained from the bitstream, and based on this, a series of processes (e.g., prediction, residual processing, block / picture reconstruction, in-loop filtering, etc.) can be performed for image decoding. The CU size and the TU size can be the same, or there can be multiple TUs in the CU area. At the same time, the CU size can generally represent the luminance component (sample) CB size. The TU size can generally represent the luminance component (sample) TB size. The chroma component (sample) CB or TB size can be derived based on the luminance component (sample) CB or TB size according to the chroma format (color format, such as 4:4:4, 4:2:2, 4:2:0, etc.) of the picture / image according to the component ratio. The TU size can be derived based on the maxTbSize that specifies the maximum available TB size. For example, when the CU size is larger than maxTbSize, multiple TUs (TBs) of maxTbSize can be derived from the CU, and transform / inverse transform can be performed in units of TU (TB). In addition, for example, when intra prediction is applied, the intra prediction mode / type can be derived in units of CU (or CB), and the neighboring reference sample derivation and prediction sample generation process can be performed in units of TU (or TB). In this case, one or more TUs (or TBs) may exist in a CU (or CB) area, and in this case, multiple TUs (or TBs) may share the same intra prediction mode / type.
[0117] A picture may be partitioned into a sequence of coding tree units (CTUs). Figure 4 An example of a picture being partitioned into CTUs is shown. A CTU may correspond to a coding tree block (CTB). Alternatively, a CTU may include a coding tree block of luma samples and two corresponding coding tree blocks of chroma samples. For example, for a picture containing three sample arrays, a CTU may include an N×N block of luma samples and two corresponding blocks of chroma samples.
[0118] Overview of CTU Segmentation
[0119] As described above, a coding unit may be obtained by recursively partitioning a coding tree unit (CTU) or a maximum coding unit (LCU) according to a quadtree / binary tree / ternary tree (QT / BT / TT) structure. For example, a CTU may be first partitioned into a quadtree structure. Thereafter, the leaf nodes of the quadtree structure may be further partitioned by a multi-type tree structure.
[0120] Partitioning according to the quadtree means that the current CU (or CTU) is equally divided into four. By partitioning according to the quadtree, the current CU can be partitioned into four CUs with the same width and the same height. When the current CU is no longer partitioned into a quadtree structure, the current CU corresponds to a leaf node of the quadtree structure. The CU corresponding to the leaf node of the quadtree structure may no longer be partitioned and may be used as the above-mentioned final coding unit. Alternatively, the CU corresponding to the leaf node of the quadtree structure may be further partitioned by a multi-type tree structure.
[0121] Figure 5 is a view showing an embodiment of a partition type of a block according to a multi-type tree structure. The partition according to the multi-type tree structure may include two types of partitions according to a binary tree structure and two types of partitions according to a ternary tree structure.
[0122] The two types of splits according to the binary tree structure may include vertical binary split (SPLIT_BT_VER) and horizontal binary split (SPLIT_BT_HOR). Vertical binary split (SPLIT_BT_VER) means that the current CU is equally split into two in the vertical direction. Figure 5 As shown in FIG, through vertical binary splitting, two CUs with the same height as the current CU and half the width of the current CU can be generated. Horizontal binary splitting (SPLIT_BT_HOR) means that the current CU is equally divided into two in the horizontal direction. Figure 5 As shown, through horizontal binary partitioning, two CUs with a height half of the height of the current CU and the same width as the current CU can be generated.
[0123] The two types of splits according to the triad structure may include vertical triad split (SPLIT_TT_VER) and horizontal triad split (SPLIT_TT_HOR). In vertical triad split (SPLIT_TT_VER), the current CU is split in a vertical direction at a ratio of 1:2:1. Figure 5As shown, through vertical trisection, two CUs with the same height as the current CU and a width of 1 / 4 of the width of the current CU, and a CU with the same height as the current CU and a width of half the width of the current CU can be generated. In horizontal trisection (SPLIT_TT_HOR), the current CU is split in the horizontal direction at a ratio of 1:2:1. Figure 5 As shown, through horizontal trifurcated division, two CUs whose height is 1 / 4 of the height of the current CU and whose width is the same as the current CU, and a CU whose height is half of the height of the current CU and whose width is the same as the current CU can be generated.
[0124] Figure 6 is a diagram illustrating a signaling mechanism of block partitioning information in a quadtree having a nested multi-type tree structure according to the present disclosure.
[0125] Here, the CTU is regarded as the root node of the quadtree and is first split into a quadtree structure. Information (e.g., qt_split_flag) specifying whether quadtree partitioning is performed for the current CU (CTU or node (QT_node) of the quadtree) is signaled. For example, when qt_split_flag has a first value (e.g., "1"), the current CU can be split by the quadtree. In addition, when qt_split_flag has a second value (e.g., "0"), the current CU is not quadtree split, but becomes a leaf node (QT_leaf_node) of the quadtree. Each quadtree leaf node can then be further split into a multi-type tree structure. That is, the leaf node of the quadtree can become a node (MTT_node) of a multi-type tree. In the multi-type tree structure, a first flag (e.g., Mtt_split_cu_flag) is signaled to indicate whether the current node is additionally split. If the corresponding node is additionally split (for example, if the first flag is 1), the second flag (for example, Mtt_split_cu_vertical_flag) may be signaled to indicate the split direction. For example, the split direction may be a vertical direction when the second flag is 1, and a horizontal direction when the second flag is 0. Then, a third flag (for example, Mtt_split_cu_binary_flag) may be signaled to indicate whether the split type is a binary split type or a ternary split type. For example, the split type may be a binary split type when the third flag is 1, and a ternary split type when the third flag is 0. The nodes of the multi-type tree obtained by binary splitting or ternary splitting may be further split into a multi-type tree structure. However, the nodes of the multi-type tree may not be split into a quadtree structure. If the first flag is 0, the corresponding node of the multi-type tree is no longer split, but becomes a leaf node (MTT_leaf_node) of the multi-type tree. The CU corresponding to the leaf node of the multi-type tree may be used as the above-mentioned final coding unit.
[0126] Based on mtt_split_cu_vertical_flag and mtt_split_cu_binary_flag, a multi-type tree partition mode (MttSplitMode) of a CU may be derived as shown in the following Table 1. In the following description, a multi-type tree partition mode may be referred to as a multi-tree partition type or a partition type.
[0127] [Table 1]
[0128] MttSplitMode mtt_split_cu_vertical_flag mtt_split_cu_binary_flag SPLIT_TT_HOR 0 0 SPLIT_BT_HOR 0 1 SPLIT_TT_VER 1 0 SPLIT_BT_VER 1 1
[0129] Figure 7 is a view showing an example of splitting a CTU into a plurality of CUs by applying a multi-type tree after applying a quadtree. Figure 7 , the bold block edge 710 represents a quadtree partition, while the remaining edges 720 represent a multi-type tree partition. A CU may correspond to a coding block (CB). In an embodiment, a CU may include a coding block of luma samples and two coding blocks of chroma samples corresponding to the luma samples.
[0130] The chroma component (sample) CB or TB size may be derived based on the luma component (sample) CB or TB size based on the component ratio according to the color format (chroma format, e.g., 4:4:4, 4:2:2, 4:2:0, etc.) of the picture / image. In the case of a 4:4:4 color format, the chroma component CB / TB size may be set equal to the luma component CB / TB size. In the case of a 4:2:2 color format, the width of the chroma component CB / TB may be set to half the width of the luma component CB / TB and the height of the chroma component CB / TB may be set to the height of the luma component CB / TB. In the case of a 4:2:0 color format, the width of the chroma component CB / TB may be set to half the width of the luma component CB / TB and the height of the chroma component CB / TB may be set to half the height of the luma component CB / TB.
[0131] In an embodiment, when the size of the CTU is based on a luma sample unit of 128, the size of the CU may have a size from 128x128 to 4x4, which is the same size as the CTU. In an embodiment, in the case of a 4:2:0 color format (or chroma format), the chroma CB size may have a size from 64x64 to 2x2.
[0132] Meanwhile, in an embodiment, the CU size and the TU size may be the same. Alternatively, there may be multiple TUs in a CU region. The TU size generally indicates the luma component (sample) transform block (TB) size.
[0133] The TU size can be derived based on the maximum allowed TB size maxTbSize as a predetermined value. For example, when the CU size is larger than maxTbSize, multiple TUs (TBs) with maxTbSize can be derived from the CU, and transform / inverse transform can be performed in units of TU (TB). For example, the maximum allowed luma TB size can be 64x64 and the maximum allowed chroma TB size can be 32x32. If the width or height of the CB split according to the tree structure is larger than the maximum transform width or height, the CB can be automatically (or implicitly) split until the TB size limits in the horizontal and vertical directions are met.
[0134] In addition, for example, when intra prediction is applied, the intra prediction mode / type may be derived in units of CU (or CB), and the neighboring reference sample derivation and prediction sample generation process may be performed in units of TU (or TB). In this case, there may be one or more TUs (or TBs) in a CU (or CB) region, and in this case, multiple TUs or (TBs) may share the same intra prediction mode / type.
[0135] Meanwhile, for a quadtree coding tree scheme with nested multi-type trees, the following parameters may be signaled from an encoding device to a decoding device as SPS syntax elements. For example, at least one of a CTU size as a parameter indicating a root node size of a quadtree, a MinQTSize as a parameter indicating a minimum allowed quadtree leaf node size, a MaxBtSize as a parameter indicating a maximum allowed binary tree root node size, a MaxTtSize as a parameter indicating a maximum allowed ternary tree root node size, a MaxMttDepth as a parameter indicating a maximum allowed hierarchical depth of multi-type tree partitioning from a quadtree leaf node, a MinBtSize as a parameter indicating a minimum allowed binary tree leaf node size, or a MinTtSize as a parameter indicating a minimum allowed ternary tree leaf node size is signaled.
[0136] As an embodiment using a 4:2:0 chroma format, the CTU size may be set to 128x128 luminance blocks and two 64x64 chroma blocks corresponding to these luminance blocks. In this case, MinOTSize may be set to 16x16, MaxBtSize may be set to 128x128, MaxTtSzie may be set to 64x64, MinBtSize and MinTtSize may be set to 4x4, and MaxMttDepth may be set to 4. Quadtree segmentation may be applied to the CTU to generate a quadtree leaf node. The quadtree leaf node may be referred to as a leaf QT node. The size of the quadtree leaf node may be from 16x16 size (e.g., MinOTSize) to 128x128 size (e.g., CTU size). If the leaf QT node is 128x128, it may not be additionally segmented into a binary tree / ternary tree. This is because, in this case, even if segmented, it exceeds MaxBtsize and MaxTtszie (e.g., 64x64). In other cases, the leaf QT node can be further split into a multi-type tree. Therefore, the leaf QT node is the root node of the multi-type tree, and the leaf QT node can have a multi-type tree depth (mttDepth) value of 0. If the multi-type tree depth reaches MaxMttdepth (for example, 4), further splitting can be ignored. If the width of the multi-type tree node is equal to MinBtSize and is less than or equal to 2xMinTtSize, further horizontal splitting can be ignored. If the height of the multi-type tree node is equal to MinBtSize and is less than or equal to 2xMinTtSize, further vertical splitting can be ignored. When splitting is not considered, the encoding device can skip the signaling of the splitting information. In this case, the decoding device can derive the splitting information with a predetermined value.
[0137] At the same time, one CTU may include a coding block of luma samples (hereinafter referred to as "luminance block") and two coding blocks of chroma samples corresponding thereto (hereinafter referred to as "chroma blocks"). The above coding tree scheme may be applied equally or individually to the luma block and chroma block of the current CU. Specifically, the luma block and chroma block in one CTU may be partitioned into the same block tree structure, and in this case, the tree structure is represented as SINGLE_TREE. Alternatively, the luma block and chroma block in one CTU may be partitioned into separate block tree structures, and in this case, the tree structure may be represented as DUAL_TREE. That is, when the CTU is divided into dual trees, a block tree structure for the luma block and a block tree structure for the chroma block may exist separately. In this case, the block tree structure for the luma block may be referred to as DUAL_TREE_LUMA, and the block tree structure for the chroma component may be referred to as DUAL_TREE_CHROMA. For P and B slices / block groups, the luma block and the chroma block in one CTU may be restricted to have the same coding tree structure. However, for I slices / patch groups, luma blocks and chroma blocks may have separate block tree structures from each other. If a separate block tree structure is applied, luma CTBs may be divided into CUs based on a specific coding tree structure, and chroma CTBs may be divided into chroma CUs based on another coding tree structure. That is, this means that a CU in an I slice / patch group to which a separate block tree structure is applied may include a coding block of a luma component or a coding block of two chroma components, and a CU of a P or B slice / patch group may include blocks of three color components (one luma component and two chroma components).
[0138] Although a quadtree coding tree structure with nested multi-type trees has been described, the structure for partitioning the CU is not limited thereto. For example, the BT structure and the TT structure may be interpreted as concepts included in a multi-partition tree (MPT) structure, and the CU may be interpreted as being partitioned by the QT structure and the MPT structure. In an example where the CU is partitioned by the QT structure and the MPT structure, a syntax element (e.g., MPT_split_type) including information about how many blocks a leaf node of the QT structure is partitioned into and a syntax element (e.g., MPT_split_mode) including information about which of the vertical and horizontal directions a leaf node of the QT structure is partitioned into may be signaled to determine the partition structure.
[0139] In another example, the CU may be partitioned in a manner different from the QT structure, the BT structure, or the TT structure. That is, instead of partitioning a CU of a lower depth into 1 / 4 of a CU of a higher depth according to the QT structure, partitioning a CU of a lower depth into 1 / 2 of a CU of a higher depth according to the BT structure, or partitioning a CU of a lower depth into 1 / 4 or 1 / 2 of a CU of a higher depth according to the TT structure, in some cases a CU of a lower depth may be partitioned into 1 / 5, 1 / 3, 3 / 8, 3 / 5, 2 / 3, or 5 / 8 of a CU of a higher depth, and the method of partitioning a CU is not limited thereto.
[0140] The quadtree coding block structure with multi-type trees can provide a very flexible block segmentation structure. Due to the segmentation types supported in the multi-type trees, different segmentation patterns can potentially produce the same coding block structure in some cases. In the encoding device and the decoding device, by limiting the occurrence of such redundant segmentation patterns, the amount of data of the segmentation information can be reduced.
[0141] Sprite-based image encoding / decoding
[0142] One encoding target picture may be partitioned into units of multiple CTUs, slices, tiles, or blocks, and the picture may be partitioned into units of multiple sub-pictures.
[0143] Within a picture, a sub-picture may be encoded or decoded regardless of whether the previous sub-picture is encoded or decoded. For example, different quantization or different resolutions may be applied to multiple sub-pictures.
[0144] In addition, each sub-picture can be processed like a separate picture. For example, the encoding target picture can be a projected picture or a packed picture in an omnidirectional image / video or a 360-degree image / video.
[0145] In this embodiment, a portion of the screen may be rendered or displayed based on a viewport of a user terminal (e.g., a head mounted display). Therefore, in order to achieve low latency, among the sub-pictures configuring a screen, at least one sub-picture covering the viewport may be encoded or decoded in priority to or independently of the remaining sub-pictures.
[0146] The encoding result of the sub-picture may be referred to as a sub-bitstream, a sub-stream, or simply a bitstream. A decoding device may decode the sub-picture from the sub-bitstream, the sub-stream, or the bitstream. In this case, a high-level syntax (HLS) such as a PPS, an SPS, a VPS, and / or a decoding parameter set (DPS) may be used to encode / decode the sub-picture.
[0147] In the present disclosure, a high-level syntax (HLS) may include at least one of an APS syntax, a PPS syntax, an SPS syntax, a VPS syntax, a DPS syntax, or an SH syntax. For example, an APS (APS syntax) or a PPS (PPS syntax) may include information / parameters that may be commonly applied to one or more slices or pictures. An SPS (SPS syntax) may include information / parameters that may be commonly applied to one or more sequences. A VPS (VPS syntax) may include information / parameters that may be commonly applied to multiple layers. A DPS (DPS syntax) may include information / parameters that may be commonly applied to an entire video. For example, a DPS may include information / parameters related to a concatenation of a coded video sequence (CVS).
[0148] A sub-picture can configure a rectangular area of a coded picture. The size of a sub-picture can be set differently within the picture. For all pictures belonging to a sequence, the size and position of a specific individual sub-picture can be set equally. Separate sub-picture sequences can be decoded independently. Blocks and slices (and CTBs) can be restricted to not cross sub-picture boundaries. To this end, the encoding device can perform encoding so that sub-pictures are decoded independently. To this end, semantic restrictions in the bitstream may be required. In addition, for each picture in a sequence, the arrangement of blocks, slices, and tiles in the sub-picture can be configured differently.
[0149] The sub-picture design is intended to be an abstraction or encapsulation that is smaller in scope than the picture level but larger than the slice or tile group level. Therefore, the VCL NAL units of a subset of the motion constrained tile set (MCTS) can be extracted from one VVC bitstream, and processing such as rearrangement to another VVC bitstream (e.g., modification at the VCL level) can be performed without difficulty. Here, MCTS is a coding technique that allows spatial and temporal independence between tiles. When MCTS is applied, it is impossible to refer to information about tiles that are not included in the MCTS to which the current tile belongs. When an image is segmented into MCTS and encoded, independent transmission and encoding of the MCTS can be performed.
[0150] The advantage of this sprite design is that it changes the viewing orientation with a mixed-resolution viewport relying on a 360° streaming scheme.
[0151] In the following, reference will be made to Figure 8 and Fig. 9 Describes an image encoding / decoding method using slices / tiles.
[0152] Figure 8 : is a flowchart illustrating a method in which an image encoding apparatus encodes an image using slices / tiles according to an embodiment of the present disclosure.
[0153] The image encoding apparatus may derive slices / tiles in the current picture by dividing the current picture ( S810 ).
[0154] The image encoding apparatus may encode the current picture based on the slice / tile derived in step S810 ( S820 ).
[0155] Fig. 9 The flowchart illustrates a method for decoding an image using slices / tiles by an image decoding apparatus according to an embodiment of the present disclosure.
[0156] The image decoding apparatus may acquire information about a video / image from a bitstream (S910).
[0157] In addition, the image decoding apparatus may derive a slice / patch in the current picture based on the information about the video / image acquired in step S910 (S920). Here, the information about the video / image may include information about the slice / patch.
[0158] Next, the image decoding apparatus may decode the current picture based on the slice / tile derived in step S920 ( S930 ).
[0159] exist Figure 8 and Fig. 9 In the present disclosure, the information about the slice / tile may include various information and / or syntax elements described in the present disclosure. The video / image information may include high-level syntax, and the high-level syntax may include information about the slice and / or information about the tile. The high-level syntax may include a picture header, and the information about the picture header may be included in the slice header described in the present disclosure. The information about the slice may include information specifying one or more slices, and the information about the tile may include information specifying one or more tiles. A slice including one or more tiles may exist in a picture.
[0160] High Level Syntax (HLS) Signaling
[0161] As described above, high-level syntax may be encoded / signaled for video / image coding. Hereinafter, signaling and syntax elements in a picture header and a slice header according to the present disclosure will be described.
[0162] Picture header and slice header
[0163] A coded picture may consist of one or more slices. The parameters of a coded picture are signaled in a picture header (PH), and the parameters of a slice are signaled in a slice header (SH). The PH is carried in its own NAL unit type. The SH may be present at the beginning of a NAL unit containing the payload of a slice (i.e., slice data). In the following, reference will be made to Fig.10 and Fig.11 Describes the syntax elements of PH and SH and the semantics of the syntax elements.
[0164] Fig.10 is a diagram of an example of the present disclosure showing signaling and syntax elements in a picture header.
[0165] picture_header_rbsp() contains information common to all slices of a coded picture associated with a picture header (PH). For example, picture_header_rbsp() may include a reference picture flag (non_reference_picture_flag), GDR picture identification information (gdr_pic_flag), no_output_of_prior_pics_flag, recovery_poc_cnt, ph_pic_parameter_set_id, etc. Here, when gdr_pic_flag is 1, recovery_poc_cnt is signaled in picture_header_rbsp().
[0166] The first value of non_reference_picture_flag (eg, 1) specifies that the picture associated with the PH is not used as a reference picture. The second value of non_reference_picture_flag (eg, 0) specifies that the picture associated with the PH may or may not be used as a reference picture.
[0167] A first value of gdr_pic_flag (eg, 1) specifies that the picture associated with the PH is a GDR picture. A second value of gdr_pic_flag (eg, 0) specifies that the picture associated with the PH is not a GDR picture.
[0168] no_output_of_prior_pics_flag affects the output of previously decoded pictures in the DPB after decoding of a coded layer video sequence (CLVSS) picture that is not the first picture in the bitstream.
[0169] recovery_poc_cnt specifies the recovery point of the decoded picture in output order.
[0170] ph_pic_parameter_set_id specifies the value of pps_pic_parameter_set_id for the picture parameter set (PPS) in use. pps_pic_parameter_set_id is a value used to identify a PPS to be referenced by another syntax.
[0171] Included in Fig.10The syntax elements in the picture_header_rbsp() syntax structure may be included in the picture_header_structure() syntax structure and signaled. In this case, the picture_header_structure() syntax structure may be included in the picture_header_rbsp() syntax structure and signaled.
[0172] Fig.11 is a diagram showing a syntax structure of a slice header according to an embodiment of the present disclosure.
[0173] like Fig.11 As shown, picture_header_in_slice_header_flag, picture_header_structure(), slice_subpic_id, slice_address, num_tiles_in_slice_minus1, etc. can be notified by slice head signal.
[0174] exist Fig.11 In the example shown, picture_header_in_slice_header_flag specifies whether a picture header syntax structure is present in a slice header syntax structure. A first value of picture_header_in_slice_header_flag (e.g., 1 or true) specifies that a picture header is present in a slice header, and a second value of picture_header_in_slice_header_flag (e.g., 0 or false) specifies that a picture header is not present in a slice header.
[0175] picture_header_structure() may be acquired based on picture_header_in_slice_header_flag. For example, when picture_header_in_slice_header_flag has a first value, picture_header_structure() may be signaled. When picture_header_in_slice_header_flag has a second value, picture_header_structure() may not be included in the slice header, but may be included in a separate NAL unit and signaled.
[0176] slice_subpic_id may be information about a sub-picture identifier for identifying a sub-picture that includes the current slice. slice_subpic_id may be acquired based on subpics_present_flag. For example, when subpics_present_flag is 1, slice_subpic_id may be signaled. subpics_present_flag may specify whether a sub-picture exists in the current picture or whether information about a sub-picture exists in the bitstream. For example, a first value of subpics_present_flag (e.g., 1 or true) may specify that information about a sub-picture exists in the bitstream or that one or more sub-pictures exist in the current picture. A second value of subpics_present_flag (e.g., 0 or false) may specify that information about a sub-picture does not exist in the bitstream or that a sub-picture does not exist in the current picture.
[0177] slice_address may specify the address in the current picture of the current slice. slice_address may be obtained based on rect_slice_flag and / or NumTilesInPic. For example, when rect_slice_flag is a first value (e.g., 1 or true) or NumTilesInPic is greater than 1, slice_address may be signaled in the slice header. At this time, rect_slice_flag may indicate an indicator of whether the slice included in the current picture is a rectangular slice. For example, rect_slice_flag may be signaled at a picture level (PPS or picture header). In addition, NumTilesInPic may specify the number of tiles included in the current picture.
[0178] num_tiles_in_slice_minus1 may specify the number of tiles included in the current slice. num_tiles_in_slice_minus1 may be obtained based on rect_slice_flag and NumTilesInPic. For example, when rect_slice_flag is a second value (eg, 0 or false) and NumTilesInPic is greater than 1, num_tiles_in_slice_minus1 may be signaled in the slice header.
[0179] exist Fig.11 In the illustrated embodiment, the requirements for bitstream consistency associated with picture_header_in_slice_header_flag may include the following.
[0180] To meet bitstream consistency, the value of picture_header_in_slice_header_flag is required to be the same in all slices of CLVS.
[0181] In addition, when picture_header_in_slice_header_flag is the first value (eg, 1), in order to satisfy bitstream consistency, it is required that there is no NAL unit with the NAL unit type equal to PH_NUT in the CLVS.
[0182] In addition, when picture_header_in_slice_header_flag is the second value (eg, 0), in order to satisfy bitstream consistency, a NAL unit with a NAL unit type equal to PH_NUT is required to exist before the first VCL NAL unit of the PU in the PU.
[0183] Fig.12 This example illustrates parsing and decoding Fig.11 Flowchart of the method of slicing header.
[0184] First, the image decoding apparatus may acquire a first flag (picture_header_in_slice_header_flag) included in a slice header (S1210).
[0185] The first flag may specify whether a picture header exists in the slice header. In addition, the first flag may specify whether the current picture includes only one slice.
[0186] When the first flag is a first value (e.g., 1 or true) (step S1220-yes), the image decoding apparatus may obtain a picture header from a slice header (S1230). When the first flag is a second value (e.g., 0 or false) (step S1220-no), the picture header may be obtained from a picture header NAL unit instead of a slice header (not shown).
[0187] Thereafter, it may be determined in step S1240 whether subpics_present_flag is a first value (e.g., 1 or true). subpics_present_flag may specify whether the current picture includes a subpicture. In addition, subpics_present_flag may specify whether information about a subpicture is included in the bitstream. subpics_present_flag may be signaled at a higher level of the slice. For example, subpics_present_flag may be included in a sequence parameter set and signaled.
[0188] When subpics_present_flag is a first value (e.g., 1 or true) (step S1240-yes), the image decoding device may obtain slice_subpic_id from the slice header (S1250). When subpics_present_flag is a second value (e.g., 0 or false) (step S1240-no), the image decoding device may omit (skip) parsing slice_subpic_id from the slice header.
[0189] Thereafter, in step S1260, it may be determined whether rect_slice_flag is a first value (e.g., 1 or true) and / or whether NumTilesInPic is greater than 1. rect_slice_flag may be an indicator indicating whether a slice included in the current picture is a rectangular slice. For example, rect_slice_flag may be signaled at a picture level (PPS or picture header). In addition, NumTilesInPic may specify the number of tiles included in the current picture.
[0190] When rect_slice_flag is a first value (e.g., 1 or true) or NumTilesInPic is greater than 1 (step S1260-yes), the image decoding device may obtain slice_address from the slice header (S1270). When rect_slice_flag is a second value (e.g., 0 or false) and NumTilesInPic is not greater than 1 (step S1260-no), the image decoding device may omit (skip) parsing slice_address from the slice header.
[0191] Thereafter, in step S1280 , it may be determined whether rect_slice_flag is a first value (eg, 1 or true) and / or whether NumTilesInPic is greater than 1.
[0192] When rect_slice_flag is a first value (e.g., 1 or true) or NumTilesInPic is not greater than 1 (step S1280-No), the image decoding device may omit (skip) parsing num_tiles_in_slice_minus1 from the slice header. When rect_slice_flag is a second value (e.g., 0 or false) and NumTilesInPic is greater than 1 (step S1280-Yes), the image decoding device may obtain num_tiles_in_slice_minus1 from the slice header (S1290).
[0193] Thereafter, the image decoding apparatus may decode the slice header by parsing subsequent syntax elements (not shown) from the slice header.
[0194] Fig.13 is an example of Fig.11 Flowchart of a method for encoding a slice header.
[0195] First, the image encoding apparatus may determine a value of a first flag (picture_header_in_slice_header_flag) and encode the first flag in a slice header (S1310).
[0196] When the first flag is a first value (e.g., 1 or true) (step S1320-yes), the image encoding device may encode the picture header in the slice header (S1330). When the first flag is a second value (e.g., 0 or false) (step S1320-no), the picture header is not encoded in the slice header, but may be included in a picture header NAL unit (not shown) and signaled.
[0197] Thereafter, in step S1340, it may be determined whether subpics_present_flag is a first value (eg, 1 or true). subpics_present_flag may be determined and signaled at a higher level of a slice. For example, subpics_present_flag may be included in a sequence parameter set and signaled.
[0198] When subpics_present_flag is a first value (e.g., 1 or true) (step S1340-yes), the image encoding apparatus may encode slice_subpic_id in the slice header (S1350). When subpics_present_flag is a second value (e.g., 0 or false) (step S1340-no), the image encoding apparatus may omit (skip) encoding slice_subpic_id in the slice header.
[0199] Thereafter, in step S1360 , it may be determined whether rect_slice_flag is a first value (eg, 1 or true) and / or whether NumTilesInPic is greater than 1.
[0200] When rect_slice_flag is a first value (e.g., 1 or true) or when NumTilesInPic is greater than 1 (step S1360-yes), the image encoding device may encode slice_address in the slice header (S1370). When rect_slice_flag is a second value (e.g., 0 or false) and NumTilesInPic is not greater than 1 (step S1360-no), the image encoding device may omit (skip) encoding slice_address in the slice header.
[0201] Thereafter, in step S1380 , it may be determined whether rect_slice_flag is a first value (eg, 1 or true) and / or whether NumTilesInPic is greater than 1.
[0202] When rect_slice_flag is a first value (e.g., 1 or true) or when NumTilesInPic is not greater than 1 (step S1380-No), the image encoding device may omit (skip) encoding num_tiles_in_slice_minus1 in the slice header. When rect_slice_flag is a second value (e.g., 0 or false) and NumTilesInPic is greater than 1 (step S1380-Yes), the image encoding device may encode num_tiles_in_slice_minus1 in the slice header (S1390).
[0203] Thereafter, the image encoding apparatus may encode the slice header by encoding subsequent syntax elements (not shown) in the slice header.
[0204] In reference Fig.12 and Fig.13 In the described examples, some steps may be changed or omitted. For example, the conditions related to the encoding / decoding of slice_address and / or num_tiles_in_slice_minus1 may be changed.
[0205] Hereinafter, a method for improving image encoding / decoding based on sub-picture will be described. Figures 11 to 13 Methods of embodiments are described.
[0206] The image encoding apparatus may encode the current picture based on the sub-picture. Alternatively, the image encoding apparatus may encode at least one sub-picture configuring the current picture and generate a bitstream including encoding information of the encoded at least one sub-picture.
[0207] The image decoding apparatus may decode at least one sub picture included in the current picture based on a bit stream including encoding information of the at least one sub picture.
[0208] As described above, picture_header_in_slice_header_flag may specify whether a picture header is present in a slice header. Additionally, picture_header_in_slice_header_flag may be used to specify whether the current picture includes only one slice or multiple slices. When the current picture includes only one slice, some syntax elements in the slice header have fixed values because the slice is the only slice in the current picture. In this case, it may be efficient not to signal some syntax elements with fixed values.
[0209] Hereinafter, various configurations of the present disclosure for performing efficient signaling will be described. The following configurations may be applied to embodiments of the present disclosure alone or in combination.
[0210] Configuration 1
[0211] When the current picture includes only one slice, the signaling of some syntax elements in the slice header may be skipped (omitted). The values of the syntax elements whose signaling is omitted may be derived or inferred by the image encoding device and / or the image decoding device.
[0212] Whether the current picture includes only one slice may be indicated by a predetermined indicator. Therefore, when the indicator indicates that the current picture includes only one slice, some syntax elements may not be included in the slice header, and their values may be inferred or derived. At this time, the indicator may be used as a condition indicating whether some syntax elements are included in the slice header.
[0213] Configuration 2
[0214] For example, the indicator described in configuration 1 may be picture_header_in_slice_header_flag.
[0215] Configuration 3
[0216] The syntax elements in the slice header whose signaling can be omitted according to the value of picture_header_in_slice_header_flag may include at least one of the following (a) or (b).
[0217] Syntax elements that specify a sub-picture that includes a slice
[0218] The reason why the signaling of the syntax element (a) can be omitted is because when each picture includes only one slice, the sub-picture is obviously not specified. For example, when picture_header_in_slice_header_flag indicates that the current picture includes only one slice, since the current picture is not encoded / decoded based on the sub-picture, the signaling of the information about the sub-picture can be omitted.
[0219] Syntax element for specifying the address of a slice
[0220] The reason why the signaling of syntax element (b) can be omitted is because it is obvious that the slice is the only first slice in the picture. For example, when picture_header_in_slice_header_fla indicates that the current picture includes only one slice, since the current slice is the only slice in the current picture, the signaling of the address of the current slice can be omitted.
[0221] Configuration 4
[0222] When each picture in the sequence has only one slice, subpictures are not used either. For example, subpics_present_flag or subpic_info_present_flag, which is a syntax element for subpictures, may be restricted to a second value (e.g., 0 or false). subpics_present_flag or subpic_info_present_flag may specify whether a subpicture exists in the current picture or whether information about a subpicture exists in the bitstream. For example, subpics_present_flag or subpic_info_present_flag may be included in a sequence parameter set and signaled.
[0223] Similarly, when subpics_present_flag or subpic_info_present_flag is a first value (e.g., 1 or true), a flag indicating whether each picture in a sequence includes only one slice or a flag indicating whether a picture header exists in a slice header (e.g., picture_header_in_slice_header_flag) may not indicate that the current picture includes only one slice, and may not indicate that a picture header exists in a slice header. Therefore, for example, when subpics_present_flag or subpic_info_present_flag is a first value (e.g., 1 or true), picture_header_in_slice_header_flag may be constrained to have a second value (e.g., 0 or false).
[0224] Configuration 5
[0225] When the picture header does not exist in the picture header NAL unit but exists in the slice header, in the CLVS of a specific layer (layer A), the picture headers of all layers that refer to layer A (i.e., dependent layers of layer A) and all layers referenced by layer A can be constrained to exist in the slice header instead of the picture header NAL unit. The above constraints are imposed to simplify the detection of picture boundaries within an access unit in the case of a multi-layer bitstream.
[0226] Fig.14 is a diagram showing a syntax structure of a slice header according to another embodiment of the present disclosure.
[0227] Because according to Fig.14 Slice header structure and according to the embodiment of Fig.11 The descriptions of the same syntax elements and the same signaling conditions in the slice header structure of the implementation scheme are the same, so the repeated descriptions will be omitted.
[0228] according to Fig.14 In an implementation manner, the condition for signaling slice_subpic_id may be changed. Specifically, the slice header may include slice_subpic_id based on subpics_present_flag and picture_header_in_slice_header_flag. For example, when subpics_present_flag is a first value (e.g., 1 or true) and picture_header_in_slice_header_flag is a second value (e.g., 0 or false), slice_subpic_id may be signaled in the slice header. This is because, as described above, when picture_header_in_slice_header_flag has a first value, the current picture includes only one slice and sub-picture-based encoding / decoding is not performed, and signaling of information about the sub-picture is unnecessary.
[0229] In addition, according to Fig.14In an embodiment of the present invention, the condition for signaling slice_address may be changed. Specifically, the slice header may include slice_address based on rect_slice_flag, NumTilesInPic, and picture_header_in_slice_header_flag. For example, when rect_slice_flag is a first value (e.g., 1 or true) or NumTilesInPic is greater than 1, and picture_header_in_slice_header_flag is a second value (e.g., 0 or false), slice_address may be signaled in the slice header. This is because, as described above, when picture_header_in_slice_header_flag has a first value, since the current picture includes only one slice, signaling of information about the address of the slice is unnecessary.
[0230] exist Fig.14 In an implementation manner, the bitstream consistency requirement of picture_header_in_slice_header_flag can be improved as follows.
[0231] First, the value of picture_header_in_slice_header_flag is required to be the same in all slices in CLVS.
[0232] In addition, when picture_header_in_slice_header_flag is the first value (eg, 1), it is required that NAL units with NAL unit type equal to PH_NUT do not exist in CLVS. This is because the picture header is included in the slice header and signaled, so a separate NAL unit for sending the picture header is not required.
[0233] In addition, when picture_header_in_slice_header_flag is the second value (e.g., 0), a NAL unit with a NAL unit type equal to PH_NUT is required to exist in the PU, before the first VCL NAL unit of the PU. That is, the current PU is required to have a PH NAL unit. This is because a separate NAL unit is required for sending a picture header.
[0234] In addition, when subpics_present_flag or subpic_info_present_flag is the first value (eg, 1), picture_header_in_slice_header_flag is required not to be the first value (eg, 1). In this case, picture_header_in_slice_header_flag may be constrained to have a second value (eg, 0).
[0235] exist Fig.14 In the example of , slice_subpic_id indicates an identifier of a sub-picture including a slice. When slice_subpic_id exists, the variable SubPicIdx is derived such that SubpicIdList[SubPicIdx] is equal to slice_subpic_id. When slice_subpic_id does not exist, the variable SubPicIdx may be derived to be equal to 0.
[0236] exist Fig.14 In the example of , the length (bit length) of slice_subpic_id can be derived as follows.
[0237] If sps_subpic_id_signalling_present_flag is equal to 1, the length of slice_subpic_id is derived to be equal to sps_subpic_id_len_minus1+1. Here, sps_subpic_id_signalling_present_flag may specify whether to signal the identifier of a subpicture in a sequence parameter set. sps_subpic_id_len_minus1 is length information of a subpicture identifier and may be included in a sequence parameter set and signaled.
[0238] Otherwise (if sps_subpic_id_signalling_present_flag is not 1), if ph_subpic_id_signalling_present_flag is 1, the length of slice_subpic_id may be derived to be equal to ph_subpic_id_len_minus1+1. Here, ph_subpic_id_signalling_present_flag may specify whether to signal the identifier of the subpicture in the picture header. ph_subpic_id_len_minus1 is the length information of the subpicture identifier and may be included in the picture header and signaled.
[0239] Otherwise (if both sps_subpic_id_signalling_present_flag and ph_subpic_id_signalling_present_flag are not 1), if pps_subpic_id_signalling_present_flag is 1, the length of slice_subpic_id may be derived to be equal to pps_subpic_id_len_minus1+1. Here, pps_subpic_id_signalling_present_flag may specify whether to signal the identifier of the subpicture in the picture parameter set. pps_subpic_id_len_minus1 is the length information of the subpicture identifier and may be included in the picture parameter set and signaled.
[0240] Otherwise (if all sps_subpic_id_signalling_present_flag, ph_subpic_id_signalling_present_flag and pps_subpic_id_signalling_present_flag are not 1), the length of slice_subpic_id may be derived to be equal to Ceil(Log2(Sps_num_subpics_minus1+1)). sps_num_subpics_minus1 is the number of subpictures of each picture in the CLVS and may be included and signaled in the sequence parameter set.
[0241] slice_address specifies the slice address of the current slice. When slice_address does not exist, the value of slice_address is inferred to be equal to 0.
[0242] picture_header_structure() may include references Fig.10 At least one syntax element included in the described picture_header_rbsp().
[0243] Fig.15 This example illustrates parsing and decoding Fig.14 Flowchart of the method of slicing header.
[0244] Fig.15 Steps S1510 to S1530 are respectively equal to Fig.12 The steps S1210 to S1230 are described above, and thus a repeated description thereof will be omitted.
[0245] Fig.15Steps S1540 to S1570 may correspond to Fig.12 Therefore, repeated description of the common parts will be omitted.
[0246] according to Fig.15 In an embodiment, in step S1540, it can be determined whether subpics_present_flag is a first value (e.g., 1 or true) and whether the first flag is a second value (e.g., 0 or false).
[0247] When subpics_present_flag is the first value and the first flag is the second value (step S1540-yes), the image decoding device may obtain slice_subpic_id from the slice header (S1550). When subpics_present_flag is the second value (e.g., 0 or false) or the first flag is the first value (e.g., 1 or true) (step S1540-no), the image decoding device may omit (skip) parsing slice_subpic_id from the slice header.
[0248] Thereafter, in step S1560 , it may be determined whether rect_slice_flag is a first value (eg, 1 or true) or whether NumTilesInPic is greater than 1 and the first flag is a second value (eg, 0 or false).
[0249] When rect_slice_flag is the first value or NumTilesInPic is greater than 1 and the first flag is the second value (step S1560-yes), the image decoding device can obtain slice_address from the slice header (S1570). When rect_slice_flag is the second value (e.g., 0 or false) and NumTilesInPic is not greater than 1 or the first flag is the first value (e.g., 1 or true) (step S1560-no), the image decoding device can omit (skip) parsing slice_address from the slice header.
[0250] Fig.15 Steps S1580 to S1590 are respectively equal to Fig.12 Steps S1280 to S1290 are described below, and thus repeated descriptions thereof will be omitted.
[0251] As reference Fig.12 As described above, the image decoding apparatus may decode the slice header by parsing subsequent syntax elements (not shown) from the slice header.
[0252] Fig.16 is an example of Fig.14 Flowchart of a method for encoding a slice header.
[0253] Fig.16 Steps S1610 to S1630 are respectively equal to Fig.13 The steps S1310 to S1330 are described above, and thus a repeated description thereof will be omitted.
[0254] Fig.16 Steps S1640 to S1670 may correspond to Fig.16 Therefore, repeated description of the common parts will be omitted.
[0255] according to Fig.16 In an implementation manner, in step S1640, it can be determined whether subpics_present_flag is a first value (e.g., 1 or true) and whether the first flag is a second value (e.g., 0 or false).
[0256] When subpics_present_flag is the first value and the first flag is the second value (step S1640-Yes), the image encoding device may encode slice_subpic_id in the slice header (S1650). When subpics_present_flag is the second value (e.g., 0 or false) or the first flag is the first value (e.g., 1 or true) (step S1640-No), the image encoding device may omit (skip) encoding slice_subpic_id in the slice header.
[0257] Thereafter, in step S1660 , it may be determined whether rect_slice_flag is a first value (eg, 1 or true) or whether NumTilesInPic is greater than 1 and the first flag is a second value (eg, 0 or false).
[0258] When rect_slice_flag is the first value or NumTilesInPic is greater than 1 and the first flag is the second value (step S1660-Yes), the image encoding device may encode slice_address in the slice header (S1670). When rect_slice_flag is the second value (e.g., 0 or false) and NumTilesInPic is not greater than 1 or the first flag is the first value (e.g., 1 or true) (step S1660-No), the image encoding device may omit (skip) encoding slice_address in the slice header.
[0259] Fig.16 Steps S1680 to S1690 are respectively equal to Fig.13 Steps S1380 to S1390 are described below, and thus repeated descriptions thereof will be omitted.
[0260] As reference Fig.13 As described above, the image encoding apparatus may encode the slice header by encoding subsequent syntax elements (not shown) in the slice header.
[0261] In reference Fig.15 and Fig.16 In the described examples, some steps may be changed or omitted. For example, the conditions related to the encoding / decoding of slice_address and / or num_tiles_in_slice_minus1 may be changed.
[0262] As a reference Figures 14 to 16 A modified example of the described implementation, the improved constraint on picture_header_in_slice_header_flag applies to Fig.11 The embodiment shown. In this case, at least some problems of the conventional method can be solved. Specifically, for example, the value of picture_header_in_slice_header_flag can be constrained based on information about the sub-picture (subpics_present_flag or subpic_info_present_flag) signaled at a higher level of the slice header. More specifically, when subpics_present_flag or subpic_info_present_flag is a first value, picture_header_in_slice_header_flag can be constrained to have a second value. Therefore, when subpics_present_flag (or subpic_info_present_flag) is a first value (when information about the sub-picture is present in the bitstream or the current picture includes a sub-picture), picture_header_in_slice_header_flag can indicate that the picture header is not present in the slice header or that the current picture includes more than one slice. Fig.14In the illustrated embodiment, when subpics_present_flag is a first value and picture_header_in_slice_header_flag is a second value, slice_subpic_id may be obtained from the slice header. However, when subpics_present_flag is a first value, since picture_header_in_slice_header_flag is constrained to have a second value, it may be sufficient to check subpics_present_flag as a parsing condition for slice_subpic_id. That is, according to this modified example, in steps S1540 and S1640, the determination as to whether the first flag is a second value may be omitted. According to this modified example, when subpics_present_flag or subpic_info_present_flag is a first value, the image encoding device may encode picture_header_in_slice_header_flag having a second value. In addition, when subpics_present_flag or subpic_info_present_flag is the first value, the image decoding device may acquire picture_header_in_slice_header_flag having the second value.
[0263] Fig.17 is a diagram showing a syntax structure of a slice header according to another embodiment of the present disclosure.
[0264] Because according to Fig.17 Slice header structure and according to the embodiment of Fig.14 The descriptions of the same syntax elements and the same signaling conditions in the slice header structure of the implementation scheme are the same, so the repeated descriptions will be omitted.
[0265] according to Fig.17In an implementation manner, the condition for signaling picture_header_in_slice_header_flag may be changed. Specifically, the slice header may include picture_header_in_slice_header_flag based on subpics_present_flag. For example, when subpics_present_flag is a first value (e.g., 1 or true), picture_header_in_slice_header_flag may not be signaled in the slice header. For example, when subpics_present_flag is a second value (e.g., 0 or false), picture_header_in_slice_header_flag may be signaled in the slice header. This is because, as described above, when subpics_present_flag has a first value, picture_header_in_slice_header_flag has a fixed value (second value) since the current picture cannot contain only one slice. Therefore, signaling of picture_header_in_slice_header_flag is unnecessary. In this case, the picture header signaled if picture_header_in_slice_header_flag is the first value may not be signaled through the slice header.
[0266] In addition, according to Fig.17 In an implementation manner, when subpics_present_flag is a first value, slice_subpic_id may be signaled in the slice header.
[0267] In the following, for the description of slice_address and num_tiles_in_slice_minus1, refer to Fig.14 .
[0268] exist Fig.17 In the implementation of the picture_header_in_slice_header_flag, the bitstream consistency requirements can be the same as those in the reference Fig.14 Same as those described.
[0269] Fig.18 This example illustrates parsing and decoding Fig.17 Flowchart of the method of slicing header.
[0270] according to Fig.18 Methods and basis Fig.15The methods differ in some conditions and order of parsing syntax elements, and the descriptions of the commonly disclosed syntax elements may be the same.
[0271] according to Fig.18 In an implementation manner, the image decoding device may determine in step S1810 whether the value of subpics_present_flag is a second value (eg, 0 or false).
[0272] When the value of subpics_present_flag is the first value (e.g., 1 or true) in step S1810, since sub-picture-based encoding / decoding is performed, the current picture includes more than one slice. Therefore, in this case, the image decoding apparatus may not acquire picture_header_in_slice_header_flag and the picture header from the slice header, but may acquire slice_subpic_id (S1850).
[0273] When the value of subpics_present_flag is the second value in step S1810, the image decoding device obtains the first flag (picture_header_in_slice_header_flag) from the slice header (S1820). The image decoding device may determine whether the first flag is the first value (S1830), and obtain the picture header from the slice header when the first flag is the first value (S1840). When the first flag is the second value, the image decoding device does not obtain the picture header from the slice header, and in this case, the image decoding device may obtain the picture header through a separate NAL unit. When the value of subpics_present_flag is the second value in step S1810, since sub-picture-based encoding / decoding is not performed, the image decoding device may not obtain information about the sub-picture (slice_subpic_id).
[0274] Fig.18 Steps S1860 to S1890 are respectively equal to Fig.15 Steps S1560 to S1590 are described below, and thus repeated descriptions thereof will be omitted.
[0275] As reference Fig.12 As described above, the image decoding apparatus may decode the slice header by parsing subsequent syntax elements (not shown) from the slice header.
[0276] Fig.19 is an example of Fig.17 Flowchart of a method for encoding a slice header.
[0277] according to Fig.19 Methods and basis Fig.16The methods differ in some conditions and order of encoding syntax elements, and the descriptions of the commonly disclosed syntax elements may be the same.
[0278] according to Fig.19 In an implementation manner, the image encoding device may determine whether the value of subpics_present_flag is a second value (eg, 0 or false) in step S1910.
[0279] When the value of subpics_present_flag is the first value (e.g., 1 or true) in step S1910, since sub-picture-based encoding / decoding is performed, the current picture includes more than one slice. Therefore, in this case, the image encoding apparatus may not encode picture_header_in_slice_header_flag and the picture header in the slice header, but may encode slice_subpic_id (S1950).
[0280] When the value of subpics_present_flag is the second value in step S1910, the image encoding apparatus may determine the value of the first flag (picture_header_in_slice_header_flag) and encode the first flag in the slice header (S1920). The image encoding apparatus may determine whether the first flag is the first value (S1930), and encode the picture header in the slice header when the first flag is the first value (S1940). When the first flag is the second value, the image encoding apparatus may not encode the picture header in the slice header, and in this case, the image encoding apparatus may signal the picture header through a separate NAL unit. When the value of subpics_present_flag is the second value in step S1910, since sub-picture-based encoding / decoding is not performed, the image encoding apparatus may not encode information about the sub-picture (slice_subpic_id) in the slice header.
[0281] Fig.19 Steps S1960 to S1990 are respectively equal to Fig.16 Steps S1660 to S1690 are described above, and thus repeated descriptions thereof will be omitted.
[0282] As reference Fig.13 As described above, the image encoding apparatus may encode the slice header by encoding subsequent syntax elements (not shown) in the slice header.
[0283] In reference Fig.18 and Fig.19In the example described, some steps may be changed or omitted. For example, the conditions related to encoding / decoding of slice_address and / or num_tiles_in_slice_minus1 may be changed.
[0284] According to an embodiment of the present disclosure, information on whether a picture header is present in a slice header and / or information on whether a picture includes only one slice may be signaled more efficiently.
[0285] In addition, according to an embodiment of the present disclosure, since information on whether a picture header exists in a slice header is signaled based on whether sub-picture based encoding / decoding is performed, signaling of unnecessary information can be prevented.
[0286] The names of the syntax elements described in the present disclosure may include information about the location of the corresponding syntax element being signaled. For example, a syntax element starting with "sps_" may mean that the corresponding syntax element is signaled in a sequence parameter set (SPS). In addition, a syntax element starting with "pps_", "ph_", "sh_", etc. means that the corresponding syntax element is signaled in a picture parameter set (PPS), a picture header, and a slice header, respectively.
[0287] Although the exemplary method of the present disclosure described above is represented as a series of operations for the sake of clarity of description, it is not intended to limit the order of executing the steps, and these steps can be performed simultaneously or in different orders when necessary. In order to implement the method according to the present disclosure, the steps described may further include other steps, may include the remaining steps except some steps, or may include other additional steps except some steps.
[0288] In the present disclosure, an image encoding device or an image decoding device that performs a predetermined operation (step) may perform an operation (step) of confirming an execution condition or situation of the corresponding operation (step). For example, if it is described that a predetermined operation is performed when a predetermined condition is met, the image encoding device or the image decoding device may perform the predetermined operation after determining whether the predetermined condition is met.
[0289] The various embodiments of the present disclosure are not a list of all possible combinations and are intended to describe representative aspects of the present disclosure, and matters described in the various embodiments may be applied independently or in combination of two or more.
[0290] Various embodiments of the present disclosure may be implemented in hardware, firmware, software or a combination thereof. In the case of implementing the present disclosure in hardware, the present disclosure may be implemented in an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a general purpose processor, a controller, a microcontroller, a microprocessor, etc.
[0291] In addition, the image decoding device and the image encoding device of the embodiment of the present disclosure can be included in multimedia broadcast transmission and reception equipment, mobile communication terminals, home theater video equipment, digital theater video equipment, surveillance cameras, video chat equipment, real-time communication equipment such as video communication, mobile streaming equipment, storage media, cameras, video on demand (VoD) service providing equipment, OTT video (over the top video) equipment, Internet streaming service providing equipment, three-dimensional (3D) video equipment, video phone video equipment, medical video equipment, etc., and can be used to process video signals or data signals. For example, OTT video equipment can include game consoles, Blu-ray players, Internet access TVs, home theater systems, smart phones, tablet PCs, digital video recorders (DVRs), etc.
[0292] Fig. 20 is a diagram showing a content streaming system to which an embodiment of the present disclosure can be applied.
[0293] like Fig. 20 As shown in , a content streaming transmission system to which the embodiments of the present disclosure are applied may mainly include an encoding server, a streaming transmission server, a network server, a media storage, a user device, and a multimedia input device.
[0294] The encoding server compresses the content input from a multimedia input device such as a smartphone, camera, camcorder, etc. into digital data to generate a bitstream and sends the bitstream to the streaming server. As another example, when a multimedia input device such as a smartphone, camera, camcorder, etc. directly generates a bitstream, the encoding server can be omitted.
[0295] A bitstream may be generated by an image encoding method or an image encoding device to which an embodiment of the present disclosure is applied, and a streaming server may temporarily store the bitstream in the process of transmitting or receiving the bitstream.
[0296] The streaming server sends multimedia data to the user device based on the user's request through the network server, and the network server serves as a medium to notify the user of the service. When the user requests the required service from the network server, the network server can deliver it to the streaming server, and the streaming server can send the multimedia data to the user. In this case, the content streaming system may include a separate control server. In this case, the control server serves as a command / response between devices in the control content streaming system.
[0297] The streaming server may receive content from a media storage and / or encoding server. For example, when content is received from an encoding server, the content may be received in real time. In this case, in order to provide a smooth streaming service, the streaming server may store the bitstream for a predetermined time.
[0298] Examples of user devices may include mobile phones, smart phones, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, tablet PCs, tablet computers, ultrabooks, wearable devices (e.g., smart watches, smart glasses, head-mounted displays), digital televisions, desktop computers, digital signage, etc.
[0299] Each server in the content streaming system may operate as a distributed server, in which case data received from each server may be distributed.
[0300] The scope of the present disclosure includes software or executable commands (e.g., operating systems, applications, firmware, programs, etc.) for enabling operations according to various embodiments of the methods to be performed on a device or computer, and a non-transitory computer-readable medium having such software or commands stored thereon and executable on a device or computer.
[0301] Industrial Applicability
[0302] The embodiments of the present disclosure may be used to encode or decode an image.
Claims
1. A method for decoding an image, the method comprising: The following steps are involved: Get a first flag of whether information about a sub-picture exists in a specified bitstream; Get the second flag of whether the picture header information exists in the specified slice header; obtaining slice address information specifying a slice address of the slice, the obtaining of the slice address information being associated with the first flag specifying the presence of the information about the sub-picture in the bitstream, wherein when the first flag is a first value of true, the second flag is constrained to have a second value of false; and decoding the bit stream based on the first flag, the second flag and the slice address information, Wherein, based on the first flag specifying that the information about the sub-picture exists in the bitstream, the second flag has a value specifying that the picture header information does not exist in the slice header.
2. The image decoding method according to claim 1, in, The slice header includes an identifier of a sub-picture based on the first flag specifying that information about the sub-picture exists in the bitstream, and the sub-picture includes a slice related to the slice header.
3. The image decoding method according to claim 1, further comprising: include: The picture header information is acquired from the slice header based on the second flag specifying that the picture header information exists in the slice header.
4. The image decoding method according to claim 1, in, The second flag has the same value with respect to all slices in the coding layer video sequence CLVS.
5. The image decoding method according to claim 1, in, Based on the second flag specifying that the picture header information exists in the slice header, there is no network abstraction layer NAL unit for sending the picture header information in the coding layer video sequence CLVS.
6. The image decoding method according to claim 1, in, Based on the second flag specifying that the picture header information does not exist in the slice header, the picture header information is obtained from a NAL unit whose network abstraction layer NAL unit type is equal to PH_NUT.
7. The image decoding method according to claim 1, in, The first flag is signaled at a higher level of the slice, and The second flag is included in the slice header and signaled.
8. A method for encoding an image, the method comprising: The following steps are involved: encoding a first flag specifying whether information about a sub-picture exists in the bitstream; Encoding a second flag indicating whether picture header information exists in a specified slice header; encoding slice address information specifying a slice address of the slice, encoding the slice address information being associated with the first flag specifying the presence of the information about the sub-picture in the bitstream, wherein when the first flag is a first value of true, the second flag is constrained to have a second value of false; and encoding the bitstream based on the first flag, the second flag and the slice address information, Wherein, based on the first flag specifying that the information about the sub-picture exists in the bitstream, the second flag has a value specifying that the picture header information does not exist in the slice header.
9. The image encoding method according to claim 8, in, The slice header includes an identifier of a sub-picture based on the first flag specifying that information about the sub-picture exists in the bitstream, and the sub-picture includes a slice related to the slice header.
10. The image encoding method according to claim 8, further comprising: include: The picture header information is encoded in the slice header based on the second flag specifying that the picture header information exists in the slice header.
11. The image encoding method according to claim 8, in, The second flag has the same value with respect to all slices in the coding layer video sequence CLVS.
12. The image encoding method according to claim 8, in, Based on the second flag specifying that the picture header information is not present in the slice header, the picture header information is signaled through a NAL unit with a network abstraction layer NAL unit type equal to PH_NUT.
13. The image encoding method according to claim 8, in, The first flag is signaled at a higher level of the slice, and The second flag is included in the slice header and signaled.
14. A method for transmitting a bit stream generated by an image encoding method, the image encoding method The following steps are involved: encoding a first flag specifying whether information about a sub-picture exists in the bitstream; Encoding a second flag indicating whether picture header information exists in a specified slice header; encoding slice address information specifying a slice address of the slice, encoding the slice address information being associated with the first flag specifying the presence of the information about the sub-picture in the bitstream, wherein when the first flag is a first value of true, the second flag is constrained to have a second value of false; and encoding the bitstream based on the first flag, the second flag and the slice address information, Wherein, based on the first flag specifying that the information about the sub-picture exists in the bitstream, the second flag has a value specifying that the picture header information does not exist in the slice header.
15. A non-transitory computer-readable recording medium storing a computer program, which, when executed by a processor, implements an image encoding method, the image encoding method The following steps are involved: encoding a first flag specifying whether information about a sub-picture exists in the bitstream; Encoding a second flag indicating whether picture header information exists in a specified slice header; encoding slice address information specifying a slice address of the slice, encoding the slice address information being associated with the first flag specifying the presence of the information about the sub-picture in the bitstream, wherein when the first flag is a first value of true, the second flag is constrained to have a second value of false; and encoding the bitstream based on the first flag, the second flag and the slice address information, Wherein, based on the first flag specifying that the information about the sub-picture exists in the bitstream, the second flag has a value specifying that the picture header information does not exist in the slice header.