Image encoding / decoding device and bit stream transmission device
By acquiring signaling flags and decoding sub-picture information in the image decoding device, the problem of low encoding/decoding efficiency in high-resolution image transmission is solved, and more efficient image transmission and storage is achieved.
Patent Information
- Application Number
- CN202510498346.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-23
- Filing Date
- 2020-09-03
- Publication Date
- 2025-07-25
AI Technical Summary
The prior art has low encoding/decoding efficiency when transmitting high resolution and high quality images, resulting in increased transmission and storage costs.
The image decoding device obtains signaling flags from the picture parameter set, determines whether to send the identifier of the sub-screen using the picture parameter set, and decodes the current screen based on the identifier information. The image encoding device divides the screen into the sub-screen and generates a picture parameter set including the signaling flag and the identifier.
Improve image encoding/decoding efficiency and reduce transmission and storage costs.
Smart Images

Figure CN120378611A_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application with the original application number 202080074435.4 (International Application No.: PCT / KR2020 / 011836, Application Date: September 3, 2020, Invention Title: Method and Apparatus for Image Encoding / Decoding Using Sub-Pictures and Bitstream Transmission Method). Technical Field
[0002] The present disclosure relates to an image encoding / decoding method and apparatus, and more particularly, to an image encoding / decoding method and apparatus using sub-pictures, and a method of transmitting a bitstream generated by the image encoding method / apparatus of the present disclosure. Background Art
[0003] Recently, the demand for high-resolution and high-quality images, such as high-definition (HD) images and ultra-high-definition (UHD) images, is increasing in various fields. As the resolution and quality of image data are improved, the amount of information or bits to be transmitted increases relatively compared to existing image data. The increase in the amount of information or bits to be transmitted leads to an increase in transmission cost and storage cost.
[0004] Therefore, an efficient image compression technology is needed to effectively transmit, store, and reproduce information on high-resolution and high-quality images. Summary of the Invention
[0005] Technical Problem
[0006] An object of the present disclosure is to provide an image encoding / decoding method and apparatus having improved encoding / decoding efficiency.
[0007] An object of the present disclosure is to provide an image encoding / decoding method and apparatus for improving encoding / decoding efficiency by effectively signaling sub-picture information.
[0008] Another object of the present disclosure is to provide a method of transmitting a bitstream generated by an image encoding method or apparatus according to the present disclosure.
[0009] Another object of the present disclosure is to provide a recording medium storing a bitstream generated by an image encoding method or apparatus according to the present disclosure.
[0010] Another object of the present disclosure is to provide a recording medium storing a bitstream received, decoded, and used for reconstructing an image by an image decoding apparatus according to the present disclosure.
[0011] The technical problems solved by the present disclosure are not limited to the above technical problems, and other technical problems not described herein will be apparent to those skilled in the art from the following description.
[0012] Technical Solution
[0013] An image decoding method performed by an image decoding device according to an aspect of the present disclosure may include: obtaining a first signaling flag from a picture parameter set, the first signaling flag specifying whether to use the picture parameter set to signal an identifier of a sub-picture that divides a current picture, obtaining sub-picture identifier information from the picture parameter set based on the first signaling flag specifying that the identifier of the sub-picture is signaled using the picture parameter set, and decoding the current picture by decoding the current picture identified based on the identifier information.
[0014] An image decoding device according to an aspect of the present disclosure may include a memory and at least one processor. The at least one processor may obtain a first signaling flag from a picture parameter set, the first signaling flag specifying whether to use the picture parameter set to signal an identifier of a sub-picture that divides a current picture, obtain sub-picture identifier information from the picture parameter set based on the first signaling flag specifying that the identifier of the sub-picture is signaled using the picture parameter set, and decode the current picture by decoding the current sub-picture identified based on the identifier information.
[0015] An image encoding method performed by an image encoding device according to an aspect of the present disclosure may include: determining a sub-picture by dividing a current picture, determining whether to use a picture parameter set to signal an identifier of the sub-picture, and generating a picture parameter set including a first signaling flag and sub-picture identifier information based on using the picture parameter set to signal the identifier of the sub-picture, the first signaling flag specifying whether to use the picture parameter set to signal the identifier of the sub-picture.
[0016] In addition, a transmission method according to another aspect of the present disclosure may transmit a bitstream generated by the image encoding device or image encoding method of the present disclosure.
[0017] In addition, a computer-readable recording medium according to another aspect of the present disclosure may store a bitstream generated by the image encoding device or image encoding method of the present disclosure.
[0018] The features of the brief overview of the present disclosure above are merely exemplary aspects of the following detailed description of the present disclosure and do not limit the scope of the present disclosure.
[0019] Advantageous Effects
[0020] According to the present disclosure, it is possible to provide an image encoding / decoding method and device having improved encoding / decoding efficiency.
[0021] In addition, according to the present disclosure, it is possible to provide an image encoding / decoding method and device that improve encoding / decoding efficiency by effectively signaling sub-picture information.
[0022] In addition, according to the present disclosure, a method of transmitting a bitstream generated by an image encoding method or apparatus according to the present disclosure can be provided.
[0023] In addition, according to the present disclosure, a recording medium storing a bitstream generated by an image encoding method or apparatus according to the present disclosure can be provided.
[0024] In addition, according to the present disclosure, a recording medium storing a bitstream received, decoded, and used for reconstructing an image by an image decoding apparatus according to the present disclosure can be provided.
[0025] Those skilled in the art will understand that the effects achievable through the present disclosure are not limited to what has been specifically described above, and other advantages of the present disclosure will be more clearly understood from the detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 is a view schematically showing a video coding system to which an embodiment of the present disclosure is applicable.
[0027] Figure 2 is a view schematically showing an image encoding apparatus to which an embodiment of the present disclosure is applicable.
[0028] Figure 3 is a view schematically showing an image decoding apparatus to which an embodiment of the present disclosure is applicable.
[0029] Figure 4 is a view showing a segmentation structure of an image according to an embodiment.
[0030] Figure 5 is a view showing an embodiment of a segmentation type of a block according to a multi-type tree structure.
[0031] Figure 6 is a view showing a signaling mechanism of block partition information in a quadtree having a nested multi-type tree structure according to the present disclosure.
[0032] Figure 7 is a view showing an embodiment of dividing a CTU into a plurality of CUs.
[0033] Figure 8 is a view illustrating an embodiment of a syntax for signaling sub-picture syntax elements in an SPS.
[0034] Figure 9 is a view illustrating an embodiment of an algorithm for deriving a predetermined variable such as SubPictTop.
[0035] Figure 10 is a view illustrating a method of encoding an image using a sub-picture by an encoding apparatus according to an embodiment.
[0036] Figure 11 It is a view illustrating a method of decoding an image using a sub - picture by a decoding device according to an embodiment.
[0037] Figure 12 It is a view illustrating an embodiment of identifying the size and position of a sub - picture using a sub - picture grid index defined in the SPS.
[0038] Figure 13 It is a view illustrating an example of the SPS syntax structure in this embodiment.
[0039] Figure 14 It is a view illustrating a method of identifying a sub - picture grid based on slices by a decoding device according to an embodiment.
[0040] Figure 15 It is a view illustrating for Figure 12 an example of the PPS syntax structure of sub - picture index signaling.
[0041] Figure 16 It is a view illustrating signaling of syntax elements specifying constraints on pps_subpics_present_flag.
[0042] Figure 17 It is a view showing the SPS syntax structure for signaling a sub - picture grid.
[0043] Figure 18 It is a view illustrating an embodiment of an algorithm for deriving a predetermined variable such as SubPicIdx.
[0044] Figure 19 It is a view illustrating an embodiment of an algorithm for deriving CtbToSubPicIdx.
[0045] Figure 20 It is a view illustrating an embodiment of the SPS syntax according to Embodiment 3.
[0046] Figure 21 It is a view illustrating, according to Embodiment 4, Figure 9 a modified embodiment of the algorithm.
[0047] Figures 22 to 24 It is a view illustrating, according to Embodiment 5, Figures 17 to 19 a modified embodiment of the example.
[0048] Figures 25 to 26 It is a view illustrating the operations of a decoding device and an encoding device according to an embodiment.
[0049] Figure 27It is a view showing a content streaming system to which embodiments of the present disclosure can be applied. Detailed Embodiments
[0050] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings to facilitate implementation by those skilled in the art. However, the present disclosure can be implemented in various different forms and is not limited to the embodiments described herein.
[0051] When describing the present disclosure, if it is determined that the detailed description of related known functions or configurations makes the scope of the present disclosure unnecessarily ambiguous, the detailed description thereof will be omitted. In the drawings, parts irrelevant to the description of the present disclosure are omitted, and similar reference numerals are assigned to similar parts.
[0052] In the present disclosure, when a component is "connected", "coupled" or "linked" to another component, it may include not only a direct connection relationship but also an indirect connection relationship in which an intermediate component exists. Additionally, when a component "includes" or "has" other components, unless otherwise specified, it means that other components may also be included, rather than excluding other components.
[0053] In the present disclosure, terms such as first and second are only used for the purpose of distinguishing one component from other components and do not limit the order or importance of the components, unless otherwise specified. Accordingly, within the scope of the present disclosure, the first component in one embodiment may be referred to as the second component in another embodiment, and similarly, the second component in one embodiment may be referred to as the first component in another embodiment.
[0054] In the present disclosure, components that are mutually distinguishable are intended to clearly describe each feature and do not mean that the components must be separated. That is, multiple components can be implemented integrated in one hardware or software unit, or one component can be distributed and implemented in multiple hardware or software units. Therefore, even without specific description, these embodiments of integration or distribution of components are included within the scope of the present disclosure.
[0055] In the present disclosure, the components described in each embodiment are not necessarily essential components, and some components may be optional components. Therefore, embodiments consisting of a subset of the components described in the embodiments are also included within the scope of the present disclosure. In addition, embodiments that include other components in addition to the components described in the various embodiments are included within the scope of the present disclosure.
[0056] The present disclosure relates to the encoding and decoding of images. Unless redefined in the present disclosure, the terms used in the present disclosure may have the general meanings commonly used in the technical field to which the present disclosure belongs.
[0057] In the present disclosure, a "picture" generally refers to a unit representing an image within a specific time period, while a slice / tile is an encoding unit that forms part of a picture, and a picture can be composed of one or more slices / tiles. Additionally, a slice / tile can include one or more coding tree units (CTUs).
[0058] In the present disclosure, a "pixel" or "pel" can mean the smallest unit that constitutes a picture (or image). Additionally, "sample" can be used as a term corresponding to a pixel. A sample generally can represent a pixel or the value of a pixel, or can represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component.
[0059] In the present disclosure, a "unit" can represent a basic unit of image processing. The unit can include at least one of a specific area of a picture and information related to that area. In some cases, the unit can be used interchangeably with terms such as "sample array", "block", or "region". Generally, an M×N block can include a set (or array) of samples (or sample arrays) or transform coefficients in M columns and N rows.
[0060] In the present disclosure, "current block" can mean one of "current coding block", "current coding unit", "encoding target block", "decoding target block", or "processing target block". When performing prediction, "current block" can mean "current prediction block" or "prediction target block". When performing transform (inverse transform) / quantization (dequantization), "current block" can mean "current transform block" or "transform target block". When performing filtering, "current block" can mean "filtering target block".
[0061] Additionally, in the present disclosure, unless explicitly stated as a chrominance block, "current block" can mean "luminance block of the current block". "Chrominance block of the current block" can be expressed by including an explicit description of a chrominance block such as "chrominance block" or "current chrominance block".
[0062] In the present disclosure, the term " / " or "," can be interpreted as indicating "and / or". For example, "A / B" and "A,B" can mean "A and / or B". Additionally, "A / B / C" and "A,B,C" can mean "at least one of A, B, and / or C".
[0063] In the present disclosure, the term "or" should be interpreted to indicate "and / or". For example, the expression "A or B" can include 1) only "A", 2) only "B", or 3) both "A and B". In other words, in the present disclosure, "or" should be interpreted to indicate "additionally or alternatively".
[0064] Overview of video encoding system
[0065] Figure 1 is a view schematically showing a video coding system according to the present disclosure.
[0066] The video coding system according to an embodiment may include an encoding device 10 and a decoding device 20. The encoding device 10 may deliver encoded video and / or image information or data in the form of a file or a stream to the decoding device 20 via a digital storage medium or a network.
[0067] The encoding device 10 according to an embodiment may include a video source generator 11, an encoding unit 12, and a transmitter 13. The decoding device 20 according to an embodiment may include a receiver 21, a decoding unit 22, and a renderer 23. The encoding unit 12 may be referred to as a video / image encoding unit, and the decoding unit 22 may be referred to as a video / image decoding unit. The transmitter 13 may be included in the encoding unit 12. The receiver 21 may be included in the decoding unit 22. The renderer 23 may include a display, and the display may be configured as a separate device or an external component.
[0068] The video source generator 11 may obtain video / images through processes such as capturing, synthesizing, or generating video / images. The video source generator 11 may include a video / image capturing device and / or a video / image generating device. The video / image capturing device may include, for example, one or more cameras, a video / image archive including previously captured video / images, etc. The video / image generating device may include, for example, a computer, a tablet computer, and a smart phone, and may (electronically) generate video / images. For example, virtual video / images may be generated by a computer or the like. In this case, the video / image capturing process may be replaced by a process of generating relevant data.
[0069] The encoding unit 12 may encode the input video / images. For compression and encoding efficiency, the encoding unit 12 may perform a series of processes such as prediction, transformation, and quantization. The encoding unit 12 may output encoded data (encoded video / image information) in the form of a bitstream.
[0070] The transmitter 13 may transmit the encoded video / image information or data output in the form of a bitstream to the receiver 21 of the decoding device 20 in the form of a file or a stream via a digital storage medium or a network. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter 13 may include an element for generating a media file in a predetermined file format and may include an element for transmitting via a broadcast / communication network. The receiver 21 may extract / receive the bitstream from the storage medium or the network and transmit the bitstream to the decoding unit 22.
[0071] The decoding unit 22 can decode video / images by performing a series of processes corresponding to the operations of the encoding unit 12, such as dequantization, inverse transformation, and prediction.
[0072] The renderer 23 can render the decoded video / images. The rendered video / images can be displayed via a display.
[0073] Overview of image encoding device
[0074] Figure 2 is a view schematically showing an image encoding device to which embodiments of the present disclosure can be applied.
[0075] As Figure 2 shown, the image encoding device 100 can include an image splitter 110, a subtractor 115, a transformer 120, a quantizer 130, a dequantizer 140, an inverse transformer 150, an adder 155, a filter 160, a memory 170, an inter-frame prediction unit 180, an intra-frame prediction unit 185, and an entropy encoder 190. The inter-frame prediction unit 180 and the intra-frame prediction unit 185 can be collectively referred to as a "prediction unit". The transformer 120, the quantizer 130, the dequantizer 140, and the inverse transformer 150 can be included in a residual processor. The residual processor may further include the subtractor 115.
[0076] In some embodiments, all or at least some of the multiple components configuring the image encoding device 100 can be configured by one hardware component (e.g., an encoder or a processor). Additionally, the memory 170 can include a decoded picture buffer (DPB) and can be configured by a digital storage medium.
[0077] The image splitter 110 may split an input image (or picture or frame) input to the image encoding device 100 into one or more processing units. For example, a processing unit may be referred to as a coding unit (CU). Coding units may be obtained by recursively splitting a coding tree unit (CTU) or a largest coding unit (LCU) according to a quadtree binary tree ternary tree (QT / BT / TT) structure. For example, a coding unit may be split into multiple coding units of a deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. For the splitting of a coding unit, a quadtree structure may be applied first, and then a binary tree structure and / or a ternary tree structure may be applied. Encoding processing according to the present disclosure may be performed based on the final coding units that are no longer splittable. The largest coding unit may be used as the final coding unit, or coding units of a deeper depth obtained by splitting the largest coding unit may be used as the final coding units. Here, the encoding processing may include processes of prediction, transformation, and reconstruction that will be described later. As another example, the processing unit of the encoding processing may be a prediction unit (PU) or a transformation unit (TU). The prediction unit and the transformation unit may be divided or split from the final coding unit. The prediction unit may be a sample prediction unit, and the transformation unit may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.
[0078] The prediction unit (inter-frame prediction unit 180 or intra-frame prediction unit 185) may perform prediction on a block to be processed (current block) and generate a prediction block including prediction samples of the current block. The prediction unit may determine whether to apply intra-frame prediction or inter-frame prediction based on the current block or CU. The prediction unit may generate various information related to the prediction of the current block and transmit the generated information to the entropy encoder 190. Information about the prediction may be encoded in the entropy encoder 190 and output in the form of a bitstream.
[0079] The intra-frame prediction unit 185 may predict the current block by referring to samples in the current picture. According to the intra-frame prediction mode and / or intra-frame prediction technique, the reference samples may be located among the neighbors of the current block or may be placed separately. The intra-frame prediction mode may include multiple non-directional modes and multiple directional modes. The non-directional modes may include, for example, the DC mode and the planar mode. According to the detail level of the prediction direction, the directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes. However, this is merely an example, and more or fewer directional prediction modes may be used according to the setting. The intra-frame prediction unit 185 may determine the prediction mode applied to the current block by using the prediction mode applied to neighboring blocks.
[0080] The inter - frame prediction unit 180 may derive a prediction block of a current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, in order to reduce the amount of motion information transmitted in the inter - frame prediction mode, the motion information may be predicted in units of blocks, sub - blocks, or samples based on the correlation of the motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may also include inter - frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter - frame prediction, neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block may be referred to as a collocated reference block, a collocated CU (colCU), etc. The reference picture including the temporal neighboring block may be referred to as a collocated picture (colPic). For example, the inter - frame prediction unit 180 may configure a motion information candidate list based on neighboring blocks and generate information specifying which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter - frame prediction may be performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the inter - frame prediction unit 180 may use the motion information of neighboring blocks as the motion information of the current block. In the case of the skip mode, different from the merge mode, the residual signal may not be transmitted. In the case of the motion vector prediction (MVP) mode, the motion vector of a neighboring block may be used as a motion vector predictor, and the motion vector of the current block may be signaled by encoding the motion vector difference and an indicator of the motion vector predictor. The motion vector difference may mean the difference between the motion vector of the current block and the motion vector predictor.
[0081] The prediction unit may generate a prediction signal based on various prediction methods and prediction techniques described below. For example, the prediction unit may not only apply intra - frame prediction or inter - frame prediction, but also apply both intra - frame prediction and inter - frame prediction simultaneously to predict the current block. The prediction method of applying both intra - frame prediction and inter - frame prediction simultaneously to predict the current block may be referred to as combined intra - and inter - frame prediction (CIIP). In addition, the prediction unit may perform intra - block copy (IBC) to predict the current block. Intra - block copy may be used for content image / video coding such as games, for example, screen content coding (SCC). IBC is a method of predicting the current picture using a previously reconstructed reference block in the current picture at a position separated from the current block by a predetermined distance. When IBC is applied, the position of the reference block in the current picture may be encoded as a vector (block vector) corresponding to the predetermined distance. IBC basically performs prediction in the current picture, but may be performed similarly to inter - frame prediction because the reference block is derived within the current picture. That is, IBC may use at least one of the inter - frame prediction techniques described in the present disclosure.
[0082] The prediction signal generated by the prediction unit can be used to generate a reconstructed signal or a residual signal. The subtractor 115 can generate a residual signal (residual block or residual sample array) by subtracting the prediction signal (prediction block or prediction sample array) output from the prediction unit from the input image signal (original block or original sample array). The generated residual signal can be transmitted to the transformer 120.
[0083] The transformer 120 can generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique can include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a karhunen-loève transform (KLT), a graph-based transform (GBT), or a conditional non-linear transform (CNT). Here, GBT refers to a transform obtained from a graph when the relationship information between pixels is represented by a graph. CNT refers to a transform obtained based on a prediction signal generated using all previously reconstructed pixels. In addition, the transform processing can be applied to a square pixel block having the same size or can be applied to a block having a variable size rather than a square.
[0084] The quantizer 130 can quantize the transform coefficients and transmit them to the entropy encoder 190. The entropy encoder 190 can encode the quantized signal (information about the quantized transform coefficients) and output a bitstream. The information about the quantized transform coefficients can be referred to as residual information. The quantizer 130 can rearrange the block-form quantized transform coefficients into a one-dimensional vector form based on the coefficient scan order and generate information about the quantized transform coefficients based on the one-dimensional vector form of the quantized transform coefficients.
[0085] The entropy encoder 190 can perform various encoding methods, such as exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. The entropy encoder 190 can encode information required for video / image reconstruction other than the quantized transform coefficients (e.g., values of syntax elements, etc.) together or separately. The encoded information (e.g., encoded video / image information) can be transmitted or stored in the form of a bitstream in units of network abstraction layer (NAL). The video / image information can also include information about various parameter sets, such as an adaptive parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information can also include general constraint information. The information signaled, transmitted, and / or syntax elements described in the present disclosure can be encoded by the above encoding process and included in the bitstream.
[0086] The bitstream can be transmitted over a network or stored in a digital storage medium. The network can include a broadcast network and / or a communication network, and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) for transmitting the signal output from the entropy encoder 190 and / or a storage unit (not shown) for storing the signal can be included as internal / external elements of the image encoding device 100. Alternatively, a transmitter can be provided as a component of the entropy encoder 190.
[0087] The quantized transform coefficients output from the quantizer 130 can be used to generate a residual signal. For example, the residual signal (residual block or residual sample) can be reconstructed by applying dequantization and inverse transformation to the quantized transform coefficients through the dequantizer 140 and the inverse transform unit 150.
[0088] The adder 155 adds the reconstructed residual signal to the prediction signal output from the inter-frame prediction unit 180 or the intra-frame prediction unit 185 to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). If the block to be processed has no residual, for example, in the case of applying the skip mode, the predicted block can be used as the reconstructed block. The adder 155 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current picture and can be used for inter-frame prediction of the next picture through filtering as described below.
[0089] The filter 160 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 160 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture and store the modified reconstructed picture in the memory 170, specifically, in the DPB of the memory 170. Various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filtering, bilateral filtering, etc. The filter 160 can generate various information related to the filtering and transmit the generated information to the entropy encoder 190, as described later in the description of each filtering method. The information related to the filtering can be encoded by the entropy encoder 190 and output in the form of a bitstream.
[0090] The modified reconstructed picture transmitted to the memory 170 can be used as a reference picture in the inter-frame prediction unit 180. When inter-frame prediction is applied by the image encoding device 100, prediction mismatch between the image encoding device 100 and the image decoding device can be avoided and the encoding efficiency can be improved.
[0091] The DPB of the memory 170 may store the modified reconstructed picture to be used as a reference picture in the inter - prediction unit 180. The memory 170 may store the motion information of the blocks from which the motion information in the current picture is derived (or encoded) and / or the motion information of the blocks that have been reconstructed in the picture. The stored motion information may be transmitted to the inter - prediction unit 180 and used as the motion information of spatially neighboring blocks or temporally neighboring blocks. The memory 170 may store the reconstructed samples of the reconstructed blocks in the current picture and may transmit the reconstructed samples to the intra - prediction unit 185.
[0092] Overview of image decoding device
[0093] Figure 3 is a view schematically showing an image decoding device to which an embodiment of the present disclosure is applicable.
[0094] As Figure 3 shown, the image decoding device 200 may include an entropy decoder 210, a de - quantizer 220, an inverse transformer 230, an adder 235, a filter 240, a memory 250, an inter - prediction unit 260, and an intra - prediction unit 265. The inter - prediction unit 260 and the intra - prediction unit 265 may be collectively referred to as a "prediction unit". The de - quantizer 220 and the inverse transformer 230 may be included in a residual processor.
[0095] According to an embodiment, all or at least some of the multiple components configuring the image decoding device 200 may be configured by hardware components (e.g., a decoder or a processor). In addition, the memory 250 may include a decoded picture buffer (DPB) or may be configured by a digital storage medium.
[0096] The image decoding device 200 that has received a bitstream including video / image information may reconstruct an image by performing a process corresponding to the process performed by Figure 2 the image encoding device 100. For example, the image decoding device 200 may perform decoding using the processing units applied in the image encoding device. Therefore, the decoding processing units may be, for example, encoding units. The encoding units may be obtained by dividing a coding tree unit or a maximum coding unit. The reconstructed image signal decoded and output by the image decoding device 200 may be reproduced by a reproducing device (not shown).
[0097] The image decoding device 200 may receive in the form of a bitstream from Figure 2The signal output by the image encoding device. The received signal can be decoded by the entropy decoder 210. For example, the entropy decoder 210 can parse the bitstream to derive the information (e.g., video / image information) required for image reconstruction (or picture reconstruction). The video / image information can also include information about various parameter sets, such as Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). In addition, the video / image information can also include general constraint information. The image decoding device can also decode the picture based on the information about the parameter set and / or the general constraint information. The information and / or syntax elements signaled / received described in this disclosure can be decoded through the decoding process and obtained from the bitstream. For example, the entropy decoder 210 decodes the information in the bitstream based on an encoding method such as Exponential Golomb coding, CAVLC, or CABAC, and outputs the values of the syntax elements required for image reconstruction and the quantization values of the transform coefficients of the residuals. More specifically, the CABAC entropy decoding method can receive the bins corresponding to each syntax element in the bitstream, determine the context model using the decoding target syntax element information, neighboring blocks, and the decoding information of the decoding target block or the information of the symbols / bins decoded in the previous stage, perform arithmetic decoding on the bins by predicting the occurrence probability of the bins according to the determined context model, and generate the symbols corresponding to the values of each syntax element. In this case, after determining the context model, the CABAC entropy decoding method can update the context model by using the information of the decoded symbols / bins for the context model of the next symbol / bin. The information related to prediction in the information decoded by the entropy decoder 210 can be provided to the prediction units (inter-frame prediction unit 260 and intra-frame prediction unit 265), and the residual values for which entropy decoding is performed in the entropy decoder 210, that is, the quantized transform coefficients and the related parameter information, can be input to the dequantizer 220. In addition, the information about filtering among the information decoded by the entropy decoder 210 can be provided to the filter 240. Furthermore, the receiver (not shown) for receiving the signal output by the image encoding device can be further configured as an internal / external component of the image decoding device 200, or the receiver can be a component of the entropy decoder 210.
[0098] In addition, the image decoding device according to the present disclosure can be referred to as a video / image / picture decoding device. The image decoding device can be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder can include the entropy decoder 210. The sample decoder can include at least one of the dequantizer 220, inverse transformer 230, adder 235, filter 240, memory 250, inter-frame prediction unit 260, or intra-frame prediction unit 265.
[0099] The dequantizer 220 may dequantize the quantized transform coefficients and output the transform coefficients. The dequantizer 220 may rearrange the quantized transform coefficients in the form of two-dimensional blocks. In this case, the rearrangement may be performed based on the coefficient scan order executed in the image coding device. The dequantizer 220 may perform dequantization on the quantized transform coefficients by using quantization parameters (e.g., quantization step information) and obtain the transform coefficients.
[0100] The inverse transformer 230 may perform an inverse transform on the transform coefficients to obtain a residual signal (residual block, residual sample array).
[0101] The prediction unit may perform prediction on the current block and generate a prediction block including the prediction samples of the current block. The prediction unit may determine whether to apply intra prediction or inter prediction to the current block based on the information about the prediction output from the entropy decoder 210, and may determine a specific intra / inter prediction mode (prediction technique).
[0102] Similar to that described in the prediction unit of the image coding device 100, the prediction unit may generate a prediction signal based on various prediction methods (techniques) described later.
[0103] The intra prediction unit 265 may predict the current block by referring to the samples in the current picture. The description of the intra prediction unit 185 is equally applicable to the intra prediction unit 265.
[0104] The inter prediction unit 260 may derive a prediction block of the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between adjacent blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may also include information about the inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the adjacent blocks may include spatially adjacent blocks present in the current picture and temporally adjacent blocks present in the reference picture. For example, the inter prediction unit 260 may configure a motion information candidate list based on the adjacent blocks, and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. The inter prediction may be performed based on various prediction modes, and the information about the prediction may include information specifying the inter prediction mode of the current block.
[0105] The adder 235 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the prediction unit (including the inter-frame prediction unit 260 and / or the intra-frame prediction unit 265). If the block to be processed has no residual, for example when the skip mode is applied, the prediction block can be used as the reconstructed block. The description of the adder 155 equally applies to the adder 235. The adder 235 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current picture and can be used for inter-frame prediction of the next picture through filtering as described below.
[0106] The filter 240 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 240 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture and store the modified reconstructed picture in the memory 250, specifically, in the DPB of the memory 250. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filtering, bilateral filtering, etc.
[0107] The (modified) reconstructed picture stored in the DPB of the memory 250 can be used as a reference picture in the inter-frame prediction unit 260. The memory 250 can store the motion information of the block from which the motion information in the current picture is derived (or decoded) and / or the motion information of the blocks that have been reconstructed in the picture. The stored motion information can be transmitted to the inter-frame prediction unit 260 to be used as the motion information of spatially neighboring blocks or temporally neighboring blocks. The memory 250 can store the reconstructed samples of the reconstructed blocks in the current picture and transmit the reconstructed samples to the intra-frame prediction unit 265.
[0108] In the present disclosure, the embodiments described in the filter 160, the inter-frame prediction unit 180, and the intra-frame prediction unit 185 of the image encoding device 100 can be equally or correspondingly applied to the filter 240, the inter-frame prediction unit 260, and the intra-frame prediction unit 265 of the image decoding device 200.
[0109] Overview of image segmentation
[0110] The video / image encoding method according to the present disclosure can be performed based on an image segmentation structure as follows. Specifically, prediction, residual processing ((inverse) transformation, (de) quantization, etc.), syntax element encoding, and filtering processes to be described later can be performed based on CTUs, CUs (and / or TUs, PUs) derived according to the image segmentation structure. The image can be segmented by block units, and the block segmentation process can be performed in the image segmenter 110 of the encoding device. The segmentation-related information can be encoded by the entropy encoder 190 and sent to the decoding device in the form of a bitstream. The entropy decoder 210 of the decoding device can derive the block segmentation structure of the current picture based on the segmentation-related information obtained from the bitstream, and based on this, a series of processes (e.g., prediction, residual processing, block / picture reconstruction, in-loop filtering, etc.) can be performed for image decoding.
[0111] A picture can be segmented into a sequence of coding tree units (CTUs). Figure 4 An example in which a picture is segmented into CTUs is shown. A CTU can correspond to a coding tree block (CTB). Alternatively, a CTU can include a coding tree block of luminance samples and two coding tree blocks of corresponding chrominance samples. For example, for a picture including three sample arrays, a CTU can include one N×N block of luminance samples and two corresponding blocks of chrominance samples.
[0112] Overview of CTU segmentation
[0113] As described above, coding units can be obtained by recursively segmenting coding tree units (CTUs) or largest coding units (LCUs) according to a quadtree / binary tree / trinary tree (QT / BT / TT) structure. For example, a CTU can be first segmented into a quadtree structure. Thereafter, the leaf nodes of the quadtree structure can be further segmented by a multi-type tree structure.
[0114] Segmentation according to a quadtree means that the current CU (or CTU) is equally segmented into four. By segmentation according to a quadtree, the current CU can be segmented into four CUs having the same width and the same height. When the current CU is no longer segmented into a quadtree structure, the current CU corresponds to a leaf node of the quadtree structure. The CU corresponding to the leaf node of the quadtree structure can no longer be segmented and can be used as the final coding unit described above. Alternatively, the CU corresponding to the leaf node of the quadtree structure can be further segmented by a multi-type tree structure.
[0115] Figure 5 is a view showing an embodiment of the segmentation type of a block according to a multi-type tree structure. Segmentation according to a multi-type tree structure can include two types of division according to a binary tree structure and two types of division according to a trinary tree structure.
[0116] The two types of division according to the binary tree structure can include vertical binary division (SPLIT_BT_VER) and horizontal binary division (SPLIT_BT_HOR). Vertical binary division (SPLIT_BT_VER) means that the current CU is equally divided into two in the vertical direction. As Figure 4 shown, through vertical binary division, two CUs with the same height as the current CU and a width half of the width of the current CU can be generated. Horizontal binary division (SPLIT_BT_HOR) means that the current CU is equally divided into two in the horizontal direction. As Figure 5 shown, through horizontal binary division, two CUs with a height half of the height of the current CU and the same width as the current CU can be generated.
[0117] The two types of division according to the ternary tree structure can include vertical ternary division (SPLIT_TT_VER) and horizontal ternary division (SPLIT_TT_HOR). In vertical ternary division (SPLIT_TT_VER), the current CU is divided in the vertical direction at a ratio of 1:2:1. As Figure 5 shown, through vertical ternary division, two CUs with the same height as the current CU and a width 1 / 4 of the width of the current CU and one CU with the same height as the current CU and a width half of the width of the current CU can be generated. In horizontal ternary division (SPLIT_TT_HOR), the current CU is divided in the horizontal direction at a ratio of 1:2:1. As Figure 5 shown, through horizontal ternary division, two CUs with a height 1 / 4 of the height of the current CU and the same width as the current CU and one CU with a height half of the height of the current CU and the same width as the current CU can be generated.
[0118] Figure 6 is a view showing a signaling mechanism of block division information in a quadtree having a nested multi-type tree structure according to the present disclosure.
[0119] Here, the CTU is regarded as the root node of the quadtree and is first divided into a quadtree structure. Information (e.g., qt_split_flag) indicating whether to perform quadtree partitioning on the current CU (CTU of the quadtree or node (QT_node)) is signaled. For example, when the qt_split_flag has a first value (e.g., "1"), the current CU can be quadtree partitioned. Additionally, when the qt_split_flag has a second value (e.g., "0"), the current CU is not quadtree partitioned but becomes a leaf node (QT_leaf_node) of the quadtree. Then each quadtree leaf node can be further divided into a multi-type tree structure. That is, the leaf node of the quadtree can become a node (MTT_node) of the multi-type tree. In the multi-type tree structure, a first flag (e.g., Mtt_split_cu_flag) is signaled to specify whether the current node is additionally partitioned. If the corresponding node is additionally partitioned (e.g., if the first flag is 1), a second flag (e.g., Mtt_split_cu_vertical_flag) can be signaled to specify the partitioning direction. For example, the partitioning direction can be the vertical direction when the second flag is 1 and the horizontal direction when the second flag is 0. Then, a third flag (e.g., Mtt_split_cu_binary_flag) can be signaled to specify whether the partitioning type is a binary partitioning type or a ternary partitioning type. For example, the partitioning type can be a binary partitioning type when the third flag is 1 and a ternary partitioning type when the third flag is 0. The nodes of the multi-type tree obtained by binary or ternary partitioning can be further divided into a multi-type tree structure. However, the nodes of the multi-type tree may not be divided into a quadtree structure. If the first flag is 0, the corresponding node of the multi-type tree is no longer partitioned but becomes a leaf node (MTT_leaf_node) of the multi-type tree. The CU corresponding to the leaf node of the multi-type tree can be used as the above-mentioned final coding unit.
[0120] Based on mtt_split_cu_vertical_flag and mtt_split_cu_binary_flag, the multi-type tree partitioning mode (MttSplitMode) of the CU can be derived as shown in Table 1 below. In the following description, the multi-type tree partitioning mode may be referred to as the multi-tree partitioning type or the partitioning type.
[0121] [Table 1]
[0122] MttSplitMode mtt_split_cu_vertical_flag mtt_split_cu_binary_flag SPLIT_TT_HOR 0 0 SPLIT_BT_HOR 0 1 SPLIT_TT_VER 1 0 SPLIT_BT_VER 1 1
[0123] Figure 7 is a view showing an example of dividing the CTU into multiple CUs by applying a multi-type tree after applying the quadtree. In Figure 7Among them, the bold block edges 710 represent quadtree partitioning, and the remaining edges 720 represent multi-type tree partitioning. A CU may correspond to a coding block (CB). In an embodiment, a CU may include one coding block of luminance samples and two coding blocks of chrominance samples corresponding to the luminance samples. The chrominance component (sample) CB or TB size may be derived based on the luminance component (sample) CB or TB size according to the component ratio of the picture / image color format (chrominance format, e.g., 4:4:4, 4:2:2, 4:2:0, etc.). In the case of the 4:4:4 color format, the chrominance component CB / TB size may be set to be equal to the luminance component CB / TB size. In the case of the 4:2:2 color format, the width of the chrominance component CB / TB may be set to half of the width of the luminance component CB / TB, and the height of the chrominance component CB / TB may be set to the height of the luminance component CB / TB. In the case of the 4:2:0 color format, the width of the chrominance component CB / TB may be set to half of the width of the luminance component CB / TB, and the height of the chrominance component CB / TB may be set to half of the height of the luminance component CB / TB.
[0124] In an embodiment, when the size of the CTU is 128 based on luminance sample units, the size of the CU may have sizes ranging from 128×128 to 4×4, which is the same size as the CTU. In one embodiment, in the case of the 4:2:0 color format (or chrominance format), the chrominance CB size may have sizes ranging from 64×64 to 2×2.
[0125] Furthermore, in an embodiment, the CU size and the TU size may be the same. Alternatively, there may be multiple TUs in the CU region. The TU size generally represents the luminance component (sample) transform block (TB) size.
[0126] The TU size may be derived based on the maximum allowable TB size maxTbSize, which is a predetermined value. For example, when the CU size is greater than maxTbSize, multiple TUs (TBs) with maxTbSize may be derived from the CU, and the transform / inverse transform may be performed in units of TUs (TBs). For example, the maximum allowable luminance TB size may be 64×64, and the maximum allowable chrominance TB size may be 32×32. If the width or height of the CB divided according to the tree structure is greater than the maximum transform width or height, the CB may be automatically (or implicitly) divided until the TB size limits in the horizontal and vertical directions are met.
[0127] In addition, for example, when intra prediction is applied, an intra prediction mode / type may be derived on a CU (or CB) basis, and neighboring reference sample derivation and predicted sample generation processing may be performed on a TU (or TB) basis. In this case, there may be one or more TUs (or TBs) in one CU (or CB) region, and in this case, multiple TUs or (TBs) may share the same intra prediction mode / type.
[0128] Furthermore, for a quadtree coding tree scheme with a nested multi-type tree, the following parameters may be signaled from an encoding device to a decoding device as SPS syntax elements. For example, at least one of the CTU size as a parameter representing the root node size of the quadtree, the MinQTSize as a parameter representing the minimum allowable quadtree leaf node size, the MaxBtSize as a parameter representing the maximum allowable binary tree root node size, the MaxTtSize as a parameter representing the maximum allowable ternary tree root node size, the MaxMttDepth as a parameter representing the maximum allowable hierarchical depth of multi-type tree partitioning starting from quadtree leaf nodes, the MinBtSize as a parameter representing the minimum allowable binary tree leaf node size, or the MinTtSize as a parameter representing the minimum allowable ternary tree leaf node size.
[0129] As an embodiment using the 4:2:0 chroma format, the CTU size can be set to 128×128 luma blocks and two 64×64 chroma blocks corresponding to these luma blocks. In this case, MinOTSize can be set to 16×16, MaxBtSize can be set to 128×128, MaxTtSzie can be set to 64×64, MinBtSize and MinTtSize can be set to 4×4, and MaxMttDepth can be set to 4. Quadtree segmentation can be applied to the CTU to generate quadtree leaf nodes. The quadtree leaf nodes can be referred to as leaf QT nodes. The size of the quadtree leaf nodes can range from a 16×16 size (e.g., MinOTSize) to a 128×128 size (e.g., CTU size). If the leaf QT node is 128×128, it may not be additionally split into a binary tree / trinary tree. This is because, in this case, even if split, it exceeds MaxBtsize and MaxTtszie (e.g., 64×64). In other cases, the leaf QT node can be further split into a multi-type tree. Thus, the leaf QT node is the root node of the multi-type tree, and the leaf QT node can have a multi-type tree depth (mttDepth) value of 0. If the multi-type tree depth reaches MaxMttdepth (e.g., 4), further splitting may not be considered. If the width of the multi-type tree node is equal to MinBtSize and less than or equal to 2xMinTtSize, further horizontal splitting may not be considered. If the height of the multi-type tree node is equal to MinBtSize and less than or equal to 2xMinTtSize, further vertical splitting may not be considered. When splitting is not considered, the encoding device can skip signaling of the splitting information. In this case, the decoding device can derive the splitting information with a predetermined value.
[0130] In addition, a CTU may include an encoded block of luminance samples (hereinafter referred to as "luminance block") and two encoded blocks of chrominance samples corresponding thereto (hereinafter referred to as "chrominance blocks"). The above-described coding tree scheme may be equally or separately applied to the luminance block and the chrominance block of the current CU. Specifically, the luminance block and the chrominance block in a CTU may be divided into the same block tree structure, and in this case, the tree structure is represented as SINGLE_TREE. Alternatively, the luminance block and the chrominance block in a CTU may be divided into separate block tree structures, and in this case, the tree structure may be represented as DUAL_TREE. That is, when a CTU is divided into a dual tree, the block tree structure for the luminance block and the block tree structure for the chrominance block may exist separately. In this case, the block tree structure for the luminance block may be referred to as DUAL_TREE_LUMA, and the block tree structure for the chrominance component may be referred to as DUAL_TREE_CHROMA. For P and B slices / tile groups, the luminance block and the chrominance block in a CTU may be restricted to having the same coding tree structure. However, for I slices / tile groups, the luminance block and the chrominance block may have separate block tree structures from each other. If a separate block tree structure is applied, the luminance CTB may be divided into CUs based on a specific coding tree structure, and the chrominance CTB may be divided into chrominance CUs based on another coding tree structure. That is, this means that the CUs in an I slice / tile group to which a separate block tree structure is applied may include an encoded block of the luminance component or two encoded blocks of the chrominance components, and the CUs of P or B slices / tile groups may include blocks of three color components (one luminance component and two chrominance components).
[0131] Although a quadtree coding tree structure having a nested multi-type tree has been described, the structure for dividing a CU is not limited thereto. For example, the BT structure and the TT structure may be interpreted as concepts included in a multi-partition tree (MPT) structure, and a CU may be interpreted as being divided by a QT structure and an MPT structure. In an example of dividing a CU by a QT structure and an MPT structure, a syntax element (e.g., MPT_split_type) including information on how many blocks a leaf node of the QT structure is divided into and a syntax element (e.g., MPT_split_mode) including information on which of the vertical direction and the horizontal direction the leaf node of the QT structure is divided into may be signaled to determine the division structure.
[0132] In another example, the CU can be split in a way different from the QT structure, the BT structure, or the TT structure. That is, different from splitting a lower-depth CU into 1 / 4 of a higher-depth CU according to the QT structure, splitting a lower-depth CU into 1 / 2 of a higher-depth CU according to the BT structure, or splitting a lower-depth CU into 1 / 4 or 1 / 2 of a higher-depth CU according to the TT structure, in some cases, a lower-depth CU can be split into 1 / 5, 1 / 3, 3 / 8, 3 / 5, 2 / 3, or 5 / 8 of a higher-depth CU, and the method of splitting the CU is not limited to this.
[0133] The quadtree coding block structure with multi-type trees can provide a very flexible block splitting structure. Due to the splitting types supported in the multi-type trees, in some cases, different splitting patterns can potentially result in the same coding block structure. In the encoding device and the decoding device, by restricting the occurrence of such redundant splitting patterns, the amount of data for splitting information can be reduced.
[0134] Quantization / dequantization
[0135] As described above, the quantizer of the encoding device can derive the quantized transform coefficients by applying quantization to the transform coefficients, and the dequantizer of the encoding device or the dequantizer of the decoding device can derive the transform coefficients by applying dequantization to the quantized transform coefficients.
[0136] In the encoding and decoding of moving images / still images, the quantization rate can be changed and the changed quantization rate can be used to adjust the compression rate. From an implementation perspective, considering the complexity, the quantization parameter (QP) can be used instead of directly using the quantization rate. For example, a quantization parameter with an integer value from 0 to 63 can be used, and each quantization parameter value can correspond to an actual quantization rate. In addition, the quantization parameter QP for the luminance component (luminance samples) and Y the quantization parameter QP for the chrominance component (chrominance samples) C .
[0137] In the quantization process, the transform coefficient C can be received as an input, divided by the quantization rate Qstep, and the quantized transform coefficient C' can be obtained based on this. In this case, considering the computational complexity, the quantization rate is multiplied by a scale to form an integer, and a shift operation can be performed by a value corresponding to the scale value. Based on the product of the quantization rate and the scale value, the quantization scale can be derived. That is, the quantization scale can be derived according to the QP. By applying the quantization scale to the transform coefficient C, the quantized transform coefficient C' can be derived based on this.
[0138] Inverse quantization is the inverse process of quantization, and can multiply the quantized transform coefficient C' by the quantization rate Qstep, and obtain the reconstructed transform coefficient C” based on this. In this case, a level scale can be derived according to the quantization parameter, the level scale can be applied to the quantized transform coefficient C', and the reconstructed transform coefficient C” can be derived based on this. Due to losses in the transform and / or quantization process, the reconstructed transform coefficient C” may be slightly different from the original transform coefficient C. Therefore, even the encoding device can perform inverse quantization in the same way as the decoding device.
[0139] At the same time, an adaptive frequency-weighted quantization technique that adjusts the quantization intensity according to the frequency can be applied. The adaptive frequency-weighted quantization technique is a method of applying the quantization intensity differently according to the frequency. In adaptive frequency-weighted quantization, a predefined quantization scaling matrix can be used to apply the quantization intensity differently according to the frequency. That is, the above quantization / inverse quantization process can be further performed based on the quantization scaling matrix. For example, different quantization scaling matrices can be used according to the size of the current block and / or whether the prediction mode applied to the current block to generate the residual signal of the current block is inter prediction or intra prediction. The quantization scaling matrix can also be referred to as a quantization matrix or a scaling matrix. The quantization scaling matrix can be predefined. In addition, the frequency quantization scale information of the quantization scaling matrix for frequency adaptive scaling can be constructed / encoded by the encoding device and signaled to the decoding device. The frequency quantization scale information can be referred to as quantization scaling information. The frequency quantization scale information can include scaling list data scaling_list_data. Based on the scaling list data, the (modified) quantization scaling matrix can be derived. In addition, the frequency quantization scale information can include presence flag information specifying whether there is scaling list data. Alternatively, when the scaling list data is signaled at a higher level (e.g., SPS), information specifying whether the scaling list data is modified at a lower level (e.g., PPS or slice group header, etc.) can be further included.
[0140] Overview of sub-picture
[0141] As described above, quantization and inverse quantization of the luminance component and the chrominance component can be performed based on the quantization parameter. In addition, one picture to be encoded can be divided into units of multiple CTUs, slices, tiles, or blocks, and further, the picture can be divided into units of multiple sub-pictures.
[0142] Within a picture, a sub-picture can be encoded or decoded regardless of whether the previous sub-picture has been encoded or decoded. For example, different quantization or different resolutions can be applied to multiple sub-pictures.
[0143] In addition, the sub-picture can be processed as a separate picture. For example, the picture to be encoded can be a projected picture or a packed picture in a 360-degree image / video or an omnidirectional image / video.
[0144] In this embodiment, a part of the picture can be rendered or displayed based on the viewport of a user terminal (e.g., a head-mounted display). Therefore, in order to achieve low latency, at least one sub-picture covering the viewport among the sub-pictures constituting a picture can be encoded or decoded preferentially or independently of the remaining sub-pictures.
[0145] The encoding result of the sub-picture can be referred to as a sub-bitstream, a sub-stream, or simply a bitstream. The decoding device can decode the sub-picture from the sub-bitstream, the sub-stream, or the bitstream. In this case, a high-level syntax (HLS) such as PPS, SPS, VPS, and / or DPS (decoding parameter set) can be used to encode / decode the sub-picture.
[0146] In the present disclosure, the high-level syntax (HLS) can include at least one of APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, or slice header syntax. For example, APS (APS syntax) or PPS (PPS syntax) can include information / parameters that are generally applicable to one or more slices or pictures. SPS (SPS syntax) can include information / parameters that are generally applicable to one or more sequences. VPS (VPS syntax) can include information / parameters that are generally applicable to multiple layers. DPS (DPS syntax) can include information / parameters that are generally applicable to the entire video. For example, DPS can include information / parameters related to the concatenation of coded video sequences (CVS).
[0147] Definition of sub-picture
[0148] The sub-pictures can form a rectangular area of the encoded picture. The size of the sub-pictures can be set differently within the picture. For all pictures belonging to one sequence, the size and position of a specific individual sub-picture can be set equally. A separate sub-picture sequence can be decoded independently. Tiles and slices (and CTB) can be restricted from crossing the sub-picture boundary. To this end, the encoding device can perform encoding so that the sub-pictures are decoded independently. For this purpose, semantic constraints in the bitstream may be required. In addition, for each picture in one sequence, the arrangement of tiles, slices, and blocks in the sub-pictures can be constructed differently.
[0149] Design purpose of sub-picture
[0150] The sub-picture design aims to abstract or encapsulate a range smaller than the picture level or larger than the slice or tile group level. Therefore, it is possible to extract the VCL NAL units of a subset of the Motion-constrained Tile Sets (MCTS) from a VVC bitstream and relocate them to another VVC bitstream without difficulty, such as making modifications at the VCL level. Here, MCTS is an encoding technique that implements spatial and temporal independence between tiles. When MCTS is applied, the information of tiles not included in the MCTS to which the current tile belongs cannot be referenced. When an image is segmented into MCTS and encoded, independent transmission and decoding of MCTS are possible.
[0151] This sub-picture design has advantages in changing the viewing direction in the viewport-related 360° stream transmission scheme.
[0152] Use cases of sub-picture
[0153] Sub-pictures are required in the viewport-related 360° scheme to provide an extended real-space resolution on the viewport. For example, the scheme of tiles covering the viewport derived from a 6K (6144×3072) ERP (equirectangular projection) picture or cube map projection (CMP) resolution (with its equivalent 4K decoding performance (HEVC level 5.1)) is included in Sections D6.3 and D.6.4 of OMAF and adopted in the VR Industry Forum guidelines. This resolution is known to be suitable for head-mounted displays using a quad high definition (2560x1440) display panel.
[0154] Encoding: Content can be encoded using two spatial resolutions, including a resolution with a cubic face size of 1656x1536 and a resolution with a cubic face size of 768x768. In both bitstreams, a 6x4 tile grid can be used, and MCTS can be encoded at each tile position.
[0155] MCTS selection for streaming: 12 MCTS can be selected from the high-resolution bitstream, and 12 additional MCTS can be obtained from the low-resolution bitstream. Therefore, streaming content for a hemisphere (180°×180°) can be generated from the high-resolution bitstream.
[0156] Decoding using the combination of MCTS and bitstreams: Receive the MCTS of a single time instance, which can be combined into an encoded picture with a resolution of 1920x4608 conforming to HEVC level 5.1. In another option for the combined picture, the width value of four tile columns is 768, the width value of two tile columns is 384, and the height value of three tile rows is 768, thus constructing a picture consisting of 3840x2304 luminance samples. Here, the width and height units can be the units of the number of luminance samples.
[0157] Sub-picture signaling
[0158] The signaling of sub - pictures can be performed at the SPS level as Figure 8 shown. Figure 8 The syntax for signaling sub - picture syntax elements in the SPS is shown. In the following, the Figure 8 syntax elements will be described.
[0159] The syntax element pic_width_max_in_luma_samples can specify the maximum width of each decoded picture of the reference SPS in terms of luma samples. The value of pic_width_max_in_luma_samples is greater than 0 and can take a value that is an integer multiple of andMinCbSizeY. Here, MinCbSizeY is a variable that specifies the minimum size of the luma - component coding block.
[0160] The syntax element pic_height_max_in_luma_samples can specify the maximum height of each decoded picture of the reference SPS in terms of luma samples. pic_height_max_in_luma_samples is greater than 0 and may have a value that is an integer multiple of MinCbSizeY.
[0161] The syntax element subpic_grid_col_width_minus1 can be used to specify the width of each element of the sub - picture identifier grid. For example, subpic_grid_col_width_minus1 can specify the width of each element of the sub - picture identifier grid in units of 4 samples, and the value obtained by adding 1 to subpic_grid_col_width_minus1 can specify the width of a single element of the sub - picture identifier grid in units of 4 samples. The length of the syntax element can be Ceil(Log2(pic_width_max_in_luma_samples / 4)) bits long.
[0162] Therefore, the variable NumSubPicGridCols that specifies the number of columns in the sub - picture grid can be derived as follows.
[0163] [Equation 1]
[0164] NumSubPicGridCols = (pic_width_max_in_luma_samples+subpic_grid_col_width_minus1*4 + 3) / (subpic_grid_col_width_minus1*4 + 4)
[0165] The syntax element subpic_grid_row_height_minus1 can be used to specify the height of each element of the sub-picture identifier grid. For example, subpic_grid_row_height_minus1 can specify the height of each element of the sub-picture identifier grid in units of 4 samples. The value obtained by adding 1 to subpic_grid_row_height_minus1 can specify the height of each element of the sub-picture identifier grid in units of 4 samples. The length of the syntax element can be Ceil(Log2(pic_height_max_in_luma_samples / 4)) bits long.
[0166] Therefore, the variable NumSubPicGridRows that specifies the number of rows in the sub-picture grid can be derived as follows.
[0167] [Formula 2]
[0168] NumSubPicGridRows = (pic_height_max_in_luma_samples + subpic_grid_row_height_minus1 * 4 + 3) / (subpic_grid_row_height_minus1 * 4 + 4)
[0169] The syntax element subpic_grid_idx[i][j] can specify the sub-picture index at the grid position (i, j). The length of the syntax element can be Ceil(Log2(max_subpics_minus1 + 1)) bits.
[0170] The variables SubPicTop[subpic_grid_idx[i][j]], SubPicLeft[subpic_grid_idx[i][j]], SubPicWidth[subpic_grid_idx[i][j]], SubPicHeight[subpic_grid_idx[i][j]] and NumSubPics can be derived as in the Figure 9 algorithm.
[0171] The syntax element subpic_treated_as_pic_flag[i] can specify whether a subpicture is treated as the same as a normal picture in the decoding process other than the loop filter process. For example, the first value of subpic_treated_as_pic_flag[i] (e.g., 0) can specify that the i-th subpicture of each coded picture in the CVS is not treated as a picture in the decoding process other than the loop filter process. The second value of subpic_treated_as_pic_flag[i] (e.g., 1) can specify that the i-th subpicture of each coded picture in the CVS is treated as a picture in the decoding process other than the loop filter process. When the value of subpic_treated_as_pic_flag[i] is not obtained from the bitstream, the value of subpic_treated_as_pic_flag[i] can be derived as the first value (e.g., 0).
[0172] The syntax element loop_filter_across_subpic_enabled_flag[i] can specify whether loop filtering is performed across the boundaries of the i-th subpicture of a separate coded picture belonging to the CVS. For example, the first value of loop_filter_across_subpic_enabled_flag[i] (e.g., 0) can specify that loop filtering is not performed across the boundaries of the i-th subpicture of a separate coded picture belonging to the CVS. The second value of loop_filter_across_subpic_enabled_flag[i] (e.g., 1) can specify that loop filtering can be performed across the boundaries of the i-th subpicture of a separate coded picture belonging to the CVS. When the value of loop_filter_across_subpic_enabled_flag[i] is not obtained from the bitstream, the value of loop_filter_across_subpic_enabled_flag[i] can be derived as the second value.
[0173] Meanwhile, for bitstream consistency, the following constraints can be applied. For any two subpictures subpicA and subpicB, when the index of subpicA is less than the index of subpicB, all coded NAL units of subpicA may have a decoding order lower than that of all coded NAL units of subpicB. Alternatively, after decoding is performed, the shape of the subpicture needs to have a perfect left boundary and a perfect upper boundary, which form the image boundary or the boundary of a previously decoded subpicture.
[0174] Overview of sub-picture based encoding and decoding
[0175] The following disclosure relates to the encoding / decoding of the above-mentioned picture and / or sub-picture. The encoding device may encode the current picture based on the sub-picture structure. Alternatively, the encoding device may encode at least one sub-picture that constitutes the current picture and output a (sub-)bitstream including (encoded) information about the at least one (encoded) sub-picture.
[0176] The decoding device may decode at least the sub-pictures belonging to the current picture based on the (sub-)bitstream including (encoded) information about the at least one sub-picture.
[0177] Figure 10 It is a view illustrating a method of encoding an image using sub-pictures by an encoding device according to an embodiment. The encoding device may divide an input picture into a plurality of sub-pictures (S1010). The encoding device may generate information about the sub-pictures (S1020). The encoding device may encode at least one sub-picture using the information about the sub-pictures. For example, each sub-picture may be independently separated and encoded using the information about the sub-pictures. Next, the encoding device may output a bitstream by encoding the image information including the information about the sub-pictures (S1030). Here, the bitstream of the sub-picture may be referred to as a sub-stream or a sub-bitstream.
[0178] The information about the sub-pictures will be variously described in this disclosure, and for example, there may be information about whether loop filtering can be performed across the boundaries of the sub-pictures, information about the sub-picture regions, information about the grid spacing for using the sub-pictures, etc.
[0179] Figure 11 It is a view showing a method of decoding an image using sub-pictures by a decoding device according to an embodiment. The decoding device may obtain information about the sub-pictures from the bitstream (S1110). Next, the decoding device may derive at least one sub-picture (S1120) and decode the at least one sub-picture (S1130).
[0180] In this way, the decoding device may decode at least one sub-picture, thereby outputting at least one decoded sub-picture or the current picture including at least one sub-picture. The bitstream may include a sub-stream or a sub-bitstream of the sub-picture.
[0181] As described above, the information about the sub-pictures may be constructed in the HLS of the bitstream. The decoding device may derive at least one sub-picture based on the information about the sub-pictures. The decoding device may decode the sub-pictures based on methods such as the CABAC method, the prediction method, the residual processing method (transformation, quantization), the loop filtering method, etc.
[0182] When outputting the decoded sub - pictures, the decoded sub - pictures can be output together in the form of an OPS (Output Sub - picture Set). For example, when the current picture is related to a 360° or omnidirectional image / video and is partially rendered, only some sub - pictures can be decoded, and some or all of the decoded sub - pictures can be rendered according to the user's viewport.
[0183] When the information specifying whether loop filtering across sub - picture boundaries is available specifies availability, the decoding device can perform loop filtering (e.g., de - blocking filtering) on the sub - picture boundaries located between two sub - pictures. At the same time, when the sub - picture boundary is equal to the picture boundary, the loop filtering process for the sub - picture boundary can be not applied.
[0184] Embodiment 1
[0185] In the above description, examples of syntax elements for signaling sub - pictures at the SPS level have been described. For example, information about the size and position of sub - blocks can be signaled at the SPS level. Thus, as Figure 12 shown in the example, a sub - picture can be defined as a set of grid granulations with the same sub - picture grid index (e.g., sub - picture grid index). Figure 12 Illustrates an implementation of using the sub - picture grid index defined in the SPS to identify the size and position of sub - pictures. In Figure 12 , the sub - picture grid granulation is represented by a dashed line, the tile boundary Tile Bdry is represented by a solid line, and the slice boundary Slice Bdry is represented by a dotted line.
[0186] However, in Figure 12 the example, the boundary of the sub - picture defined as a set of grid granulations with a sub - picture grid index of 2 (represented by the Figure 12 shading in) has a sub - picture boundary that does not coincide with the tile boundary and the slice boundary. The boundary of the sub - picture identified according to the sub - picture grid information provided at the SPS may not coincide with the boundary of the slice and / or tile. Thus, a slice or tile with a boundary having a boundary across the sub - picture can be set.
[0187] In this case, since the slice or tile is determined within the sub - picture, there is a problem that the boundary of the tile or slice can be determined in a form not considered during encoding. Alternatively, during encoding and decoding, there is also a problem of reference to a part of a non - existent slice or a part of a tile.
[0188] To solve the above problems, the correspondence between slices and sub-pictures can be signaled at the PPS. Thus, the sub-picture grid can be aligned according to the positions of the slices. Therefore, the sub-picture syntax can be provided at the PPS instead of the SPS. Hereinafter, a method of signaling the identifier of the sub-picture using the PPS will be described.
[0189] Figure 13 is a view showing an example of the SPS syntax structure of the present embodiment. In the SPS syntax structure, signaling of at least one syntax element 1310 of the sub-picture can be omitted.
[0190] Figure 14 is a view showing an example of a method for a decoding device to identify a sub-picture grid based on slices according to an embodiment. The operation of the decoding device will be described because an encoding device can use a corresponding method to generate a bitstream.
[0191] First, the decoding device can determine whether sub-pictures of the current picture exist in the PPS (S1410). The decoding device can determine this based on the syntax element pps_subpics_present_flag described below.
[0192] Next, when sub-pictures of the current picture exist in the PPS, the decoding device can determine the maximum number of sub-pictures (S1420). The decoding device can determine the maximum number of sub-pictures based on the syntax element max_subpics_minus1 described below.
[0193] Next, the decoding device can determine the number of slices in the current picture (S1430). The decoding device can determine the number of slices in the current picture based on the syntax element num_slices_in_pic_minus1 described below.
[0194] Next, the decoding device can determine the ID of the slice belonging to the current picture and the ID of the sub-picture corresponding to the slice (S1440). The decoding device can determine the ID of the slice belonging to the current picture and the ID of the sub-picture corresponding to the slice based on the syntax elements slice_id[i] and slice_to_subpic_id[i] described below.
[0195] Figure 15 is an illustration for Figure 12View of an example of the PPS syntax structure for sub - picture index signaling. The syntax element pps_subpics_present_flag can specify whether the sub - picture information of the current picture exists in the PPS. For example, the first value of pps_subpics_present_flag (e.g., 0) can specify that the sub - picture information of the current picture is not included in the PPS. In this case, there may be only one sub - picture in the current picture. The second value of pps_subpics_present_flag (e.g., 1) can specify that the sub - picture information of the current picture is included in the PPS. In this case, there may be at least one sub - picture in the current picture.
[0196] When the value of pps_subpics_present_flag is the second value (e.g., 1), the syntax element max_subpics_minus1 can be obtained from the bitstream. max_subpics_minus1 can be used to signal the maximum number of sub - pictures included in the current picture. For example, the decoding device can determine the maximum number of sub - pictures included in the current picture as the value obtained by adding 1 to max_subpics_minus1.
[0197] The syntax element rect_slice_flag can specify whether the slices in the current picture are split into rectangles. The first value of rect_slice_flag (e.g., 0) can specify that the slices are split into a raster - scan shape. The second value of rect_slice_flag (e.g., 1) can specify that the slices are split into rectangles.
[0198] When rect_slice_flag is the second value (e.g., 1), the syntax element num_slices_in_pic_minus1 can be obtained from the PPS. num_slices_in_pic_minus1 can specify the value obtained by subtracting 1 from the number of slices belonging to the current picture. Thus, the decoding device can determine the number of slices belonging to the current picture as the value obtained by adding 1 to num_slices_in_pic_minus1.
[0199] The syntax element signaled_slice_id_flag can specify whether the slice identifier is signaled explicitly. For example, the first value of signaled_slice_id_flag (e.g., 0) can specify that the slice identifier is not signaled explicitly. The second value of signaled_slice_id_flag (e.g., 1) can specify that the slice identifier is signaled explicitly.
[0200] In addition, when the value of pps_subpics_present_flag is the second value (e.g., 1), the syntax elements slice_id[i] and slice_to_subpic_id[i] can be obtained from the bitstream. slice_id[i] can specify the identifier of the slice. slice_to_subpic_id[i] can specify the subpicture identifier corresponding to the slice. The length of slice_to_subpic_id[i] can be Ceil(Log2(max_sub_pics_minus1 + 1)) bits long.
[0201] Through the above processing, the maximum number of subpictures in the sequence and the enable flag of the subpicture tool can be defined in the PPS. Through this processing, it can be assumed that the subpictures are not used for the raster scan slice mode but only for the rectangular slice mode. In addition, when a flag such as signaled_slice_id_flag is enabled, the slice ID can be explicitly signaled. To use the presence of subpictures and the rectangular slice mode, explicit signaling of the slice ID can be considered. In addition, new syntax elements (such as slice_to_subpic_id[i]) for matching the slice and the subpicture grid can be further considered, which is similar to the syntax element subpic_grid_idx obtained from the SPS.
[0202] Figure 16 It is a view showing the signaling of the syntax element specifying the constraint on the above pps_subpics_present_flag. The syntax element no_subpictures_constraint_flag 1610 can be used to signal whether the pps_subpics_present_flag is constrained. For example, the first value of no_subpictures_constraint_flag (e.g., 0) can specify that the constraint does not apply to the value of pps_subpics_present_flag. The second value of no_subpictures_constraint_flag (e.g., 1) can specify that the value of pps_subpics_present_flag is constrained to the first value (e.g., 0).
[0203] Embodiment 2
[0204] The subpictures in the picture may have the same width and height. However, to improve the coding efficiency, the subpictures may have different widths and heights. To signal the width and height of the subpictures, the subpicture grid information can be signaled.
[0205] In one embodiment, sub-picture grid information having a consistent width and height can be signaled. Alternatively, the width and height of each element of the sub-picture grid can be explicitly signaled separately. By this, greater flexibility can be demonstrated in the encoding and decoding processes using sub-pictures.
[0206] The encoding device and the decoding device can reduce the amount of the bitstream by using a flag that specifies whether the sub-picture grid is signaled uniformly. Figure 17 is a view illustrating an SPS syntax structure for signaling a sub-picture grid. Hereinafter, it will be described with reference to Figure 17 as follows.
[0207] The syntax element subpics_present_flag can specify whether sub-picture parameters are provided in the current SPS RBSP syntax. For example, a first value (e.g., 0) of subpics_present_flag can specify that sub-picture parameters are not provided in the current SPS RBSP syntax. A second value (e.g., 1) of subpics_present_flag can specify that sub-picture parameters are provided in the current SPS RBSP syntax.
[0208] Meanwhile, when a bitstream is generated by sub-bitstream acquisition processing and the bitstream includes only a sub-picture subset of the input bitstream for the sub-bitstream acquisition processing, the value of subpics_present_flag can be set to 1 in the RVSP of the SPS.
[0209] The syntax element max_subpics_minus1 can specify a value obtained by subtracting 1 from the maximum number of sub-pictures provided in the current encoded picture sequence. For example, the value obtained by adding 1 to max_subpics_minus1 can specify the maximum number of sub-pictures. The value of max_subpics_minus1 can be from 0 to 254.
[0210] The syntax element uniform_subpic_grid_flag can specify whether the sub-picture grid is uniformly distributed on the picture. In addition, uniform_subpic_grid_flag can specify the signaling of the syntax elements subpic_grid_col_width_minus1 and subpic_grid_row_height_minus1.
[0211] For example, a first value of the uniform_subpic_grid_flag (e.g., 0) may specify that sub-picture column boundaries and sub-picture row boundaries may be distributed on the picture in a non-uniform manner. Further, the first value of the uniform_subpic_grid_flag (e.g., 0) may further specify that the syntax elements num_subpic_grid_col_minus1 and num_subpic_grid_row_minus1, and the syntax elements subpic_grid_row_height_minus1[i] and subpic_col_width_minus1[i] are signaled as a list of syntax element pairs.
[0212] Meanwhile, a second value of the uniform_subpic_grid_flag (e.g., 1) may specify that sub-picture column boundaries and sub-picture row boundaries are distributed on the picture in a uniform manner. Further, the uniform_subpic_grid_flag may specify that the syntax elements subpic_grid_col_width_minus1 and subpic_grid_row_height_minus1 are signaled.
[0213] When the value of the uniform_subpic_grid_flag is not obtained from the bitstream, the value of the uniform_subpic_grid_flag may be derived as 1.
[0214] The syntax element subpic_grid_col_width_minus1 may specify the width of a single element in the sub-picture identifier grid. For example, the value obtained by adding 1 to subpic_grid_col_width_minus1 may specify the width of each element in the sub-picture identifier grid in units of 4 samples. subpic_grid_col_width_minus1 may have a length of Ceil(Log2(pic_width_max_in_luma_samples / 4)) bits. Meanwhile, when the value of the uniform_subpic_grid_flag is the second value (e.g., 1), the value of the variable NumSubPicGridCols that specifies the number of sub-picture grid columns may be derived as follows.
[0215] [Formula 3]
[0216] NumSubPicGridCols = (pic_width_max_in_luma_samples + subpic_grid_col_width_minus1 * 4 + 3) / (subpic_grid_col_width_minus1 * 4 + 4)
[0217] The syntax element subpic_grid_row_height_minus1 can specify the height of each element in the sub-picture identifier grid. For example, the value obtained by adding 1 to subpic_grid_row_height_minus1 can specify the height of each element in the sub-picture identifier grid in units of 4 samples. subpic_grid_row_height_minus1 may have a length of Ceil(Log2(pic_height_max_in_luma_samples / 4)) bits. Meanwhile, when the value of uniform_subpic_grid_flag is the second value (e.g., 1), the value of the variable NumSubPicGridRows that specifies the number of sub-picture grid rows can be derived as follows.
[0218] [Formula 4]
[0219] NumSubPicGridRows = (pic_height_max_in_luma_samples + subpic_grid_row_height_minus1 * 4 + 3) / (subpic_grid_row_height_minus1 * 4 + 4)
[0220] The syntax element num_subpic_grid_col_minus1 can be used to signal the value of NumSubPicGridCols. For example, the value obtained by adding 1 to num_subpic_grid_col_minus1 can specify the value of NumSubPicGridCols.
[0221] The syntax element num_subpic_grid_row_minus1 can be used to signal the value of NumSubPicGridRows. For example, the value obtained by adding 1 to num_subpic_grid_row_minus1 can specify the value of NumSubPicGridRows.
[0222] The syntax element subpic_grid_idx[i][j] can specify the sub - picture index at the grid position (i, j). subpic_grid_idx[i][j] may have a length of Ceil(Log2(max_subpics_minus1 + 1)) bits.
[0223] The variables SubPicTop[subpic_grid_idx[i][j]], SubPicLeft[subpic_grid_idx[i][j]], SubPicWidth[subpic_grid_idx[i][j]], SubPicHeight[subpic_grid_idx[i][j]] and NumSubPics can be derived using Figure 9 the algorithm.
[0224] The syntax element subpic_treated_as_pic_flag[i] can be used to specify whether the i - th sub - picture of a single coded picture in the coded picture sequence is treated as a picture in the decoding process other than the loop - filtering process. For example, the first value of subpic_treated_as_pic_flag[i] (e.g., 0) can specify that the i - th sub - picture of a single coded picture in the coded picture sequence is not treated as a picture in the decoding process other than the loop - filtering process.
[0225] The second value of subpic_treated_as_pic_flag[i] (e.g., 1) can specify that the i - th sub - picture of a single coded picture in the coded picture sequence is treated as a picture in the decoding process other than the loop - filtering process.
[0226] When subpic_treated_as_pic_flag[i] is not obtained from the bitstream, the value of subpic_treated_as_pic_flag[i] can be derived as the first value.
[0227] The syntax element loop_filter_across_subpic_enabled_flag[i] can be used to specify whether to perform the loop - filtering operation across the boundaries of the i - th sub - picture of a separate coded picture belonging to the coded picture sequence.
[0228] For example, the first value of loop_filter_across_subpic_enabled_flag[i] (e.g., 0) can specify that the loop - filtering operation is not performed across the boundaries of the i - th sub - picture of a single coded picture belonging to the coded picture sequence.
[0229] The second value of loop_filter_across_subpic_enabled_flag[i] (e.g., 1) may specify that loop filtering operations are performed across the boundaries of the i-th sub-picture of separate coded pictures belonging to an encoded picture sequence.
[0230] When loop_filter_across_subpic_enabled_flag[i] is not obtained from the bitstream, the value of loop_filter_across_subpic_enabled_flag[i] may be derived as the second value.
[0231] To comply with bitstream consistency, the following constraints may be applied. For any two sub-pictures subpicA and subpicB, when the index of subpicA is less than the index of subpicB, all coded NAL units of subpicA may have a decoding order lower than that of all coded NAL units of subpicB. Additionally, after decoding, the shape of the sub-picture needs to have a perfect left boundary and a perfect upper boundary, which form the picture boundary or the boundary of a previously decoded sub-picture.
[0232] Meanwhile, the variables SubPicIdx, SubPicLeftBoundaryPos, SubPicTopBoundaryPos, SubPicRightBoundaryPos, and SubPicBotBoundaryPos may be derived based on the value of uniform_subpic_grid_flag according to Figure 18 the algorithm.
[0233] Additionally, a list CtbToSubPicIdx[ctbAddrRs] of sub-picture indices that specify the CTB address ctbAddrRs according to the raster scan order of the picture may be derived based on the value of uniform_subpic_grid_flag according to Figure 19 the algorithm. Here, the value of ctbAddrRs may be from 0 to PicSizeInCtbsY - 1.
[0234] Embodiment 3
[0235] In the above embodiment, the sub-picture grid may have a 4-pixel resolution. The following table shows the storage cost in a sub-picture grid using a 4-sample resolution.
[0236] [Table 2]
[0237]
[0238] As shown in the above table, allocating the sub-picture grid in a 4K video sequence requires 553 - kB of memory. Similarly, when the resolution increases to 8K, the memory requirement in the worst case increases to 2.21 MB. In this embodiment, a new syntax element for controlling the grid size is disclosed. Figure 20 shows the SPS syntax according to an embodiment. As Figure 20 shown in the example of, the syntax element grid_spacing can be added to the SPS syntax.
[0239] The syntax element grid_spacing_minus4 2010 can be used to specify the grid spacing value. For example, the value obtained by adding 4 to grid_spacing_minus4 can specify the grid spacing of the current sub-picture grid.
[0240] Therefore, the above NumSubPicGridCols can be derived as follows.
[0241] [Formula 5]
[0242] NumSubPicGridCols = (pic_width_max_in_luma_samples + subpic_grid_col_width_minus1 * 4 + 3) / (subpic_grid_col_width_minus1 * (grid_spacing_minus4 + 4) + 4)
[0243] In addition, the above NumSubPicGridRows can be derived as follows.
[0244] [Formula 6]
[0245] NumSubPicGridRows = (pic_height_max_in_luma_samples + subpic_grid_row_height_minus1 * 4 + 3) / (subpic_grid_row_height_minus1 * (grid_spacing_minus4 + 4) + 4)
[0246] By limiting the grid spacing as described above, the memory in the worst case can be controlled more effectively. The above-described embodiment corresponds to the case where the value of the minimum grid interval to be used is 4, but the scope of the present disclosure is not limited thereto. For example, SPS syntax signaling can be performed for the case where the minimum grid spacing is 8, 16, 32, 64, or even the CTU. For this purpose, an expression such as grid_spacing_minusx can be used to represent the above-described syntax element grid_spacing. Here, "x" can be 8, 16, 32, 64, 128, and / or any other value.
[0247] Therefore, in order to obtain greater flexibility in sub-picture design, user-defined grid spacing can be supported. In this case, in order to control the memory requirements in the worst case, constraints can be applied to make the user-defined grid spacing size equal to or greater than the minimum grid spacing size.
[0248] In this regard, constraints on syntax elements (e.g., grid_spacing_minusx) can be applied as follows. In order to signal the minimum grid spacing, the syntax element grid_spacing_minus8 can be used to signal the value obtained by subtracting 8 from the minimum grid spacing. In a similar manner, the syntax element grid_spacing_minus16 can be used to signal the value obtained by subtracting 16 from the minimum grid spacing, the syntax element grid_spacing_minus32 can be used to signal the value obtained by subtracting 32 from the minimum grid spacing, the syntax element grid_spacing_minus64 can be used to signal the value obtained by subtracting 64 from the minimum grid spacing, or the syntax element grid_spacing_minus128 can be used to signal the value obtained by subtracting 128 from the minimum grid spacing.
[0249] The following table shows the memory load according to the grid spacing size. In the following table, it can be observed that the memory load decreases as the grid spacing increases. This is more obvious when the resolution increases.
[0250] [Table 3]
[0251] Grid spacing 4K (4096x2160) 8K (8192x4320) 4x 4 553kB 2.21MB 8x 8 138kB 553kB 16x 16 34kB 138kB 32x 32 8.6kB 34kB 64x 64 2.2kB 8.6kB 128x 128 0.5kB 2.2kB
[0252] Embodiment 4
[0253] In Figure 9In the above-described embodiment, when the width and height of the sub-picture are both 1, the width and height of the last column of the sub-picture and the last column and the last row of the sub-picture are not clearly generated. Therefore, this embodiment discloses modifications to SubPicHeight and SubPicWidth to clearly generate the width and height of the sub-picture.
[0254] Figure 21 An algorithm for this change is shown in Figure 21 . By calculating SubPicHeight and SubPicWidth, the area of the sub-picture can be calculated more accurately.
[0255] Embodiment 5
[0256] In Embodiment 2 above, a method of modifying the grid spacing of the sub-picture to provide greater flexibility to enable easier signaling of large sub-pictures is disclosed.
[0257] In the present disclosure, a method of signaling a non-uniform sub-picture grid is disclosed. Making the sub-picture grid information uniform or non-uniform may be more effective because it may be more aligned with the tile structure of the PPS and fewer sub-picture grid indices are signaled compared to large sub-pictures.
[0258] Hereinafter, only the changes compared to Embodiment 2 above will be described. According to Figure 17 , the SPS syntax can be modified as shown in Figure 22 . In the Figure 22 syntax, the syntax element subpic_grid_rows_height_minus1[i] can be used to specify the height of the i-th row of the sub-picture grid element. For example, the value obtained by adding 1 to subpic_grid_rows_height_minus1[i] can specify the height of the i-th row. The syntax element subpic_grid_cols_width_minus1[i] can be used to specify the width of the i-th column of the sub-picture grid element. For example, the value obtained by adding 1 to subpic_grid_cols_width_minus1[i] can specify the width of the i-th column.
[0259] When the value of uniform_subpic_grid_flag is the first value (e.g., 0) that specifies the width and height of the sub-picture grid elements are inconsistent, the individual width and height of the sub-picture grid elements can be signaled by subpic_grid_cols_width_minus1[i] and subpic_grid_rows_height_minus1[i]. In addition, when the value of uniform_subpic_grid_flag is the second value (e.g., 1) that specifies the width and height of the sub-picture grid elements are consistent, the individual width and height of the sub-picture grid elements can be signaled by subpic_grid_col_width_minus1 and subpic_grid_row_height_minus1.
[0260] As the expressions of the syntax elements change according to the value of uniform_subpic_grid_flag, the variables SubPicIdx, SubPicLeftBoundaryPos, SubPicTopBoundaryPos, SubPicRightBoundaryPos, and SubPicBotBoundaryPos can be derived according to Figure 23 the algorithm, and a list CtbToSubPicIdx[ctbAddrRs] of sub-picture indices specifying the CTB address ctbAddrRs according to the raster scan order of the picture can be derived according to Figure 24 the algorithm.
[0261] Encoding and decoding method
[0262] Hereinafter, an image decoding method performed by an image decoding device will be described with reference to Figure 25 The image decoding device according to an embodiment may include a memory and a processor, and the processor may perform the following operations.
[0263] First, the decoding device may obtain a first signaling flag (e.g., pps_subpics_present_flag in PPS) from the picture parameter set, and this first signaling flag specifies whether to use the picture parameter set to signal the identifier of the sub-picture that divides the current picture (S2510).
[0264] Next, when the first signaling flag specifies using the picture parameter set to signal the identifier of the sub-picture, the decoding device may obtain the sub-picture identifier information (e.g., slice_to_subpic_id[i]) from the picture parameter set (S2520).
[0265] More specifically, an identifier number information specifying a number of sub-picture identifier information obtained from a picture parameter set (e.g., num_slices_in_pic_minus1) can be obtained from the picture parameter set. Additionally, the decoding device can obtain the sub-picture identifier information from the picture parameter set based on the identifier number information. The identifier number information can specify a value obtained by subtracting 1 from the number of sub-picture identifier information obtained from the picture parameter set.
[0266] Next, the decoding device can decode the current picture by decoding the current sub-picture identified based on the identifier information (S2530).
[0267] Meanwhile, a sub-picture constraint flag can be obtained from the bitstream. When the sub-picture constraint flag specifies that the value of the sub-picture signaling flag is constrained to a predetermined value, the decoding device can determine the value of the sub-picture signaling flag as the value signaled by the picture parameter set not using the identifier of the specified sub-picture.
[0268] Furthermore, the decoding device can obtain a second signaling flag (e.g., subpics_present_flag in SPS) from the sequence parameter set, which specifies whether to use the identifier of the sub-picture that divides the current picture signaled by the sequence parameter set.
[0269] Additionally, when the second signaling flag specifies that the identifier of the sub-picture is signaled by the sequence parameter set, the decoding device can obtain dimension number information (e.g., num_subpic_grid_col_minus1, num_subpic_grid_row_minus1) specifying a number of dimension information of the sub-picture included in the sequence parameter set. When the width of the sub-picture is inconsistently determined (e.g., uniform_subpic_grid_flag == 0), the dimension number information can be obtained from the sequence parameter set. Moreover, the dimension information can also be obtained from the sequence parameter set according to the number identified by the dimension number information.
[0270] More specifically, the size information may include width information (e.g., subpic_grid_col_width_minus1[i]) and height information (e.g., subpic_grid_row_height_minus1[i]) of the sub-picture. The width of the sub-picture can be obtained for at least one sub-picture that divides the width of the current picture. Additionally, the height of the sub-picture can be obtained for at least one sub-picture that divides the height of the current block. Meanwhile, when the width of the sub-picture is consistently determined (e.g., uniform_subpic_grid_flag == 1), one size information can be obtained for each of the width and height (e.g., subpic_grid_col_width_minus1, subpic_grid_row_height_minus1).
[0271] Accordingly, the decoding device can obtain the size information of the sub-picture from the sequence parameter set based on the size number information.
[0272] In addition, the decoding device can determine the size of the sub-picture based on the size information and the spacing information obtained from the sequence parameter set. Here, the spacing information can specify the sample unit of the value specified by the size information. The sample unit specified by the spacing information can have a sample unit greater than the 4-sample unit.
[0273] Hereinafter, reference will be made to Figure 26 Describe an image encoding method performed by an image encoding device. The image encoding device according to an embodiment may include a memory and a processor, and the processor may perform operations corresponding to the operations of the above decoding device.
[0274] For example, the encoding device can determine sub-pictures by dividing the current picture (S2610). Next, the encoding device can determine whether to signal the identifier of the sub-picture using the picture parameter set (S2620). Next, when signaling the identifier of the sub-picture using the picture parameter set, the encoding device can generate a picture parameter set including a first signaling flag pps_subpics_present_flag that specifies that the identifier of the sub-picture is signaled using the picture parameter set and the sub-picture identifier information slice_to_subpic_id[i] (S2630).
[0275] More specifically, the encoding device may generate a picture parameter set including identifier number information (e.g., num_slices_in_pic_minus1) and sub-picture identifier information, and the identifier number information specifies the number of sub-picture identifier information included in the picture parameter set. The identifier number information may specify a value obtained by subtracting 1 from the number of sub-picture identifier information obtained from the picture parameter set.
[0276] Meanwhile, the encoding device may generate a bitstream including a sub-picture constraint flag. The encoding device may set the value of the sub-picture constraint flag to a value in which the value of the specified sub-picture signaling flag is constrained to a predetermined value. For example, the encoding device may set the value of the sub-picture signaling flag to a value that specifies not to signal the identifier of the sub-picture using the picture parameter set.
[0277] In addition, the encoding device may generate a sequence parameter set, which includes a second signaling flag (e.g., subpics_present_flag in SPS), and the second signaling flag specifies whether to signal the identifier of the sub-picture that divides the current picture using the sequence parameter set.
[0278] Furthermore, the encoding device may also include dimension number information in the sequence parameter set that specifies the number of dimension information of the sub-picture. When the width of the sub-picture is inconsistently determined, the dimension number information may be included in the sequence parameter set. More specifically, the dimension information may include the width information and height information of the sub-picture. The width of the sub-picture may be included in the sequence parameter set for at least one sub-picture that divides the width of the current picture. In addition, the height of the sub-picture may be included in the sequence parameter set for at least one sub-picture that divides the height of the current picture. Meanwhile, when the width of the sub-picture is consistently determined, only one dimension information may be included in the sequence parameter set for each of the width and height.
[0279] In addition, the encoding device may signal the dimension information and the interval information using the sequence parameter set to signal the dimension of the sub-picture. Here, the interval information may specify the sample unit of the value specified by the dimension information. The sample unit specified by the interval information may have a sample unit greater than 4-sample units.
[0280] Application embodiment
[0281] Although for clarity of description, the above-described exemplary method of the present disclosure is represented as a series of operations, it is not intended to limit the order of execution of the steps, and these steps may be performed simultaneously or in a different order when necessary. To implement the method according to the present invention, the described steps may further include other steps, may include the remaining steps except some steps, or may include other additional steps except some steps.
[0282] In the present disclosure, an image encoding apparatus or an image decoding apparatus that performs a predetermined operation (step) may perform an operation (step) of confirming an execution condition or situation of the corresponding operation (step). For example, if it is described that a predetermined operation is performed when a predetermined condition is satisfied, the image encoding apparatus or the image decoding apparatus may perform the predetermined operation after determining whether the predetermined condition is satisfied.
[0283] The various embodiments of the present disclosure are not a list of all possible combinations and are intended to describe representative aspects of the present disclosure, and the matters described in the various embodiments may be applied independently or in combinations of two or more.
[0284] The various embodiments of the present disclosure may be implemented in hardware, firmware, software, or a combination thereof. In the case where the present disclosure is implemented by hardware, the present disclosure may be implemented by an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a general-purpose processor, a controller, a microcontroller, a microprocessor, or the like.
[0285] In addition, an image decoding apparatus and an image encoding apparatus to which embodiments of the present disclosure are applied may be included in a multimedia broadcast transmission and reception apparatus, a mobile communication terminal, a home theater video apparatus, a digital cinema video apparatus, a surveillance camera, a video chat apparatus, a real-time communication apparatus such as video communication, a mobile streaming apparatus, a storage medium, a camera, a video on demand (VoD) service providing apparatus, an over the top video (OTT video) apparatus, an Internet streaming service providing apparatus, a three-dimensional (3D) video apparatus, a video phone video apparatus, a medical video apparatus, etc., and may be used to process video signals or data signals. For example, an OTT video apparatus may include a game machine, a Blu-ray player, an Internet access TV, a home theater system, a smart phone, a tablet PC, a digital video recorder (DVR), etc.
[0286] Figure 27 is a view showing a content streaming system to which embodiments of the present disclosure can be applied.
[0287] As Figure 27 shown, a content streaming system to which embodiments of the present disclosure are applied may mainly include an encoding server, a streaming server, a network server, a media storage device, a user device, and a multimedia input device.
[0288] The encoding server compresses the content input from multimedia input devices such as smart phones, cameras, video cameras, etc. into digital data to generate a bitstream and sends the bitstream to the streaming media server. As another example, when multimedia input devices such as smart phones, cameras, video cameras, etc. directly generate a bitstream, the encoding server can be omitted.
[0289] The bitstream can be generated by an image encoding method or an image encoding device applying an embodiment of the present disclosure, and the streaming media server can temporarily store the bitstream in the process of sending or receiving the bitstream.
[0290] The streaming media server sends multimedia data to the user device based on a request from the user through the network server, and the network server serves as a medium for informing the user of the service. When the user requests a required service from the network server, the network server can forward it to the streaming media server, and the streaming media server can send multimedia data to the user. In this case, the content streaming system may include a separate control server. In this case, the control server is used to control commands / responses between devices in the content streaming system.
[0291] The streaming media server can receive content from the media storage device and / or the encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming media server can store the bitstream within a predetermined time.
[0292] Examples of user devices may include mobile phones, smart phones, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smart watches, smart glasses, head-mounted displays), digital TVs, desktop computers, digital signage, etc.
[0293] Each server in the content streaming system can operate as a distributed server, and in this case, the data received from each server can be distributed.
[0294] The scope of the present disclosure includes software or machine-executable commands (e.g., operating systems, applications, firmware, programs, etc.) for enabling the operations of methods according to various embodiments to be executed on a device or computer, and non-transitory computer-readable media having such software or commands stored thereon and executable on a device or computer.
[0295] Industrial Applicability
[0296] Embodiments of the present disclosure can be used for encoding or decoding images.
Claims
1. An image decoding device, the image decoding device comprising: a memory; and at least one processor, the at least one processor being connected to the memory and configured to: obtain a first signaling flag from a picture parameter set, the first signaling flag specifying whether to use the picture parameter set to signal an identifier of a sub-picture included in a current picture; based on the first signaling flag specifying to use the picture parameter set to signal the identifier of the sub-picture, obtain sub-picture identifier information of the sub-picture from the picture parameter set; and decode the current picture by decoding a current sub-picture identified based on the sub-picture identifier information, wherein the image decoding device is further configured to: obtain a second signaling flag from a sequence parameter set, the second signaling flag specifying whether to use the sequence parameter set to signal the identifier of the sub-picture included in the current picture; based on the second signaling flag specifying to use the sequence parameter set to signal the identifier of the sub-picture, obtain size quantity information specifying a quantity of size information of the sub-picture included in the sequence parameter set; and based on the size quantity information, obtain the size information of the sub-picture from the sequence parameter set.
2. The image decoding device according to claim 1, Among them, obtain identifier quantity information specifying a quantity of the sub-picture identifier information obtained from the picture parameter set, and wherein the sub-picture identifier information is further obtained from the picture parameter set based on the identifier quantity information.
3. The image decoding device according to claim 2, wherein, The identifier quantity information specifies a value obtained by subtracting 1 from a quantity of the sub-picture identifier information obtained from the picture parameter set.
4. The image decoding device according to claim 1, Among them, a sub-picture constraint flag is obtained from a bitstream, and wherein based on the sub-picture constraint flag, a value of the first signaling flag is constrained to a predetermined value, and the first signaling flag is determined to be a value specifying not to use the picture parameter set to signal the identifier of the sub-picture.
5. The image decoding device according to claim 1, wherein, Based on the width of the sub-picture being inconsistently determined, obtain the size quantity information from the sequence parameter set.
6. The image decoding device according to claim 1, wherein, Obtain the size information from the sequence parameter set by a number identified according to the size quantity information.
7. The image decoding device according to claim 1, wherein The size information includes width information and height information of the sub-picture.
8. The image decoding device according to claim 7, wherein, Obtain the width of the sub-picture for at least one sub-picture dividing the width of the current picture.
9. The image decoding device according to claim 7, wherein, Obtain the height of the sub-picture for at least one sub-picture dividing the height of the current picture.
10. The image decoding device according to claim 1, wherein, Based on the width of the sub-picture being consistently determined, obtain one size information for each of width and height.
11. The image decoding device according to claim 1, Among them, the size of the sub-picture is determined based on the size information and spacing information obtained from the sequence parameter set, wherein the spacing information specifies a sample unit of a value specified by the size information, and wherein the sample unit specified by the spacing information is greater than a 4-sample unit.
12. An image encoding device, the image encoding device comprising: A memory; And At least one processor, the at least one processor being connected to the memory and configured to: Determine sub - pictures by splitting a current picture; Determine whether to signal the identifier of the sub - picture using a picture parameter set; and Based on signaling the identifier of the sub - picture using the picture parameter set, generate a picture parameter set including a first signaling flag and sub - picture identifier information of the sub - picture, the first signaling flag specifying whether to signal the identifier of the sub - picture using the picture parameter set, Wherein, the image encoding device is further configured to: Determine whether to signal the identifier of the sub - picture using a sequence parameter set; Based on signaling the identifier of the sub - picture using the sequence parameter set, determine the number of size information of the sub - picture; and Based on the number of the size information of the sub - picture, encode the size information of the sub - picture into the sequence parameter set, Wherein, a second signaling flag and the size number information of the sub - picture are encoded into the sequence parameter set, the second signaling flag specifying whether to signal the identifier of the sub - picture using the sequence parameter set, and the size number information specifying the number of the size information of the sub - picture.
13. A device for transmitting a bitstream for an image, the device comprising: At least one processor, the at least one processor being configured to obtain the bitstream generated by an image encoding device; And A transmitter, the transmitter being configured to transmit the bitstream, Wherein, the bitstream is generated by the following operations: Determine sub - pictures by splitting a current picture; Determine whether to signal the identifier of the sub - picture using a picture parameter set; and Based on signaling the identifier of the sub - picture using the picture parameter set, generate a picture parameter set including a first signaling flag and sub - picture identifier information of the sub - picture, the first signaling flag specifying whether to signal the identifier of the sub - picture using the picture parameter set, Wherein, the device for transmitting the bitstream is further configured to: Determine whether to signal the identifier of the sub - picture using a sequence parameter set; Based on signaling the identifier of the sub - picture using the sequence parameter set, determine the number of size information of the sub - picture; and Based on the number of the size information of the sub - picture, encode the size information of the sub - picture into the sequence parameter set, Wherein, a second signaling flag and the size number information of the sub - picture are encoded into the sequence parameter set, the second signaling flag specifying whether to signal the identifier of the sub - picture using the sequence parameter set, and the size number information specifying the number of the size information of the sub - picture.