Method and apparatus for signaling slice-related information
By signaling slice-related information in the image encoding system, the problem of high-resolution and high-quality images is solved, and more efficient image compression and segmentation is achieved.
Patent Information
- Application Number
- CN202080093827.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-11-27
- Filing Date
- 2020-11-26
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2040-11-26
AI Technical Summary
When sending and storing high-resolution and high-quality images, the prior art has problems that the amount of information increases and leads to high costs, and it is necessary to improve image encoding efficiency.
The segmentation structure of the current picture is derived and decoded by signaling slice-related information in the image encoding system, including a flag related to the existence of the sub-picture and a flag on whether each sub-picture includes only one slice.
The overall image compression efficiency and segmentation efficiency are improved, and the image encoding process is optimized.
Smart Images

Figure CN115004709B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to image coding technology, and more particularly, to a method and apparatus for signaling slice-related information in an image coding system. Background Art
[0002] Recently, demand for high-resolution, high-quality images, such as high-definition (HD) and ultra-high-definition (UHD), has been growing in various fields. Because image data has high resolution and high quality, the amount of information or bits to be transmitted has increased compared to conventional image data. Consequently, when image data is transmitted using media such as conventional wired / wireless broadband lines or stored using existing storage media, the transmission and storage costs increase.
[0003] Therefore, there is a need for efficient image compression technology that can effectively transmit, store, and reproduce information of high-resolution and high-quality images. Summary of the Invention
[0004] Technical issues
[0005] The present disclosure provides a method and apparatus for improving image coding efficiency.
[0006] The present disclosure also provides a method and apparatus for signaling segmentation information of a picture.
[0007] The present disclosure also provides a method and device for decoding a current picture based on segmentation information of the current picture.
[0008] Technical Solution
[0009] In one aspect, a method for decoding an image is provided. The method includes obtaining image information including segmentation information of a current picture from a bitstream; deriving a segmentation structure of the current picture including at least one slice based on the segmentation information of the current picture; and decoding the current picture based on the segmentation structure, wherein the image information includes a first flag related to the presence of sub-picture information and a second flag related to whether each sub-picture in the current picture includes only one slice, wherein the number of sub-pictures in the current picture is derived to be equal to 1 based on the value of the first flag being equal to 0, and wherein whether the number of slices included in the current picture is equal to 1 is determined by the value of the second flag based on the value of the first flag being equal to 0.
[0010] In another aspect, an image encoding method performed by an encoding device is provided. The method includes: partitioning a current picture into at least one slice; deriving a partition structure of the current picture including the at least one slice; generating partition information of the current picture based on the partition structure; and encoding image information of the current picture including the partition information, wherein the image information includes a first flag related to the presence of sub-picture information and a second flag related to whether each sub-picture in the current picture includes only one slice, wherein the number of sub-pictures in the current picture is derived to be equal to 1 based on the value of the first flag being equal to 0, and wherein whether the number of slices included in the current picture is equal to 1 is determined by the value of the second flag based on the value of the first flag being equal to 0.
[0011] In another aspect, a non-transitory computer-readable storage medium storing a bitstream containing image information that enables an image decoding method to be performed is provided. The image decoding method includes: obtaining image information including segmentation information of a current picture from a bitstream; deriving a segmentation structure of the current picture including at least one slice based on the segmentation information of the current picture; and decoding the current picture based on the segmentation structure, wherein the image information includes a first flag related to the presence of sub-picture information and a second flag related to whether each sub-picture in the current picture includes only one slice, wherein the number of sub-pictures in the current picture is derived to be equal to 1 based on the value of the first flag being equal to 0, and wherein whether the number of slices included in the current picture is equal to 1 is determined by the value of the second flag based on the value of the first flag being equal to 0.
[0012] Beneficial effects
[0013] According to the present disclosure, the overall image / video compression efficiency can be improved.
[0014] According to the present invention, the efficiency of segmentation can be improved.
[0015] According to the present disclosure, the efficiency of segmentation can be improved based on the segmentation information of the current picture. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 An example of a video / image coding system to which the embodiments of this document can be applied is schematically illustrated.
[0017] Figure 2 is a schematic diagram illustrating a configuration of a video / image encoding device to which the embodiments of this document can be applied.
[0018] Figure 3 is a schematic diagram illustrating a configuration of a video / image decoding device to which the embodiments of this document can be applied.
[0019] Figure 4 An exemplary hierarchical structure of encoded data is illustrated.
[0020] Figure 5 is a diagram illustrating an example of dividing a picture.
[0021] Figure 6 is a flowchart illustrating a picture encoding process according to an embodiment.
[0022] Figure 7 is a flowchart illustrating a picture decoding process according to an embodiment.
[0023] Figure 8 is a block diagram showing the configuration of an encoding apparatus according to an embodiment.
[0024] Figure 9 is a block diagram showing the configuration of a decoding device according to an embodiment.
[0025] Figure 10 is a flowchart illustrating the operation of an encoding device according to an embodiment.
[0026] Figure 11 is a block diagram showing the configuration of an encoding apparatus according to an embodiment.
[0027] Figure 12 is a flowchart illustrating the operation of a decoding device according to an embodiment.
[0028] Figure 13 is a block diagram showing the configuration of a decoding device according to an embodiment.
[0029] Figure 14 An example of a content streaming system to which the embodiments disclosed in this document can be applied is shown. DETAILED DESCRIPTION
[0030] This document can be modified in various ways and has various embodiments, and specific embodiments will be described in detail and illustrated in the accompanying drawings. However, this is not intended to limit this document to specific embodiments. The general terms in this specification are only used to describe specific embodiments and are not used to limit the technical spirit of this document. Singular expressions include plural expressions unless the context clearly indicates otherwise. Terms such as "including" and "having" in this specification should be understood to indicate the presence of characteristics, numbers, steps, operations, elements, parts or combinations thereof described in the specification, and do not exclude the possibility of the presence or addition of one or more other characteristics, numbers, steps, operations, elements, parts or combinations thereof.
[0031] In addition, the elements in the drawings described in this document are independently illustrated for the convenience of describing different characteristic functions. This does not mean that each element is implemented as different hardware or different software. For example, at least two of the elements can be combined to form a single element, and a single element can also be divided into multiple elements. Without departing from the main purpose of this document, embodiments in which elements are combined and / or separated are also included in the scope of the rights of this document.
[0032] In the present disclosure, "A or B" may mean "only A", "only B", or "both A and B". In other words, in the present disclosure, "A or B" may be interpreted as "A and / or B". For example, in the present disclosure, "A, B or C" may mean "only A", "only B", "only C", or "any combination of A, B, and C".
[0033] As used in this disclosure, a slash ( / ) or a comma may mean "and / or". For example, "A / B" may mean "A and / or B". Thus, "A / B" may mean "only A", "only B", or "both A and B". For example, "A, B, C" may mean "A, B, or C".
[0034] In the present disclosure, “at least one of A and B” may mean “only A”, “only B”, or “both A and B”. In addition, in the present specification, the expression “at least one of A or B” or “at least one of A and / or B” may be interpreted as “at least one of A and B”.
[0035] In addition, in the present disclosure, “at least one of A, B, and C” may mean “only A,” “only B,” “only C,” or “any combination of A, B, and C.” In addition, “at least one of A, B, or C” or “at least one of A, B, and / or C” may mean “at least one of A, B, and C.”
[0036] In addition, the brackets used in the present disclosure may mean "for example." Specifically, when it is indicated as "prediction (intra-frame prediction)", it may mean that "intra-frame prediction" is proposed as an example of "prediction." In other words, "prediction" in the present disclosure is not limited to "intra-frame prediction", and "intra-frame prediction" may be proposed as an example of "prediction". In addition, when it is indicated as "prediction (i.e., intra-frame prediction)", it may also mean that "intra-frame prediction" is proposed as an example of "prediction".
[0037] In the present disclosure, technical features separately described in one drawing may be implemented separately or simultaneously.
[0038] Hereinafter, preferred embodiments of the present document will be described in more detail with reference to the accompanying drawings. Hereinafter, in the accompanying drawings, the same reference numerals are used for the same elements, and redundant descriptions of the same elements may be omitted.
[0039] Figure 1 An example of a video / image coding system to which the embodiments of this document can be applied is schematically illustrated.
[0040] Reference Figure 1 The video / image coding system may include a first device (source device) and a second device (receiver device). The source device may transmit the coded video / image information or data in the form of a file or stream transmission to the receive device via a digital storage medium or a network.
[0041] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.
[0042] A video source may acquire video / images through a process of capturing, synthesizing, or generating video / images. A video source may include a video / image capture device and / or a video / image generation device. For example, a video / image capture device may include one or more cameras, a video / image archive including previously captured video / images, etc. For example, a video / image generation device may include a computer, a tablet computer, and a smartphone, and may (electronically) generate video / images. For example, a virtual video / image may be generated by a computer, etc. In this case, the video / image capture process may be replaced by a process of generating relevant data.
[0043] An encoding device encodes input video / images. For compression and coding efficiency, the encoding device performs a series of processes such as prediction, transformation, and quantization. The encoded data (encoded video / image information) is output as a bitstream.
[0044] The transmitter can transmit the encoded image / image information or data, output as a bitstream, in the form of a file or stream to a receiver of a receiving device via a digital storage medium or network. Digital storage media may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter may include components for generating a media file in a predetermined file format and may also include components for transmission via a broadcast / communication network. The receiver may receive / extract the bitstream and transmit the received bitstream to a decoding device.
[0045] The decoding device may decode a video / image by performing a series of processes corresponding to the operations of the encoding device, such as inverse quantization, inverse transformation, and prediction.
[0046] The renderer may render the decoded video / image, and the rendered video / image may be displayed on a display.
[0047] This document relates to video / image coding. For example, the methods / implementations disclosed in this document can be applied to methods disclosed in Versatile Video Coding (VVC), EVC (Essential Video Coding) standards, AOMedia Video 1 (AV1) standards, the second-generation Audio Video Coding standard (AVS2), or next-generation video / image coding standards (e.g., H.267 or H.268, etc.).
[0048] This document proposes various embodiments of video / image coding, and unless otherwise mentioned, these embodiments can be performed in combination with each other.
[0049] In the present disclosure, a video may refer to a collection of a series of images over time. Generally, a picture refers to a unit of an image representing a specific time region, and a slice / tile is a unit constituting a portion of a picture. A slice / tile may include one or more coding tree units (CTUs). A picture may include one or more slices / tiles.
[0050] A tile is a rectangular area of CTUs within a specific tile column and a specific tile row in a picture. A tile column is a rectangular area of CTUs whose height is equal to the height of the picture and whose width is specified by the syntax elements in the picture parameter set. A tile row is a rectangular area of CTUs whose height is specified by the syntax elements in the picture parameter set and whose width is equal to the picture width. Tile scanning is a specific sequential ordering of the CTUs of a partitioned picture in which the CTUs are sequentially ordered by a CTU raster scan in the tile and the tiles in the picture are sequentially ordered by a raster scan of the tiles of the picture. A slice may include multiple intact tiles or multiple consecutive CTU columns in the tiles of a picture that may be included in one NAL unit. In the present disclosure, tile groups may be used interchangeably with slices. For example, in the present disclosure, a tile group / tile group header may be referred to as a slice / slice header.
[0051] In addition, a picture can be divided into two or more sub-pictures. A sub-picture can be a rectangular area of one or more slices within a picture.
[0052] A pixel or a picture element (pel) may refer to the smallest unit constituting a picture (or image). In addition, "sample" may be used as a term corresponding to a pixel. A sample may generally represent a pixel or a pixel value, and may represent only a pixel / pixel value of a luminance component or a pixel / pixel value of a chrominance component.
[0053] A unit may represent a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to the region. A unit may include a luma block and two chroma (e.g., CB, CR) blocks. In some cases, a unit may be used interchangeably with terms such as block or region. In general, an M×N block may include M columns and N rows of samples (or sample arrays) or a set (or array) of transform coefficients.
[0054] Figure 2 is a diagram schematically illustrating a configuration of a video / image encoding device to which an embodiment of this document can be applied. Hereinafter, a video encoding device may include an image encoding device.
[0055] Reference Figure 2 , the encoding device 200 includes an image segmenter 210, a predictor 220, a residual processor 230 and an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, an inverse quantizer 234, and an inverse transformer 235. The residual processor 230 may also include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstructed block generator. According to an embodiment, the image segmenter 210, the predictor 220, the residual processor 230, the entropy encoder 240, the adder 250, and the filter 260 may be configured by at least one hardware component (e.g., an encoder chipset or processor). In addition, the memory 270 may include a decoded picture buffer (DPB) or may be configured by a digital storage medium. The hardware component may also include the memory 270 as an internal / external component.
[0056] The image splitter 210 may split the input image (or picture or frame) input to the encoding device 200 into one or more processors. For example, a processor may be referred to as a coding unit (CU). In this case, the coding unit may be recursively split from a coding tree unit (CTU) or a maximum coding unit (LCU) according to a quadtree, binary tree, and ternary tree (QTBTTT) structure. For example, a coding unit may be split into multiple coding units of greater depth based on a quadtree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quadtree structure may be applied first, followed by the binary tree structure and / or ternary structure. Alternatively, the binary tree structure may be applied first. The encoding process according to the present disclosure may be performed based on the final coding unit that is no longer split. In this case, the maximum coding unit may be used as the final coding unit based on image characteristics, coding efficiency, etc., or, if necessary, the coding unit may be recursively split into coding units of greater depth, and the coding unit of the optimal size may be used as the final coding unit. Here, the encoding process may include prediction, transformation, and reconstruction processes (described later). As another example, the processor may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may be split or divided from the final coding unit. The prediction unit may be a unit for sample prediction, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.
[0057] In some cases, the term "unit" may be used interchangeably with terms such as "block" or "region." In general, an M×N block may represent a set of samples or transform coefficients consisting of M columns and N rows. A sample may generally represent a pixel or pixel value, and may represent only the pixel / pixel value of the luma component or only the pixel / pixel value of the chroma component. A sample may be used as a term corresponding to a picture (or image) of a pixel or picture element.
[0058] In the encoding device 200, the prediction signal (prediction block, prediction sample array) output from the inter predictor 221 or the intra predictor 222 is subtracted from the input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is sent to the transformer 232. In this case, as shown, the part of the encoding device 200 that subtracts the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) can be referred to as the subtractor 231. The predictor can perform prediction on a block to be processed (hereinafter referred to as the current block) and generate a prediction block including prediction samples of the current block. The predictor can determine whether to apply intra prediction or inter prediction based on the current block or CU. As described later in the description of each prediction mode, the predictor can generate various information related to prediction, such as prediction mode information, and send the generated information to the entropy encoder 240. The information about the prediction can be encoded in the entropy encoder 240 and output in the form of a bitstream.
[0059] The intra-frame predictor 222 can predict the current block by referring to samples in the current picture. Depending on the prediction mode, the referenced samples may be located near the current block or may be spaced apart. In intra-frame prediction, the prediction mode may include multiple non-directional modes and multiple directional modes. For example, the non-directional mode may include a DC mode and a planar mode. For example, depending on the level of detail of the prediction direction, the directional mode may include 33 directional prediction modes or 65 directional prediction modes. However, this is merely an example, and more or fewer directional prediction modes may be used depending on the settings. The intra-frame predictor 222 may use the prediction mode applied to the neighboring blocks to determine the prediction mode applied to the current block.
[0060] The inter-frame predictor 221 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector in a reference picture. To reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. It can also include information about the inter-frame prediction direction (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, neighboring blocks can include spatially neighboring blocks in the current picture and temporally neighboring blocks in the reference picture. The reference picture containing the reference block and the reference picture containing the temporally neighboring block can be the same or different. Temporally neighboring blocks can be referred to as collocated reference blocks, collocated CUs (colCUs), etc., and the reference picture containing temporally neighboring blocks can be referred to as collocated pictures (colPics). For example, the inter-frame predictor 221 can configure a motion information candidate list based on the neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index for the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the inter-frame predictor 221 can use the motion information of the neighboring block as the motion information of the current block. In skip mode, unlike merge mode, a residual signal may not be transmitted. In the case of motion vector prediction (MVP) mode, the motion vector of the neighboring block may be used as a motion vector predictor, and the motion vector of the current block may be indicated by signaling a motion vector difference.
[0061] The predictor 220 can generate a prediction signal based on various prediction methods described below. For example, the predictor can not only apply intra prediction or inter prediction to predict a block, but also apply both intra prediction and inter prediction at the same time. This can be referred to as a combination of inter and intra prediction (CIIP). In addition, the predictor can predict a block based on an intra-block copy (IBC) prediction mode or a palette mode. The IBC prediction mode or palette mode can be used for content image / video coding of games, etc., for example, screen content coding (SCC). IBC basically performs prediction in the current picture, but can be performed similarly to inter prediction in terms of deriving a reference block in the current picture. That is, IBC can use at least one of the inter prediction techniques described in this document. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, the sample values within the picture can be signaled based on information about the palette table and the palette index.
[0062] The prediction signal generated by the predictor (including the inter-frame predictor 221 and / or the intra-frame predictor 222) can be used to generate a reconstruction signal or to generate a residual signal. The transformer 232 can generate a transform coefficient by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loève transform (KLT), a graph-based transform (GBT), or a conditional nonlinear transform (CNT). Here, GBT means a transform obtained from a graph when the relationship information between pixels is represented by a graph. CNT refers to a transform generated based on a prediction signal generated using all previously reconstructed pixels. In addition, the transform process can be applied to square pixel blocks of the same size, or can be applied to blocks of variable size rather than square.
[0063] The quantizer 233 can quantize the transform coefficients and send them to the entropy encoder 240. The entropy encoder 240 can encode the quantized signal (information about the quantized transform coefficients) and output a bitstream. The information about the quantized transform coefficients can be called residual information. The quantizer 233 can rearrange the quantized transform coefficients of the block type into a one-dimensional vector form based on the coefficient scanning order, and generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. Information about the transform coefficients can be generated. The entropy encoder 240 can perform various encoding methods such as (for example) exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. The entropy encoder 240 can encode information necessary for video / image reconstruction in addition to the quantized transform coefficients (for example, syntax element values, etc.) together or separately. The encoded information (for example, the encoded video / image information) can be transmitted or stored in the form of a bitstream in units of NAL (Network Abstraction Layer). The video / image information may also include information about various parameter sets such as an adaptive parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may also include general constraint information. In this document, information and / or syntax elements transmitted / signaled from the encoding device to the decoding device may be included in the video / image information. The video / image information may be encoded through the above-mentioned encoding process and included in the bitstream. The bitstream may be transmitted over a network or stored in a digital storage medium. The network may include a broadcast network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) that transmits the signal output from the entropy encoder 240 or a storage device (not shown) that stores the signal may be included as an internal / external element of the encoding device 200, and alternatively, the transmitter may be included in the entropy encoder 240.
[0064] The quantized transform coefficients output from the quantizer 233 can be used to generate a prediction signal. For example, the residual signal (residual block or residual sample) can be reconstructed by applying inverse quantization and inverse transformation to the quantized transform coefficients via the inverse quantizer 234 and the inverse transformer 235. The adder 250 adds the reconstructed residual signal to the prediction signal output from the inter-frame predictor 221 or the intra-frame predictor 222 to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). If there is no residual of the block to be processed (such as when skip mode is applied), the prediction block can be used as a reconstructed block. The adder 250 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current picture, and as described below, can be used for inter-frame prediction of the next picture through filtering.
[0065] Furthermore, luma mapping with chroma scaling (LMCS) may be applied during the picture encoding and / or reconstruction process.
[0066] The filter 260 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 260 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and store the modified reconstructed picture in the memory 270, specifically, in the DPB of the memory 270. Various filtering methods may include, for example, deblocking filtering, sample adaptive offset (SAO), adaptive loop filter, bilateral filter, etc. The filter 260 can generate various types of information related to filtering and transmit the generated information to the entropy encoder 290, as described in the subsequent description of each filtering method. The information related to filtering can be encoded by the entropy encoder 290 and output in the form of a bitstream.
[0067] The modified reconstructed picture transmitted to the memory 270 may be used as a reference picture in the inter predictor 221. When inter prediction is applied by the encoding apparatus, prediction mismatch between the encoding apparatus 200 and the decoding apparatus may be avoided, and encoding efficiency may be improved.
[0068] The DPB of the memory 270 can store the modified reconstructed picture for use as a reference picture in the inter-frame predictor 221. The memory 270 can store the motion information of the block from which the motion information in the current picture is derived (or encoded) and / or the motion information of the block in the reconstructed picture. The stored motion information can be transmitted to the inter-frame predictor 221 to be used as the motion information of the spatially adjacent block or the motion information of the temporally adjacent block. The memory 270 can store the reconstructed samples of the reconstructed block in the current picture and can transmit the reconstructed samples to the intra-frame predictor 222.
[0069] In addition, in the present disclosure, at least one of quantization / dequantization and / or transform / inverse transform may be omitted. When quantization / dequantization is omitted, the quantized transform coefficient may be referred to as a transform coefficient. When transform / inverse transform is omitted, the transform coefficient may be referred to as a coefficient or a residual coefficient, or, for consistency of expression, may still be referred to as a transform coefficient.
[0070] In addition, in the present disclosure, the quantized transform coefficient and the transform coefficient may be referred to as a transform coefficient and a scaled transform coefficient, respectively. In this case, the residual information may include information about the transform coefficient, and the information of the transform coefficient may be signaled through the residual coding syntax. The transform coefficient may be derived based on the residual information (or information about the transform coefficient), and the scaled transform coefficient may be derived by inverse transforming (scaling) the transform coefficient. Residual samples may be derived based on the inverse transform (transform) of the scaled transform coefficient. This may be applied / expressed in the same manner in other parts of the present disclosure.
[0071] Figure 3 is a diagram schematically illustrating the configuration of a video / image decoding device to which the disclosure of this document can be applied.
[0072] Reference Figure 3 , the decoding device 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-frame predictor 331 and an intra-frame predictor 332. The residual processor 320 may include an inverse quantizer 321 and an inverse transformer 321. According to an embodiment, the entropy decoder 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 may be configured by hardware components (e.g., a decoder chipset or processor). In addition, the memory 360 may include a decoded picture buffer (DPB) or may be configured by a digital storage medium. The hardware components may also include the memory 360 as an internal / external component.
[0073] When a bit stream including video / image information is input, the decoding apparatus 300 can reconstruct the bit stream corresponding to the bit stream in FIG. Figure 2 The image corresponding to the processing of the video / image information in the encoding device. For example, the decoding device 300 can derive the unit / block based on the block segmentation related information obtained from the bit stream. The decoding device 300 can use the processor applied in the encoding device to perform decoding. Therefore, for example, the decoding processor can be a coding unit, and the coding unit can be divided from the coding tree unit or the maximum coding unit according to the quadtree structure, the binary tree structure and / or the ternary tree structure. One or more transform units can be derived from the coding unit. The reconstructed image signal decoded and output by the decoding device 300 can be reproduced by the reproduction device.
[0074] The decoding device 300 may receive the data from the Figure 2 The received signal is output by the encoding device, and the entropy decoder 310 can decode the received signal. For example, the entropy decoder 310 can parse the bitstream to derive information necessary for image reconstruction (or picture reconstruction) (e.g., video / image information). The video / image information may also include information about various parameter sets such as the Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). In addition, the video / image information may also include conventional constraint information. The decoding device may also decode the picture based on the information about the parameter set and / or conventional constraint information. The signaled / received information and / or syntax elements subsequently described in this document can be decoded through a decoding process and obtained from the bitstream. For example, the entropy decoder 310 can decode the information in the bitstream based on a coding method such as Exponential Golomb coding, CABAC, or CAVLC, and output syntax elements required for image reconstruction and quantized values of the transform coefficients for the residual. More specifically, the CABAC entropy decoding method can receive the bin corresponding to each syntax element in the bitstream, use the decoding target syntax element information, the decoding information of the decoding target block, or the information of the symbol / bin decoded in the previous stage to determine the context model, and perform arithmetic decoding on the bin by predicting the probability of occurrence of the bin according to the determined context model, and generate a symbol corresponding to the value of each syntax element. In this case, the CABAC entropy decoding method can update the context model by using the information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model. The information related to prediction among the information decoded by the entropy decoder 310 can be provided to the predictor (inter-frame predictor 332 and intra-frame predictor 331), and the residual value (i.e., quantized transform coefficient and related parameter information) entropy decoded in the entropy decoder 310 can be input to the residual processor 320. The residual processor 320 can derive a residual signal (residual block, residual sample, residual sample array). In addition, information about filtering among the information decoded by the entropy decoder 310 can be provided to the filter 350. In addition, a receiver (not shown) for receiving a signal output from the encoding device may be configured as an internal / external element of the decoding device 300, or the receiver may be a component of the entropy decoder 310. In addition, the decoding device according to this document may be referred to as a video / image / picture decoding device, and the decoding device may be classified into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoder 310, and the sample decoder may include at least one of the inverse quantizer 321, the inverse transformer 322, the adder 340, the filter 350, the memory 360, the inter-frame predictor 332, and the intra-frame predictor 331.
[0075] The inverse quantizer 321 may inversely quantize the quantized transform coefficients and output the transform coefficients. The inverse quantizer 321 may rearrange the quantized transform coefficients in a two-dimensional block format. In this case, the rearrangement may be performed based on the coefficient scanning order performed in the encoding device. The inverse quantizer 321 may inversely quantize the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain the transform coefficients.
[0076] The inverse transformer 322 inversely transforms the transform coefficients to obtain a residual signal (residual block, residual sample array).
[0077] The predictor may perform prediction on the current block and generate a prediction block including prediction samples of the current block. The predictor may determine whether to apply intra prediction or inter prediction to the current block based on the information on prediction output from the entropy decoder 310 and may determine a specific intra / inter prediction mode.
[0078] The predictor 330 can generate a prediction signal based on various prediction methods described below. For example, the predictor can not only apply intra prediction or inter prediction to predict a block, but also apply intra prediction and inter prediction at the same time. This can be called a combination of inter and intra prediction (CIIP). In addition, the predictor can predict a block based on an intra-block copy (IBC) prediction mode or a palette mode. The IBC prediction mode or palette mode can be used for content image / video encoding of games, etc., for example, screen content coding (SCC). IBC basically performs prediction in the current picture, but can be performed similarly to inter prediction in terms of deriving a reference block in the current picture. That is, IBC can use at least one of the inter prediction techniques described in this document. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, the sample value within the picture can be signaled based on information about the palette table and the palette index.
[0079] The intra-frame predictor 331 can predict the current block by referencing samples in the current picture. Depending on the prediction mode, the referenced samples can be located near the current block or can be located separately. In intra-frame prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. The intra-frame predictor 331 can determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.
[0080] The inter-frame predictor 332 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. Motion information can include a motion vector and a reference picture index. Motion information can also include information on the inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, neighboring blocks can include spatial neighboring blocks in the current picture and temporal neighboring blocks in the reference picture. For example, the inter-frame predictor 332 can configure a motion information candidate list based on neighboring blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and information about the prediction can include information indicating the inter-frame prediction mode for the current block.
[0081] The adder 340 can generate a reconstructed signal (reconstructed picture, reconstructed block, or reconstructed sample array) by adding the obtained residual signal to the prediction signal (prediction block or prediction sample array) output from the predictor 330. If there is no residual of the processing target block (such as when a skip mode is applied), the prediction block can be used as the reconstructed block.
[0082] The adder 340 may be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal may be used for intra-frame prediction of the next block to be processed in the current picture, and as described later, may also be output through filtering or may also be used for inter-frame prediction of the next picture.
[0083] In addition, luma mapping with chroma scaling (LMCS) can also be applied to the picture decoding process.
[0084] The filter 350 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 350 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image, and store the modified reconstructed image in the memory 360, specifically, in the DPB of the memory 360. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0085] The (modified) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter-frame predictor 331. The memory 360 can store the motion information of the block from which the motion information in the current picture is derived (or decoded) and / or the motion information of the block in the reconstructed picture. The stored motion information can be transmitted to the inter-frame predictor 331 to be used as the motion information of the spatially adjacent block or the motion information of the temporally adjacent block. The memory 360 can store the reconstructed samples of the reconstructed block in the current picture and transmit the reconstructed samples to the intra-frame predictor 332.
[0086] In the present disclosure, the embodiments described in the filter 260, the inter-frame predictor 221, and the intra-frame predictor 222 of the encoding device 200 may be the same as the filter 350, the inter-frame predictor 332, and the intra-frame predictor 331 of the decoding device 300 or may be applied respectively to correspond to the filter 350, the inter-frame predictor 332, and the intra-frame predictor 331 of the decoding device 300. The same may also be applied to the inter-frame predictor 332 and the intra-frame predictor 331.
[0087] As described above, when performing video encoding, prediction is performed to enhance compression efficiency. A prediction block including prediction samples of a current block (i.e., a target coding block) can be generated by prediction. In this case, the prediction block includes prediction samples in the spatial domain (or pixel domain). The prediction block is derived identically in the encoding device and the decoding device. The encoding device can enhance image coding efficiency by signaling information (residual information) about the residual between the original block (rather than the original sample values of the original block themselves) and the prediction block to the decoding device. The decoding device can derive a residual block including residual samples based on the residual information, can generate a reconstructed block including reconstructed samples by adding the residual block and the prediction block, and can generate a reconstructed picture including the reconstructed block.
[0088] Residual information can be generated through a transformation process and a quantization process. For example, the encoding device can derive a residual block between the original block and the prediction block, derive a transform coefficient by performing a transformation process on the residual samples (residual sample array) included in the residual block, derive a quantized transform coefficient by performing a quantization process on the transform coefficient, and can signal the relevant residual information (through a bitstream) to the decoding device. In this case, the residual information may include value information such as the quantized transform coefficient, position information, transform scheme, transform kernel, and quantization parameter. The decoding device can perform an inverse quantization / inverse transform process based on the residual information and derive residual samples (or residual blocks). The decoding device can generate a reconstructed picture based on the prediction block and the residual block. In addition, for inter-frame prediction reference of subsequent pictures, the encoding device can derive a residual block by performing inverse quantization / inverse transform on the quantized transform coefficient, and can generate a reconstructed picture.
[0089] Figure 4 An exemplary hierarchical structure of encoded data is illustrated.
[0090] Reference Figure 4 , the coded data can be divided into the video coding layer (VCL) that manipulates the coding process of the video / image and the video / image itself, and the network abstraction layer (NAL) that stores and sends the coded video / image data and is between the video coding layer (VCL) and lower systems.
[0091] The VCL can generate a Supplemental Enhancement Information (SEI) message, which is supplemented by the header corresponding to the sequence and picture and the parameter set (Picture Parameter Set (PPS), Sequence Parameter Set (SPS), Video Parameter Set (VPS), etc.) of the video / image in the encoding process. The SEI message is separated from the information for the video / image (slice data). The VCL including the information for the video / image includes slice data and a slice header. In addition, the slice header may be called a tile group header, and the slice data may be called tile group data.
[0092] In NAL, a NAL unit can be generated by adding header information (NAL unit header) to the raw byte sequence payload (RBSP) generated in the VCL. In this case, RBSP refers to slice data, parameter sets, SEI messages, etc. generated in the VCL. The NAL unit header may include NAL unit type information specified according to the RBSP data included in the NAL unit.
[0093] The NAL unit, which is a basic unit of NAL, performs a function of mapping a coded image to a bit sequence of a lower system such as a file format, a real-time transport protocol (RTP), a transport stream (TS), etc. according to a predetermined specification.
[0094] As shown in the figure, according to the RBSP generated in the VCL, the NAL unit can be divided into a VCL NAL unit and a non-VCL NAL unit. The VCL NAL unit may refer to a NAL unit including information (slice data) for an image, and the non-VCL NAL unit may refer to a NAL unit including information required for decoding the image (parameter set or SEI message).
[0095] The VCL NAL units and non-VCL NAL units described above can be transmitted over the network by attaching header information according to the data specifications of the lower system. For example, the NAL unit can be converted into a predetermined data format such as the H.266 / VVC file format, RTP (Real-time Transport Protocol), TS (Transport Stream), etc.
[0096] As described above, for a NAL unit, a NAL unit type may be specified according to an RBSP data structure included in the NAL unit, and information of the NAL unit type may be stored in a NAL unit header and signaled.
[0097] For example, NAL units can be classified into VCL NAL unit types and non-VCL NAL unit types depending on whether the NAL unit includes information for the image (slice data). VCL NAL unit types can be classified according to the nature and type of the picture included in the VCL NAL unit, and non-VCL NAL unit types can be classified according to the type of parameter set.
[0098] The following describes an example of a NAL unit type specified according to the type of parameter set included in a non-VCL NAL unit type. The NAL unit type can be specified according to the type of parameter set. For example, the NAL unit type can be specified as one of an APS (adaptation parameter set) NAL unit (including the type of NAL unit of the APS), a DPS (decoding parameter set) NAL unit (including the type of NAL unit of the DPS), a VPS (video parameter set) NAL unit (including the type of NAL unit of the VPS), an SPS (sequence parameter set) NAL unit (including the type of NAL unit of the SPS), and a PPS (picture parameter set) NAL unit (including the type of NAL unit of the PPS).
[0099] The above-mentioned NAL unit type may have syntax information for the NAL unit type, and the syntax information may be stored in the NAL unit header and signaled. For example, the syntax information may be nal_unit_type, and the NAL unit type may be specified as a nal_unit_type value.
[0100] In addition, as described above, a picture can have multiple slices, and a slice can include a slice header and slice data. In this case, in addition to multiple slices (a collection of slice headers and slice data), a picture header can also be added to a picture. The picture header (picture header syntax) can include information / parameters that can be commonly applied to the picture. The slice header (slice header syntax) can include information / parameters that can be commonly applied to the slice. The APS (APS syntax) or PPS (PPS syntax) can include information / parameters that can be commonly applied to one or more slices or pictures. The SPS (SPS syntax) can include information / parameters that can be commonly applied to one or more sequences. The VPS (VPS syntax) can include information / parameters that can be commonly applied to multiple layers. The DPS (DPS syntax) can include information / parameters that can be commonly applied to the entire video. The DPS can include information / parameters related to the concatenation of CVS (Coded Video Sequence). In the present disclosure, the high-level syntax (HLS) can include at least one of the APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, picture header syntax, or slice header syntax.
[0101] In the present disclosure, the image / video information encoded from the encoding device to the decoding device and signaled in a bitstream format may include information included in a slice header, information included in a picture header, information included in an APS, information included in a PPS, information included in an SPS, information included in a VPS, and / or information included in a DPS, as well as information related to segmentation in a picture, intra-frame / inter-frame prediction information, residual information, loop filtering information, etc. In addition, the image / video information may also include information in a NAL unit header.
[0102] Figure 5 is a diagram illustrating an example of dividing a picture.
[0103] A picture may be partitioned into coding tree units (CTUs), and a CTU may correspond to a coding tree block (CTB). A CTU may include a coding tree block for luma samples and two corresponding coding tree blocks for chroma samples. In addition, the maximum available size of a CTU for coding and prediction may be different from the maximum available size of a CTU for transform.
[0104] A tile may correspond to a series of CTUs covering a rectangular area, and one picture may be partitioned into one or more tile rows and one or more tile columns.
[0105] In addition, a slice can include an integer number of intact tiles or an integer number of consecutive intact CTU columns. In this case, two types of slicing modes can be supported, including raster scan slicing mode and rectangular slicing mode.
[0106] In raster scan slicing mode, a slice can include a series of intact tiles in a tile raster scan of a picture. In rectangular slicing mode, a slice can include multiple intact tiles that together form a rectangular area of the picture. Alternatively, in rectangular slicing mode, a slice can include multiple consecutive CTU columns in tiles that together form a rectangular area of the picture. The tiles in a rectangular slice can be scanned in tile raster scan order within the rectangular area corresponding to the respective slice.
[0107] Furthermore, a sub-picture may include one or more slices covering a rectangular area of a picture.
[0108] Figure 5 (a) is a diagram illustrating an example in which a picture is divided into raster scan slices. For example, a picture can be divided into 12 tiles and 3 raster scan slices.
[0109] in addition, Figure 5 (b) is a diagram illustrating an example in which a picture is divided into rectangular slices. For example, a picture can be divided into 24 tiles (6 tile rows and 4 tile columns) and 9 rectangular slices.
[0110] also, Figure 5 (c) is a diagram illustrating an example of a picture being divided into tiles and rectangular slices. For example, a picture can be divided into 24 tiles (2 tile rows and 2 tile columns) and 4 rectangular slices.
[0111] Figure 6 is a flowchart illustrating a picture encoding process according to an embodiment.
[0112] In one embodiment, picture segmentation may be performed by the image segmentor 210 of the encoding apparatus (step S600 ), and picture encoding may be performed by the entropy encoder 240 of the encoding apparatus (step S610 ).
[0113] The encoding device according to the embodiment can derive the slices and / or tiles included in the current picture (step S600). For example, the encoding device can perform picture segmentation to encode the input current picture. For example, the encoding device can derive the slices and / or tiles included in the current picture. The encoding device can segment the picture into various formats by considering the image properties and coding efficiency of the current picture, and generate information indicating the segmentation format with the best coding efficiency, and then signal the information to the decoding device.
[0114] According to an embodiment, the encoding device may perform encoding on the current picture based on the derived slices and / or tiles (step S610). For example, the encoding device may encode video / image information including information about the slices and / or tiles and output the information in a bitstream format. The output bitstream may be forwarded to a decoding device via a digital storage medium or a network.
[0115] Figure 7 is a flowchart illustrating a picture decoding process according to an embodiment.
[0116] In one embodiment, the step of obtaining video / image information from the bitstream (step S710) and the step of deriving slices and / or tiles in the current picture (step S720) can be performed by the entropy decoder 310 of the decoding device, and the step of reconstructing the current picture based on the slices and / or tiles can be performed by the adder 340 of the decoding device.
[0117] The decoding device according to the embodiment can obtain video / image information from the received bit stream (step S710). The video / image information may include HLS, and the HLS may include slice information or tile information. The slice information may include information for specifying one or more slices in the current picture, and the tile information may include information for specifying one or more tiles in the current picture. The slice information or tile information can be obtained through various parameter sets, picture headers and / or slice headers.
[0118] Furthermore, the current picture may include a tile including one or more slices or a slice including one or more tiles.
[0119] The decoding apparatus according to the embodiment may derive a slice and / or a tile based on video / image information including information of the slice and / or tile in the current picture (step S720 ).
[0120] The decoding apparatus according to the embodiment may reconstruct (decode) the current picture based on slices and / or tiles (step S730 ).
[0121] Figure 8 is a block diagram showing a configuration of an encoding device according to an embodiment, and Figure 9 is a block diagram showing the configuration of a decoding device according to an embodiment.
[0122] Figure 8 An example of a block diagram of an encoding device is shown. Figure 8 The encoding device 800 shown includes a partitioner 810 and an entropy encoder 820. The partitioner 810 may perform the same Figure 2 The image segmenter 210 of the encoding device shown in FIG. 1 may perform the same and / or similar operations as the image segmenter 210 of the encoding device shown in FIG. 1 , and the entropy encoder 820 may perform the same and / or similar operations as the image segmenter 210 of the encoding device shown in FIG. Figure 2 The encoding device 800 may perform the same and / or similar operations as the entropy encoder 240 of the encoding device shown. For example, the segmenter 810 may derive at least one slice and / or at least one tile included in the current picture. For example, the encoding device may perform picture segmentation for encoding the input current picture. The input video may be segmented in the segmenter 810 and encoded in the entropy encoder 820. After encoding, the encoded video may be output from the encoding device 800.
[0123] Figure 9 An example of a block diagram of a decoding device is shown. Figure 9 The decoding device 900 shown includes an entropy decoder 910 and a reconstruction processor 920. The entropy decoder 910 can perform the same Figure 3 The reconstruction processor 920 may include the same and / or similar operations as the entropy decoder 310 of the decoding device shown. Figure 3 At least one component other than the entropy decoder 310 shown. The entropy decoder 910 may decode the input received from the encoding apparatus 800 and derive information of the tile. A processing unit may be determined based on the decoded information, and a reconstruction processor 920 may perform decoding based on the processing unit and may generate a reconstructed sample.
[0124] Furthermore, as described above, a picture can be divided into sub-pictures, tiles, and slices. Sub-picture information can be signaled via the SPS, and tile and rectangular slice information can be signaled via the PPS. Furthermore, raster scan slice information can be signaled via the slice header.
[0125] For example, the SPS syntax including sub-picture information can be expressed as shown in the following table.
[0126] [Table 1]
[0127]
[0128] For example, the PPS syntax including information of tiles and rectangular slices can be represented as shown in the following table.
[0129] [Table 2]
[0130]
[0131] In addition, for example, the syntax of the slice header including information of the raster scan slice can be expressed as shown in the following table.
[0132] [Table 3]
[0133]
[0134] In addition, the slice information and tile information in the current picture may include a flag related to whether each sub-picture in the current picture includes a single slice. This flag may be referred to as single_slice_per_subpic_flag or pps_single_slice_per_subpic_flag, but is not limited thereto. In addition, the sub-picture information may include a flag related to the presence of sub-picture information, and this flag may be referred to as subpics_present_flag or sps_subpic_info_present_flag, but is not limited thereto. For example, the sub-picture information may be included in a parameter set. For example, the sub-picture information may be included in an SPS.
[0135] Conventionally, when the value of the flag related to the presence of sub-picture information is zero, the flag value is restricted so that the value of the flag related to whether the sub-picture includes only one slice becomes zero. That is, when the value of the flag related to the presence of sub-picture information is zero, the sub-picture is determined to be unavailable, and the value of the flag related to whether the sub-picture includes only one slice is restricted to zero. However, this condition is very strict. For example, even when there is no sub-picture information, the current picture can be divided into two or more tiles, and all tiles can be included in a single slice. In this case, the current picture includes only one slice.
[0136] Therefore, the embodiment of the present disclosure proposes a method for eliminating the restriction that the value of the flag related to whether a sub-picture includes only one slice becomes zero when the value of the flag related to the presence of sub-picture information is zero. In this case, the flag related to whether a sub-picture includes only one slice can indicate that the current picture includes only one slice even when the sub-picture information does not exist.
[0137] For example, according to an embodiment, even when the coded layer video sequence (CLVS) does not have sub-picture information, a flag related to whether the sub-picture includes only one slice may be present. That is, even when the CLVS does not have sub-picture information, the flag related to whether the sub-picture includes only one slice may have a value of 0 or 1.
[0138] For example, when the value of the flag related to the presence of sub-picture information is zero and the value of the flag related to whether the sub-picture includes only one slice is 1, the current picture may include only one slice. That is, when there is no signaled sub-picture and the value of the flag related to whether the sub-picture includes only one slice is 1, the number of slices in the picture can be inferred to be 1.
[0139] In addition, when the value of the flag related to the presence of sub-picture information is zero, the number of sub-pictures in the current picture may be 1. For example, when the value of the flag related to the presence of sub-picture information is zero, the number of sub-pictures present in each of all pictures of the SPS of the reference image information may be 1.
[0140] In addition, a flag related to the number of slices included in the current picture may be included in the PPS of the image information. The flag related to the number of slices included in the current picture may be referred to as num_slices_in_pic_minus1 or pps_num_slices_in_pic_minus1, but is not limited thereto. In addition, a flag related to the number of sub-pictures included in the current picture may be included in the SPS of the image information. The flag related to the number of sub-pictures included in the current picture may be referred to as sps_num_subpics_minus1, but is not limited thereto.
[0141] When sub-picture information does not exist and the value of the flag related to whether the sub-picture includes only one slice is 1, the flag related to the number of slices included in the current picture can be inferred to have a value of 0. In addition, when sub-picture information does not exist and the value of the flag related to whether the sub-picture includes only one slice is 1, the flag related to the number of slices included in the current picture and the flag related to the number of sub-pictures included in the current picture can be inferred to have the same value.
[0142] In addition, in a case where a value of a flag related to whether a sub-picture includes only one slice is 1, all CTUs in the picture may belong to only one slice included in the picture.
[0143] The semantics of the syntax element including a flag related to whether a sub-picture includes only one slice and a flag related to the number of slices included in the current picture can be represented as in the following table.
[0144] [Table 4]
[0145]
[0146] Referring to the above table, when the value of single_slice_per_subpic_flag corresponding to the flag related to whether a sub-picture includes only one slice is 1, each sub-picture may include a single rectangular slice. In addition, when the value of single_slice_per_subpic_flag is zero, each sub-picture may include one or more rectangular slices. When the value of single_slice_per_subpic_flag is 1, the value of num_slices_in_pic_minus1 corresponding to the flag related to the number of slices included in the current picture can be inferred to have the value of sps_num_subpics_minus1 corresponding to the flag related to the number of sub-pictures included in the current picture.
[0147] Also, in a case where a value of single_slice_per_subpic_flag is 1 and a value of subpics_present_flag corresponding to a flag related to the presence of sub-picture information is zero, a picture referring to the PPS may have a single slice per picture.
[0148] In addition, the scanning process, which is the order in which tiles in a picture are decoded, can be determined according to the following table.
[0149] [Table 5]
[0150]
[0151]
[0152] Figure 10 is a flowchart illustrating the operation of an encoding device according to an embodiment, and Figure 11 is a block diagram showing the configuration of an encoding apparatus according to an embodiment.
[0153] Figure 10 The method shown can be Figure 2 or Figure 11 The encoding device shown performs. Figure 10 Steps S1010 to S1030 shown can be performed by Figure 2 The image segmenter 210 shown in FIG. 1 is executed, and step S1040 can be performed by Figure 2 The entropy encoder 240 shown in FIG. 1 is executed. In addition, the operations according to steps S1010 to S1040 are based on the reference Figures 1 to 9 Therefore, the description and reference are omitted or simplified. Figures 1 to 9 The description of the specification overlaps with the detailed description.
[0154] Reference Figure 10According to an embodiment, the encoding device may divide the current picture into at least one slice (step S1010). For example, the encoding device may divide the current picture into at least one sub-picture, at least one slice, and / or at least one tile. For example, the encoding device may perform picture division to encode the input current picture.
[0155] The encoding apparatus according to an embodiment may derive a partition structure of a current picture including at least one slice (step S1020 ).
[0156] According to an embodiment, the encoding device may generate segmentation information of the current picture based on the segmentation structure (step S1030). For example, the image segmentor 210 of the encoding device may generate segmentation information of the current picture based on at least one sub-picture, at least one slice, and / or at least one tile.
[0157] The encoding device according to the embodiment may encode image information including segmentation information (step S1040). For example, the image information may include at least one of segmentation information of the current picture or prediction information of the current block. Alternatively, the image information may include prediction samples derived by the predictor 220 of the encoding device and residual information generated from original samples in the residual processor 230 of the encoding device.
[0158] For example, the image information may include a first flag related to the presence of sub-picture information and a second flag related to whether the sub-picture includes only one slice. For example, the value of the first flag may correspond to zero, and the value of the second flag may correspond to 1.
[0159] For example, in a case where a value of a first flag related to the presence of sub-picture information is zero and a value of a second flag is 1, the number of slices included in the current picture may be derived to be equal to 1.
[0160] For example, the first flag related to the presence of sub-picture information may be included in an SPS (Sequence Parameter Set) of image information.
[0161] For example, the second flag related to whether the sub-picture includes only one slice may be included in the PPS (Picture Parameter Set) of the image information.
[0162] For example, the image information may include a third flag related to the number of slices included in the current picture, and the third flag may be included in the PPS of the image information.
[0163] In addition, for example, the image information may include a fourth flag related to the number of sub-pictures included in the current picture, and the fourth flag may be included in the SPS of the image information.
[0164] In addition, for example, in the case where the value of the first flag is zero, the number of sub-pictures present in each of all pictures of the SPS of the reference image information can be derived to be equal to 1.
[0165] In addition, the image information may include prediction information of the current picture. The prediction information may include information on the inter-frame prediction mode or intra-frame prediction mode performed in the current picture. The encoding device may generate the prediction information of the current picture and encode it.
[0166] In addition, the bitstream can be sent to the decoding device via a network or a (digital) storage medium. Here, the network can include a broadcast network and / or a communication network. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.
[0167] Figure 12 is a flowchart illustrating the operation of a decoding device according to an embodiment, and Figure 13 is a block diagram showing the configuration of a decoding device according to an embodiment.
[0168] Figure 12 The method shown can be Figure 3 or Figure 13 Steps S1210 and S1220 can be performed by Figure 3 The entropy decoder 310 shown in FIG. 12 is executed. In addition, step S1230 can be performed by Figure 3 The adder 340 shown in FIG. 1 is executed. In addition, the operations according to steps S1210 to S1230 are based on the reference Figures 1 to 9 Therefore, the description and reference are omitted or simplified. Figures 1 to 9 The description of the specification overlaps with the detailed description.
[0169] The decoding device according to the embodiment may obtain image information including segmentation information of the current picture from the bitstream (step S1210). For example, the entropy decoder 310 of the decoding device may obtain image information including segmentation information of the current picture from the bitstream. For example, the segmentation information may include sub-picture information, slice information, and / or tile information of the current picture. The sub-picture information includes information about at least one sub-picture included in the current picture, the slice information may include information about at least one slice included in the current picture, and the tile information may include information about at least one tile included in the current picture.
[0170] In addition, the image information may include at least a portion of prediction-related information or residual-related information. For example, the prediction-related information may include inter-frame prediction mode information or inter-frame prediction type information.
[0171] The decoding device according to the embodiment may derive a partition structure of a current picture including at least one slice (step S1220). For example, the entropy decoder 310 of the decoding device may derive at least one slice included in the current picture based on slice information included in the partition information. For example, the entropy decoder 310 of the decoding device may derive a partition structure of the current picture based on slice information of the current picture.
[0172] The decoding device according to the embodiment may decode the current picture based on the partition structure (step S1230). For example, the entropy decoder 310 of the decoding device may generate a reconstructed block or a reconstructed picture based on at least one slice. For example, the decoding device may derive prediction samples by performing an inter-frame prediction mode or an intra-frame prediction mode on the current picture based on prediction information received through the bitstream, and generate a reconstructed block by adding the prediction samples and residual samples.
[0173] For example, the image information may include a first flag related to the presence of sub-picture information and a second flag related to whether the sub-picture includes only one slice. For example, the value of the first flag may correspond to zero, and the value of the second flag may correspond to 1.
[0174] For example, in a case where the value of the first flag is zero and the value of the second flag is 1, the number of slices included in the current picture may be derived to be equal to 1.
[0175] For example, the first flag related to the presence of sub-picture information may be included in an SPS (Sequence Parameter Set) of image information.
[0176] For example, the second flag related to whether the sub-picture includes only one slice may be included in the PPS (Picture Parameter Set) of the image information.
[0177] For example, the image information may include a third flag related to the number of slices included in the current picture, and the third flag may be included in the PPS of the image information.
[0178] In addition, for example, the image information may include a fourth flag related to the number of sub-pictures included in the current picture, and the flag related to the number of sub-pictures included in the current picture may be included in the SPS of the image information.
[0179] In addition, for example, in the case where the value of the first flag is zero, the number of sub-pictures present in each of all pictures of the SPS of the reference image information can be derived to be equal to 1.
[0180] Although the method has been described based on a flowchart that lists steps or blocks in sequence in the above embodiments, the steps of this document are not limited to a specific order, and specific steps may be performed in different steps or in a different order or simultaneously with respect to the above steps. In addition, it will be understood by those skilled in the art that the steps in the flowchart are not exclusive and that another step may be included therein or one or more steps in the flowchart may be deleted without affecting the scope of the present disclosure.
[0181] The above-mentioned method according to the present disclosure may be in the form of software, and the encoding device and / or decoding device according to the present disclosure may be included in an apparatus for performing image processing (e.g., TV, computer, smart phone, set-top box, display device, etc.).
[0182] When the embodiments of the present disclosure are implemented with software, the above-mentioned methods can be implemented with modules (processing or functions) that perform the above-mentioned functions. The modules can be stored in a memory and executed by a processor. The memory can be installed inside or outside the processor and can be connected to the processor via various well-known devices. The processor may include an application-specific integrated circuit (ASIC), other chipsets, logic circuits and / or data processing devices. The memory may include a read-only memory (ROM), a random access memory (RAM), a flash memory, a memory card, a storage medium and / or other storage devices. In other words, according to the embodiments of the present disclosure, it can be implemented and executed on a processor, a microprocessor, a controller or a chip. For example, the functional units illustrated in the corresponding figures can be implemented and executed on a computer, a processor, a microprocessor, a controller or a chip. In this case, information about the implementation (for example, information about instructions) or the algorithm can be stored in a digital storage medium.
[0183] In addition, the decoding device and encoding device of the embodiment of the present document can be included in a multimedia broadcast transceiver, a mobile communication terminal, a home theater video device, a digital theater video device, a surveillance camera, a video chat device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camera, a video on demand (VoD) service provider, an over-the-top (OTT) video device, an Internet streaming service provider, a 3D video device, a virtual reality (VR) device, an augmented reality (AR) device, an image phone video device, a vehicle terminal (e.g., a vehicle (including an autonomous vehicle) terminal, an aircraft terminal, or a ship terminal) and a medical video device; and can be used to process image signals or data. For example, the OTT video device may include a game console, a Blueray player, a networked TV, a home theater system, a smartphone, a tablet PC, and a digital video recorder (DVR).
[0184] In addition, the processing method of the present document can be generated in the form of a program executed by a computer and can be stored in a computer-readable recording medium. The multimedia data with a data structure according to the present document can also be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices stored with computer-readable data. The computer-readable recording medium may include, for example, a Blu-ray disc (BD), a universal serial bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. The computer-readable recording medium also includes a medium implemented in the form of a carrier wave (e.g., transmission on the Internet). In addition, the bit stream generated by the encoding method can be stored in a computer-readable recording medium, or can be transmitted through a wired or wireless communication network.
[0185] In addition, the embodiments of this document can be implemented as a computer program product using a program code, and the program code can be executed by a computer according to the embodiments of this document. The program code can be stored on a computer-readable carrier.
[0186] Figure 14 An example of a content streaming system to which the embodiments disclosed in this document can be applied is shown.
[0187] Reference Figure 14 The content streaming system to which the embodiments of this document are applied may basically include an encoding server, a streaming server, a network server, a media storage, a user device, and a multimedia input device.
[0188] The encoding server is used to compress content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data, generate a bitstream, and transmit it to the streaming server. As another example, if the multimedia input device such as smartphones, cameras, and camcorders directly generates the bitstream, the encoding server can be omitted.
[0189] The bitstream may be generated by the encoding method or the bitstream generation method to which the embodiments of this document are applied, and the streaming server may temporarily store the bitstream during the process of transmitting or receiving the bitstream.
[0190] The streaming server transmits multimedia data to the user device via a network server based on the user's request. The network server serves as a tool for notifying the user of available services. When the user requests a desired service, the network server transfers the request to the streaming server, which then transmits the multimedia data to the user. In this regard, the content streaming system may include a separate control server, in which case the control server is used to control commands and responses between the various devices in the content streaming system.
[0191] The streaming server may receive content from a media storage and / or encoding server. For example, when receiving content from an encoding server, the content may be received in real time. In this case, the streaming server may store the bitstream for a predetermined period of time to provide a smooth streaming service.
[0192] For example, user devices may include mobile phones, smart phones, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation, tablet PCs, tablet PCs, ultrabooks, wearable devices (e.g., watch-type terminals (smart watches), glasses-type terminals (smart glasses), head-mounted displays (HMDs)), digital TVs, desktop computers, digital signs, etc.
[0193] Each server in the content streaming system may be operated as a distributed server, and in such case, data received by each server may be processed in a distributed manner.
[0194] The claims of this disclosure may be combined in various ways. For example, the technical features in the method claims of this disclosure may be combined to be implemented or performed in a device, and the technical features in the device claims may be combined to be implemented or performed in a method. Furthermore, the technical features in the method claims and the device claims may be combined to be implemented or performed in a device. Furthermore, the technical features in the method claims and the device claims may be combined to be implemented or performed in a method.
Claims
1. An image decoding method performed by a decoding device, the image decoding method comprising the following steps: receiving image information including segmentation information for a current picture from a bitstream; deriving a partition structure of the current picture including at least one slice based on the partition information for the current picture; as well as Decoding the current picture based on the segmentation structure, The segmentation information includes a first flag related to the existence of sub-picture information and a second flag related to whether each sub-picture in the current picture includes only one slice. Wherein, based on the value of the first flag being equal to 0, the number of sub-pictures in the current picture is derived to be equal to 1, Wherein, based on the value of the first flag being equal to 0, determining whether the number of slices included in the current picture is equal to 1 by the value of the second flag, and Wherein, based on the value of the first flag being equal to 0 and the value of the second flag being equal to 1, it is determined that the number of the slices included in the current picture is equal to 1.
2. A method for encoding an image, the method comprising the steps of: Split the current image into at least one slice; deriving a partitioning structure of the current picture including at least one slice; Generating segmentation information of the current picture based on the segmentation structure; as well as encoding the image information for the current picture including the segmentation information, The segmentation information includes a first flag related to the existence of sub-picture information and a second flag related to whether each sub-picture in the current picture includes only one slice. Wherein, based on the value of the first flag being equal to 0, the number of sub-pictures in the current picture is derived to be equal to 1, Wherein, based on the value of the first flag being equal to 0, determining whether the number of slices included in the current picture is equal to 1 by the value of the second flag, and Wherein, based on the value of the first flag being equal to 0 and the value of the second flag being equal to 1, it is determined that the number of the slices included in the current picture is equal to 1.
3. A method for transmitting image data, the method comprising the following steps: Obtaining a bitstream generated by a method, wherein the method comprises the steps of: partitioning a current picture into at least one slice, deriving a partition structure of the current picture including the at least one slice, generating partition information of the current picture based on the partition structure, and generating the bitstream by encoding image information for the current picture including the partition information; and sending said data comprising said bitstream, The segmentation information includes a first flag related to the existence of sub-picture information and a second flag related to whether each sub-picture in the current picture includes only one slice. Wherein, based on the value of the first flag being equal to 0, the number of sub-pictures in the current picture is derived to be equal to 1, Wherein, based on the value of the first flag being equal to 0, determining whether the number of slices included in the current picture is equal to 1 by the value of the second flag, and Wherein, based on the value of the first flag being equal to 0 and the value of the second flag being equal to 1, it is determined that the number of the slices included in the current picture is equal to 1.
Citation Information
Patent Citations
Method and apparatus for processing a video signal based on adaptive block patitioning
KR1020180033030A