Image decoding method, image encoding method, and data transmission method

By segmenting and deducing symbols from images, the efficiency of image encoding and decoding is improved, solving the problem of high transmission and storage costs for high-resolution, high-quality images, and achieving efficient image compression.

CN120897067APending Publication Date: 2025-11-04LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511337841.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-11-27
Filing Date
2020-11-26
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

The transmission and storage of high-resolution, high-quality images are costly, necessitating effective image compression techniques to improve coding efficiency.

Method used

By segmenting the image, we can deduce the existence of sub-image information and determine whether each sub-image contains only one slice, thereby improving the efficiency of image encoding and decoding.

Benefits of technology

It improves the overall image/video compression and segmentation efficiency, and performs efficient decoding based on the segmentation information of the current image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120897067A_ABST
    Figure CN120897067A_ABST
Patent Text Reader

Abstract

The invention relates to an image decoding method, an image encoding method, and a data transmission method. In a method for decoding an image by a decoding device according to the present disclosure, a current picture is configured to include a single slice based on a flag indicating whether information about a sub-picture exists and a flag indicating whether the sub-picture includes the single slice.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the original application No. 202080093704.1 (International Application No. PCT / KR2020 / 016944, filed on November 26, 2020, entitled "Method and apparatus for signaling picture partitioning information") for an invention patent application. TECHNICAL FIELD

[0002] The disclosure relates to an image coding technology, and most particularly, to a method and apparatus for signaling picture partitioning information in an image coding system. BACKGROUND

[0003] Recently, in various fields, the demand for high-resolution, high-quality images such as high-definition (HD) images and ultra-high-definition (UHD) images is increasing. Because image data has high resolution and high quality, the amount of information or bits to be transmitted increases relative to conventional image data. Therefore, when transmitting image data using a medium such as a conventional wired / wireless broadband line or storing image data using an existing storage medium, the transmission cost and storage cost thereof increase.

[0004] Therefore, there is a need for an efficient image compression technology that effectively transmits, stores, and reproduces information of high-resolution, high-quality images. SUMMARY

[0005] TECHNICAL PROBLEM

[0006] The disclosure provides a method and apparatus for improving image coding efficiency.

[0007] The disclosure also provides a method and apparatus for signaling partitioning information of a picture.

[0008] The disclosure also provides a method and apparatus for decoding a current picture based on partitioning information of the current picture.

[0009] TECHNICAL SOLUTION

[0010] In an aspect, an image decoding method performed by a decoding apparatus is provided. The method includes the steps of obtaining image information about a current picture from a bitstream; and decoding the current picture based on the image information, wherein the image information includes a first flag related to the presence of sub-picture information and a second flag related to whether each sub-picture includes only one slice, and wherein based on the first flag and the second flag, the number of slices included in the current picture is derived to be equal to 1.

[0011] In another aspect, there is provided an image encoding method performed by an encoding device. The method comprises the steps of deriving at least one slice by partitioning a current picture; and encoding image information of the current picture based on the at least one slice, wherein the image information comprises a first flag related to a presence of sub-picture information and a second flag related to whether each sub-picture in the current picture comprises only one slice, and wherein based on the first flag and the second flag, a number of slices included in the current picture is derived to be equal to 1.

[0012] In yet another aspect, there is provided a non-transitory computer-readable storage medium storing a bitstream comprising image information causing an image decoding method to be performed. The image decoding method comprises the steps of obtaining image information about a current picture from the bitstream; and decoding the current picture based on the image information, wherein the image information comprises a first flag related to a presence of sub-picture information and a second flag related to whether each sub-picture comprises only one slice, and wherein based on the first flag and the second flag, a number of slices included in the current picture is derived to be equal to 1.

[0013] Advantageous effects

[0014] According to the present disclosure, the overall image / video compression efficiency can be improved.

[0015] According to the present disclosure, the efficiency of partitioning can be improved.

[0016] According to the present disclosure, the efficiency of partitioning can be improved based on the partitioning information of a current picture. BRIEF DESCRIPTION OF DRAWINGS

[0017] Figure 1 An example of a video / image encoding system to which embodiments of the present document can be applied is schematically illustrated.

[0018] Figure 2 is a schematic diagram illustrating a configuration of a video / image encoding device to which embodiments of the present document can be applied.

[0019] Figure 3 is a schematic diagram illustrating a configuration of a video / image decoding device to which embodiments of the present document can be applied.

[0020] Figure 4 An exemplary hierarchical structure of encoded data is illustrated.

[0021] Figure 5 is a diagram illustrating an example of partitioning a picture.

[0022] Figure 6 is a flowchart illustrating a picture encoding process according to an embodiment.

[0023] Figure 7 FIG. 5 is a flowchart illustrating a picture decoding process according to an embodiment.

[0024] Figure 8 FIG. 6 is a flowchart illustrating an operation of an encoding apparatus according to an embodiment.

[0025] Figure 9 FIG. 7 is a block diagram illustrating a configuration of an encoding apparatus according to an embodiment.

[0026] Figure 10 FIG. 8 is a flowchart illustrating an operation of a decoding apparatus according to an embodiment.

[0027] Figure 11 FIG. 9 is a block diagram illustrating a configuration of a decoding apparatus according to an embodiment.

[0028] Figure 12 FIG. 10 illustrates an example of a content streaming system to which embodiments disclosed in the present document can be applied. DETAILED DESCRIPTION

[0029] The present document can be modified in various ways and has various embodiments, and a specific embodiment will be described in detail and illustrated in the accompanying drawings. However, this is not intended to limit the present document to a specific embodiment. The terms used in the present specification are merely used to describe a specific embodiment, and are not intended to limit the technical spirit of the present document. The singular expression includes the plural expression unless the context clearly indicates otherwise. Terms such as "include" and "have" in the present specification should be understood to indicate the presence of features, numbers, steps, operations, elements, parts, or combinations thereof described in the specification, and do not exclude the presence or addition of one or more other features, numbers, steps, operations, elements, parts, or combinations thereof.

[0030] In addition, the elements in the drawings described in the present document are independently illustrated for convenience of description related to different characteristic functions. This does not mean that each element is implemented as a different hardware or a different software. For example, at least two of the elements can be combined to form a single element, and a single element can also be divided into a plurality of elements. Embodiments in which elements are combined and / or separated are also included in the scope of rights of the present document without departing from the gist of the present document.

[0031] In the present disclosure, "A or B" can mean "only A", "only B", or "both A and B". In other words, in the present disclosure, "A or B" can be interpreted as "A and / or B". For example, in the present disclosure, "A, B, or C" can mean "only A", "only B", "only C", or "any combination of A, B, and C".

[0032] A slash ( / ) or a comma used in the present disclosure can mean "and / or". For example, "A / B" can mean "A and / or B". Thus, "A / B" can mean "only A", "only B", or "both A and B". For example, "A, B, C" can mean "A, B, or C".

[0033] In the present disclosure, "at least one of A and B" can mean "only A", "only B", or "both A and B". Also, in the present specification, the expression "at least one of A or B" or "at least one of A and / or B" can be interpreted as "at least one of A and B".

[0034] Also, in the present disclosure, "at least one of A, B, and C" can mean "only A", "only B", "only C", or "any combination of A, B, and C". Also, "at least one of A, B, or C" or "at least one of A, B, and / or C" can mean "at least one of A, B, and C".

[0035] Also, the parentheses used in the present disclosure can mean "for example". Specifically, when indicated as "prediction (intra prediction)", it can mean that "intra prediction" is proposed as an example of "prediction". In other words, the "prediction" of the present disclosure is not limited to "intra prediction", and "intra prediction" can be proposed as an example of "prediction". Also, when indicated as "prediction (i.e., intra prediction)", it can also mean that "intra prediction" is proposed as an example of "prediction".

[0036] In the present disclosure, technical features individually illustrated in the accompanying drawings can be implemented alone or can be simultaneously implemented.

[0037] Hereinafter, preferred embodiments of the present document are described more specifically with reference to the accompanying drawings. Hereinafter, in the accompanying drawings, the same reference numerals are used for the same elements, and redundant descriptions for the same elements can be omitted.

[0038] Figure 1 An example of a video / image encoding system to which embodiments of the present document can be applied is schematically illustrated.

[0039] Referring to Figure 1 , the video / image encoding system can include a first apparatus (a source apparatus) and a second apparatus (a receiving apparatus). The source apparatus can transmit encoded video / image information or data in a file or stream form to the receiving apparatus via a digital storage medium or a network.

[0040] A source device can include a video source, an encoding apparatus, and a transmitter. A receiving device can include a receiver, a decoding apparatus, and a renderer. The encoding apparatus can be referred to as a video / image encoding apparatus, and the decoding apparatus can be referred to as a video / image decoding apparatus. The transmitter can be included in the encoding apparatus. The receiver can be included in the decoding apparatus. The renderer can include a display, and the display can be configured as a separate device or an external component.

[0041] A video source can acquire a video / image through a process of capturing, synthesizing, or generating a video / image. The video source can include a video / image capturing device, and / or a video / image generating device. For example, the video / image capturing device can include one or more cameras, a video / image archive including previously captured video / images, etc. For example, the video / image generating device can include a computer, a tablet, and a smartphone, and can generate a video / image (electronically). For example, a virtual video / image can be generated through a computer, etc. In this case, the video / image capturing process can be replaced by a process of generating related data.

[0042] An encoding apparatus can encode an input video / image. For compression and encoding efficiency, the encoding apparatus can perform a series of processes such as prediction, transformation, and quantization. The encoded data (encoded video / image information) can be output in the form of a bitstream.

[0043] A transmitter can transmit the encoded image / image information or data output in the form of a bitstream in the form of a file or a stream to a receiver of a receiving device through a digital storage medium or a network. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter can include an element for generating a media file through a predetermined file format, and can include an element for transmission through a broadcasting / communication network. The receiver can receive / extract a bitstream and transmit the received bitstream to a decoding apparatus.

[0044] A decoding apparatus can decode a video / image by performing a series of processes such as dequantization, inverse transformation, and prediction corresponding to the operations of the encoding apparatus.

[0045] A renderer can render a decoded video / image. The rendered video / image can be displayed through a display.

[0046] This document relates to video / image encoding. For example, the methods / embodiments disclosed in this document can be applied to methods disclosed in the Versatile Video Coding (VVC), EVC (Elementary Video Coding) standard, AOMedia Video 1 (AV1) standard, second generation Audio Video Coding standard (AVS2), or next generation video / image encoding standards (e.g., H.267 or H.268, etc.).

[0047] Various embodiments of video / image encoding are presented in this document, and can be performed in combination with each other unless otherwise mentioned.

[0048] In this disclosure, a video can mean a set of a series of images over time. Typically, a picture means a unit of an image representing a specific time region, and a slice / tile is a unit of a part constituting a picture. A slice / tile can include one or more coding tree units (CTU). One picture can include one or more slices / tiles.

[0049] A tile is a rectangular region of CTUs within a specific tile column and a specific tile row in a picture. A tile column is a rectangular region of CTUs having a height equal to a height of a picture and a width specified by a syntax element in a picture parameter set. A tile row is a rectangular region of CTUs having a height specified by a syntax element in a picture parameter set and a width equal to a picture width. Tile scanning is a specific order of partitioning CTUs of a picture in which CTUs are continuously ordered in a CTU raster scan in a tile, and tiles in a picture are continuously ordered in a raster scan of tiles of the picture. A slice can include a plurality of intact tiles or a plurality of consecutive CTU columns in a tile of a picture that can be included in one NAL unit. In this disclosure, a tile group can be used interchangeably with a slice. For example, in this disclosure, a tile group / tile group header can be referred to as a slice / slice header.

[0050] In addition, a picture can be divided into two or more sub-pictures. A sub-picture can be a rectangular region of one or more slices within a picture.

[0051] A pixel or pel can mean a minimum unit constituting one picture (or image). In addition, a "sample" can be used as a term corresponding to a pixel. A sample can typically represent a pixel or a value of a pixel, and can represent only a pixel / pixel value of a luminance component or a pixel / pixel value of a chrominance component.

[0052] A unit can represent a basic unit of image processing. A unit can include at least one of a specific region of a picture and information related to the region. One unit can include one luminance block and two chrominance (e.g., cb, cr) blocks. In some cases, a unit can be used interchangeably with terms such as a block or a region. In general, an MxN block can include a set (or array) of M columns and N rows of samples (or sample arrays) or transform coefficients. Figure 2 FIG. 1 is a diagram schematically illustrating a configuration of a video / image encoding apparatus to which embodiments of the present document can be applied. Hereinafter, a video encoding apparatus can include an image encoding apparatus.

[0053] Referring to Figure 2The encoding apparatus 200 includes an image partitioner 210, a predictor 220, a residual processor 230, and an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 can include an inter-predictor 221 and an intra-predictor 222. The residual processor 230 can include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 can further include a subtractor 231. The adder 250 can be referred to as a reconstructor or a reconstructed block generator. According to embodiments, the image partitioner 210, the predictor 220, the residual processor 230, the entropy encoder 240, the adder 250, and the filter 260 can be configured by at least one hardware component (e.g., an encoder chipset or a processor). In addition, the memory 270 can include a decoded picture buffer (DPB), or can be configured by a digital storage medium. The hardware component can further include the memory 270 as an internal / external component.

[0054] The image partitioner 210 can partition an input image (or picture or frame) input to the encoding apparatus 200 into one or more processors. For example, the processor can be referred to as a coding unit (CU). In this case, the coding unit can be recursively partitioned from a coding tree unit (CTU) or a largest coding unit (LCU) according to a quad tree binary tree ternary (QTBT TT) structure. For example, one coding unit can be partitioned into a plurality of coding units at a deeper depth based on a quad tree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quad tree structure can be applied first, and the binary tree structure and / or the ternary structure can be applied later. Alternatively, the binary tree structure can be applied first. The encoding process according to the disclosure can be performed based on a final coding unit that is no longer partitioned. In this case, the largest coding unit can be used as the final coding unit based on coding efficiency, etc., according to the image characteristics, or if necessary, the coding unit can be recursively partitioned into coding units at a deeper depth and a coding unit having an optimal size can be used as the final coding unit. Here, the encoding process can include a process of prediction, transformation, and reconstruction (to be described later). As another example, the processor can further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit can be split or partitioned from the final coding unit described above. The prediction unit can be a unit of sample prediction, and the transform unit can be a unit for deriving a transform coefficient and / or a unit for deriving a residual signal from the transform coefficient.

[0055] In some cases, a unit can be used interchangeably with a term such as a block or a region. In general, an M x N block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, can represent only a pixel / pixel value of a luminance component, or can represent only a pixel / pixel value of a chrominance component. A sample can be used as a term corresponding to one picture (or image) of pixels or picture elements.

[0056] In the encoding apparatus 200, a prediction signal (prediction block, prediction sample array) output from the inter-predictor 221 or the intra-predictor 222 is subtracted from an input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is sent to the transformer 232. In this case, as illustrated, the portion in which the prediction signal (prediction block, prediction sample array) is subtracted from the input image signal (original block, original sample array) in the encoding apparatus 200 can be referred to as a subtractor 231. The predictor can perform prediction on a block to be processed (hereinafter, referred to as a current block) and generate a prediction block including prediction samples of the current block. The predictor can determine whether to apply intra-prediction or inter-prediction on a basis of the current block or the CU. As described later in the description of each prediction mode, the predictor can generate various information related to prediction such as prediction mode information, and send the generated information to the entropy encoder 240. The information about prediction can be encoded in the entropy encoder 240 and output in the form of a bitstream.

[0057] The intra-predictor 222 can predict the current block by referring to samples in the current picture. Depending on the prediction mode, the referred samples can be located in the vicinity of the current block or can be spaced apart. In intra-prediction, the prediction mode can include a plurality of non-directional modes and a plurality of directional modes. For example, the non-directional modes can include a DC mode and a planar mode. For example, depending on the level of detail of the prediction direction, the directional modes can include 33 directional prediction modes or 65 directional prediction modes. However, this is merely an example, and more or less directional prediction modes can be used according to settings. The intra-predictor 222 can determine the prediction mode applied to the current block using the prediction mode applied to the neighboring block.

[0058] The inter predictor 221 can derive a prediction block of a current block based on a reference block (a reference sample array) specified by a motion vector on a reference picture. Here, to reduce the amount of motion information transmitted in the inter prediction mode, motion information can be predicted in units of a block, a sub-block, or a sample based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the neighboring blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block can be the same or different. The temporal neighboring block can be referred to as a collocated reference block, a collocated CU (colCU), etc., and the reference picture including the temporal neighboring block can be referred to as a collocated picture (colPic). For example, the inter predictor 221 can configure a motion information candidate list based on the neighboring blocks and generate information indicating which candidate is used to derive a motion vector and / or a reference picture index of the current block. Inter prediction can be performed based on various prediction modes. For example, in the case of a skip mode and a merge mode, the inter predictor 221 can use motion information of the neighboring blocks as motion information of the current block. In the skip mode, unlike the merge mode, a residual signal can not be transmitted. In the case of a motion vector prediction (MVP) mode, a motion vector of the neighboring block can be used as a motion vector predictor, and a motion vector of the current block can be indicated by signaling a motion vector difference.

[0059] The predictor 220 can generate a prediction signal based on various prediction methods described below. For example, the predictor can not only apply intra prediction or inter prediction to predict one block, but also simultaneously apply both intra prediction and inter prediction. This can be referred to as combined inter and intra prediction (CIIP). In addition, the predictor can predict a block based on an intra block copy (IBC) prediction mode or a palette mode. The IBC prediction mode or the palette mode can be used for content image / video encoding of a game or the like, for example, screen content coding (SCC). The IBC basically performs prediction in the current picture, but can be similar to inter prediction in terms of deriving a reference block in the current picture. That is, the IBC can use at least one of the inter prediction techniques described in the present document. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, sample values within a picture can be signaled based on information about a palette table and a palette index.

[0060] The prediction signal generated by the predictor (including the inter-predictor 221 and / or the intra-predictor 222) can be used to generate a reconstructed signal or to generate a residual signal. The transformer 232 can generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique can include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-loève transform (KLT), a graph-based transform (GBT), or a conditional nonlinear transform (CNT). Here, the GBT means a transform obtained from a graph when a relationship information between pixels is expressed with a graph. The CNT refers to a transform generated based on a prediction signal generated using all previously reconstructed pixels. Also, the transform process can be applied to a square pixel block having the same size, or can be applied to a block having a variable size other than a square.

[0061] The quantizer 233 can quantize the transform coefficients and transmit them to the entropy encoder 240, and the entropy encoder 240 can encode and output a bitstream with respect to the quantized transform coefficients. The information with respect to the quantized transform coefficients can be referred to as residual information. The quantizer 233 can rearrange the quantized transform coefficients of a block type into a one-dimensional vector form based on a coefficient scan order, and generate information with respect to the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. Information with respect to the transform coefficients can be generated. The entropy encoder 240 can perform various encoding methods such as, for example, exponential Golomb, context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), and the like. The entropy encoder 240 can encode information necessary for video / image reconstruction, other than the quantized transform coefficients (e.g., values of syntax elements, etc.), together or separately. The encoded information (e.g., encoded video / image information) can be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer). The video / image information can further include information with respect to various parameter sets such as an adaptive parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information can further include regular constraint information. In this document, information and / or syntax elements transmitted / signaled from the encoding apparatus to the decoding apparatus can be included in the video / picture information. The video / picture information can be encoded through the above-described encoding process and included in the bitstream. The bitstream can be transmitted through a network, or can be stored in a digital storage medium. The network can include a broadcasting network and / or a communication network, and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) that transmits a signal output from the entropy encoder 240 or a storage (not shown) that stores the signal can be included as an internal / external element of the encoding apparatus 200, and alternatively, the transmitter can be included in the entropy encoder 240.

[0062] The quantized transform coefficients output from the quantizer 233 can be used to generate a prediction signal. For example, a residual signal (a residual block or a residual sample array) can be reconstructed by applying dequantization and inverse transform to the quantized transform coefficients via the dequantizer 234 and the inverse transformer 235. The adder 250 adds the reconstructed residual signal to the prediction signal output from the inter-predictor 221 or the intra-predictor 222 to generate a reconstructed signal (a reconstructed picture, a reconstructed block, a reconstructed sample array). If there is no residual of a block to be processed, such as in a case where a skip mode is applied, the prediction block can be used as the reconstructed block. The adder 250 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-prediction of a next block to be processed in the current picture, and can be used for inter-prediction of a next picture by filtering as described below.

[0063] Further, luma mapping with chroma scaling (LMCS) can be applied during the picture encoding and / or reconstruction process.

[0064] The filter 260 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 260 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and store the modified reconstructed picture in the memory 270, specifically, in the DPB of the memory 270. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset (SAO), adaptive loop filter, bilateral filter, etc. The filter 260 can generate various types of information related to filtering, and transmit the generated information to the entropy encoder 290 as described in the description of each filtering method below. The information related to filtering can be encoded by the entropy encoder 290 and output in the form of a bitstream.

[0065] The modified reconstructed picture transmitted to the memory 270 can be used as a reference picture in the inter-predictor 221. When inter-prediction is applied by the encoding apparatus, prediction mismatch between the encoding apparatus 200 and a decoding apparatus can be avoided, and encoding efficiency can be improved.

[0066] The DPB of the memory 270 can store the modified reconstructed picture to be used as a reference picture in the inter-predictor 221. The memory 270 can store motion information of a block from which motion information in the current picture is derived (or encoded) and / or motion information of a block in a picture that has been reconstructed. The stored motion information can be transmitted to the inter-predictor 221 to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory 270 can store reconstructed samples of a reconstructed block in the current picture, and can transmit the reconstructed samples to the intra-predictor 222.

[0067] Figure 3FIG. 1 is a diagram schematically illustrating a configuration of a video / image decoding apparatus to which the disclosure of the present document can be applied.

[0068] Referring to Figure 3 , the decoding apparatus 300 can include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 can include an intra predictor 331 and an inter predictor 332. The residual processor 320 can include a dequantizer 321 and an inverse transformer 322. According to an embodiment, the entropy decoder 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 can be configured by hardware components (e.g., a decoder chipset or a processor). In addition, the memory 360 can include a decoded picture buffer (DPB) or can be configured by a digital storage medium. The hardware components can further include the memory 360 as an internal / external component.

[0069] When a bitstream including video / image information is input, the decoding apparatus 300 can reconstruct an image corresponding to a process of processing video / image information in the encoding apparatus. Figure 2 For example, the decoding apparatus 300 can derive a unit / block based on block partitioning-related information obtained from the bitstream. The decoding apparatus 300 can perform decoding using a processor applied in the encoding apparatus. Accordingly, for example, the processor for decoding can be an encoding unit, and the encoding unit can be partitioned from a coding tree unit or a largest coding unit according to a quad tree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units can be derived from the encoding unit. The reconstructed image signal decoded and output by the decoding apparatus 300 can be reproduced by a reproduction apparatus.

[0070] The decoding apparatus 300 can receive a bitstream from Figure 2The signal output from the encoding apparatus can be received, and the received signal can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can parse a bitstream to derive information (e.g., video / image information) necessary for image reconstruction (or picture reconstruction). The video / image information can further include information on various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information can further include general constraint information. The decoding apparatus can also decode a picture based on the information on the parameter sets and / or the general constraint information. The information and / or the syntax elements signaled / received described later in this document can be decoded by the decoding process and obtained from the bitstream. For example, the entropy decoder 310 can decode information in the bitstream based on an encoding method such as exponential Golomb coding, CABAC, or CAVLC, and output syntax elements necessary for image reconstruction and quantized values of transform coefficients for a residual. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element in the bitstream, determine a context model using decoded target syntax element information, decoding information of a decoded target block, or information of decoded symbols / bins in a previous stage, and perform arithmetic decoding on the bins by predicting a probability of occurrence of the bins according to the determined context model, and generate a symbol corresponding to a value of each syntax element. In this case, the CABAC entropy decoding method can update the context model by using information of the decoded symbols / bins for a context model of a next symbol / bin after determining the context model. Prediction-related information among the information decoded by the entropy decoder 310 can be provided to the predictors (inter-predictor 332 and intra-predictor 331), and the residual values (i.e., quantized transform coefficients and related parameter information) entropy-decoded in the entropy decoder 310 can be input to the residual processor 320. The residual processor 320 can derive a residual signal (a residual block, residual samples, a residual sample array). In addition, filtering-related information among the information decoded by the entropy decoder 310 can be provided to the filter 350. Further, a receiver (not shown) for receiving a signal output from the encoding apparatus can be additionally configured as an internal / external element of the decoding apparatus 300, or the receiver can be a component of the entropy decoder 310. Furthermore, the decoding apparatus according to the present document can be referred to as a video / image / picture decoding apparatus, and the decoding apparatus can be classified into an information decoder (a video / image / picture information decoder) and a sample decoder (a video / image / picture sample decoder). The information decoder can include the entropy decoder 310, and the sample decoder can include at least one of the inverse quantizer 321, the inverse transformer 322, the adder 340, the filter 350, the memory 360, the inter-predictor 332, and the intra-predictor 331.

[0071] The dequantizer 321 can dequantize the quantized transform coefficients and output the transform coefficients. The dequantizer 321 can rearrange the quantized transform coefficients in a two-dimensional block form. In this case, the rearrangement can be performed based on a coefficient scanning order performed in the encoding apparatus. The dequantizer 321 can perform dequantization on the quantized transform coefficients using a quantization parameter (e.g., quantization step length information) and obtain the transform coefficients.

[0072] The inverse transformer 322 inverse-transforms the transform coefficients to obtain a residual signal (a residual block, a residual sample array).

[0073] The predictor can perform prediction on the current block and generate a prediction block including prediction samples of the current block. The predictor can determine whether to apply intra prediction or inter prediction to the current block based on information on prediction output from the entropy decoder 310 and can determine a specific intra / inter prediction mode.

[0074] The predictor 330 can generate a prediction signal based on various prediction methods described below. For example, the predictor can not only apply intra prediction or inter prediction to predict one block, but also simultaneously apply intra prediction and inter prediction. This can be referred to as combined inter and intra prediction (CIIP). In addition, the predictor can predict a block based on an intra block copy (IBC) prediction mode or a palette mode. The IBC prediction mode or the palette mode can be used for content image / video encoding of games, etc., for example, screen content coding (SCC). IBC basically performs prediction in the current picture, but can be similar to inter prediction in terms of deriving a reference block in the current picture. That is, IBC can use at least one of the inter prediction techniques described in the present document. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, sample values within a picture can be signaled based on information on a palette table and a palette index.

[0075] The intra predictor 331 can predict the current block by referring to samples in the current picture. Depending on the prediction mode, the referred samples can be located in the vicinity of the current block, or can be located apart. In intra prediction, the prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The intra predictor 331 can determine a prediction mode applied to the current block by using a prediction mode applied to a neighboring block.

[0076] The inter predictor 332 can derive a prediction block of the current block based on a reference block (a reference sample array) designated by a motion vector on a reference picture. In this case, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of a block, a sub-block, or a sample based on the correlation of the motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the neighboring blocks can include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. For example, the inter predictor 332 can configure a motion information candidate list based on the neighboring blocks and derive a motion vector and / or a reference picture index of the current block based on received candidate selection information. The inter prediction can be performed based on various prediction modes, and information about the prediction can include information indicating a mode of the inter prediction for the current block.

[0077] The adder 340 can generate a reconstructed signal (a reconstructed picture, a reconstructed block, or a reconstructed sample array) by adding the obtained residual signal to a prediction signal (a prediction block or a prediction sample array) output from the predictor 330. If there is no residual of the target block to be processed, such as in the case of applying a skip mode, the prediction block can be used as the reconstructed block.

[0078] The adder 340 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra prediction of the next block in the current picture, and as described later, can also be output through filtering or can also be used for inter prediction of the next picture.

[0079] In addition, luma mapping with chroma scaling (LMCS) can also be applied to the picture decoding process.

[0080] The filter 350 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 350 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture and store the modified reconstructed picture in the memory 360, specifically, in the DPB of the memory 360. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.

[0081] The (modified) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter prediction 332. The memory 360 can store motion information of a block from which motion information in the current picture is derived (or decoded) and / or motion information of a block in the already reconstructed picture. The stored motion information can be transferred to the inter prediction 332 to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory 360 can store reconstructed samples of a reconstructed block in the current picture and transfer the reconstructed samples to the intra prediction 331.

[0082] In the present disclosure, the embodiments described in the filter 260, the inter prediction 221, and the intra prediction 222 of the encoding device 200 can be the same as or respectively applied to correspond to the filter 350, the inter prediction 332, and the intra prediction 331 of the decoding device 300. The same can also apply to the inter prediction 332 and the intra prediction 331.

[0083] As described above, in performing video encoding, prediction is performed to enhance compression efficiency. A prediction block including prediction samples of a current block (i.e., a target coding block) can be generated by prediction. In this case, the prediction block includes prediction samples in a spatial domain (or pixel domain). The prediction block is derived identically in the encoding device and the decoding device. The encoding device can enhance image encoding efficiency by signaling information (residual information) about a residual between an original block (not the original sample values of the original block itself) and the prediction block to the decoding device. The decoding device can derive a residual block including residual samples based on the residual information, can generate a reconstructed block including reconstructed samples by adding the residual block and the prediction block, and can generate a reconstructed picture including the reconstructed block.

[0084] The residual information can be generated through a transform process and a quantization process. For example, the encoding device can derive a residual block between an original block and a prediction block, can derive transform coefficients by performing a transform process on residual samples (a residual sample array) included in the residual block, can derive quantized transform coefficients by performing a quantization process on the transform coefficients, and can signal the relevant residual information (through a bitstream) to the decoding device. In this case, the residual information can include value information of the quantized transform coefficients, position information, a transform scheme, a transform kernel, and a quantization parameter, etc. The decoding device can perform a dequantization / inverse transform process based on the residual information and can derive residual samples (or a residual block). The decoding device can generate a reconstructed picture based on the prediction block and the residual block. Furthermore, for inter prediction reference of a subsequent picture, the encoding device can derive a residual block by dequantizing / inversely transforming the quantized transform coefficients and can generate a reconstructed picture.

[0085] Figure 4 An exemplary hierarchy of coded data is illustrated.

[0086] Referring to Figure 4 , coded data can be divided into a video / image manipulation encoding process and a video coding layer (VCL) of a video / image itself and a network abstraction layer (NAL) that stores and transmits the encoded video / image and is between the video coding layer (VCL) and a lower system.

[0087] The VCL can generate a supplemental enhancement information (SEI) message, which is supplemental information required in an encoding process of a header corresponding to a sequence and a picture and a parameter set (picture parameter set (PPS), sequence parameter set (SPS), video parameter set (VPS), etc.) of a video / image. The SEI message is separate from information (slice data) for a video / image. The VCL including the information for the video / image includes slice data and a slice header. In addition, the slice header can be referred to as a tile group header, and the slice data can be referred to as tile group data.

[0088] In the NAL, a NAL unit can be generated by adding a header information (NAL unit header) to a raw byte sequence payload (RBSP) generated in the VCL. In this case, the RBSP is referred to as slice data, a parameter set, an SEI message, etc. generated in the VCL. The NAL unit header can include NAL unit type information designated according to RBSP data included in the NAL unit.

[0089] The NAL unit, which is a basic unit of the NAL, performs a function of mapping an encoded image to a bit sequence of a lower system such as a file format, a real-time transport protocol (RTP), a transport stream (TS), etc. according to a predetermined specification.

[0090] As illustrated in the drawing, the NAL unit can be divided into a VCL NAL unit and a non-VCL NAL unit according to the RBSP generated in the VCL. The VCL NAL unit can mean a NAL unit including information (slice data) for an image, and the non-VCL NAL unit can mean a NAL unit including information (a parameter set or an SEI message) required to decode an image.

[0091] The VCL NAL unit and the non-VCL NAL unit described above can be transmitted through a network by attaching header information according to a data specification of a lower system. For example, the NAL unit can be transformed into a data format of a predetermined specification such as an H.266 / VVC file format, an RTP (real-time transport protocol), a TS (transport stream), etc.

[0092] As described above, for a NAL unit, a NAL unit type can be specified according to a RBSP data structure included in the NAL unit, and information of the NAL unit type can be stored in a NAL unit header and signaled.

[0093] For example, according to whether a NAL unit includes information (slice data) for a picture, the NAL unit can be classified into a VCL NAL unit type and a non-VCL NAL unit type. The VCL NAL unit type can be classified according to properties and types of a picture included in the VCL NAL unit, and the non-VCL NAL unit type can be classified according to types of parameter sets.

[0094] Examples of a NAL unit type specified according to a type of a parameter set included in a non-VCL NAL unit type are described below. The NAL unit type can be specified according to a type of a parameter set. For example, the NAL unit type can be specified as one of an APS (adaptive parameter set) NAL unit (a type of a NAL unit including an APS), a DPS (decoding parameter set) NAL unit (a type of a NAL unit including a DPS), a VPS (video parameter set) NAL unit (a type of a NAL unit including a VPS), an SPS (sequence parameter set) NAL unit (a type of a NAL unit including an SPS), and a PPS (picture parameter set) NAL unit (a type of a NAL unit including a PPS).

[0095] The above-described NAL unit type can have syntax information for the NAL unit type, and the syntax information can be stored in a NAL unit header and signaled. For example, the syntax information can be nal_unit_type, and the NAL unit type can be specified as a nal_unit_type value.

[0096] Also, as described above, one picture can have a plurality of slices, and one slice can include a slice header and slice data. In this case, in addition to the plurality of slices (a set of slice header and slice data), a picture header can be added in one picture. The picture header (picture header syntax) can include information / parameters that can be commonly applied to a picture. The slice header (slice header syntax) can include information / parameters that can be commonly applied to a slice. The APS (APS syntax) or the PPS (PPS syntax) can include information / parameters that can be commonly applied to one or more slices or pictures. The SPS (SPS syntax) can include information / parameters that can be commonly applied to one or more sequences. The VPS (VPS syntax) can include information / parameters that can be commonly applied to a plurality of layers. The DPS (DPS syntax) can include information / parameters that can be commonly applied to an entire video. The DPS can include information / parameters related to a concatenation of CVS (coded video sequence). In the disclosure, the high-level syntax (HLS) can include at least one of the APS syntax, the PPS syntax, the SPS syntax, the VPS syntax, the DPS syntax, the picture header syntax, or the slice header syntax.

[0097] In the disclosure, the image / video information encoded from the encoding apparatus to the decoding apparatus and signaled in a bitstream format can include information included in a slice header, information included in a picture header, information included in an APS, information included in a PPS, information included in an SPS, information included in a VPS, and / or information included in a DPS, and information related to partitioning in a picture, intra / inter prediction information, residual information, loop filtering information, etc. In addition, the image / video information can further include information of a NAL unit header.

[0098] Figure 5 FIG. 1 is a diagram illustrating an example of partitioning a picture.

[0099] A picture can be partitioned into coding tree units (CTUs), and a CTU can correspond to a coding tree block (CTB). A CTU can include a coding tree block of luma samples and corresponding two coding tree blocks of chroma samples. Also, a maximum allowed size of a CTU for coding and prediction can be different from a maximum allowed size of a CTU for transform.

[0100] A tile can correspond to a series of CTUs covering a rectangular region, and one picture can be partitioned into one or more tile rows and one or more tile columns.

[0101] Also, a slice can include an integer number of complete tiles or an integer number of consecutive complete CTU columns. In this case, two types of slice modes, including a raster scan slice mode and a rectangular slice mode, can be supported.

[0102] In the raster scan slice mode, a slice can include a series of intact tiles in a tile raster scan of a picture. In the rectangular slice mode, a slice can include a plurality of intact tiles that together form a rectangular region of a picture. Alternatively, in the rectangular slice mode, a slice can include a plurality of consecutive CTU columns in a tile that together form a rectangular region of a picture. The tiles in a rectangular slice can be scanned in a tile raster scan order in the rectangular region corresponding to the respective slice.

[0103] Further, a sub-picture can include one or more slices covering a rectangular region of a picture.

[0104] Figure 5 (a) of FIG. 1 is a diagram illustrating an example in which a picture is partitioned into raster scan slices. For example, a picture can be partitioned into 12 tiles and 3 raster scan slices.

[0105] Further, Figure 5 (b) of FIG. 1 is a diagram illustrating an example in which a picture is partitioned into rectangular slices. For example, a picture can be partitioned into 24 tiles (6 tile rows 4 tile columns) and 9 rectangular slices.

[0106] Further, Figure 5 (c) of FIG. 1 is a diagram illustrating an example in which a picture is partitioned into tiles and rectangular slices. For example, a picture can be partitioned into 24 tiles (2 tile rows 2 tile columns) and 4 rectangular slices.

[0107] Figure 6 is a flowchart illustrating a picture encoding process according to an embodiment.

[0108] In one embodiment, picture partitioning (step S600) can be performed by the image partitioner 210 of the encoding device, and picture encoding (step S610) can be performed by the entropy encoder 240 of the encoding device.

[0109] The encoding device according to an embodiment can derive slices and / or tiles included in a current picture (step S600). For example, the encoding device can perform picture partitioning to encode an input current picture. For example, the encoding device can derive slices and / or tiles included in a current picture. The encoding device can partition a picture in various formats by considering image properties and encoding efficiency of the current picture, and generate information representing a partitioning format having the best encoding efficiency, and then, the information can be signaled to a decoding device.

[0110] The encoding apparatus according to the embodiment can perform encoding on the current picture based on the derived slice and / or tile (step S610). For example, the encoding apparatus can encode video / image information including information of the slice and / or tile, and output the information in a bitstream format. The output bitstream can be forwarded to a decoding apparatus through a digital storage medium or a network.

[0111] Figure 7 is a flowchart illustrating a picture decoding process according to the embodiment.

[0112] In the embodiment, the steps of obtaining video / image information from a bitstream (step S710) and deriving a slice and / or tile in a current picture (step S720) can be performed by an entropy decoder 310 of a decoding apparatus, and the step of reconstructing the current picture based on the slice and / or tile can be performed by an adder 340 of the decoding apparatus.

[0113] The decoding apparatus according to the embodiment can obtain video / image information from a received bitstream (step S710). The video / image information can include an HLS, and the HLS can include information of a slice or information of a tile. The information of the slice can include information for specifying one or more slices in the current picture, and the information of the tile can include information for specifying one or more tiles in the current picture. The information of the slice or the information of the tile can be obtained through various parameter sets, a picture header, and / or a slice header.

[0114] Further, the current picture can include a tile including one or more slices or a slice including one or more tiles.

[0115] The decoding apparatus according to the embodiment can derive a slice and / or tile based on the video / image information including information of the slice and / or tile in the current picture (step S720).

[0116] The decoding apparatus according to the embodiment can reconstruct (decode) the current picture based on the slice and / or tile (step S730).

[0117] Further, as described above, a picture can be partitioned into subpictures, tiles, and slices. Information of a subpicture can be signaled through an SPS, and information of a tile and a rectangular slice can be signaled through a PPS. In addition, information of a raster-scan slice can be signaled through a slice header.

[0118] For example, an SPS syntax including information of a subpicture can be expressed as in the following table.

[0119] [Table 1]

[0120]

[0121] For example, the PPS syntax including information of a tile and a rectangular slice can be expressed as the following table.

[0122] [Table 2]

[0123]

[0124] In addition, for example, the slice header syntax including information of a raster-scan slice can be expressed as the following table.

[0125] [Table 3]

[0126]

[0127] In addition, the information of a slice in the current picture and the information of a tile can include a flag related to whether each subpicture in the current picture includes a single slice. The flag can be referred to as single_slice_per_subpic_flag or pps_single_slice_per_subpic_flag, but can not be limited thereto. In addition, the information of a subpicture can include a flag related to the presence of subpicture information, and the flag can be referred to as subpics_present_flag or sps_subpic_info_present_flag, but can not be limited thereto. For example, the information of a subpicture can be included in a parameter set. For example, the information of a subpicture can be included in an SPS.

[0128] Conventionally, in a case where the value of the flag related to the presence of subpicture information is zero, the value of the flag is limited such that the value of the flag related to whether a subpicture includes only one slice becomes zero. That is, in a case where the value of the flag related to the presence of subpicture information is zero, it is determined that the subpicture is not available, and the value of the flag related to whether the subpicture includes only one slice is limited to zero. However, this condition is very strict. For example, even in the absence of subpicture information, the current picture can be divided into two or more tiles, and all the tiles can be included in a single slice. In this case, the current picture includes only one slice.

[0129] Accordingly, an embodiment of the disclosure proposes a method of removing the limitation that the value of the flag related to whether a subpicture includes only one slice becomes zero in a case where the value of the flag related to the presence of subpicture information is zero. In this case, the flag related to whether a subpicture includes only one slice can indicate a case where the current picture includes only one slice even in the absence of subpicture information.

[0130] For example, according to the embodiment, even in a case where subpicture information is not present in a coded layer video sequence (CLVS), a flag related to whether a subpicture includes only one slice can be present. That is, even in a case where subpicture information is not present in a CLVS, the flag related to whether a subpicture includes only one slice can have a value of zero or 1.

[0131] For example, in a case where a value of a flag related to presence of subpicture information is zero and a value of a flag related to whether a subpicture includes only one slice is 1, the current picture can include only one slice. That is, in a case where no signaled subpicture is present and a value of the flag related to whether a subpicture includes only one slice is 1, a number of slices in the picture can be inferred to be 1.

[0132] In addition, in a case where a value of a flag related to presence of subpicture information is zero, a number of subpictures in the current picture can be 1. For example, in a case where a value of a flag related to presence of subpicture information is zero, a number of subpictures present in each of all pictures of an SPS referring to picture information can be 1.

[0133] In addition, a flag related to a number of slices included in the current picture can be included in a PPS of picture information. The flag related to the number of slices included in the current picture can be referred to as num_slices_in_pic_minus1 or pps_num_slices_in_pic_minus1, but can not be limited thereto. In addition, a flag related to a number of subpictures included in the current picture can be included in an SPS of picture information. The flag related to the number of subpictures included in the current picture can be referred to as sps_num_subpics_minus1, but can not be limited thereto.

[0134] In a case where subpicture information is not present and a value of a flag related to whether a subpicture includes only one slice is 1, a flag related to a number of slices included in the current picture can be inferred to have a value of zero. In addition, in a case where subpicture information is not present and a value of a flag related to whether a subpicture includes only one slice is 1, a flag related to a number of slices included in the current picture and a flag related to a number of subpictures included in the current picture can be inferred to have the same value.

[0135] In addition, in a case where a value of a flag related to whether a subpicture includes only one slice is 1, all CTUs in the picture can belong to only one slice included in the picture.

[0136] The semantics of the syntax elements including the flag related to whether a subpicture includes only one slice and the flag related to the number of slices included in the current picture can be expressed as in the following table.

[0137] [Table 4]

[0138]

[0139] Referring to the above table, in a case where the value of single_slice_per_subpic_flag corresponding to the flag related to whether a subpicture includes only one slice is 1, each subpicture can include a single rectangular slice. Also, in a case where the value of single_slice_per_subpic_flag is zero, each subpicture can include one or more rectangular slices. In a case where the value of single_slice_per_subpic_flag is 1, the value of num_slices_in_pic_minus1 corresponding to the flag related to the number of slices included in the current picture can be inferred to have the value of sps_num_subpics_minus1 corresponding to the flag related to the number of subpictures included in the current picture.

[0140] Further, in a case where the value of single_slice_per_subpic_flag is 1 and the value of subpics_present_flag corresponding to the flag related to the presence of subpicture information is zero, the picture referring to the PPS can have a single slice per picture.

[0141] Further, the scanning process which is the order of decoding the tiles in a picture can be determined according to the following table.

[0142] [Table 5]

[0143]

[0144]

[0145] Figure 8 is a flowchart illustrating an operation of an encoding apparatus according to an embodiment, and Figure 9 is a block diagram illustrating a configuration of an encoding apparatus according to an embodiment.

[0146] Figure 8 The method shown in Figure 2 or Figure 9 The method shown in Figure 9 The step S810 shown in Figure 2 The step S810 shown in Figure 2The entropy encoder 240 illustrated in FIG. 2 performs. Further, the operations according to steps S810 and S820 are based on a reference Figures 1 to 7 A part of the description described above. Therefore, the detailed description overlapping with the reference Figures 1 to 7 A part of the description described above. Therefore, the detailed description overlapping with the reference

[0147] Referring to Figure 8 The encoding apparatus according to the embodiment can split a current picture and derive at least one slice (step S810). For example, the image splitter 210 of the encoding apparatus can generate split information of the current picture based on the at least one slice.

[0148] The encoding apparatus according to the embodiment can encode image information of the current picture based on the at least one slice (step S810). The image information can include split information generated based on the at least one slice.

[0149] For example, the image information can include a first flag related to presence of sub-picture information and a second flag related to whether a sub-picture includes only one slice. For example, based on the first flag and the second flag, a number of slices included in the current picture can be derived to be equal to 1.

[0150] For example, in a case where a value of the first flag related to presence of sub-picture information is zero and a value of the second flag is 1, the number of slices included in the current picture can be derived to be equal to 1.

[0151] For example, in a case where a value of the first flag related to presence of sub-picture information is zero, a number of sub-pictures present in the current picture can be 1.

[0152] For example, the first flag related to presence of sub-picture information can be included in SPS (Sequence Parameter Set) of the image information.

[0153] For example, the second flag related to whether a sub-picture includes only one slice can be included in PPS (Picture Parameter Set) of the image information.

[0154] For example, the image information can include a third flag related to a number of slices included in the current picture, and the third flag can be included in PPS of the image information.

[0155] In addition, for example, the image information can include a fourth flag related to a number of sub-pictures included in the current picture, and the fourth flag can be included in SPS of the image information.

[0156] In addition, for example, in a case where a value of the first flag is zero, a number of sub-pictures present in each of all pictures referring to the SPS of the image information can be derived to be equal to 1.

[0157] Further, the picture information can include prediction information of the current picture. The prediction information can include information of an inter prediction mode or an intra prediction mode performed in the current picture. The encoding apparatus can generate and encode the prediction information of the current picture.

[0158] Further, the bitstream can be transmitted to a decoding apparatus through a network or a (digital) storage medium. Here, the network can include a broadcasting network and / or a communication network. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.

[0159] Figure 10 is a flowchart illustrating an operation of an encoding apparatus according to an embodiment, and Figure 11 is a block diagram illustrating a configuration of a decoding apparatus according to an embodiment.

[0160] Figure 10 The method disclosed in Figure 3 may be performed by a decoding apparatus shown in Figure 11 . In particular, steps S1010 and S1020 can be performed by an entropy decoder 310 shown in Figure 3 . Further, the operations according to steps S1010 to S1020 are based on a part of the specification described with reference to Figures 1 to 7 . Therefore, the detailed description overlapping with the specification described with reference to Figures 1 to 7 is omitted or briefly described.

[0161] The decoding apparatus according to an embodiment can obtain picture information of a current picture from a bitstream (step S1010). For example, the entropy decoder 310 of the decoding apparatus can obtain the picture information including segmentation information of the current picture from the bitstream. For example, the segmentation information can include slice information of the current picture. In addition, the picture information can include at least a part of prediction-related information or residual-related information. For example, the prediction-related information can include inter prediction mode information or inter prediction type information.

[0162] The decoding apparatus according to an embodiment can decode the current picture based on at least the picture information (step S1020). For example, the entropy decoder 310 of the decoding apparatus can derive a segmentation structure of the current picture based on the slice information of the current picture.

[0163] For example, the picture information can include a first flag related to existence of sub-picture information and a second flag related to whether a sub-picture includes only one slice. For example, based on the first flag and the second flag, the number of slices included in the current picture can be derived to be equal to 1.

[0164] For example, in a case where the value of the first flag is zero and the value of the second flag is 1, the number of slices included in the current picture can be derived to be equal to 1.

[0165] For example, in a case where the value of the flag related to the presence of sub-picture information is zero, the number of sub-pictures present in the current picture can be 1.

[0166] For example, the first flag related to the presence of sub-picture information can be included in the SPS (Sequence Parameter Set) of the picture information.

[0167] For example, the second flag related to whether the sub-picture includes only one slice can be included in the PPS (Picture Parameter Set) of the picture information.

[0168] For example, the picture information can include a third flag related to the number of slices included in the current picture, and the third flag can be included in the PPS of the picture information.

[0169] In addition, for example, the picture information can include a fourth flag related to the number of sub-pictures included in the current picture, and the flag related to the number of sub-pictures included in the current picture can be included in the SPS of the picture information.

[0170] In addition, for example, in a case where the value of the first flag is zero, the number of sub-pictures present in each of all pictures referring to the SPS of the picture information can be derived to be equal to 1.

[0171] Although the method has been described based on a flowchart in which steps or blocks are sequentially listed in the above-described embodiments, the steps of the present document are not limited to a specific order, and a specific step can be performed in a different step or in a different order or simultaneously with respect to the above-described steps. In addition, those of ordinary skill in the art will understand that the steps in the flowchart are not exclusive, and another step can be included therein or one or more steps in the flowchart can be deleted without affecting the scope of the present disclosure.

[0172] The above-mentioned method according to the present disclosure can be in the form of software, and the encoding apparatus and / or the decoding apparatus according to the present disclosure can be included in an apparatus for performing image processing (e.g., a TV, a computer, a smart phone, a set-top box, a display apparatus, etc.).

[0173] When the embodiments of the disclosure are implemented with software, the above-described methods can be implemented with a module (process or function) that performs the above-described functions. The module can be stored in a memory and executed by a processor. The memory can be installed in the inside or outside of the processor, and can be connected to the processor via various well-known means. The processor can include an application-specific integrated circuit (ASIC), other chip sets, logic circuits, and / or data processing devices. The memory can include read-only memory (ROM), random access memory (RAM), flash memory, memory cards, storage media, and / or other storage devices. In other words, according to the embodiments of the disclosure, a processor, a microprocessor, a controller, or a chip can be implemented and executed. For example, the functional units illustrated in the respective drawings can be implemented and executed on a computer, a processor, a microprocessor, a controller, or a chip. In this case, information about the implementation (for example, information about instructions) or an algorithm can be stored in a digital storage medium.

[0174] In addition, the decoding apparatus and the encoding apparatus to which the embodiments of the disclosure are applied can be included in a multimedia broadcast transceiver, a mobile communication terminal, a home theater video device, a digital theater video device, a surveillance camera, a video chat device, a real-time communication device such as a video communication, a mobile streaming device, a storage medium, a camcorder, a video on demand (VoD) service provider, an over-the-top (OTT) video device, an Internet streaming service provider, a 3D video device, a virtual reality (VR) device, an augmented reality (AR) device, a picture phone video device, a vehicle terminal (for example, a vehicle (including an autonomous vehicle) terminal, an airplane terminal, or a ship terminal), and a medical video device; and can be used to process an image signal or data. For example, the OTT video device can include a game console, a Blu-ray player, a networked TV, a home theater system, a smart phone, a tablet PC, and a digital video recorder (DVR).

[0175] In addition, the processing method to which the disclosure is applied can be generated in the form of a program executed by a computer, and can be stored in a computer-readable recording medium. Multimedia data having a data structure according to the disclosure can also be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distribution devices that store computer-readable data. The computer-readable recording medium can include, for example, a Blu-ray disc (BD), a universal serial bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. The computer-readable recording medium also includes a medium implemented in the form of a carrier wave (for example, transmission over the Internet). In addition, a bitstream generated by the encoding method can be stored in a computer-readable recording medium or can be transmitted through a wired or wireless communication network.

[0176] In addition, the embodiments of the present document can be implemented as a computer program product using a program code, and the program code can be executed by a computer according to the embodiments of the present document. The program code can be stored on a computer readable carrier.

[0177] Figure 12 An example of a content streaming system to which the embodiments disclosed in the present document can be applied is illustrated.

[0178] Referring to Figure 12 A content streaming system to which the embodiments of the present document are applied can basically include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.

[0179] The encoding server serves to compress content input from a multimedia input device such as a smart phone, a camera, a camcorder, or the like, into digital data, generate a bitstream, and transmit the same to the streaming server. As another example, in the case where a bitstream is directly generated by a multimedia input device such as a smart phone, a camera, a camcorder, or the like, the encoding server can be omitted.

[0180] A bitstream can be generated by an encoding method or a bitstream generation method to which the embodiments of the present document can be applied. Also, the streaming server can temporarily store a bitstream in the process of transmitting or receiving the bitstream.

[0181] The streaming server transmits multimedia data to a user device through a web server that serves as a tool for informing a user of what services exist based on a request of the user. When the user requests a service that the user wants, the web server transfers the request to the streaming server, and the streaming server transmits multimedia data to the user. In this regard, the content streaming system can include a separate control server, and in this case, the control server serves to control commands / responses between the respective devices in the content streaming system.

[0182] The streaming server can receive content from a media storage and / or an encoding server. For example, in the case where content is received from the encoding server, the content can be received in real time. In this case, the streaming server can store a bitstream for a predetermined period of time to provide a streaming service smoothly.

[0183] For example, the user device can include a mobile phone, a smart phone, a laptop computer, a digital broadcasting terminal, a personal digital assistant (PDA), a portable multimedia player (PMP), a navigation, a slate PC, a tablet PC, an ultrabook, a wearable device (for example, a watch-type terminal (a smart watch), a glasses-type terminal (smart glasses), a head-mounted display (HMD)), a digital TV, a desktop computer, a digital signage, or the like.

[0184] Each server in the content streaming system can be operated as a distributed server, and in this case, data received by each server can be processed in a distributed manner.

[0185] The claims of the present disclosure can be combined in various ways. For example, the technical features in the method claims of the present disclosure can be combined to be implemented or executed in a device, and the technical features in the device claims can be combined to be implemented or executed in a method. Further, the technical features in the method claims and the device claims can be combined to be implemented or executed in a device. Further, the technical features in the method claims and the device claims can be combined to be implemented or executed in a method.

Claims

1. An image decoding method performed by a decoding device, the image decoding method comprising the steps of: obtaining, from a bitstream, image information about a current picture; and decoding the current picture based on the image information, wherein the image information comprises a first flag related to the presence of subpicture information and a second flag related to whether each subpicture comprises only one slice, wherein, based on the first flag and the second flag, the number of slices comprised in the current picture is derived to be equal to 1, wherein, based on the value of the first flag being equal to 0 and the value of the second flag being equal to 1, the number of slices comprised in the current picture is derived to be equal to 1, wherein the value of the second flag being equal to 1 indicates that each subpicture consists of one and only one slice, and wherein the first flag is comprised in a sequence parameter set.

2. An image encoding method performed by an encoding device, the image encoding method comprising the steps of: deriving at least one slice by partitioning a current picture; and encoding image information for the current picture based on the at least one slice, wherein the image information comprises a first flag related to the presence of subpicture information and a second flag related to whether each subpicture comprises only one slice, wherein, based on the first flag and the second flag, the number of slices comprised in the current picture is derived to be equal to 1, wherein, based on the value of the first flag being equal to 0 and the value of the second flag being equal to 1, the number of slices comprised in the current picture is derived to be equal to 1, wherein the value of the second flag being equal to 1 indicates that each subpicture consists of one and only one slice, and wherein the first flag is comprised in a sequence parameter set.

3. A method of transmission of data for an image, the method of transmission comprising the steps of: obtaining a bitstream generated by a method, wherein the method comprises performing: deriving at least one slice by partitioning a current picture and generating the bitstream by encoding image information for the current picture based on the at least one slice; and transmitting the data containing the bitstream, wherein the image information comprises a first flag related to the presence of subpicture information and a second flag related to whether each subpicture comprises only one slice, wherein, based on the first flag and the second flag, the number of slices comprised in the current picture is derived to be equal to 1, wherein, based on the value of the first flag being equal to 0 and the value of the second flag being equal to 1, the number of slices comprised in the current picture is derived to be equal to 1, wherein the value of the second flag being equal to 1 indicates that each subpicture consists of one and only one slice, and wherein the first flag is comprised in a sequence parameter set.