Method and apparatus for signaling picture segmentation information
By signaling picture partitioning information using flags to determine the number of slices in a current picture, the method addresses the need for efficient video coding, enhancing compression efficiency and reducing costs for high-resolution video data.
Patent Information
- Application Number
- JP2024118630
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-11-27
- Filing Date
- 2024-07-24
- Publication Date
- 2025-06-09
- Estimated Expiration
- 2040-11-26
AI Technical Summary
The increasing demand for high-resolution and high-quality videos has led to a need for more efficient video coding technologies to reduce transmission and storage costs.
A method and apparatus for signaling picture partitioning information in a video coding system, which includes using flags to determine the number of slices in a current picture, thereby improving video coding efficiency and decoding performance.
The proposed solution enhances video coding efficiency by optimizing picture partitioning and decoding processes, leading to improved compression efficiency and reduced costs for transmitting and storing high-resolution video data.
Smart Images

Figure 0007690095000007 
Figure 0007690095000008 
Figure 0007690095000009
Abstract
Description
Technical Field
[0001] The present disclosure relates to video coding technology, and to a method and apparatus for signaling picture partitioning information in a video coding system.
Background Art
[0002] Recently, the demand for high-resolution and high-quality videos such as HD (High Definition) videos and UHD (Ultra High Definition) videos has been increasing in various fields. As video data becomes higher in resolution and quality, the amount of information or bits transmitted relatively increases compared to existing video data. Therefore, when transmitting video data using a medium such as an existing wired or wireless broadband line, or storing video data using an existing storage medium, the transmission cost and storage cost increase.
[0003] Therefore, in order to effectively transmit, store, and reproduce high-resolution and high-quality video information, a highly efficient video compression technology is required.
Summary of the Invention
Problems to be Solved by the Invention
[0004] A technical problem of the present disclosure is to provide a method and apparatus for increasing video coding efficiency.
[0005] Another technical problem of the present disclosure is to provide a method and apparatus for signaling picture partitioning information.
[0006] Still another technical problem of the present disclosure is to provide a method and apparatus for performing decoding on a current picture based on partitioning information for the current picture.
Means for Solving the Problems
[0007] According to one embodiment of the present disclosure, a video decoding method executed by a decoding device is provided. The method includes obtaining video information for a current picture from a bitstream and performing decoding for the current picture based on the video information. The video information includes a first flag regarding the presence or absence of sub-picture information and a second flag regarding whether the sub-picture includes only one slice. Based on the first flag and the second flag, it is derived that the number of slices included in the current picture is one.
[0008] According to another embodiment of the present disclosure, a video encoding method executed by an encoding device is provided. The method includes dividing a current picture to derive at least one slice and encoding video information for the current picture based on the at least one slice. The video information includes a first flag regarding the presence or absence of sub-picture information and a second flag regarding whether the sub-picture includes only one slice. Based on the first flag and the second flag, it is derived that the number of slices included in the current picture is one.
[0009] According to still another embodiment of the present disclosure, a computer-readable digital storage medium storing encoded video information that causes a decoding device to execute a video decoding method is provided. The decoding method according to the one embodiment includes obtaining video information for a current picture from a bitstream and performing decoding for the current picture based on the video information. The video information includes a first flag regarding the presence or absence of sub-picture information and a second flag regarding whether the sub-picture includes only one slice. Based on the first flag and the second flag, it is derived that the number of slices included in the current picture is one.
Advantages of the Invention
[0010] According to this specification, the overall video / video compression efficiency can be improved.
[0011] According to this specification, the efficiency of picture partitioning can be improved.
[0012] According to this specification, based on the partitioning information for the current picture, the efficiency of picture partitioning can be improved.
Brief Description of the Drawings
[0013]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Best Mode for Carrying Out the Invention
[0014] This document can be modified in various ways and can have various embodiments. Specific embodiments will be illustrated in the drawings and described in detail. However, this is not intended to limit this document to specific embodiments. The terms commonly used in this specification are merely used to describe specific embodiments and are not used with the intention of limiting the technical idea of this document. Singular expressions include plural expressions unless the context clearly indicates otherwise. Terms such as "including" or "having" in this specification are intended to specify the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and it should be understood that the presence or addition possibility of one or more other features, numbers, steps, operations, components, parts, or combinations thereof, etc., is not precluded in advance.
[0015] On the other hand, each configuration in the drawings described in this document is independently illustrated for the convenience of explaining different characteristic functions, and it does not mean that each configuration is implemented by separate hardware or separate software. For example, among the configurations, two or more configurations can be combined to form one configuration, and one configuration can also be divided into multiple configurations. Embodiments in which each configuration is integrated and / or separated are included in the scope of rights of this document as long as they do not deviate from the essence of this document.
[0016] In this specification, "A or B" can mean "only A", "only B", or "both A and B". In other words, in this document, "A or B" can be interpreted as "A and / or B". For example, in this specification, "A, B or C" can mean "only A", "only B", "only C", or "any combination of A, B and C".
[0017] The slashes ( / ) and commas used in this document can mean "and / or". For example, "A / B" can mean "A and / or B". Thus, "A / B" can mean "only A", "only B", or "both A and B". For example, "A, B, C" can mean "A, B or C".
[0018] In this specification, "at least one of A and B" can mean "only A", "only B", or "both A and B". Also, in this specification, expressions such as "at least one of A or B" and "at least one of A and / or B" can be interpreted in the same way as "at least one of A and B".
[0019] Also, in this specification, "at least one of A, B, and C" can mean "only A", "only B", "only C", or "any combination of A, B, and C". Also, "at least one of A, B, or C" and "at least one of A, B, and / or C" can mean "at least one of A, B, and C".
[0020] Also, the parentheses used in this specification can mean "for example". Specifically, when expressed as "prediction (intra prediction)", "intra prediction" is proposed as an example of "prediction". As another expression, "prediction" in this specification is not limited to "intra prediction", but "intra prediction" is proposed as an example of "prediction". Also, when expressed as "prediction (i.e., intra prediction)", "intra prediction" is proposed as an example of "prediction".
[0021] In this specification, the technical features individually described within one drawing can be embodied individually or simultaneously.
[0022] Hereinafter, with reference to the accompanying drawings, the preferred embodiments of the present disclosure will be described in more detail. Hereinafter, the same reference numerals will be used for the same components in the drawings, and redundant descriptions for the same components can be omitted.
[0023] FIG. 1 schematically shows an example of a video / video coding system to which the present disclosure is applicable.
[0024] Referring to FIG. 1, a video / image coding system can include a first device (source device) and a second device (receiving device). The source device can transmit encoded video / image information or data in file or streaming form through a digital storage medium or network to the receiving device.
[0025] The source device can include a video source, an encoding device, and a transmitting unit. The receiving device can include a receiving unit, a decoding device, and a renderer. The encoding device can be called a video / image encoding device, and the decoding device can be called a video / image decoding device. A transmitter can be included in the encoding device. A receiver can be included in the decoding device. The renderer can also include a display unit, and the display unit can also be composed of a separate device or an external component.
[0026] The video source can obtain video / image through processes such as video / image capture, synthesis, or generation. The video source can include a video / image capture device and / or a video / image generation device. The video / image capture device can include, for example, one or more cameras, a video / image archive including previously captured video / image, etc. The video / image generation device can include, for example, a computer, a tablet, and a smartphone, etc., and can (electronically) generate video / image. For example, virtual video / image can be generated through a computer, etc., in which case the video / image capture process can be replaced by a process in which related data is generated.
[0027] The encoding device can encode the input video / image. The encoding device can execute a series of procedures such as prediction, transformation, quantization, etc. for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0028] The transmitting unit can transmit the encoded video / image information or data output in the form of a bitstream to the receiving unit of the receiving device via a digital storage medium or network in file or streaming form. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit can include elements for generating a media file via a predetermined file format and can include elements for transmission via a broadcast / communication network. The receiving unit can receive / extract the bitstream and transmit it to the decoding device.
[0029] The decoding device can decode the video / image by executing a series of procedures such as inverse quantization, inverse transformation, prediction, etc. corresponding to the operation of the encoding device.
[0030] The renderer can render the decoded video / image. The rendered video / image can be displayed via the display unit.
[0031] This document relates to video / video coding. For example, the methods / embodiments disclosed in this document can be applied to the methods disclosed in the VVC (Versatile Video Coding) standard, EVC (essential video coding) standard, AV1 (AOMedia Video 1) standard, AVS2 (2nd generation of audio video coding standard) or next-generation video / video coding standards (e.g., H.267 or H.267).
[0032] This document presents various embodiments related to video / video coding, and unless otherwise mentioned, the embodiments can also be executed in combination with each other.
[0033] In this document, video can mean a collection of a series of images over time. A picture generally means a unit representing one image at a specific time period, and a slice / tile is a unit that constitutes a part of a picture in coding. A slice / tile can include one or more CTUs (coding tree units). One picture can be composed of one or more slices / tiles.
[0034] A tile is a rectangular region of CTUs within a particular tile column and particular tile row in a picture. The tile column is a rectangular region of CTUs having a height equal to the height of the picture and a width specified by syntax elements in the picture parameter set. The tile row is a rectangular region of CTUs having a height specified by syntax elements in the picture parameter set and a width equal to the width of the picture.A tile scan can indicate a specific sequential ordering of CTUs partitioning a picture, where the CTUs can be ordered consecutively in a CTU raster scan within a tile, and tiles in a picture can be ordered consecutively in a raster scan of the tiles of the picture. A slice can contain multiple consecutive CTU rows within one tile of a picture that can be included in a number of complete tiles or one NAL unit. In this document, tile groups and slices can be used interchangeably. For example, in this document, a tile group / tile group header can be referred to as a slice / slice header.
[0035] On the other hand, one picture can be divided into two or more sub-pictures. A sub-picture can be an rectangular region of one or more slices within a picture.
[0036] A pixel or pel can mean the smallest unit that constitutes one picture (or video). Also, the term "sample" can be used as a term corresponding to a pixel. A sample can generally indicate a pixel or a pixel value, and can also indicate only the pixel / pixel value of the luma component or only the pixel / pixel value of the chroma component.
[0037] A unit can represent the basic unit of video processing. A unit can include at least one of a specific area of a picture and information related to that area. One unit can include one luma block and two chroma (e.g., cb, cr) blocks. A unit can, in some cases, be used interchangeably with terms such as block or area. In general, an M×N block can include a sample (or, sample array) consisting of M columns and N rows, or a set (or, array) of transform coefficients.
[0038] Figure 2 is a diagram schematically explaining the configuration of a video / video encoding device to which this document can be applied. Hereinafter, the video encoding device can include a video encoding device.
[0039] As shown in FIG. 2, the encoding apparatus 200 can be configured to include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 can include an inter-predictor 221 and an intra-predictor 222. The residual processor 230 can include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 can further include a subtractor 231. The adder 250 can be called a reconstructor or a reconstructed block generator. The above-described image partitioner 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 can be configured by one or more hardware components (e.g., an encoder chipset or a processor) according to an embodiment. Also, the memory 270 can include a DPB (decoded picture buffer) and can also be configured by a digital storage medium. The hardware component can further include the memory 270 as an internal / external component.
[0040] The image segmentation unit 210 can divide an input image (or picture, frame) input to the encoding device 200 into one or more processing units. As an example, the processing unit can be called a coding unit (CU). In this case, the coding unit can be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) by a QTBTTT (Quad-tree binary-tree ternary-tree) structure. For example, one coding unit can be divided into multiple coding units with a deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure can be applied first, and then the binary-tree structure and / or the ternary structure can be applied. Or, the binary-tree structure can also be applied first. The coding procedure according to the present disclosure can be performed based on the final coding unit that is no longer divided. In this case, based on coding efficiency according to image characteristics, etc., the largest coding unit can be used as the final coding unit, or, if necessary, the coding unit can be recursively divided into coding units with a deeper depth so that the coding unit with the optimal size can be used as the final coding unit. Here, the coding procedure can include procedures such as prediction, transformation, and restoration described later. As another example, the processing unit can further include a prediction unit (PU: Prediction Unit) or a transform unit (TU: Transform Unit). In this case, the prediction unit and the transform unit can each be divided or partitioned from the final coding unit described above.The prediction unit can be a unit of sample prediction, and the conversion unit can be a unit for deriving a conversion coefficient and / or a unit for deriving a residual signal from the conversion coefficient.
[0041] The term "unit" can, in some cases, be used interchangeably with terms such as "block" or "area". In general, an M×N block can represent a set of samples or transform coefficients, etc., consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luma component, or only the pixel / pixel value of the chroma component. A sample can be used as a term corresponding to a pixel or a pel for one picture (or image).
[0042] The encoding device 200 can subtract a prediction signal (predicted block, predicted sample array) output from the inter prediction unit 221 or the intra prediction unit 222 from an input video signal (original block, original sample array) to generate a residual signal (residual signal, residual block, residual sample array), and the generated residual signal is transmitted to the conversion unit 232. In this case, as shown in the figure, the unit that subtracts the prediction signal (predicted block, predicted sample array) from the input video signal (original block, original sample array) within the encoding device 200 can be called the subtraction unit 231. The prediction unit can perform prediction on a block to be processed (hereinafter referred to as the current block) and generate a predicted block including predicted samples for the current block. The prediction unit can determine whether intra prediction or inter prediction is to be applied in units of the current block or CU. The prediction unit can generate various pieces of information related to prediction, such as prediction mode information, as described later, in the description of each prediction mode, and transmit them to the entropy encoding unit 240. The information related to prediction can be encoded by the entropy encoding unit 240 and output in the form of a bit stream.
[0043] The intra prediction unit 222 can predict the current block by referring to samples within the current picture. The samples to be referred to can be located adjacent to the current block or remotely located depending on the prediction mode. In intra prediction, the prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes can include, for example, the DC mode and the Planar mode. The directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the level of detail of the prediction direction. However, this is an example, and more or fewer directional prediction modes can be used depending on the setting. The intra prediction unit 222 can also determine the prediction mode to be applied to the current block using the prediction mode applied to the adjacent block.
[0044] The inter prediction unit 221 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between adjacent blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the adjacent blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block can be the same or different. The temporal neighboring blocks can be called by names such as collocated reference blocks and collocated CUs (col CUs), and the reference picture including the temporal neighboring blocks can also be called a collocated picture (colPic). For example, the inter prediction unit 221 can construct a motion information candidate list based on adjacent blocks, and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction can be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the inter prediction unit 221 can use the motion information of adjacent blocks as the motion information of the current block. In the case of skip mode, unlike merge mode, a residual signal may not be transmitted.In the case of the motion information prediction (motion vector prediction, MVP) mode, the motion vector of an adjacent block is used as a motion vector predictor, and by signaling the motion vector difference, the motion vector of the current block can be indicated.
[0045] The prediction unit 220 can generate a prediction signal based on various prediction methods described later. For example, the prediction unit can apply intra prediction or inter prediction for predicting a block, and can also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). In addition, the prediction unit can also be based on the intra block copy (IBC) prediction mode or the palette mode for predicting a block. The IBC prediction mode or the palette mode can be used for content video / motion video coding such as games, for example, like SCC (screen content coding). IBC basically performs prediction within the current picture, but can be executed in a manner similar to inter prediction in terms of deriving a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this document. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, the sample values in the picture can be signaled based on information regarding the palette table and the palette index.
[0046] The prediction signal generated via the prediction unit (including the inter prediction unit 221 and / or the intra prediction unit 222) can be used to generate a restored signal or can be used to generate a residual signal. The conversion unit 232 can generate transform coefficients by applying a conversion technique to the residual signal. For example, the conversion technique can include at least one of DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), KLT (Karhunen-Loeve Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT means the conversion obtained from this graph when representing the relationship information between pixels with a graph. CNT means the conversion obtained based on generating a prediction signal using all previously reconstructed pixels. Also, the conversion process can be applied to pixel blocks having the same size of a square and can also be applied to blocks of variable size that are not square.
[0047] The quantization unit 233 quantizes the transform coefficients and transmits them to the entropy encoding unit 240. The entropy encoding unit 240 can encode the quantized signal (information regarding the quantized transform coefficients) and output it as a bitstream. The information regarding the quantized transform coefficients may be referred to as residual information. The quantization unit 233 can reorder the quantized transform coefficients in block form into a one-dimensional vector form based on the coefficient scan order, and can also generate the information regarding the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. The entropy encoding unit 240 can execute various encoding methods such as, for example, exponential Golomb, CAVLC (context-adaptive variable length coding), CABAC (context-adaptive binary arithmetic coding), etc. The entropy encoding unit 240 can encode, together or separately, information necessary for video / image restoration (e.g., values of syntax elements, etc.) in addition to the quantized transform coefficients. The encoded information (e.g., encoded video / video information) can be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer) units. The video / video information can further include information regarding various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS), etc. Also, the video / video information can further include general constraint information. In this document, the information and / or syntax elements transmitted / signaled from the encoding device to the decoding device can be included in the video / video information.The video / video information can be encoded through the encoding procedure described above and included in the bitstream. The bitstream can be transmitted via a network or stored in a digital storage medium. Here, the network can include a broadcast network and / or a communication network, etc., and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The signal output from the entropy encoding unit 240 can be configured as an internal / external element of the encoding device 200 by a transmission unit (not shown) for transmission and / or a storage unit (not shown) for storage, or the transmission unit can also be included in the entropy encoding unit 240.
[0048] The quantized transform coefficients output from the quantization unit 233 can be used to generate a prediction signal. For example, a residual signal (residual block or residual sample) can be restored by applying inverse quantization and inverse transformation to the quantized transform coefficients via the inverse quantization unit 234 and the inverse transform unit 235. The addition unit 155 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the restored residual signal to the prediction signal output from the inter prediction unit 221 or the intra prediction unit 222. When there is no residual for the block to be processed, as in the case where the skip mode is applied, the predicted block can be used as the reconstructed block. The addition unit 250 can be called a restoration unit or a reconstructed block generation unit. The generated reconstructed signal can be used for intra prediction of the next block to be processed within the current picture and, as will be described later, can also be used for inter prediction of the next picture after passing through filtering.
[0049] On the other hand, LMCS (luma mapping with chrom ascaling) can also be applied in the picture encoding and / or restoration process.
[0050] The filtering unit 260 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 260 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be stored in the memory 270, specifically, in the DPB of the memory 270. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset (SAO), adaptive loop filter, bilateral filter, etc. As will be described later in the description of each filtering method, the filtering unit 260 can generate various information related to filtering and transmit it to the entropy encoding unit 290. The information related to filtering can be encoded by the entropy encoding unit 290 and output in the form of a bitstream.
[0051] The modified restored picture transmitted to the memory 270 can be used as a reference picture in the inter prediction unit 280. When inter prediction is applied through this, the encoding device can avoid prediction mismatches between the encoding device 200 and the decoding device, and can also improve the encoding efficiency.
[0052] The DPB of memory 270 can store the corrected restored picture for use as a reference picture in the inter prediction unit 221. Memory 270 can store the motion information of the blocks for which the motion information in the current picture has been derived (or encoded) and / or the motion information of the blocks in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 221 for utilization as the motion information of spatially adjacent blocks or temporally adjacent blocks. Memory 270 can store the restored samples of the restored blocks in the current picture and can transmit them to the intra prediction unit 222.
[0053] FIG. 3 is a diagram schematically explaining the configuration of a video / video decoding apparatus to which this document can be applied.
[0054] As shown in FIG. 3, the decoding apparatus 300 can be configured to include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filtering unit 350, and a memory 360. The predictor 330 can include an inter prediction unit 331 and an intra prediction unit 332. The residual processor 320 can include a dequantizer 321 and an inverse transformer 321. The above-described entropy decoder 310, residual processor 320, predictor 330, adder 340, and filtering unit 350 can be configured by one hardware component (for example, a decoder chipset or a processor) according to an embodiment. Also, the memory 360 can include a DPB (decoded picture buffer) and can also be configured by a digital storage medium. The hardware component can further include the memory 360 as an internal / external component.
[0055] If a bitstream including video / image information is input, the decoding device 300 can restore an image corresponding to the process in which the video / image information is processed by the encoding device in FIG. 3. For example, the decoding device 300 can derive units / blocks based on the block division related information obtained from the bitstream. The decoding device 300 can perform decoding using the processing units applied in the encoding device. Therefore, the processing unit for decoding can be, for example, a coding unit, and the coding unit can be divided according to a quad-tree structure, a binary tree structure, and / or a ternary tree structure from a coding tree unit or a maximum coding unit. One or more transform units can be derived from the coding unit. Then, the restored image signal decoded and output via the decoding device 300 can be reproduced via a reproducing device.
[0056] The decoding device 300 can receive the signal output from the encoding device in FIG. 3 in the form of a bitstream, and the received signal can be decoded via the entropy decoding unit 310. For example, the entropy decoding unit 310 can parse the bitstream to derive information (e.g., video / video information) necessary for video restoration (or picture restoration). The video / video information can further include information regarding various parameter sets such as an Adaptation Parameter Set (APS), a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), or a Video Parameter Set (VPS). Also, the video / video information can further include general constraint information. The decoding device can further decode a picture based on the information regarding the parameter set and / or the general constraint information. The signaling / received information and / or syntax elements described later in this document can be decoded via the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 310 can decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the value of the syntax element necessary for video restoration and the quantized value of the transform coefficient regarding the residual. More specifically, the CABAC entropy decoding method receives the bin corresponding to each syntax element in the bitstream, determines a context model using the syntax element information to be decoded, the information adjacent to the decoding target block, and the information of the symbol / bin decoded in the previous step or the decoding information of the decoding target block, predicts the occurrence probability of the bin based on the determined context model, and performs arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element.At this time, the CABAC entropy decoding method can update the context model by using the information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model. Among the information decoded by the entropy decoding unit 310, the information related to prediction is provided to the prediction unit (inter prediction unit 332 and intra prediction unit 331), and the residual value for which entropy decoding is executed by the entropy decoding unit 310, that is, the quantized transform coefficient and related parameter information can be input to the residual processing unit 320. The residual processing unit 320 can derive a residual signal (residual block, residual sample, residual sample array). Also, among the information decoded by the entropy decoding unit 310, the information related to filtering can be provided to the filtering unit 350. On the other hand, a receiving unit (not shown) that receives a signal output from the encoding device can be further configured as an internal / external element of the decoding device 300, or the receiving unit can also be a component of the entropy decoding unit 310. On the other hand, the decoding device according to this document can be called a video / video / picture decoding device, and the decoding device can also be classified into an information decoder (video / video / picture information decoder) and a sample decoder (video / video / picture sample decoder). The information decoder can include the entropy decoding unit 310, and the sample decoder can include at least one of the inverse quantization unit 321, inverse transform unit 322, addition unit 340, filtering unit 350, memory 360, inter prediction unit 332, and intra prediction unit 331.
[0057] In the inverse quantization unit 321, the quantized transform coefficients can be inverse quantized to output transform coefficients. The inverse quantization unit 321 can reorder the quantized transform coefficients in a two-dimensional block form. In this case, the reordering can be performed based on the coefficient scan order performed by the encoding device. The inverse quantization unit 321 can perform inverse quantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) to obtain transform coefficients.
[0058] In the inverse transform unit 322, the transform coefficients are inverse transformed to obtain a residual signal (residual block, residual sample array).
[0059] The prediction unit can perform prediction on the current block to generate a predicted block including predicted samples for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied to the current block based on the information regarding the prediction output from the entropy decoding unit 310, and can determine a specific intra / inter prediction mode.
[0060] The prediction unit 330 can generate a prediction signal based on various prediction methods described below. For example, the prediction unit can apply intra prediction or inter prediction for the prediction of one block, and can also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). Also, the prediction unit can be based on the intra block copy (IBC) prediction mode or the palette mode for the prediction of a block. The IBC prediction mode or the palette mode can be used for content video / moving picture coding such as games, for example, like SCC (screen content coding). IBC basically performs prediction within the current picture, but can be executed in a way similar to inter prediction in terms of deriving a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this document. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, information regarding the palette table and the palette index can be included in and signaled in the video / video information.
[0061] The intra prediction unit 331 can predict the current block by referring to samples within the current picture. The samples to be referred to can be located adjacent to or away from the current block depending on the prediction mode. The prediction modes in intra prediction can include a plurality of non-directional modes and a plurality of directional modes. The intra prediction unit 331 can also determine the prediction mode to be applied to the current block by using the prediction mode applied to the adjacent blocks.
[0062] The inter prediction unit 332 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between adjacent blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, adjacent blocks can include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit 332 can configure a motion information candidate list based on adjacent blocks, and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction can be performed based on various prediction modes, and information regarding the prediction can include information indicating a mode of inter prediction for the current block.
[0063] The addition unit 340 can generate a restored signal (restored picture, restored block, restored sample array) by adding the obtained residual signal to the predicted signal (predicted block, predicted sample array) output from the prediction unit 330. When there is no residual for the processing target block as in the case where the skip mode is applied, the predicted block can be used as the restored block.
[0064] The addition unit 340 can be called a restoration unit or a restored block generation unit. The generated restored signal can be used for intra prediction of the next processing target block within the current picture, and can also be output after filtering as described later, or can be used for inter prediction of the next picture.
[0065] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied in the picture decoding process.
[0066] The filtering unit 350 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 350 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be transmitted to the memory 60, specifically, the DPB of the memory 360. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0067] The (modified) restored picture stored in the DPB of the memory 360 can be used as a reference picture by the inter prediction unit 331. The memory 360 can store the motion information of the blocks for which the motion information in the current picture has been derived (or decoded) and / or the motion information of the blocks in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 331 for utilization as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 360 can store the restored samples of the restored blocks in the current picture and transmit them to the intra prediction unit 332.
[0068] In this specification, the embodiments described in the filtering unit 260, the inter prediction unit 221, and the intra prediction unit 222 of the encoding device 100 can be applied in the same or corresponding manner to the filtering unit 350, the inter prediction unit 332, and the intra prediction unit 331 of the decoding device 300, respectively.
[0069] As described above, prediction is performed to increase the compression efficiency when performing video coding. Thereby, a predicted block including prediction samples for a current block which is a block to be coded can be generated. Here, the predicted block includes prediction samples in the spatial domain (or pixel domain). The predicted block is derived in the same way in an encoding device and a decoding device, and the encoding device can increase the image coding efficiency by signaling information (residual information) regarding a residual between the original block and the predicted block, which is not the original sample value of the original block, to the decoding device. The decoding device can derive a residual block including residual samples based on the residual information, add the residual block and the predicted block to generate a restored block including restored samples, and generate a restored picture including the restored block.
[0070] The residual information can be generated through conversion and quantization procedures. For example, an encoding device can derive a residual block between the original block and the predicted block, execute a conversion procedure on the residual samples (residual sample array) included in the residual block to derive conversion coefficients, execute a quantization procedure on the conversion coefficients to derive quantized conversion coefficients, and thus signal the relevant residual information (via a bitstream) to a decoding device. Here, the residual information can include information such as the value information, position information, conversion technique, conversion kernel, quantization parameter, etc. of the quantized conversion coefficients. The decoding device can execute an inverse quantization / inverse conversion procedure based on the residual information to derive residual samples (or a residual block). The decoding device can generate a reconstructed picture based on the predicted block and the residual block. Also, the encoding device can inverse quantize / inverse convert the quantized conversion coefficients for reference in inter prediction of subsequent pictures to derive a residual block and generate a reconstructed picture based on this.
[0071] FIG. 4 exemplarily shows a hierarchical structure for coded data.
[0072] Referring to FIG. 4, the coded data can be divided into a NAL (Network abstraction layer) between a VCL (video coding layer) that handles video / image coding processing and itself, and a lower system that stores and transmits the coded video / image data.
[0073] VCL can generate parameter sets (such as Picture Parameter Set (PPS), Sequence Parameter Set (SPS), Video Parameter Set (VPS), etc.) corresponding to headers such as sequences and pictures, and Supplemental Enhancement Information (SEI) messages that are additionally required in the video / image coding process. The SEI messages are separated from the information (slice data) for the video / image. VCL that contains information for the video / image consists of slice data and slice headers. On the other hand, the slice header can be referred to as a tile group header, and the slice data can be referred to as tile group data.
[0074] In NAL, a NAL unit can be generated by adding header information (NAL unit header) to the Raw Byte Sequence Payload (RBSP) generated by VCL. At this time, RBSP refers to slice data, parameter sets, SEI messages, etc. generated by VCL. The NAL unit header can contain NAL unit type information specified by the RBSP data included in the NAL unit.
[0075] The NAL unit, which is the basic unit of NAL, plays a role in mapping the coded video to a bit sequence of a lower-level system such as a file format according to a predetermined standard, Real-time Transport Protocol (RTP), Transport Stream (TS), etc.
[0076] As shown in the figure, NAL units can be divided into VCL NAL units and Non-VCL NAL units by the RBSP generated in the VCL. A VCL NAL unit can mean a NAL unit that contains information (slice data) for video, and a Non-VCL NAL unit can mean a NAL unit that contains information (parameter set or SEI message) necessary for decoding the video.
[0077] As described above, the aforementioned VCL NAL units and Non-VCL NAL units can be transmitted via a network with header information attached according to the data standard of the lower system. For example, NAL units can be transformed into data forms of a predetermined standard such as the H.266 / VVC file format, RTP (Real-time Transport Protocol), TS (Transport Stream), etc., and transmitted via various networks.
[0078] As described above, the NAL unit type can be specified by the RBSP data structure (structure) included in the NAL unit, and information about such a NAL unit type can be stored in the NAL unit header and signaled.
[0079] For example, depending on whether the NAL unit contains information (slice data) for video, it can be roughly classified into a VCL NAL unit type and a Non-VCL NAL unit type. The VCL NAL unit type can be classified according to the nature and type of the picture included in the VCL NAL unit, etc., and the Non-VCL NAL unit type can be classified according to the type of the parameter set, etc.
[0080] The following is an example of a NAL unit type specified by the type of parameter set included in the Non-VCL NAL unit type, etc. The NAL unit type can be specified by the type of parameter set, etc. For example, the NAL unit type is the APS (Adaptation Parameter Set) NAL unit which is the type for the NAL unit including APS, the DPS (Decoding Parameter Set) NAL unit which is the type for the NAL unit including DPS, the VPS (Video Parameter Set) NAL unit which is the type for the NAL unit including VPS, the SPS (Sequence Parameter Set) NAL unit which is the type for the NAL unit including SPS, and the PPS (Picture Parameter Set) NAL unit which is the type for the NAL unit including PPS, and can be specified as any one of them.
[0081] The above-mentioned NAL unit type has syntax information for the NAL unit type, and the syntax information can be stored in the NAL unit header and signaled. For example, the syntax information can be nal_unit_type, and the NAL unit type can be specified by the nal_unit_type value.
[0082] On one hand, as described above, one picture can include a plurality of slices, and one slice can include a slice header and slice data. In this case, for the plurality of slices (slice header and slice data set) within one picture, one picture header can be further added. The said picture header (picture header syntax) can include information / parameters that can be commonly applied to the said picture. The said slice header (slice header syntax) can include information / parameters that can be commonly applied to the said slice. The said APS (APS syntax) or PPS (PPS syntax) can include information / parameters that can be commonly applied to one or more slices or pictures. The said SPS (SPS syntax) can include information / parameters that can be commonly applied to one or more sequences. The said VPS (VPS syntax) can include information / parameters that can be commonly applied to multi-layers. The said DPS (DPS syntax) can include information / parameters that can be commonly applied to the entire video. The said DPS can include information / parameters regarding the concatenation of CVS (coded video sequence). In this document, the high level syntax (HLS) refers to at least one of the said APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, a picture header syntax, and slice header syntax.
[0083] In this document, video / video information encoded by an encoding device and signaled in the form of a bitstream by a decoding device can include not only information related to intra-picture partitioning, intra / inter prediction information, residual information, in-loop filtering information, etc., but also information included in the slice header, information included in the picture header, information included in the APS, information included in the PPS, information included in the SPS, information included in the VPS, and / or information included in the DPS. Further, the video / video information can further include information of the NAL unit header.
[0084] FIG. 5 is a drawing showing an example of partitioning a picture.
[0085] A picture can be divided into coding tree units (CTUs), and a CTU can correspond to a coding tree block (CTB). A CTU can include a coding tree block of luma samples and two coding tree blocks of corresponding chroma samples. On the other hand, the maximum allowable size of a CTU for coding and prediction, etc., may be different from the maximum allowable size of a CTU for transformation.
[0086] A tile can correspond to a series of CTUs covering a rectangular area of a picture, and a picture can be divided into one or more tile rows and one or more tile columns.
[0087] On the other hand, a slice can be composed of an integer number of complete tiles or an integer number of consecutive complete CTU rows. At this time, two slice modes including a raster-scan slice mode and a rectangular slice mode can be supported.
[0088] In the raster scan slice mode, a slice can include a series of complete tiles in a tile raster scan of a picture. In the square slice mode, a slice can include a plurality of complete tiles that collectively form a square region of the picture. Alternatively, in the square slice mode, a slice can include a number of consecutive CTU rows within one tile that collectively form a square region of the picture. Tiles within a square slice can be scanned in tile raster scan order within the square region corresponding to the slice.
[0089] On the other hand, a subpicture can include one or more slices that cover a square region of a picture.
[0090] FIG. 5(a) is a drawing showing an example of dividing a picture into raster scan slices. For example, a picture can be divided into 12 tiles and 3 raster scan slices.
[0091] Also, FIG. 5(b) is a drawing showing an example of dividing a picture into square slices. For example, a picture can be divided into 24 tiles (6 tile columns and 4 tile rows) and 9 square slices.
[0092] Also, FIG. 5(c) is a drawing showing an example of dividing a picture into tiles and square slices. For example, a picture can be divided into 24 tiles (2 tile columns and 2 tile rows) and 4 square slices.
[0093] FIG. 6 is a flowchart showing a picture encoding procedure according to an embodiment.
[0094] In one embodiment, picture partitioning (S600) can be executed by the video division unit 210 of the encoding device, and picture encoding (S610) can be executed by the entropy encoding unit 240 of the encoding device.
[0095] An encoding device according to an embodiment can derive slices and / or tiles included in a current picture (S600). For example, the encoding device can perform picture partitioning for encoding the input current picture. For example, the encoding device can derive slices and / or tiles included in the current picture. The encoding device can partition the picture into various forms in consideration of the video characteristics and coding efficiency of the current picture, generate information indicating a partitioning form having optimal coding efficiency, and signal it to the decoding device.
[0096] An encoding device according to an embodiment can perform encoding on a current picture based on the derived slices and / or tiles (S610). For example, the encoding device can encode video / picture information including information related to slices and / or tiles and output it in the form of a bitstream. The output bitstream can be transmitted to the decoding device via a digital storage medium or a network.
[0097] FIG. 7 is a flowchart showing a picture decoding procedure according to an embodiment.
[0098] In one embodiment, the step of obtaining video / picture information from a bitstream (S710) and the step of deriving slices and / or tiles within a current picture (S720) can be executed by the entropy decoding unit 310 of the decoding device, and the step of restoring the current picture based on the slices and / or tiles can be executed by the addition unit 340 of the decoding device.
[0099] A decoding device according to an embodiment can acquire video / image information from a received bitstream (S710). The video / image information can include HLS, and the HLS can include information related to slices or information related to tiles. The information related to slices can include information identifying one or more slices in the current picture, and the information related to tiles can include information identifying one or more tiles in the current picture. The information related to slices or the information related to tiles can be acquired through various parameter sets, picture headers, and / or slice headers.
[0100] On the other hand, the current picture can include a tile including one or more slices or a slice including one or more tiles.
[0101] A decoding device according to an embodiment can derive slices and / or tiles within the current picture based on video / image information including information related to slices and / or tiles (S720).
[0102] A decoding device according to an embodiment can restore (decode) the current picture based on the slices and / or tiles (S730).
[0103] On the other hand, as described above, a picture can be divided into sub-pictures, tiles, and slices. Information related to sub-pictures can be signaled through SPS, and information related to tiles and rectangular slices can be signaled through PPS. Also, information related to raster-scan slices can be signaled through a slice header.
[0104] For example, the SPS syntax including information related to sub-pictures can be as follows in the following table.
[0105]
Table 1
[0106] For example, the PPS syntax including information on tiles and rectangular slices may be as shown in the following table.
[0107]
Table 2
[0108] Also, for example, the slice header syntax including information on raster-scan slices may be as shown in the following table.
[0109]
Table 3
[0110] On the other hand, the information on slices within the current picture and the information on tiles can include a flag regarding whether each sub-picture within the current picture contains a single slice. The flag may be referred to as single_slice_per_subpic_flag or pps_single_slice_per_subpic_flag, but is not limited thereto. Also, the information on sub-pictures can include a flag regarding the presence or absence of sub-picture information, and the flag may be referred to as subpics_present_flag or sps_subpic_info_present_flag, but is not limited thereto. For example, the information on sub-pictures can be included in a parameter set. For example, the information on sub-pictures can be included in the SPS.
[0111] Conventionally, when the value of the flag regarding the presence or absence of sub-picture information is 0, it was restricted such that the value of the flag regarding whether the sub-picture contains a single slice becomes 0. That is, when the value of the flag regarding the presence or absence of sub-picture information is 0, it was determined that the sub-picture is not available, and the value of the flag regarding whether the sub-picture contains a single slice was restricted to 0. However, such a condition is very restrictive. For example, even when there is no sub-picture information, the current picture can be divided into two or more tiles, and all the tiles can be included within one slice. In such a case, the current picture comes to include a single slice.
[0112] Therefore, one embodiment of this document proposes a solution to eliminate the restriction such that when the value of the flag regarding the presence or absence of sub-picture information is 0, the value of the flag regarding whether the sub-picture contains only one slice becomes 0. In such a case, the flag regarding whether the sub-picture contains a single slice can indicate the case where the current picture includes a single slice even when there is no sub-picture information.
[0113] For example, according to the above-described embodiment, even when there is no sub-picture information for a CLVS (coded layer video sequence), a flag regarding whether the sub-picture contains a single slice can exist. That is, even when there is no sub-picture information for a CLVS, the flag regarding whether the sub-picture contains a single slice can have a value of 0 or 1.
[0114] For example, when the value of the flag regarding the presence or absence of sub-picture information is 0 and the value of the flag regarding whether the sub-picture contains a single slice is 1, the current picture can include a single slice. That is, when there is no signaled sub-picture and the value of the flag regarding whether the sub-picture contains a single slice is 1, it can be inferred that the number of slices within the picture is 1.
[0115] Also, when the value of the flag regarding the presence or absence of sub-picture information is 0, the number of sub-pictures within the current picture can be 1. For example, when the value of the flag regarding the presence or absence of sub-picture information is 0, the number of sub-pictures existing in each of all pictures referring to the SPS of the video information can be 1.
[0116] On the other hand, the flag regarding the number of slices included in the current picture can be included in the PPS of the video information. The flag regarding the number of slices included in the current picture can be referred to as num_slices_in_pic_minus1 or pps_num_slices_in_pic_minus1, but is not limited thereto. Also, the flag regarding the number of sub-pictures included in the current picture can be included in the SPS of the video information. The flag regarding the number of sub-pictures included in the current picture can be referred to as sps_num_subpics_minus1, but is not limited thereto.
[0117] When there is no signaled sub-picture information and the value of the flag regarding whether the sub-picture includes a single slice is 1, it can be inferred that the flag regarding the number of slices included in the current picture has a value of 0. Also, when there is no signaled sub-picture information and the value of the flag regarding whether the sub-picture includes a single slice is 1, it can be inferred that the flag regarding the number of slices included in the current picture and the flag regarding the number of sub-pictures included in the current picture have the same value.
[0118] Also, when there is no signaled sub-picture and the value of the flag regarding whether the sub-picture includes a single slice is 1, all CTUs within the picture can belong to the single slice included in the picture.
[0119] The semantics for a syntax element that includes a flag regarding whether the sub - picture according to the foregoing embodiment includes a single slice and a flag regarding the number of slices included in the current picture can be shown as in the following table.
[0120]
Table 4
[0121] Referring to the said table, when the value of single_slice_per_subpic_flag corresponding to the flag regarding whether the sub - picture includes a single slice is 1, each sub - picture can be composed of one rectangular slice. Also, when the value of single_slice_per_subpic_flag is 0, each sub - picture can be composed of one or more rectangular slices. When the value of single_slice_per_subpic_flag is 1, it can be inferred that the value of num_slices_in_pic_minus1 corresponding to the flag regarding the number of slices included in the current picture has the same value as SPS_num_subpics_minus1 corresponding to the flag regarding the number of sub - pictures included in the current picture.
[0122] Also, when the value of single_slice_per_subpic_flag is 1 and the value of subpics_present_flag corresponding to the flag regarding the existence of sub - picture information is 0, the picture referring to the PPS can have one slice per picture.
[0123] On the other hand, the scanning process, which is the procedure for decoding the tiles within a picture, can be determined by the following table.
[0124]
Table 5 - 1
[0125]
Table 5-2
[0126] FIG. 8 is a flowchart showing the operation of an encoding apparatus according to an embodiment, and FIG. 9 is a block diagram showing the configuration of the encoding apparatus according to an embodiment.
[0127] The method disclosed in FIG. 8 can be executed by the encoding apparatus disclosed in FIG. 2 or FIG. 9. S810 in FIG. 9 can be executed by the video division unit 210 disclosed in FIG. 2, and S820 can be executed by the entropy encoding unit 240 disclosed in FIG. 2. Further, since the operations by S810 to S820 are based on a part of the content described above with reference to FIGS. 1 to 7, specific content overlapping with the content described above with reference to FIGS. 1 to 7 will be omitted or simplified in the description.
[0128] Referring to FIG. 8, an encoding apparatus according to an embodiment can divide a current picture to derive at least one slice (S810). For example, the video division unit 210 of the encoding apparatus can generate division information for the current picture based on the at least one slice.
[0129] An encoding apparatus according to an embodiment can encode video information for a current picture based on at least one slice (S810). The video information can include division information generated based on the at least one slice.
[0130] For example, the video information can include a first flag regarding the presence or absence of sub-picture information and a second flag regarding whether the sub-picture includes only one slice. For example, based on the first flag and the second flag, it can be derived that the number of slices included in the current picture is one.
[0131] For example, when the value of the first flag regarding the presence or absence of the sub-picture information is 0 and the second flag is 1, it can be derived that the number of slices included in the current picture is 1.
[0132] For example, when the value of the first flag regarding the presence or absence of the sub-picture information is 0, the number of sub-pictures present in the current picture can be 1.
[0133] For example, the first flag regarding the presence or absence of the sub-picture information can be included in the SPS (Sequence Parameter Set) of the video information.
[0134] For example, the second flag regarding whether the sub-picture includes only one slice can be included in the PPS (Picture Parameter Set) of the video information.
[0135] For example, the video information includes a third flag regarding the number of slices included in the current picture, and the third flag can be included in the PPS of the video information.
[0136] Also, for example, the video information includes a fourth flag regarding the number of sub-pictures included in the current picture, and the fourth flag can be included in the SPS of the video information.
[0137] Also, for example, when the value of the first flag is 0, it can be derived that the number of sub-pictures present in each of all the pictures referring to the SPS of the video information is 1.
[0138] On the one hand, the video information can include prediction information for the current picture. The prediction information can include information for an inter prediction mode or an intra prediction mode executed on the current picture. The encoding device can generate and encode the prediction information for the current picture.
[0139] On the other hand, the bitstream can be transmitted to the decoding device via a network or a (digital) storage medium. Here, the network can include a broadcast network and / or a communication network, etc., and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.
[0140] FIG. 10 is a flowchart showing the operation of a decoding device according to an embodiment, and FIG. 11 is a block diagram showing the configuration of a decoding device according to an embodiment.
[0141] The method disclosed in FIG. 10 can be executed by the decoding device disclosed in FIG. 3 or FIG. 11. Specifically, S1010 and S1020 can be executed by the entropy decoding unit 310 disclosed in FIG. 3. Also, the operations by S1010 and S1020 are based on a part of the content described above with reference to FIGS. 1 to 7. Therefore, specific content overlapping with the content described above with reference to FIGS. 1 to 7 will be omitted or simplified in the description.
[0142] The decoding apparatus according to one embodiment can acquire video information for the current picture from a bitstream (S1010). For example, the entropy decoding unit 310 of the decoding apparatus can acquire video information including split information for the current picture from the bitstream. For example, the split information can include slice information for the current picture. Further, the video information can include at least a part of prediction-related information or residual-related information. For example, the prediction-related information can include inter prediction mode information or inter prediction type information.
[0143] The decoding apparatus according to one embodiment can execute decoding for the current picture based on the video information (S1020). For example, the entropy decoding unit 310 of the decoding apparatus can derive the split structure of the current picture based on the slice information for the current picture.
[0144] For example, the video information can include a first flag regarding the presence or absence of sub-picture information and a second flag regarding whether the sub-picture includes only one slice. For example, based on the first flag and the second flag, the number of slices included in the current picture can be derived to be one.
[0145] For example, when the value of the first flag is 0 and the value of the second flag is 1, the number of slices included in the current picture can be derived to be one.
[0146] For example, when the value of the flag regarding the presence or absence of the sub-picture information is 0, the number of sub-pictures present in the current picture can be one.
[0147] For example, the first flag regarding the presence or absence of the sub-picture information can be included in the SPS (Sequence Parameter Set) of the video information.
[0148] For example, a second flag regarding whether the sub-picture includes a single slice can be included in the PPS (Picture Parameter Set) of the video information.
[0149] For example, the video information includes a third flag regarding the number of slices included in the current picture, and the third flag can be included in the PPS of the video information.
[0150] Also, for example, the video information includes a fourth flag regarding the number of sub-pictures included in the current picture, and the flag regarding the number of sub-pictures included in the current picture can be included in the SPS of the video information.
[0151] Also, for example, when the value of the first flag is 0, it can be derived that the number of sub-pictures existing in each of all the pictures referring to the SPS of the video information is 1.
[0152] In the foregoing embodiments, the method is described based on a flowchart in a series of steps or blocks, but the corresponding embodiments are not limited to the order of the steps. A certain step can occur in a different order or simultaneously with steps different from the foregoing. Also, those skilled in the art can understand that the steps shown in the flowchart are not exclusive, other steps are included, or one or more steps of the flowchart can be deleted without affecting the scope of the embodiments of this document.
[0153] The method according to the embodiments of this document described above can be embodied in software form, and the encoding device and / or decoding device according to this document can be included in a device that executes video processing, such as a TV, a computer, a smartphone, a set-top box, a display device, etc.
[0154] In this document, when an embodiment is implemented in software, the above-described method can be implemented by modules (processes, functions, etc.) that perform the above-described functions. The modules can be stored in a memory and executed by a processor. The memory can be inside or outside the processor and can be connected to the processor by various well-known means. The processor can include an ASIC (application-specific integrated circuit), other chip sets, logic circuits, and / or data processing devices. The memory can include a ROM (read-only memory), a RAM (random access memory), a flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described in this document can be implemented and executed on a processor, a microprocessor, a controller, or a chip. For example, the functional units shown in each drawing can be implemented and executed on a computer, a processor, a microprocessor, a controller, or a chip. In this case, information for implementation (e.g., information on instructions) or an algorithm can be stored in a digital storage medium.
[0155] In addition, the decoding device and encoding device to which the present disclosure is applied can be included in a multimedia broadcast transmission / reception device, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video conferencing device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camcorder, an on-demand video (VoD) service providing device, an OTT video (Over the top video) device, an Internet streaming service providing device, a three-dimensional (3D) video device, a VR (virtual reality) device, an AR (augmented reality) device, a videophone video device, a transportation means terminal (e.g., a vehicle (including an autonomous driving vehicle) terminal, an airplane terminal, a ship terminal, etc.), and a medical video device, etc., and can be used to process a video signal or a data signal. For example, as an OTT video (Over the top video) device, it can include a game console, a Blu-ray player, an Internet-connected TV, a home theater system, a smartphone, a tablet PC, a DVR (Digital Video Recorder), etc.
[0156] In addition, the processing method to which the present disclosure is applied can be produced in the form of a program executed by a computer and can be stored in a computer-readable recording medium. Also, multimedia data having a data structure according to the example(s) of this document can be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices in which data that can be read by a computer is stored. The computer-readable recording medium can include, for example, a Blu-ray Disc (BD), a Universal Serial Bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. Also, the computer-readable recording medium includes a medium embodied in the form of a carrier wave (for example, transmission via the Internet). Also, a bitstream generated by an encoding method can be stored in a computer-readable recording medium or can be transmitted via a wired or wireless communication network.
[0157] In addition, embodiments of the present disclosure can be embodied as a computer program product by program code, and the program code can be executed by a computer according to the example(s) of this document. The program code can be stored on a carrier readable by a computer.
[0158] FIG. 12 shows an example of a content streaming system to which the disclosure of this document is applicable.
[0159] Referring to FIG. 12, a content streaming system to which the present disclosure is applied can generally include an encoding server, a streaming server, a web server, a media repository, a user device, and a multimedia input device.
[0160] The encoding server compresses the content input from multimedia input devices such as smartphones, cameras, camcorders, etc. into digital data to generate a bitstream, and plays the role of sending this to the streaming server. As another example, when multimedia input devices such as smartphones, cameras, camcorders, etc. directly generate a bitstream, the encoding server can be omitted.
[0161] The bitstream can be generated by an encoding method or a bitstream generation method applied to the embodiments of this document, and the streaming server can temporarily store the bitstream in the process of sending or receiving the bitstream.
[0162] The streaming server sends multimedia data to the user device based on a user request via a web server, and the web server plays the role of a medium to inform the user of what services are available. When the user requests a desired service from the web server, the web server transmits this to the streaming server, and the streaming server sends multimedia data to the user. At this time, the content streaming system can include a separate control server. In this case, the control server plays the role of controlling commands / responses between each device within the content streaming system.
[0163] The streaming server can receive content from a media repository and / or an encoding server. For example, when it comes to receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.
[0164] Examples of the user device include a mobile phone, a smart phone, a laptop computer, a digital broadcast terminal, a PDA (personal digital assistants), a PMP (portable multimedia player), a navigation device, a slate PC, a tablet PC, an ultrabook, a wearable device (e.g., a smartwatch, a smart glass, an HMD (head mounted display)), a digital TV, a desktop computer, and a digital signage.
[0165] Each server in the content streaming system can be operated as a distributed server. In this case, the data received by each server can be distributedly processed.
[0166] The claims described in this specification can be combined in various ways. For example, the technical features of the method claims in this specification can be combined and implemented in a device, and the technical features of the device claims in this specification can be combined and implemented in a method. Also, the technical features of the method claims in this specification and the technical features of the device claims can be combined and implemented in a device, and the technical features of the method claims in this specification and the technical features of the device claims can be combined and implemented in a method.
Claims
1. 1. A method for decoding an image performed by a decoding device, the method comprising: obtaining video information relating to a current picture from a bitstream; and decoding the current picture based on the video information. the video information includes a first flag related to the presence of sub-picture information and a second flag related to whether each sub-picture includes a single slice; based on the value of the first flag being equal to 0 and the value of the second flag being equal to 1, the number of slices included in the current picture is derived to be equal to 1; A method according to claim 1, wherein the value of the second flag equal to one indicates that each sub-picture consists of a single slice.
2. 1. A method of encoding video performed by an encoding device, the method comprising: deriving at least one slice by dividing a current picture; encoding video information for the current picture based on the at least one slice; the video information includes a first flag related to the presence of sub-picture information and a second flag related to whether each sub-picture in the current picture includes a single slice; based on the value of the first flag being equal to 0 and the value of the second flag being equal to 1, the number of slices included in the current picture is derived to be equal to 1; A method according to claim 1, wherein the value of the second flag equal to one indicates that each sub-picture consists of a single slice.
3. A method for transmitting data for a video, the method comprising: Obtaining a bitstream generated by a method, the method comprising: deriving at least one slice by dividing a current picture; and generating the bitstream by encoding video information for the current picture based on the at least one slice; transmitting the data including the bitstream; the video information includes a first flag related to the presence of sub-picture information and a second flag related to whether each sub-picture in the current picture includes a single slice; based on the value of the first flag being equal to 0 and the value of the second flag being equal to 1, the number of slices included in the current picture is derived to be equal to 1; A method of transmitting, wherein the value of the second flag equal to 1 indicates that each sub-picture consists of a single slice.