Method and apparatus for signaling picture segmentation information
By signaling picture partitioning information with flags to identify sub-pictures and slices, the method addresses the inefficiencies in high-resolution video compression, enhancing both video coding efficiency and picture partitioning.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- LG ELECTRONICS INC
- Filing Date
- 2026-02-19
- Publication Date
- 2026-05-13
AI Technical Summary
The increasing demand for high-resolution and high-quality videos has led to a surge in video data transmission and storage costs due to the increased amount of information, necessitating a more efficient video compression technology.
A method and apparatus for signaling picture partitioning information, including flags to determine the presence of sub-pictures and slices, allowing for improved video coding efficiency through precise picture partitioning and decoding.
This approach enhances overall video compression efficiency and improves the efficiency of picture partitioning by optimizing the division of video data into slices and sub-pictures.
Smart Images

Figure 2026077844000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to video coding technology, and relates to a method and an apparatus for signaling picture partitioning information in a video coding system.
Background Art
[0002] Recently, the demand for high-resolution and high-quality videos such as HD (High Definition) videos and UHD (Ultra High Definition) videos has been increasing in various fields. As video data becomes higher in resolution and quality, the amount of information or bits to be transmitted relatively increases compared to existing video data. Therefore, when transmitting video data using a medium such as an existing wired or wireless broadband line or storing video data using an existing storage medium, the transmission cost and the storage cost increase.
[0003] Therefore, in order to effectively transmit, store, and reproduce high-resolution and high-quality video information, a highly efficient video compression technology is required.
Summary of the Invention
Problems to be Solved by the Invention
[0004] A technical problem of the present disclosure is to provide a method and an apparatus for increasing video coding efficiency.
[0005] Another technical problem of the present disclosure is to provide a method and an apparatus for signaling picture partitioning information.
[0006] Still another technical problem of the present disclosure is to provide a method and an apparatus for performing decoding on a current picture based on partitioning information for the current picture.
Means for Solving the Problems
[0007] According to one embodiment of the present disclosure, a video decoding method performed by a decoding device is provided. The method includes the steps of obtaining video information for a current picture from a bitstream and performing decoding on the current picture based on the video information, wherein the video information includes a first flag indicating whether or not sub-picture information exists, and a second flag indicating whether or not the sub-picture contains only one slice, and based on the first and second flags, it is derived that the number of slices contained in the current picture is one.
[0008] In another embodiment of the present disclosure, a video encoding method is provided which is performed by an encoding device. The method includes the steps of dividing a current picture to derive at least one slice and encoding video information for the current picture based on the at least one slice, wherein the video information includes a first flag indicating the presence or absence of sub-picture information and a second flag indicating whether the sub-picture contains only one slice, and based on the first and second flags, it is derived that the number of slices contained in the current picture is one.
[0009] In another embodiment of the present disclosure, a computer-readable digital storage medium is provided for storing encoded video information that causes a decoding device to perform a video decoding method. The decoding method according to the present embodiment includes the steps of obtaining video information for a current picture from a bitstream and performing decoding on the current picture based on the video information, wherein the video information includes a first flag indicating the presence or absence of subpicture information and a second flag indicating whether the subpicture contains only one slice, and based on the first and second flags, it is derived that the number of slices contained in the current picture is one. [Effects of the Invention]
[0010] This specification allows for an improvement in overall video compression efficiency.
[0011] This specification can improve the efficiency of picture partitioning.
[0012] According to this specification, the efficiency of picture partitioning can be improved based on the division information for the picture. [Brief explanation of the drawing]
[0013] [Figure 1] An example of a video / image coding system to which this disclosure applies is schematically shown below. [Figure 2] This diagram schematically illustrates the configuration of a video / image encoding device to which this disclosure applies. [Figure 3] This diagram schematically illustrates the configuration of a video / image decoding device to which this disclosure applies. [Figure 4] This illustrates a hierarchical structure for coded data. [Figure 5] This is a diagram illustrating an example of partitioning a picture. [Figure 6] This is a flowchart illustrating a picture encoding procedure according to one embodiment. [Figure 7] This is a flowchart illustrating a picture decoding procedure according to one embodiment. [Figure 8] This is a flowchart illustrating the operation of an encoding device according to one embodiment. [Figure 9] This block diagram shows the configuration of an encoding device according to one embodiment. [Figure 10] This is a flowchart illustrating the operation of a decoding device according to one embodiment. [Figure 11] This is a block diagram showing the configuration of a decoding device according to one embodiment. [Figure 12] This document provides examples of content streaming systems to which the disclosures described herein apply. [Modes for carrying out the invention]
[0014] This document may be modified in various ways and may have various embodiments, and specific embodiments are illustrated in the drawings and described in detail. However, this is not intended to limit this document to any specific embodiment. Terms used herein are used solely to describe specific embodiments and are not intended to limit the technical ideas of this document. Singular expressions include plural expressions unless the context clearly indicates otherwise. Terms such as “includes” or “has” herein are intended to specify the existence of features, figures, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood not to preemptively exclude the existence or possibility of adding one or more other features, figures, steps, actions, components, parts, or combinations thereof.
[0015] On the other hand, each configuration shown in the diagrams described in this document is illustrated independently for the purpose of explaining its distinct characteristic functions, and does not mean that each configuration is embodied in separate hardware or separate software. For example, two or more configurations may be combined to form a single configuration, and a single configuration may be divided into multiple configurations. Embodiments in which each configuration is integrated and / or separated are also included within the scope of the rights of this document, as long as they do not deviate from the essence of this document.
[0016] In this specification, "A or B" can mean "only A", "only B", or "both A and B". In other words, in this document, "A or B" can be interpreted as "A and / or B". For example, in this specification, "A, B or C" can mean "only A", "only B", "only C", or "any combination of A, B and C".
[0017] The slashes ( / ) and commas used in this document can mean "and / or". For example, "A / B" can mean "A and / or B". Thus, "A / B" can mean "only A", "only B", or "both A and B". For example, "A, B, C" can mean "A, B or C".
[0018] In this specification, "at least one of A and B" can mean "only A", "only B", or "both A and B". Also, in this specification, expressions such as "at least one of A or B" and "at least one of A and / or B" can be interpreted in the same way as "at least one of A and B".
[0019] Also, in this specification, "at least one of A, B, and C" can mean "only A", "only B", "only C", or "any combination of A, B, and C". Also, "at least one of A, B, or C" and "at least one of A, B, and / or C" can mean "at least one of A, B, and C".
[0020] Also, the parentheses used in this specification can mean "for example". Specifically, when expressed as "prediction (intra prediction)", "intra prediction" is proposed as an example of "prediction". As another expression, the "prediction" in this specification is not limited to "intra prediction", but "intra prediction" is proposed as an example of "prediction". Also, when expressed as "prediction (that is, intra prediction)", "intra prediction" is proposed as an example of "prediction".
[0021] In this specification, the technical features separately described within one drawing can be embodied separately or simultaneously.
[0022] Hereinafter, with reference to the accompanying drawings, preferred embodiments of the present disclosure will be described in more detail. Hereinafter, the same reference numerals will be used for the same components in the drawings, and redundant descriptions for the same components can be omitted.
[0023] FIG. 1 schematically shows an example of a video / video coding system to which the present disclosure is applicable.
[0024] Referring to Figure 1, a video / image coding system may include a first device (source device) and a second device (receiving device). The source device can transmit encoded video / image information or data to the receiving device in file or streaming form via a digital storage medium or network.
[0025] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be called a video / image encoding device, and the decoding device may be called a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, which may consist of a separate device or external component.
[0026] A video source can acquire video / images through processes such as video / image capture, synthesis, or generation. A video source may include video / image capture devices and / or video / image generation devices. Video / image capture devices may include, for example, one or more cameras, or video / image archives containing previously captured video / images. Video / image generation devices may include, for example, computers, tablets, and smartphones, and can generate video / images (electronically). For example, virtual video / images may be generated via a computer, in which case the video / image capture process can be replaced by the process of generating the relevant data.
[0027] An encoding device can encode input video / image data. For compression and coding efficiency, the encoding device can perform a series of steps, including prediction, transformation, and quantization. The encoded data (encoded video / image information) can be output in bitstream format.
[0028] The transmitting unit can transmit encoded video / image information or data output in bitstream format to the receiving unit of a receiving device via a digital storage medium or network in file or streaming format. The digital storage medium can include a variety of storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The transmitting unit may include elements for generating media files via a predetermined file format and may include elements for transmission via a broadcast / communication network. The receiving unit can receive / extract the bitstream and transmit it to a decoding device.
[0029] A decoding device can decode video / images by performing a series of steps, such as inverse quantization, inverse transformation, and prediction, corresponding to the operation of an encoding device.
[0030] The renderer can render the decoded video / image. The rendered video / image can be displayed via the display unit.
[0031] This document relates to video / image coding. For example, the methods / examples disclosed herein can be applied to methods disclosed in the VVC (Versatile Video Coding) standard, EVC (essential video coding) standard, AV1 (AOMedia Video 1) standard, AVS2 (2nd generation of audio video coding standard), or next-generation video / image coding standards (e.g., H.267 or H.267).
[0032] This document presents various examples of video / image coding, and unless otherwise noted, these examples can be combined and implemented in conjunction with each other.
[0033] In this document, "video" can mean a collection of images over time. "Picture" generally refers to a single image representing a specific time period, while "slice" or "tile" is a unit that constitutes part of a picture in coding. A slice or tile can contain one or more coding tree units (CTUs). A single picture can consist of one or more slices or tiles.
[0034] A tile is a rectangular region of CTUs within a particular tile column and particular tile row in a picture. The tile column is a rectangular region of CTUs having a height equal to the height of the picture and a width specified by syntax elements in the picture parameter set. The tile row is a rectangular region of CTUs having a height specified by syntax elements in the picture parameter set and a width equal to the width of the picture.A tile scan can represent a specific sequential ordering of CTUs partitioning a picture in which the CTUs are ordered consecutively in a CTU raster scan in a tile, whereas tiles in a picture are ordered consecutively in a raster scan of the tiles of the picture. A slice can contain multiple consecutive rows of CTUs within a single tile of a picture, which may be contained in a number of complete tiles or a single NAL unit. In this document, tile groups and slices may be used interchangeably. For example, in this document, a tile group / tile group header may be referred to as a slice / slice header.
[0035] On the other hand, a single picture can be divided into two or more subpictures. A subpicture can be a rectangular region of one or more slices within a picture.
[0036] A pixel or pel can refer to the smallest unit that makes up a picture (or image). Alternatively, the term "sample" can be used as a counterpart to "pixel." A sample can generally represent a pixel or a pixel value, and can represent only the luma component pixel / pixel value, or only the chroma component pixel / pixel value.
[0037] A unit can represent a basic unit of image processing. A unit can contain at least one of a specific region of a picture and information associated with that region. A unit can contain one luma block and two chroma (e.g., cb, cr) blocks. The term unit may be used interchangeably with terms such as block or area. In general, an M×N block can contain a sample (or sample array) consisting of M columns and N rows, or a set (or array) of transform coefficients.
[0038] Figure 2 is a schematic diagram illustrating the configuration of a video / image encoding device to which this document can be applied. Hereinafter, the term "video encoding device" may include image encoding devices.
[0039] As shown in Figure 2, the encoding device 200 can be configured to include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-prediction unit 221 and an intra-prediction unit 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be called a reconstructor or a reconstructed block generator. The image segmentation unit 210, prediction unit 220, residual processing unit 230, entropy encoding unit 240, addition unit 250, and filtering unit 260 described above can be configured by one or more hardware components (e.g., an encoder chipset or processor) depending on the embodiment. The memory 270 may also include a DPB (decoded picture buffer) and may be configured by a digital storage medium. The hardware components may further include the memory 270 as an internal / external component.
[0040] The image splitting unit 210 can split an input image (or picture, frame) input to the encoding device 200 into one or more processing units. For example, the processing units may be called coding units (CUs). In this case, the coding units can be recursively split from a coding tree unit (CTU) or the largest coding unit (LCU) using a QTBTTT (Quad-tree binary-tree ternary-tree) structure. For example, one coding unit can be split into multiple coding units of deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first, followed by the binary-tree structure and / or the ternary structure. Alternatively, the binary-tree structure may be applied first. The coding procedure according to this disclosure may be performed based on the final coding unit that is not further split. In this case, based on coding efficiency due to image characteristics, the largest coding unit can be used as the final coding unit, or, if necessary, the coding unit can be recursively divided into lower-depth coding units so that the optimally sized coding unit is used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration, which will be described later. As another example, the processing unit may further comprise a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit can each be separated or partitioned from the final coding unit described above.The prediction unit may be a unit of sample prediction, and the conversion unit may be a unit for deriving conversion coefficients and / or a unit for deriving a residual signal from conversion coefficients.
[0041] The term "unit" can sometimes be used interchangeably with terms such as "block" or "area." Generally, an M×N block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and may represent only the luminance (luma) component pixel / pixel value, or only the chroma component pixel / pixel value. A sample can be used as the term corresponding to a single picture (or image) pixel or pel.
[0042] The encoding device 200 can generate a residual signal (residual block, residual sample array) by subtracting the predicted signal (predicted block, predicted sample array) output from the inter-prediction unit 221 or intra-prediction unit 222 from the input video signal (original block, original sample array), and the generated residual signal is transmitted to the conversion unit 232. In this case, as shown in the figure, the unit that subtracts the predicted signal (predicted block, predicted sample array) from the input video signal (original block, original sample array) within the encoding device 200 may be called the subtraction unit 231. The prediction unit can make predictions for the block to be processed (hereinafter referred to as the current block) and generate a predicted block that includes predicted samples for the current block. The prediction unit can decide whether intra-prediction or inter-prediction is applied on a current block or CU basis. As will be described later in the explanation for each prediction mode, the prediction unit can generate various information related to prediction, such as prediction mode information, and transmit it to the entropy encoding unit 240. Prediction information can be encoded by the entropy encoding unit 240 and output in bitstream format.
[0043] The intra-prediction unit 222 can predict the current block by referring to a sample in the current picture. The referenced sample may be located adjacent to the current block or at a distance, depending on the prediction mode. The prediction mode in intra-prediction may include multiple non-directional modes and multiple directional modes. Non-directional modes may include, for example, DC mode and planar mode. Directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes, depending on the degree of fineness of the prediction direction. However, this is illustrative, and more or fewer directional prediction modes may be used depending on the settings. The intra-prediction unit 222 may also determine the prediction mode to apply to the current block using the prediction modes applied to adjacent blocks.
[0044] The interprediction unit 221 can derive a predicted block relative to the current block based on a reference block (reference sample array) identified by motion vectors on the reference picture. In this case, in order to reduce the amount of motion information transmitted in interprediction mode, motion information can be predicted in units of blocks, subblocks, or samples based on the correlation of motion information between adjacent blocks and the current block. The motion information may include motion vectors and reference picture indices. The motion information may further include interprediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of interprediction, adjacent blocks may include spatially adjacent blocks that exist in the current picture and temporally adjacent blocks that exist in the reference picture. The reference picture containing the reference block and the reference picture containing the temporally adjacent block may be the same or different. The temporally adjacent block may be called a collocated reference block, col CU, etc., and the reference picture containing the temporally adjacent block may be called a collocated picture (colPic). For example, the inter-prediction unit 221 can construct a motion information candidate list based on adjacent blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter-prediction can be performed based on various prediction modes, and for example, in skip mode and merge mode, the inter-prediction unit 221 can use the motion information of adjacent blocks as the motion information of the current block. In skip mode, unlike merge mode, a residual signal may not be transmitted.In motion vector prediction (MVP) mode, the motion vector of an adjacent block is used as a motion vector predictor, and the motion vector difference is signaled to indicate the motion vector of the current block.
[0045] The prediction unit 220 can generate prediction signals based on various prediction methods described later. For example, the prediction unit can apply intra-prediction or inter-prediction for prediction of a single block, and can also apply intra-prediction and inter-prediction simultaneously. This can be called combined inter and intra prediction (CIIP). The prediction unit can also be based on intra-block copy (IBC) prediction mode or palette mode for prediction of a block. The IBC prediction mode or palette mode can be used for content video / movie coding such as games, for example, as in SCC (screen content coding). IBC basically performs prediction within the current picture, but can be performed similarly to inter-prediction in that it derives reference blocks within the current picture. That is, IBC can utilize at least one of the inter-prediction techniques described in this document. Palette mode can be considered an example of intra-coding or intra-prediction. When palette mode is applied, sample values within the picture can be signaled based on information about the palette table and palette index.
[0046] The prediction signal generated via the prediction unit (including the inter-prediction unit 221 and / or the intra-prediction unit 222) can be used to generate a reconstructed signal or a residual signal. The transformation unit 232 can generate transformation coefficients by applying a transformation technique to the residual signal. For example, the transformation technique may include at least one of the following: DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), KLT (Karhunen-Loeve Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT refers to a transformation obtained from a graph when relational information between pixels is represented by a graph. CNT refers to a transformation obtained by generating a prediction signal using all previously reconstructed pixels and based on that. The transformation process can also be applied to pixel blocks of the same size that are square, or to non-square blocks of variable size.
[0047] The quantization unit 233 quantizes the conversion coefficients and transmits them to the entropy encoding unit 240, which can encode the quantized signal (information about the quantized conversion coefficients) and output it as a bitstream. The information about the quantized conversion coefficients may be called residual information. The quantization unit 233 can rearrange the block-form quantized conversion coefficients into a one-dimensional vector form based on the coefficient scan order, and can also generate information about the quantized conversion coefficients based on the one-dimensional vector form of the quantized conversion coefficients. The entropy encoding unit 240 can perform various encoding methods, such as exponential Golomb, CAVLC (context-adaptive variable length coding), and CABAC (context-adaptive binary arithmetic coding). The entropy encoding unit 240 can encode information necessary for video / image restoration (e.g., values of syntax elements) together with or separately from the quantized conversion coefficients. The encoded information (e.g., encoded video / image information) can be transmitted or stored in bitstream form in units of NAL (network abstraction layer) units. The video / image information may further include information on various parameter sets, such as the adaptation parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS). The video / image information may also further include general constraint information. In this document, information and / or syntax elements transmitted / signaled from the encoding device to the decoding device may be included in the video / image information.The video / image information can be encoded through the encoding procedure described above and included in the bitstream. The bitstream can be transmitted over a network or stored in a digital storage medium. Here, the network can include broadcast networks and / or communication networks, and the digital storage medium can include a variety of storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The signal output from the entropy encoding unit 240 can be transmitted by a transmitting unit (not shown) and / or stored by a storage unit (not shown) which are configured as internal / external elements of the encoding device 200, or the transmitting unit can be included in the entropy encoding unit 240.
[0048] The quantized conversion coefficients output from the quantization unit 233 can be used to generate a prediction signal. For example, a residual signal (residual block or residual sample) can be reconstructed by applying inverse quantization and inverse transformation to the quantized conversion coefficients via the inverse quantization unit 234 and the inverse transformation unit 235. The adder 155 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter-prediction unit 221 or the intra-prediction unit 222. If there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as the reconstructed block. The adder 250 can be called the reconstruction unit or reconstructed block generation unit. The generated reconstructed signal can be used for intra-prediction of the next block to be processed in the current picture, or, as described later, for inter-prediction of the next picture after filtering.
[0049] On the other hand, LMCS (luma mapping with chromium ascaling) can also be applied during the picture encoding and / or restoration process.
[0050] The filtering unit 260 can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit 260 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be stored in the memory 270, specifically in the DPB of the memory 270. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset (SAO), adaptive loop filter, and bilateral filter. The filtering unit 260 can generate various filtering-related information and transmit it to the entropy encoding unit 290, as will be described later in the explanation of each filtering method. The filtering-related information can be encoded by the entropy encoding unit 290 and output in bitstream format.
[0051] The corrected restored picture sent to memory 270 can be used as a reference picture in the interpretation unit 280. When interpretation is applied via this, the encoding device can avoid prediction mismatches between the encoding device 200 and the decoding device, and can also improve encoding efficiency.
[0052] The DPB in memory 270 can store the corrected restored picture for use as a reference picture in the inter-prediction unit 221. Memory 270 can store motion information of blocks from which motion information in the current picture has been derived (or encoded) and / or motion information of blocks in the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit 221 for use as motion information of spatially adjacent blocks or motion information of temporally adjacent blocks. Memory 270 can store restored samples of restored blocks in the current picture and transmit them to the intra-prediction unit 222.
[0053] Figure 3 is a schematic diagram illustrating the configuration of a video / image decoding device to which this document can be applied.
[0054] As shown in Figure 3, the decoding device 300 can be configured to include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-prediction unit 331 and an intra-prediction unit 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 321. The entropy decoder 310, residual processor 320, predictor 330, adder 340, and filtering unit 350 described above can be configured by a single hardware component (e.g., a decoder chipset or processor) depending on the embodiment. The memory 360 may include a decoded picture buffer (DPB) and may also be configured by a digital storage medium. The aforementioned hardware component may also further include memory 360 as an internal / external component.
[0055] When a bitstream containing video / image information is input, the decoding device 300 can reconstruct the image in accordance with the process by which the video / image information was processed in the encoding device shown in Figure 3. For example, the decoding device 300 can derive units / blocks based on block division-related information obtained from the bitstream. The decoding device 300 can perform decoding using the processing units applied in the encoding device. Therefore, the decoding processing unit can be, for example, a coding unit, which can be divided from a coding tree unit or a maximum coding unit according to a quad-tree structure, a binary tree structure, and / or a terminally tree structure. One or more conversion units can be derived from the coding unit. The reconstructed image signal decoded and output via the decoding device 300 can then be reproduced via a playback device.
[0056] The decoding device 300 can receive the signal output from the encoding device shown in Figure 3 in bitstream form, and the received signal can be decoded via the entropy decoding unit 310. For example, the entropy decoding unit 310 can parse the bitstream to derive information necessary for video restoration (or picture restoration) (e.g., video / image information). The video / image information may further include information about various parameter sets, such as the adaptation parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS). The video / image information may also further include general constraint information. The decoding device can further decode the picture based on the parameter set information and / or the general constraint information. The signaling / received information and / or syntax elements described later in this document can be decoded via the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 310 can decode information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the values of syntax elements necessary for image restoration and the quantized values of conversion coefficients related to the residual. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element in the bitstream, determines a context model using the syntax element information to be decoded and the decoding information of adjacent and decoded blocks or the symbol / bin information decoded in a previous step, predicts the probability of bin occurrence based on the determined context model, and performs arithmetic decoding of the bins to generate symbols corresponding to the values of each syntax element.At this time, the CABAC entropy decoding method can update the context model after determining the context model by utilizing the information of the decoded symbol / bin for the context model of the next symbol / bin. Of the information decoded by the entropy decoding unit 310, information related to prediction is provided to the prediction unit (inter-prediction unit 332 and intra-prediction unit 331), and the residual values for which entropy decoding has been performed in the entropy decoding unit 310, i.e., quantized conversion coefficients and related parameter information, can be input to the residual processing unit 320. The residual processing unit 320 can derive residual signals (residual blocks, residual samples, residual sample arrays). In addition, of the information decoded by the entropy decoding unit 310, information related to filtering can be provided to the filtering unit 350. On the other hand, a receiving unit (not shown) that receives signals output from the encoding device can be further configured as an internal / external element of the decoding device 300, or the receiving unit may also be a component of the entropy decoding unit 310. On the other hand, the decoding device relating to this document may be called a video / image / picture decoding device, and the decoding device may also be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoding unit 310, and the sample decoder may include at least one of the inverse quantization unit 321, inverse transformation unit 322, addition unit 340, filtering unit 350, memory 360, inter-prediction unit 332, and intra-prediction unit 331.
[0057] The inverse quantization unit 321 can inverse quantize the quantized transformation coefficients and output the transformation coefficients. The inverse quantization unit 321 can rearrange the quantized transformation coefficients in a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scan order performed by the encoding device. The inverse quantization unit 321 can perform inverse quantization on the quantized transformation coefficients using quantization parameters (e.g., quantization step size information) and obtain the transformation coefficients.
[0058] In the inverse conversion unit 322, the conversion coefficients are inversely converted to obtain a residual signal (residual block, residual sample array).
[0059] The prediction unit can make predictions for the current block and generate a predicted block containing prediction samples for the current block. Based on the prediction information output from the entropy decoding unit 310, the prediction unit can determine whether intra-prediction or inter-prediction is applied to the current block, and can determine a specific intra / inter-prediction mode.
[0060] The prediction unit 330 can generate prediction signals based on various prediction methods described later. For example, the prediction unit can apply intra-prediction or inter-prediction for prediction of a single block, and can also apply intra-prediction and inter-prediction simultaneously. This may be called combined inter and intra prediction (CIIP). The prediction unit can also be based on an intra-block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode can be used for content video / movie coding such as games, for example, as in SCC (screen content coding). IBC basically performs prediction within the current picture, but can be performed similarly to inter-prediction in that it derives a reference block within the current picture. That is, IBC can utilize at least one of the inter-prediction techniques described in this document. Palette mode can be considered an example of intra-coding or intra-prediction. When palette mode is applied, information about the palette table and palette index can be included in the video / movie information and signaled.
[0061] The intra-prediction unit 331 can predict the current block by referring to a sample in the current picture. The referenced sample may be located adjacent to the current block or at a distance, depending on the prediction mode. The prediction mode in intra-prediction may include multiple non-directional modes and multiple directional modes. The intra-prediction unit 331 can also determine the prediction mode to be applied to the current block by utilizing the prediction modes applied to adjacent blocks.
[0062] The interprediction unit 332 can derive a predicted block for the current block based on a reference block (reference sample array) identified by motion vectors on a reference picture. In this case, to reduce the amount of motion information transmitted in interprediction mode, motion information can be predicted in blocks, subblocks, or samples based on the correlation of motion information between adjacent blocks and the current block. The motion information may include motion vectors and reference picture indices. The motion information may further include interprediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of interprediction, adjacent blocks may include spatially adjacent blocks present in the current picture and temporally adjacent blocks present in the reference picture. For example, the interprediction unit 332 can construct a motion information candidate list based on adjacent blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Interprediction can be performed based on various prediction modes, and the prediction information may include information indicating the mode of interprediction for the current block.
[0063] The summing unit 340 can generate a restored signal (restored picture, restored block, restored sample array) by adding the acquired residual signal to the predicted signal (predicted block, predicted sample array) output from the prediction unit 330. If there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as the restored block.
[0064] The addition unit 340 may be called the restoration unit or restoration block generation unit. The generated restoration signal can be used for intra-prediction of the next block to be processed in the current picture, and can be output after filtering as described later, or it can be used for intra-prediction of the next picture.
[0065] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied during the picture decoding process.
[0066] The filtering unit 350 can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit 350 can apply various filtering methods to the restored picture to generate a modified restored picture, and can transmit the modified restored picture to the memory 60, specifically to the DPB of memory 360. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, and bilateral filter.
[0067] The (modified) restored picture stored in the DPB of memory 360 can be used as a reference picture by the inter-prediction unit 331. Memory 360 can store motion information of blocks from which motion information in the current picture has been derived (or decoded) and / or motion information of blocks in the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit 331 for use as motion information of spatially adjacent blocks or motion information of temporally adjacent blocks. Memory 360 can store restored samples of restored blocks in the current picture and transmit them to the intra-prediction unit 332.
[0068] In this specification, the embodiments described for the filtering unit 260, the inter-prediction unit 221, and the intra-prediction unit 222 of the encoding device 100 can also be applied identically or in a corresponding manner to the filtering unit 350, the inter-prediction unit 332, and the intra-prediction unit 331 of the decoding device 300, respectively.
[0069] As mentioned above, prediction is performed to improve compression efficiency when performing video coding. This allows for the generation of a predicted block containing predicted samples for the current block, which is the block to be coded. Here, the predicted block contains predicted samples in the spatial domain (or pixel domain). The predicted block is similarly derived by the encoding and decoding devices, and the encoding device can improve image coding efficiency by signaling the decoding device with information about the residual between the original block and the predicted block (residual information), which is not the original sample value of the original block itself. The decoding device can derive a residual block containing residual samples based on the residual information, and can generate a restored block containing restored samples by adding the residual block and the predicted block, thereby generating a restored picture containing the restored block.
[0070] The residual information can be generated through transformation and quantization procedures. For example, an encoding device can signal the relevant residual information (via a bitstream) to a decoding device by deriving a residual block between the original block and the predicted block, performing a transformation procedure on the residual samples (residual sample array) contained in the residual block to derive transformation coefficients, and performing a quantization procedure on the transformation coefficients to derive quantized transformation coefficients. Here, the residual information may include information such as the value information, position information, transformation technique, transformation kernel, and quantization parameters of the quantized transformation coefficients. The decoding device can derive a residual sample (or residual block) by performing an inverse quantization / inverse transformation procedure based on the residual information. The decoding device can generate a reconstructed picture based on the predicted block and the residual block. The encoding device can also derive a residual block by inverse quantization / inverse transformation of the quantized transformation coefficients for reference for subsequent interpretation of the picture, and generate a reconstructed picture based on this.
[0071] Figure 4 illustrates the hierarchical structure of coded data.
[0072] Referring to Figure 4, coded data can be divided into the VCL (video coding layer), which handles the video / image coding process and itself, and the NAL (Network Abstraction Layer), which is located between the underlying systems that store and transmit the coded video / image data.
[0073] VCL can generate parameter sets corresponding to headers such as sequence and picture (e.g., picture parameter set (PPS), sequence parameter set (SPS), video parameter set (VPS)) and SEI (Supplemental Enhancement Information) messages additionally required during the video / image coding process. SEI messages are separated from information about the video / image (slice data). VCL containing information about the video / image consists of slice data and slice headers. On the other hand, the slice header may be referred to as a tile group header, and the slice data may be referred to as tile group data.
[0074] In NAL, a NAL unit can be generated by adding header information (NAL unit header) to an RBSP (Raw Byte Sequence Payload) generated by VCL. In this case, the RBSP refers to slice data, parameter sets, SEI messages, etc., generated by VCL. The NAL unit header can include NAL unit type information identified by the RBSP data contained in the NAL unit.
[0075] The NAL unit, the basic unit of NAL, is responsible for mapping coded video to a file format conforming to a predetermined standard, or to a bit sequence of a lower-level system such as RTP (Real-time Transport Protocol) or TS (Transport Storage).
[0076] As illustrated, NAL units can be divided into VCL NAL units and Non-VCL NAL units by the RBSP generated in VCL. VCL NAL units can represent NAL units that contain information about the video (slice data), while Non-VCL NAL units can represent NAL units that contain information necessary for decoding the video (parameter sets or SEI messages).
[0077] The aforementioned VCL NAL units and Non-VCL NAL units can be transmitted over a network with header information added according to the data standards of the underlying system. For example, NAL units can be transformed into data formats of predetermined standards such as H.266 / VVC file format, RTP (Real-time Transport Protocol), and TS (Transport Stream), and transmitted over various networks.
[0078] As mentioned above, the NAL unit type can be identified by the RBSP data structure contained within the NAL unit, and information about such NAL unit types can be stored in the NAL unit header and signaled.
[0079] For example, NAL units can be broadly classified into VCL NAL unit types and Non-VCL NAL unit types depending on whether or not they contain information (slice data) related to the video. VCL NAL unit types can be further classified by the nature and type of picture contained in the VCL NAL unit, while Non-VCL NAL unit types can be further classified by the type of parameter set.
[0080] The following is an example of a NAL unit type identified by the type of parameter set included in the Non-VCL NAL unit type. A NAL unit type can be identified by the type of parameter set. For example, a NAL unit type can be identified as one of the following: an APS (Adaptation Parameter Set) NAL unit, a DPS (Decoding Parameter Set) NAL unit, a VPS (Video Parameter Set) NAL unit, a SPS (Sequence Parameter Set) NAL unit, and a PPS (Picture Parameter Set) NAL unit.
[0081] The aforementioned NAL unit type has syntax information for the NAL unit type, and this syntax information can be stored in the NAL unit header and signaled. For example, the syntax information may be nal_unit_type, and the NAL unit type can be identified by the nal_unit_type value.
[0082] On the other hand, as mentioned above, a single picture can contain multiple slices, and a single slice can contain a slice header and slice data. In this case, a single picture header can be added to multiple slices (slice headers and slice data sets) within a single picture. The picture header (picture header syntax) can contain information / parameters that can be applied in common to the picture. The slice header (slice header syntax) can contain information / parameters that can be applied in common to the slice. The APS (APS syntax) or PPS (PPS syntax) can contain information / parameters that can be applied in common to one or more slices or pictures. The SPS (SPS syntax) can contain information / parameters that can be applied in common to one or more sequences. The VPS (VPS syntax) can contain information / parameters that can be applied in common to multiple layers. The DPS (DPS syntax) can contain information / parameters that can be applied in common to video in general. The DPS can contain information / parameters related to the concatenation of a CVS (coded video sequence). In this document, High-level syntax (HLS) may include at least one of the following: APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, picture header syntax, or slice header syntax.
[0083] In this document, the video information encoded by the decoding device from the encoding device and signaled in bitstream form includes not only picture partitioning-related information, intra / inter prediction information, residual information, and loop filtering information, but also information contained in the slice header, the picture header, the APS, the PPS, the SPS, the VPS, and / or the DPS. Furthermore, the video information may further include information from the NAL unit header.
[0084] Figure 5 is a diagram illustrating an example of partitioning a picture.
[0085] A picture can be divided into coding tree units (CTUs), and a CTU can correspond to a coding tree block (CTB). A CTU may contain two coding tree blocks: one for a luma sample and another for a corresponding chroma sample. On the other hand, the maximum allowable size of a CTU for coding and prediction may differ from the maximum allowable size of a CTU for transformation.
[0086] A tile can correspond to a series of CTUs that cover a rectangular area of a picture, and the picture can be divided into one or more tile rows and one or more tile columns.
[0087] On the other hand, a slice can consist of an integer number of complete tiles or an integer number of consecutive complete CTU rows. In this case, two slicing modes can be supported, including raster-scan slice mode and rectangular slice mode.
[0088] In raster scan slice mode, a slice can contain a series of complete tiles in a tile raster scan of a picture. In quadrilateral slice mode, a slice can contain multiple complete tiles that collectively form a quadrilateral region of a picture. Alternatively, in quadrilateral slice mode, a slice can contain multiple consecutive CTU rows within a single tile that collectively form a quadrilateral region of a picture. Tiles within a quadrilateral slice can be scanned in tile raster scan order within the quadrilateral region corresponding to the slice.
[0089] On the other hand, a subpicture can contain one or more slices that cover the rectangular area of the picture.
[0090] Figure 5(a) is a diagram showing an example of a picture being divided into raster scan slices. For example, the picture can be divided into 12 tiles and 3 raster scan slices.
[0091] Furthermore, Figure 5(b) is a diagram showing an example of dividing a picture into rectangular slices. For example, a picture can be divided into 24 tiles (6 tile columns and 4 tile rows) and 9 rectangular slices.
[0092] Furthermore, Figure 5(c) is a diagram showing an example of a picture being divided into tiles and rectangular slices. For example, a picture can be divided into 24 tiles (2 tile columns and 2 tile rows) and 4 rectangular slices.
[0093] Figure 6 is a flowchart showing a picture encoding procedure according to one embodiment.
[0094] In one embodiment, picture partitioning (S600) can be performed by the video splitting unit 210 of the encoding device, and picture encoding (S610) can be performed by the entropy encoding unit 240 of the encoding device.
[0095] An encoding device according to one embodiment can derive slices and / or tiles contained in the currently selected picture (S600). For example, the encoding device can perform picture partitioning for encoding the input currently selected picture. For example, the encoding device can derive slices and / or tiles contained in the currently selected picture. The encoding device can partition the picture into various forms considering the image characteristics and coding efficiency of the currently selected picture, and can generate information indicating the partitioning form having the optimal coding efficiency and signal it to the decoding device.
[0096] An encoding device according to one embodiment can perform encoding on the current picture based on the derived slices and / or tiles (S610). For example, the encoding device can encode video / image information including information about slices and / or tiles and output it in bitstream form. The output bitstream can be transmitted to a decoding device via a digital storage medium or a network.
[0097] Figure 7 is a flowchart showing a picture decoding procedure according to one embodiment.
[0098] In one embodiment, the steps of obtaining video / image information from a bitstream (S710) and deriving slices and / or tiles within the current picture (S720) can be performed by the entropy decoding unit 310 of the decoding device, and the step of restoring the current picture based on the slices and / or tiles can be performed by the addition unit 340 of the decoding device.
[0099] A decoding device according to one embodiment can acquire video / image information from a received bitstream (S710). The video / image information may include HLS, which may include information about slices or information about tiles. Information about slices may include information that identifies one or more slices in the current picture, and information about tiles may include information that identifies one or more tiles in the current picture. Information about slices or tiles can be acquired through a variety of parameter sets, picture headers and / or slice headers.
[0100] On the other hand, a picture can currently contain a tile that includes one or more slices, or a slice that includes one or more tiles.
[0101] A decoding apparatus according to one embodiment can derive the currently active slice and / or tile in a picture based on video / image information including information about slices and / or tiles (S720).
[0102] A decoding apparatus according to one embodiment can restore (decode) the current picture based on slices and / or tiles (S730).
[0103] On the other hand, as mentioned above, a picture can be divided into subpictures, tiles, and slices. Information about subpictures can be signaled through SPS, and information about tiles and rectangular slices can be signaled through PPS. In addition, information about raster-scan slices can be signaled through a slice header.
[0104] For example, SPS syntax including information about subpictures may be as shown in the following table.
[0105] [Table 1]
[0106] For example, PPS syntax including information about tiles and rectangular slices may be as shown in the following table.
[0107] [Table 2]
[0108] Furthermore, for example, the slice header syntax containing information about raster-scan slices may be as shown in the following table.
[0109] [Table 3]
[0110] On the other hand, information about slices and tiles within a picture may include a flag indicating whether each subpicture within the picture contains a single slice. This flag may be referred to as single_slice_per_subpic_flag or pps_single_slice_per_subpic_flag, but is not limited to these. Similarly, information about subpictures may include a flag indicating the existence of subpicture information, which may be referred to as subpics_present_flag or sps_subpic_info_present_flag, but is not limited to these. For example, information about subpictures may be included in a parameter set (parameter_set). For example, information about subpictures may be included in SPS.
[0111] Traditionally, if the flag value for the existence of subpicture information was 0, the flag value for whether the subpicture contains a single slice was restricted to 0. That is, if the flag value for the existence of subpicture information was 0, it was determined that the subpicture was unavailable, and the flag value for whether the subpicture contains a single slice was restricted to 0. However, this condition is very restrictive. For example, even if subpicture information does not exist, the current picture can be divided into two or more tiles, and all tiles can be contained within a single slice. In such cases, the current picture will contain a single slice.
[0112] Therefore, one embodiment of this document proposes a method to remove the restriction that if the value of the flag regarding the existence of subpicture information is 0, the value of the flag regarding whether the subpicture contains only one slice is also 0. In such a case, the flag regarding whether the subpicture contains only one slice can indicate that the current picture contains only one slice even if the subpicture information does not exist.
[0113] For example, according to the embodiment described above, even if sub-picture information does not exist for a CLVS (coded layer video sequence), a flag may exist indicating whether or not the sub-picture contains a single slice. That is, even if sub-picture information does not exist for a CLVS, the flag indicating whether or not the sub-picture contains a single slice may have a value of 0 or 1.
[0114] For example, if the flag indicating the existence of subpicture information is 0, and the flag indicating whether a subpicture contains a single slice is 1, then the picture can currently contain a single slice. That is, if there are no signaled subpictures and the flag indicating whether a subpicture contains a single slice is 1, then it can be inferred that the number of slices in the picture is 1.
[0115] Furthermore, if the flag value regarding the existence of sub-picture information is 0, the current number of sub-pictures within a picture may be 1. For example, if the flag value regarding the existence of sub-picture information is 0, the number of sub-pictures present in each picture that references the SPS of the video information may be 1.
[0116] On the other hand, a flag regarding the number of slices currently contained in a picture can be included in the PPS of the video information. This flag may be referred to as num_slices_in_pic_minus1 or pps_num_slices_in_pic_minus1, but is not limited to these terms. Similarly, a flag regarding the number of subpictures currently contained in a picture can be included in the SPS of the video information. This flag may be referred to as sps_num_subpics_minus1, but is not limited to these terms.
[0117] If there is no signaled sub-picture information and the flag indicating whether a sub-picture contains a single slice is valued at 1, it can be inferred that the flag indicating the number of slices currently contained in the picture has a value of 0. Furthermore, if there is no signaled sub-picture information and the flag indicating whether a sub-picture contains a single slice is valued at 1, it can be inferred that the flag indicating the number of slices currently contained in the picture and the flag indicating the number of sub-pictures currently contained in the picture have the same value.
[0118] Additionally, if there are no signaled subpictures and the flag indicating whether a subpicture contains a single slice is valued at 1, then all CTUs within the picture can belong to a single slice contained within the picture.
[0119] The semantics for a syntax element that includes a flag indicating whether the subpicture according to the above embodiment contains a single slice and a flag indicating the number of slices currently contained in the picture can be shown in the following table.
[0120] [Table 4]
[0121] Referring to the table above, if the value of single_slice_per_subpic_flag, which corresponds to the flag indicating whether a subpicture contains a single slice, is 1, then each subpicture can consist of one quadrilateral slice. Also, if the value of single_slice_per_subpic_flag is 0, then each subpicture can consist of one or more quadrilateral slices. If the value of single_slice_per_subpic_flag is 1, it can be inferred that the value of num_slices_in_pic_minus1, which corresponds to the flag indicating the number of slices currently contained in the picture, is the same as the value of SPS_num_subpics_minus1, which corresponds to the flag indicating the number of subpictures currently contained in the picture.
[0122] Furthermore, if the value of single_slice_per_subpic_flag is 1 and the value of subpics_present_flag, which corresponds to the flag regarding the existence of subpicture information, is 0, then a picture referencing PPS can have one slice per picture.
[0123] On the other hand, the scanning process, which is the procedure for decoding the tiles within a picture, can be determined by the table below.
[0124] [Table 5-1]
[0125] [Table 5-2]
[0126] Figure 8 is a flowchart showing the operation of an encoding device according to one embodiment, and Figure 9 is a block diagram showing the configuration of an encoding device according to one embodiment.
[0127] The method disclosed in Figure 8 can be performed by the encoding device disclosed in Figure 2 or Figure 9. S810 in Figure 9 can be performed by the video splitting unit 210 disclosed in Figure 2, and S820 can be performed by the entropy encoding unit 240 disclosed in Figure 2. Furthermore, the operations of S810 to S820 are based in part on the content described in Figures 1 to 7. Therefore, specific details that overlap with the content described in Figures 1 to 7 will be omitted or simplified in this explanation.
[0128] Referring to Figure 8, an encoding device according to one embodiment can divide the current picture and derive at least one slice (S810). For example, the video division unit 210 of the encoding device can generate division information for the current picture based on the at least one slice.
[0129] An encoding device according to one embodiment can encode video information for a current picture based on at least one slice (S810). The video information may include segmentation information generated based on the at least one slice.
[0130] For example, the video information may include a first flag indicating whether or not sub-picture information exists, and a second flag indicating whether or not the sub-picture contains only one slice. For example, based on the first and second flags, it can be derived that the number of slices contained in the current picture is one.
[0131] For example, if the value of the first flag regarding the existence or nonexistence of the subpicture information is 0 and the second flag is 1, it can be derived that the number of slices included in the current picture is 1.
[0132] For example, if the value of the first flag regarding the existence or nonexistence of the subpicture information is 0, the number of subpictures present in the current picture may be 1.
[0133] For example, the first flag regarding the existence or non-existence of the subpicture information can be included in the SPS (Sequence Parameter Set) of the video information.
[0134] For example, a second flag indicating whether or not the subpicture contains only one slice can be included in the Picture Parameter Set (PPS) of the video information.
[0135] For example, the video information includes a third flag relating to the number of slices contained in the current picture, and the third flag may be included in the PPS of the video information.
[0136] Furthermore, for example, the video information may include a fourth flag relating to the number of sub-pictures included in the current picture, and the fourth flag may be included in the SPS of the video information.
[0137] Furthermore, for example, if the value of the first flag is 0, it can be derived that the number of subpictures present in each picture that references the SPS of the video information is 1.
[0138] On the other hand, the video information may include prediction information for the current picture. The prediction information may include information for an inter-prediction mode or intra-prediction mode to be performed on the current picture. The encoding device can generate and encode the prediction information for the current picture.
[0139] On the other hand, the bitstream can be transmitted to a decoding device via a network or a (digital) storage medium. Here, the network may include broadcast networks and / or communication networks, and the digital storage medium may include a variety of storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.
[0140] Figure 10 is a flowchart showing the operation of a decoding device according to one embodiment, and Figure 11 is a block diagram showing the configuration of a decoding device according to one embodiment.
[0141] The method disclosed in Figure 10 can be performed by the decoding device disclosed in Figure 3 or Figure 11. Specifically, steps S1010 and S1020 can be performed by the entropy decoding unit 310 disclosed in Figure 3. Furthermore, the operation of steps S1010 and S1020 is based in part on the content described in Figures 1 to 7. Therefore, specific details that overlap with the content described in Figures 1 to 7 will be omitted or simplified in this explanation.
[0142] A decoding device according to one embodiment can acquire video information for the current picture from a bitstream (S1010). For example, the entropy decoding unit 310 of the decoding device can acquire video information including segmentation information for the current picture from a bitstream. For example, the segmentation information may include slice information for the current picture. The video information may also include at least a portion of prediction-related information or residual-related information. For example, the prediction-related information may include inter-prediction mode information or inter-prediction type information.
[0143] A decoding device according to one embodiment can perform decoding on the current picture based on video information (S1020). For example, the entropy decoding unit 310 of the decoding device can derive the division structure of the current picture based on slice information of the current picture.
[0144] For example, the video information may include a first flag indicating whether or not sub-picture information exists, and a second flag indicating whether or not the sub-picture contains only one slice. For example, based on the first and second flags, it can be derived that the number of slices contained in the current picture is one.
[0145] For example, if the value of the first flag is 0 and the value of the second flag is 1, it can be derived that the number of slices currently included in the picture is 1.
[0146] For example, if the value of the flag regarding the existence of the sub-picture information is 0, the number of sub-pictures present in the current picture may be 1.
[0147] For example, the first flag regarding the existence or non-existence of the subpicture information can be included in the SPS (Sequence Parameter Set) of the video information.
[0148] For example, a second flag indicating whether or not the subpicture contains a single slice can be included in the Picture Parameter Set (PPS) of the video information.
[0149] For example, the video information includes a third flag relating to the number of slices contained in the current picture, and the third flag may be included in the PPS of the video information.
[0150] Furthermore, for example, the video information includes a fourth flag relating to the number of sub-pictures included in the current picture, and the flag relating to the number of sub-pictures included in the current picture may be included in the SPS of the video information.
[0151] Furthermore, for example, if the value of the first flag is 0, it can be derived that the number of subpictures present in each picture that references the SPS of the video information is 1.
[0152] In the embodiments described above, the method is explained based on a flowchart in a series of steps or blocks, but the embodiments are not limited to the order of the steps, and some steps may occur with other steps, in a different order, or simultaneously. Furthermore, those skilled in the art will understand that the steps shown in the flowchart are not exclusive, other steps may be included, or one or more steps in the flowchart may be deleted without affecting the scope of the embodiments described herein.
[0153] The methods described in the embodiments of this document above can be implemented in software form, and the encoding and / or decoding devices described in this document can be included in devices that perform video processing, such as TVs, computers, smartphones, set-top boxes, and display devices.
[0154] In this document, when embodiments are implemented in software, the methods described above can be implemented by modules (processes, functions, etc.) that perform the functions described above. These modules are stored in memory and can be executed by a processor. The memory may be internal or external to the processor and may be connected to the processor by a variety of well-known means. The processor may include an ASIC (application-specific integrated circuit), other chipsets, logic circuits, and / or data processing devices. The memory may include ROM (read-only memory), RAM (random access memory), flash memory, memory cards, storage media, and / or other storage devices. That is, the embodiments described in this document can be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units shown in each drawing can be implemented and executed on a computer, processor, microprocessor, controller, or chip. In this case, information on instructions or algorithms for implementation may be stored on a digital storage medium.
[0155] Furthermore, decoding and encoding devices to which this disclosure applies may include multimedia broadcasting transceivers, mobile communication terminals, home cinema video equipment, digital cinema video equipment, surveillance cameras, video interaction devices, real-time communication devices such as video communications, mobile streaming devices, storage media, camcorders, video-on-demand (VoD) service providers, over-the-top (OTT) video equipment, internet streaming service providers, 3D video equipment, virtual reality (VR) equipment, argumentative reality (AR) equipment, image-phone video equipment, transportation terminals (e.g., vehicle terminals (including autonomous vehicles), airplane terminals, ship terminals, etc.), and medical video equipment, and can be used to process video signals or data signals. For example, over-the-top (OTT) video equipment may include game consoles, Blu-ray players, internet-connected TVs, home theater systems, smartphones, tablet PCs, and digital video recorders (DVRs).
[0156] Furthermore, the processing methods to which this disclosure applies can be produced in the form of programs executed on a computer and stored on a computer-readable recording medium. Multimedia data having data structures according to the embodiments(et al.) of this document can also be stored on a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices that store data that can be read by a computer. The computer-readable recording medium can include, for example, Blu-ray discs (BDs), general-purpose serial buses (USBs), ROMs, PROMs, EPROMs, EEPROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, and optical data storage devices. The computer-readable recording medium also includes media embodied in the form of carrier waves (e.g., transmission over the Internet). Furthermore, bitstreams generated by encoding methods can be stored on a computer-readable recording medium or transmitted over a wireless network.
[0157] Furthermore, embodiments of the present disclosure can be embodied in computer program products comprising program code, the program code can be executed on a computer according to the embodiments(et al.) of this document. The program code can be stored on a computer-readable carrier.
[0158] Figure 12 shows an example of a content streaming system to which the disclosures in this document apply.
[0159] Referring to Figure 12, the content streaming system to which this disclosure applies can broadly include an encoding server, a streaming server, a web server, a media repository, user equipment, and multimedia input devices.
[0160] The encoding server is responsible for compressing content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream, and then transmitting this bitstream to the streaming server. In other cases, if a multimedia input device such as a smartphone, camera, or camcorder directly generates the bitstream, the encoding server may be omitted.
[0161] The bitstream can be generated by an encoding method or bitstream generation method applicable to the embodiments of this document, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.
[0162] The streaming server transmits multimedia data to user devices based on user requests via a web server, and the web server acts as an intermediary to inform users about available services. When a user requests a desired service from the web server, the web server transmits this to the streaming server, and the streaming server transmits multimedia data to the user. In this case, the content streaming system may include a separate control server, in which case the control server controls the commands and responses between the devices within the content streaming system.
[0163] The streaming server can receive content from a media storage and / or encoding server. For example, if it starts receiving content from the encoding server, it can receive the content in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.
[0164] Examples of user devices include mobile phones, smartphones, laptop computers, digital broadcasting terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (such as smartwatches, smart glasses, HMDs (head-mounted displays)), digital TVs, desktop computers, and digital signage.
[0165] Each server within the aforementioned content streaming system can be operated as a distributed server, in which case the data received by each server can be processed in a distributed manner.
[0166] The claims described herein can be combined in various ways. For example, the technical features of the method claims herein can be combined and embodied in an apparatus, and the technical features of the apparatus claims herein can be combined and embodied in a method. Furthermore, the technical features of the method claims and the technical features of the apparatus claims herein can be combined and embodied in an apparatus, and the technical features of the method claims and the technical features of the apparatus claims herein can be combined and embodied in a method.
Claims
1. In a video decoding method performed by a decoding device, Currently, the process involves obtaining video information about the picture from the bitstream, The step includes decoding the current picture based on the aforementioned video information, The aforementioned video information includes a first flag related to the presence of subpicture information and a second flag related to whether each subpicture contains a single slice. Based on the fact that the value of the first flag is equal to 0 and the value of the second flag is equal to 1, the number of slices currently included in the picture is derived to be equal to 1. The value of the second flag being equal to 1 indicates that each subpicture consists of a single slice. The first flag is a method included in the sequence parameter set.
2. In a video encoding method performed by an encoding device, Currently, the picture is divided to derive at least one slice, The step includes encoding video information for the current picture based on at least one slice, The video information includes a first flag related to the presence of sub-picture information and a second flag related to whether each sub-picture in the current picture contains a single slice. Based on the fact that the value of the first flag is equal to 0 and the value of the second flag is equal to 1, the number of slices currently included in the picture is derived to be equal to 1. The value of the second flag being equal to 1 indicates that each subpicture consists of a single slice. The first flag is a method included in the sequence parameter set.
3. A method for transmitting video data, A step of obtaining a bitstream generated by a method, wherein the method is Currently, by dividing the picture, we can derive at least one slice, A step comprising generating the bitstream by encoding video information for the current picture based on at least one slice, The step of transmitting the data, which includes the bitstream, The video information includes a first flag related to the presence of sub-picture information and a second flag related to whether each sub-picture in the current picture contains a single slice. Based on the fact that the value of the first flag is equal to 0 and the value of the second flag is equal to 1, the number of slices currently included in the picture is derived to be equal to 1. The value of the second flag being equal to 1 indicates that each subpicture consists of a single slice. The first flag is a transmission method included in the sequence parameter set.