Syntax design method and apparatus for performing coding by using syntax

The image decoding method improves coding efficiency by using affine and sub-block TMVP flags to optimize merge modes, addressing the need for efficient compression of high-resolution images and VR/AR content.

JP2025108632AActive Publication Date: 2025-07-23LG ELECTRONICS INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025068293
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2018-10-08
Filing Date
2025-04-17
Publication Date
2025-07-23
Estimated Expiration
2039-10-08

AI Technical Summary

Technical Problem

The increasing demand for high-resolution and high-quality images/videos, including VR and AR content, has led to a need for highly efficient image/video compression technologies to reduce transmission and storage costs.

Method used

An image decoding method that uses affine flags and sub-block TMVP flags to determine whether to apply merge modes, improving motion prediction efficiency by decoding or encoding these flags based on bitstream conditions.

Benefits of technology

Enhances image coding efficiency through high-level and low-level syntax designs, particularly in motion prediction, reducing transmission and storage costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025108632000001_ABST
    Figure 2025108632000001_ABST
Patent Text Reader

Abstract

To provide an image decoding method and an image encoding method for enhancing the overall image / video compression efficiency.SOLUTION: A decoding method comprises: decoding, on the basis of a bitstream, an affine flag which indicates whether or not affine prediction is applicable to a current block, and a sub-block TMVP flag which indicates whether or not a temporal motion vector predictor based on a sub-block of the current block is usable; determining whether or not to decode a predetermined merge mode flag which indicates whether or not to apply a predetermined merge mode to the current block, on the basis of the decoded affine flag and the decoded sub-block TMVP flag; deriving prediction samples of the current block on the basis of the determination of whether or not to decode the predetermined merge mode flag; and generating reconstructed samples of the current block on the basis of the prediction samples of the current block.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to image coding technology, and more particularly, to a syntax design method and an apparatus for performing coding using syntax in an image coding system.

Background Art

[0002] In recent years, the demand for high-resolution and high-quality images / videos such as 4K or UHD (Ultra High Definition) images / videos of 8K or higher has been increasing in various fields. As the image / video data becomes higher in resolution and quality, the amount of information or bits to be transmitted relatively increases compared to the existing image / video data. Therefore, when transmitting image data using a medium such as an existing wired or wireless broadband line, or storing image / video data using an existing storage medium, the transmission cost and storage cost increase.

[0003] Also, in recent years, the interest and demand for immersive media such as VR (Virtual Reality), AR (Artificial Reality) contents, and holograms have been increasing, and the broadcast of images / videos having image characteristics different from those of real images, such as game images, has been increasing.

[0004] Accordingly, there is a need for a highly efficient image / video compression technology to effectively compress, transmit, store, and reproduce the information of high-resolution and high-quality images / videos having various characteristics as described above.

Summary of the Invention

Problems to be Solved by the Invention

[0005] The technical problem of the present disclosure is to provide a method and an apparatus for increasing image coding efficiency.

[0006] Another technical problem of the present disclosure is to provide a syntax design method and an apparatus for coding using the syntax.

[0007] Still another technical problem of the present disclosure is to provide a method and an apparatus for coding using high-level and low-level syntax design methods and syntax.

[0008] Still another technical problem of the present disclosure is to provide a method and an apparatus for using high-level and / or low-level syntax elements for motion prediction based on sub-blocks.

[0009] Still another technical problem of the present disclosure is to provide a method and an apparatus for using high-level and / or low-level syntax elements for motion prediction based on an affine model.

[0010] Still another technical problem of the present disclosure is to provide a method and an apparatus for determining whether to decode a determined merge mode flag indicating whether to apply a predetermined merge mode to a current block based on an affine flag and a sub-block TMVP flag.

Means for Solving the Problems

[0011] According to an embodiment of the present disclosure, an image decoding method performed by a decoding device is provided. The method includes decoding an affine flag indicating whether affine prediction can be applied to a current block based on a bitstream and a sub-block TMVP flag indicating whether a temporal motion vector predictor based on a sub-block of the current block can be used; determining whether to decode a pre-determined merge mode flag indicating whether to apply the pre-determined merge mode to the current block based on the decoded affine flag and the decoded sub-block TMVP flag; deriving a prediction sample for the current block based on the determination of whether to decode the pre-determined merge mode flag; and generating a restored sample for the current block based on the prediction sample for the current block. When the value of the affine flag is 1 or the value of the sub-block TMVP flag is 1, it is determined that the pre-determined merge mode flag is to be decoded.

[0012] According to another embodiment of the present disclosure, a decoding device for performing image decoding is provided. The decoding device can use an affine flag indicating whether affine prediction can be applied to a current block based on a bitstream, and a sub-block TMVP flag indicating whether a temporal motion vector predictor based on sub-blocks of the current block can be used, to decode an already-determined merge mode flag indicating whether to apply a predetermined merge mode to the current block based on the decoded affine flag and the decoded sub-block TMVP flag. An entropy decoding unit that determines whether to decode the already-determined merge mode flag, a prediction unit that derives prediction samples for the current block based on the determination regarding whether to decode the already-determined merge mode flag, and an addition unit that generates restored samples for the current block based on the prediction samples for the current block. When the value of the affine flag is 1 or the value of the sub-block TMVP flag is 1, it is determined that the already-determined merge mode flag is to be decoded.

[0013] According to still another embodiment of the present disclosure, there is provided an image encoding method performed by an encoding device. The method includes determining whether affine prediction can be applied to a current block and whether a temporal motion vector predictor based on sub-blocks of the current block can be used, determining whether to encode a determined merge mode flag indicating whether to apply a predetermined merge mode to the current block based on the determination as to whether affine prediction can be applied to the current block and whether the temporal motion vector predictor based on sub-blocks of the current block can be used, and encoding the affine flag indicating whether affine prediction can be applied to the current block, the sub-block TMVP flag indicating whether the temporal motion vector predictor based on sub-blocks of the current block can be used, and the determined merge mode flag based on the determination as to whether to encode the determined merge mode flag, wherein it is determined that the determined merge mode flag is encoded when the value of the affine flag is 1 or the value of the sub-block TMVP flag is 1.

[0014] According to still another embodiment of the present disclosure, an encoding device that performs image encoding is provided. The encoding device determines whether affine prediction can be applied to a current block and whether a temporal motion vector predictor based on sub-blocks of the current block can be used, and based on the determination as to whether affine prediction can be applied to the current block and whether the temporal motion vector predictor based on the sub-blocks of the current block can be used, a prediction unit that determines whether to encode a determined merge mode flag indicating whether to apply a predetermined merge mode to the current block, and based on the determination as to whether to encode the determined merge mode flag, an affine flag indicating whether affine prediction can be applied to the current block, a sub-block TMVP flag indicating whether the temporal motion vector predictor based on the sub-blocks of the current block can be used, and an entropy encoding unit that encodes the determined merge mode flag, and is characterized in that when the value of the affine flag is 1 or the value of the sub-block TMVP flag is 1, it is determined that the determined merge mode flag is to be encoded.

[0015] According to still another embodiment of the present disclosure, a decoder-readable storage medium that stores information regarding instructions for causing a video decoding device to perform a decoding method according to some embodiments or the like is provided.

[0016] According to still another embodiment of the present disclosure, there is provided a decoder-readable storage medium that stores information regarding an instruction for causing a video decoding device to perform a decoding method according to an embodiment. The decoding method according to the embodiment includes decoding an affine flag indicating whether affine prediction can be applied to a current block and a sub-block TMVP flag indicating whether a temporal motion vector predictor based on sub-blocks of the current block can be used, based on a bitstream; determining whether to decode a pre-determined merge mode flag indicating whether to apply the pre-determined merge mode to the current block, based on the decoded affine flag and the decoded sub-block TMVP flag; deriving prediction samples for the current block based on the determination as to whether to decode the pre-determined merge mode flag; and generating restored samples for the current block based on the prediction samples for the current block, and it is determined that the pre-determined merge mode flag is to be decoded when the value of the affine flag is 1 or the value of the sub-block TMVP flag is 1.

Advantages of the Invention

[0017] According to the present disclosure, general image / video compression efficiency can be improved.

[0018] According to the present disclosure, image coding efficiency can be improved through high-level and low-level syntax designs.

[0019] According to the present disclosure, image coding efficiency can be improved by using high-level and / or low-level syntax elements for performing motion prediction based on sub-blocks.

[0020] According to the present disclosure, image coding efficiency can be improved by using high-level and / or low-level syntax elements for performing motion prediction based on an affine model.

[0021] According to the present disclosure, by determining whether to decode a predetermined merge mode flag indicating whether to apply a predetermined merge mode to a current block based on an affinity flag and a sub-block TMVP flag, the image coding efficiency can be improved.

Brief Description of the Drawings

[0022]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Modes for Carrying Out the Invention

[0023] According to an embodiment of the present disclosure, an image decoding method performed by a decoding device is provided. The method includes decoding an affine flag indicating whether affine prediction can be applied to a current block based on a bitstream and a sub-block TMVP flag indicating whether a temporal motion vector predictor based on a sub-block of the current block can be used; determining whether to decode a predetermined merge mode flag indicating whether to apply the predetermined merge mode to the current block based on the decoded affine flag and the decoded sub-block TMVP flag; deriving a prediction sample for the current block based on the determination of whether to decode the predetermined merge mode flag; and generating a restored sample for the current block based on the prediction sample for the current block. It is characterized in that when the value of the affine flag is 1 or the value of the sub-block TMVP flag is 1, it is determined to decode the predetermined merge mode flag.

[0024] The present disclosure can be subject to various modifications and can have various embodiments. Specific embodiments will be illustrated in the drawings and described in detail. However, this is not intended to limit the present disclosure to the specific embodiments. The terms commonly used in this specification are used merely to explain specific embodiments and are not intended to limit the technical idea of the present disclosure. Singular expressions include plural expressions unless the context clearly indicates otherwise. Terms such as "including" or "having" in this specification are intended to specify the existence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and it should be understood that they do not preclude the existence or addition possibility of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0025] On the one hand, each component in the drawings described in the present disclosure is independently illustrated for the convenience of explaining different characteristic functions and the like, and it does not mean that each component is realized by separate hardware or separate software. For example, among the components, two or more components can be combined to form one component, and one component can also be divided into a plurality of components. Embodiments in which each component is integrated and / or separated are included in the scope of rights of the present disclosure as long as they do not deviate from the essence of the present disclosure.

[0026] Hereinafter, with reference to the accompanying drawings, preferred embodiments of the present disclosure will be described in more detail. Hereinafter, the same reference numerals are used for the same components in the drawings, and duplicate descriptions of the same components can be omitted.

[0027] FIG. 1 schematically shows an example of a video / image coding system to which the present disclosure can be applied.

[0028] As shown in FIG. 1, the video / image coding system can include a first device (source device) and a second device (receiver device). The source device can transmit encoded video / image information or data in a file or streaming form to the receiver device via a digital storage medium or a network.

[0029] The source device can include a video source, an encoding device, and a transmission unit. The receiver device can include a reception unit, a decoding device, and a renderer. The encoding device can be called a video / image encoding device, and the decoding device can be called a video / image decoding device. A transmitter can be provided in the encoding device. A receiver can be provided in the decoding device. The renderer can include a display unit, and the display unit can also be composed of a separate device or an external component.

[0030] The video source can obtain video / images through processes such as video / image capture, synthesis, or generation. The video source can include a video / image capture device and / or a video / image generation device. The video / image capture device can include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. The video / image generation device can include, for example, a computer, a tablet, and a smartphone, etc., and can (electronically) generate video / images. For example, virtual video / images can be generated via a computer, etc., and in this case, the video / image capture process can be replaced by the process of generating related data.

[0031] The encoding device can encode the input video / image. The encoding device can perform a series of procedures such as prediction, transformation, quantization, etc. for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.

[0032] The transmitting unit can transmit the encoded video / image information or data output in the form of a bitstream to the receiving unit of the receiving device via a digital storage medium or a network in file or streaming form. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit can include elements for generating a media file via a predetermined file format and can include elements for transmission via a broadcast / communication network. The receiving unit can receive / extract the bitstream and transmit it to the decoding device.

[0033] The decoding device can decode the video / image by performing a series of procedures such as inverse quantization, inverse transformation, prediction, etc. corresponding to the operation of the encoding device.

[0034] The renderer can render the decoded video / image. The rendered video / image can be displayed via the display unit.

[0035] This document relates to video / image coding. For example, the methods / embodiments disclosed in this document can be applied to the methods disclosed in the VVC (versatile video coding) standard, EVC (essential video coding) standard, AV1 (AOMedia Video 1) standard, AVS2 (2nd generation of audio video coding standard), or next-generation video / image coding standards (such as H.267 or H.268, etc.).

[0036] This document presents various embodiments related to video / image coding. Unless otherwise specified, the above embodiments etc. can also be combined with each other.

[0037] In this document, "video" can mean a collection of a series of images etc. according to the passage of time. "Picture" generally means a unit indicating one image in a specific time period, and "slice" / "tile" is a unit constituting a part of a picture in coding. A slice / tile can include one or more CTUs (Coding Tree Units). One picture can be composed of one or more slices / tiles. One picture can be composed of one or more tile groups. One tile group can include one or more tiles. A brick can represent a rectangular region of CTU rows within a tile in a picture. A tile can be partitioned into multiple bricks, and each brick can be composed of one or more CTU rows within the tile. A tile that is not partitioned into multiple bricks may also be referred to as a brick.A brick scan can indicate a specific sequential ordering of CTUs partitioning a picture, where the CTUs can be ordered consecutively in a CTU raster scan within a brick, bricks within a tile can be ordered consecutively in a raster scan of the bricks of the tile, and tiles in a picture can be ordered consecutively in a raster scan of the tiles of the picture. A tile is a rectangular region of CTUs within a particular tile column and a particular tile row in a picture. The tile column is a rectangular region of CTUs having a height equal to the height of the picture and a width specified by syntax elements in the picture parameter set.The tile row is a rectangular region of CTUs having a width specified by syntax elements in the picture parameter set and a height equal to the height of the picture. A tile scan is a specific sequential ordering of CTUs partitioning a picture in which the CTUs are ordered consecutively in CTU raster scan in a tile whereas tiles in a picture are ordered consecutively in a raster scan of the tiles of the picture. A slice includes an integer number of bricks of a picture that may be exclusively contained in a single NAL unit.A slice may consists of either a number of complete tiles or only a consecutive sequence of complete bricks of one tile. In this document, tile group and slice can be used interchangeably. For example, in this document, tile group / tile group header can be called slice / slice header.

[0038] A pixel or pel can mean the smallest unit that makes up one picture (or image). Also, as a term corresponding to a pixel, "sample" can be used. A sample can generally indicate a pixel or the value of a pixel, and can indicate only the pixel / pixel value of the luma component, or only the pixel / pixel value of the chroma component.

[0039] A unit can indicate the basic unit of image processing. A unit can include at least one of a specific area of a picture and information related to that area. One unit can include one luma block and two chroma (e.g., cb, cr) blocks. A unit can, in some cases, be used interchangeably with terms such as block or area. In general, an M×N block can include a set (or array) of samples (or sample array) or transform coefficients consisting of M columns and N rows.

[0040] In this document, the terms “ / ” and “,” shall be construed to mean “and / or.” For example, “A / B” shall be construed to mean “A and / or B,” and “A, B” shall be construed to mean “A and / or B.” Additionally, “A / B / C” means “at least one of A, B, and / or C.” Also, “A, B, C” also means “at least one of A, B, and / or C.”

[0041] Further, in this document, the term “or” shall be construed to mean “and / or.” For example, “A or B” can mean 1) only “A,” 2) only “B,” or 3) both “A and B.” In other words, the term “or” in this document can be construed to mean “additionally or alternatively.”

[0042] FIG. 2 is a diagram schematically illustrating the configuration of a video / image encoding apparatus to which the present disclosure can be applied. Hereinafter, the video encoding apparatus can include an image encoding apparatus.

[0043] As shown in FIG. 2, the encoding apparatus 200 can be configured to include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 can include an inter predictor 221 and an intra predictor 222. The residual processor 230 can include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 can further include a subtractor (231). The adder 250 can be called a reconstructor or a reconstructed block generator. The above-described image partitioner 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 can be configured by one or more hardware components (e.g., an encoder chipset or a processor) according to an embodiment. Also, the memory 270 can include a DPB (decoded picture buffer) and can also be configured by a digital storage medium. The hardware component can further include the memory 270 as an internal / external component.

[0044] The image segmentation unit 210 can divide the input image (or picture, frame) input to the encoding device 200 into one or more processing units. As an example, the processing unit can be called a coding unit (CU). In this case, the coding unit can be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) by a QTBTTT (Quad-tree binary-tree ternary-tree) structure. For example, one coding unit can be divided into multiple coding units with a deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure can be applied first, and then the binary-tree structure and / or the ternary structure can be applied. Or, the binary-tree structure can also be applied first. The coding procedure according to the present disclosure can be performed based on the final coding unit that is no longer divided. In this case, based on the coding efficiency according to the image characteristics, etc., the largest coding unit can be used as the final coding unit, or, if necessary, the coding unit can be recursively divided into coding units with a deeper depth so that the coding unit with the optimal size can be used as the final coding unit. Here, the coding procedure can include procedures such as prediction, transformation, and restoration described later. As another example, the processing unit can further include a prediction unit (PU: Prediction Unit) or a transform unit (TU: Transform Unit). In this case, the prediction unit and the transform unit can each be divided or partitioned from the above-described final coding unit.The prediction unit can be a unit of sample prediction, and the conversion unit can be a unit for deriving a conversion coefficient and / or a unit for deriving a residual signal from the conversion coefficient.

[0045] The unit can, in some cases, be used interchangeably with terms such as a block or an area. In a general case, an M×N block can represent a set such as samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luma component, or can also represent only the pixel / pixel value of the chroma component. A sample can be used as a term corresponding to a pixel or a pel for one picture (or image).

[0046] The encoding device 200 can subtract the predicted signal (predicted block, predicted sample array) output from the inter prediction unit 221 or the intra prediction unit 222 from the input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is transmitted to the conversion unit 232. In this case, as shown in the figure, the unit that subtracts the predicted signal (predicted block, predicted sample array) from the input image signal (original block, original sample array) within the encoder 200 can be called the subtraction unit 231. The prediction unit can perform a prediction on the block to be processed (hereinafter referred to as the current block) and generate a predicted block including predicted samples for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied in units of the current block or CU. The prediction unit can generate various pieces of information related to prediction, such as prediction mode information, and transmit it to the entropy encoding unit 240 as described later in the description of each prediction mode. The information related to prediction can be encoded by the entropy encoding unit 240 and output in the form of a bitstream.

[0047] The intra prediction unit 222 can predict the current block by referring to samples within the current picture. The samples to be referred to can be located in the neighborhood of the current block or at a distance from it depending on the prediction mode. In intra prediction, the prediction mode can include a plurality of non - directional modes and a plurality of directional modes. The non - directional modes can include, for example, the DC mode and the Planar mode. The directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the fineness of the prediction direction. However, this is an example, and more or fewer directional prediction modes can be used depending on the setting. The intra prediction unit 222 can also determine the prediction mode to be applied to the current block using the prediction mode applied to the neighboring blocks.

[0048] The inter prediction unit 221 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the peripheral block and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the peripheral block can include a spatial neighboring block existing in the current picture and a temporal neighboring block existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block can be the same or different. The temporal neighboring block can be called by names such as a collocated reference block and a collocated CU (col CU), and the reference picture including the temporal neighboring block can also be called a collocated picture (colPic). For example, the inter prediction unit 221 can construct a motion information candidate list based on the peripheral block, and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction can be performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the inter prediction unit 221 can use the motion information of the peripheral block as the motion information of the current block. In the case of the skip mode, different from the merge mode, the residual signal may not be transmitted.In the case of the motion information prediction (motion vector prediction, MVP) mode, the motion vector of a neighboring block is used as a motion vector predictor, and the motion vector difference is signaled to indicate the motion vector of the current block.

[0049] The prediction unit 220 can generate a prediction signal based on various prediction methods described later. For example, the prediction unit can apply intra prediction or inter prediction for the prediction of one block, and can also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). In addition, the prediction unit can be based on the intra block copy (IBC) prediction mode or the palette mode for the prediction of a block. The IBC prediction mode or the palette mode can be used for content image / video coding such as games, for example, like SCC (screen content coding). IBC basically performs prediction within the current picture, but can be performed in the same way as inter prediction in terms of deriving a reference block within the current picture. That is, IBC can use at least one of the inter prediction techniques described in this document. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, the sample values within the picture can be signaled based on the information regarding the palette table and the palette index.

[0050] The prediction signal generated via the prediction unit (including the inter prediction unit 221 and / or the intra prediction unit 222) can be used to generate a restored signal or can be used to generate a residual signal. The conversion unit 232 can generate transform coefficients by applying a conversion technique to the residual signal. For example, the conversion technique can include at least one of DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), KLT (Karhunen-Loeve Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT means the conversion obtained from this graph when trying to represent the relationship information between pixels with a graph. CNT means the conversion obtained based on generating a prediction signal using all previously reconstructed pixels. Also, the conversion process can be applied to a pixel block having the same size of a square and can also be applied to a block of variable size that is not square.

[0051] The quantization unit 233 quantizes the transform coefficients and transmits them to the entropy encoding unit 240. The entropy encoding unit 240 can encode the quantized signal (information regarding the quantized transform coefficients) and output it as a bitstream. The information regarding the quantized transform coefficients can be referred to as residual information. The quantization unit 233 can reorder the block-form quantized transform coefficients in a one-dimensional vector form based on the coefficient scan order, and can also generate the information regarding the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. The entropy encoding unit 240 can perform various encoding methods such as, for example, exponential Golomb, CAVLC (context-adaptive variable length coding), CABAC (context-adaptive binary arithmetic coding), etc. The entropy encoding unit 240 can encode, together or separately, information necessary for video / image restoration (e.g., values of syntax elements, etc.) in addition to the quantized transform coefficients. The encoded information (e.g., encoded video / image information) can be transmitted or stored in units of NAL (network abstraction layer) units in the form of a bitstream. The video / image information can further include information regarding various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Also, the video / image information can further include general constraint information. In this document, the information and / or syntax elements transmitted / signaled from the encoding device to the decoding device can be included in the video / image information. The video / image information can be encoded through the above-described encoding procedure and included in the bitstream.The bitstream can be transmitted via a network or stored in a digital storage medium. Here, the network can include a broadcast network and / or a communication network, etc., and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The signal output from the entropy encoding unit 240 can be configured such that a transmission unit (not shown) for transmission and / or a storage unit (not shown) for storage are internal / external elements of the encoding device 200, or the transmission unit can also be included in the entropy encoding unit 240.

[0052] The quantized transform coefficients output from the quantization unit 233 can be used to generate a prediction signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients via the inverse quantization unit 234 and the inverse transformation unit 235, a residual signal (residual block or residual sample) can be restored. The addition unit 155 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the restored residual signal to the prediction signal output from the inter prediction unit 221 or the intra prediction unit 222. When there is no residual for the block to be processed, as in the case where the skip mode is applied, the predicted block can be used as the reconstructed block. The addition unit 250 can be called a restoration unit or a reconstructed block generation unit. The generated reconstructed signal can be used for intra prediction of the next block to be processed within the current picture and, as will be described later, can also be used for inter prediction of the next picture after passing through filtering.

[0053] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied in the picture encoding and / or restoration process.

[0054] The filtering unit 260 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 260 can apply various filtering methods to the restored picture to generate a modified restored picture, and can store the modified restored picture in the memory 270, specifically, in the DPB of the memory 270. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit 260 can generate various information related to filtering and transmit it to the entropy encoding unit 240, as will be described later in the description of each filtering method. The information related to filtering can be encoded by the entropy encoding unit 240 and output in the form of a bitstream.

[0055] The modified restored picture transmitted to the memory 270 can be used as a reference picture in the inter prediction unit 221. The encoding device can avoid prediction mismatches between the encoding device 100 and the decoding device and improve the encoding efficiency when inter prediction is applied through this.

[0056] The DPB of the memory 270 can store the modified restored picture for use as a reference picture in the inter prediction unit 221. The memory 270 can store the motion information of the block where the motion information in the current picture was derived (or encoded) and / or the motion information of the block in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 221 for utilization as the motion information of the spatial neighboring block or the temporal neighboring block. The memory 270 can store the restored samples of the restored blocks in the current picture and transmit them to the intra prediction unit 222.

[0057] FIG. 3 is a diagram schematically illustrating the configuration of a video / image decoding apparatus to which the present disclosure can be applied.

[0058] As shown in FIG. 3, the decoding apparatus 300 can be configured to include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filtering unit 350, and a memory 360. The predictor 330 can include an inter-prediction unit 331 and an intra-prediction unit 332. The residual processor 320 can include a dequantizer 321 and an inverse transformer 321. The entropy decoder 310, the residual processor 320, the predictor 330, the adder 340, and the filtering unit 350 described above can be configured by one hardware component (e.g., a decoder chipset or a processor) according to an embodiment. Also, the memory 360 can include a DPB (decoded picture buffer) and can also be configured by a digital storage medium. The hardware component can further include the memory 360 as an internal / external component.

[0059] If a bitstream including video / image information is input, the decoding device 300 can restore an image corresponding to the process in which the video / image information was processed by the encoding device in FIG. 3. For example, the decoding device 300 can derive units / blocks based on the block division related information obtained from the bitstream. The decoding device 300 can perform decoding using the processing units applied in the encoding device. Therefore, the processing unit for decoding can be, for example, a coding unit, and the coding unit can be divided according to a quad-tree structure, a binary tree structure, and / or a ternary tree structure from a coding tree unit or a maximum coding unit. One or more transform units can be derived from the coding unit. Then, the restored image signal decoded and output via the decoding device 300 can be reproduced via a reproducing device.

[0060] The decoding device 300 can receive the signal output from the encoding device in FIG. 3 in the form of a bitstream, and the received signal can be decoded via the entropy decoding unit 310. For example, the entropy decoding unit 310 can parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information can further include information regarding various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Also, the video / image information can further include general constraint information. The decoding device can further decode a picture based on the information regarding the parameter set and / or the general constraint information. The signaling / received information and / or syntax elements described later in this document can be decoded via the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 310 can decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the value of the syntax element necessary for image restoration and the quantized value of the transform coefficient regarding the residual. More specifically, the CABAC entropy decoding method receives the bin corresponding to each syntax element in the bitstream, determines a context model using the syntax element information to be decoded, the information of the surrounding and decoded blocks to be decoded, or the information of the symbol / bin decoded in the previous step, predicts the occurrence probability of the bin according to the determined context model, performs arithmetic decoding of the bin, and can generate a symbol corresponding to the value of each syntax element. At this time, after determining the context model, the CABAC entropy decoding method can update the context model using the information of the symbol / bin decoded for the context model of the next symbol / bin.Of the information decoded by the entropy decoding unit 310, the information related to prediction is provided to the prediction unit (inter prediction unit 332 and intra prediction unit 331), and the residual value for which entropy decoding is performed by the entropy decoding unit 310, that is, the quantized transform coefficient and related parameter information, can be input to the residual processing unit 320. The residual processing unit 320 can derive a residual signal (residual block, residual sample, residual sample array). Also, of the information decoded by the entropy decoding unit 310, the information related to filtering can be provided to the filtering unit 350. On the other hand, a receiving unit (not shown) that receives a signal output from the encoding device can be further configured as an internal / external element of the decoding device 300, or the receiving unit can be a component of the entropy decoding unit 310. On the other hand, the decoding device according to the present document can be called a video / image / picture decoding device, and the decoding device can be classified into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder can include the entropy decoding unit 310, and the sample decoder can include at least one of the inverse quantization unit 321, inverse transform unit 322, addition unit 340, filtering unit 350, memory 360, inter prediction unit 332, and intra prediction unit 331.

[0061] In the inverse quantization unit 321, the quantized transform coefficient can be inverse quantized to output a transform coefficient. The inverse quantization unit 321 can reorder the quantized transform coefficients in a two-dimensional block form. In this case, the reordering can be performed based on the coefficient scan order performed by the encoding device. The inverse quantization unit 321 can perform inverse quantization on the quantized transform coefficient using a quantization parameter (for example, quantization step size information) to obtain a transform coefficient.

[0062] In the inverse conversion unit 322, the conversion coefficients are inversely converted to obtain a residual signal (residual block, residual sample array).

[0063] The prediction unit can perform prediction on the current block and generate a predicted block including predicted samples for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied to the current block based on the information regarding the prediction output from the entropy decoding unit 310, and can determine a specific intra / inter prediction mode.

[0064] The prediction unit 320 can generate a prediction signal based on various prediction methods described later. For example, for prediction of one block, the prediction unit can apply not only intra prediction or inter prediction, but also can apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). Also, for prediction of a block, the prediction unit can be based on the intra block copy (IBC) prediction mode, or can be based on the palette mode. The IBC prediction mode or the palette mode can be used for content image / video coding such as games, for example, like SCC (screen content coding). IBC basically performs prediction within the current picture, but can be performed in the same way as inter prediction in terms of deriving a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this document. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, information regarding the palette table and palette index can be included in and signaled in the video / image information.

[0065] The Intra prediction unit 331 can predict the current block by referring to samples within the current picture. The samples to be referred can be located in the neighborhood (neighbor) of the current block according to the prediction mode, or can be located remotely. In intra prediction, the prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The Intra prediction unit 331 can also determine the prediction mode to be applied to the current block by using the prediction mode applied to the neighboring blocks.

[0066] The Inter prediction unit 332 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on the reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the neighboring blocks can include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. For example, the Inter prediction unit 332 can construct a motion information candidate list based on the neighboring blocks, and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction can be performed based on various prediction modes, and the information related to the prediction can include information indicating the mode of inter prediction for the current block.

[0067] The adder 340 can generate a restored signal (restored picture, restored block, restored sample array) by adding the obtained residual signal to the predicted signal (predicted block, predicted sample array) output from the prediction unit (including the inter prediction unit 332 and / or the intra prediction unit 331). When there is no residual for the processing target block, as in the case where the skip mode is applied, the predicted block can be used as the restored block.

[0068] The adder 340 can be referred to as a restoration unit or a restored block generation unit. The generated restored signal can be used for intra prediction of the next processing target block in the current picture, can be output after filtering as will be described later, or can also be used for inter prediction of the next picture.

[0069] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied during the picture decoding process.

[0070] The filtering unit 350 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 350 can apply various filtering methods to the restored picture to generate a modified restored picture, and can transmit the modified restored picture to the memory 360, specifically, to the DPB of the memory 360. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.

[0071] The (corrected) restored picture stored in the DPB of the memory 360 can be used as a reference picture in the inter prediction unit 332. The memory 360 can store the motion information of the block from which the motion information in the current picture was derived (or decoded) and / or the motion information of the blocks in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 260 for utilization as the motion information of spatially neighboring blocks or temporally neighboring blocks. The memory 360 can store the restored samples of the restored blocks in the current picture and can transmit them to the intra prediction unit 331.

[0072] In this specification, the embodiments described in the filtering unit 260, the inter prediction unit 221, and the intra prediction unit 222 of the encoding device 100 can be applied to the filtering unit 350, the inter prediction unit 332, and the intra prediction unit 331 of the decoding device 300 in the same or corresponding manner, respectively.

[0073] As described above, prediction is performed to increase the compression efficiency when performing video coding. Thereby, a predicted block including prediction samples for the current block, which is the block to be coded, can be generated. Here, the predicted block includes prediction samples in the spatial domain (or pixel domain). The predicted block is derived in the same way in the encoding device and the decoding device, and the encoding device can increase the image coding efficiency by signaling information (residual information) regarding the residual between the original block and the predicted block, which is not the original sample value of the original block, to the decoding device. The decoding device can derive a residual block including residual samples based on the residual information, add the residual block and the predicted block to generate a restored block including restored samples, and generate a restored picture including the restored block.

[0074] The residual information can be generated through the conversion and quantization procedures. For example, an encoding device can derive a residual block between the original block and the predicted block, execute a conversion procedure on the residual samples (residual sample array) included in the residual block to derive conversion coefficients, execute a quantization procedure on the conversion coefficients to derive quantized conversion coefficients, and thus signal the related residual information (via a bitstream) to a decoding device. Here, the residual information can include information such as the value information, position information, conversion technique, conversion kernel, and quantization parameter of the quantized conversion coefficients. The decoding device can execute an inverse quantization / inverse conversion procedure based on the residual information to derive residual samples (or a residual block). The decoding device can generate a restored picture based on the predicted block and the residual block. Also, the encoding device can inverse quantize / inverse convert the quantized conversion coefficients for reference in the inter prediction of subsequent pictures to derive a residual block and generate a restored picture based on this.

[0075] In one embodiment, to control the motion prediction based on sub - blocks, a sub - block TMVP flag indicating whether a sub - block - based temporal motion vector predictor can be used can be employed. The sub - block TMVP flag can be signaled at the SPS (Sequence Parameter Set) level and can control the on / off of the motion prediction based on sub - blocks. The sub - block TMVP flag can be referred to as sps_sbtmvp_enabled_flag, for example, as shown in Table 1 below.

[0076] Also, in order to control the affine motion prediction method, an affine flag indicating whether affine prediction can be applied to the current block can be used. The affine flag can be signaled at the SPS level and can control the on / off of affine prediction. The affine flag can be referred to as sps_affine_enabled_flag, for example, as shown in Table 1 below. When the value of the affine flag is 1, an affine type flag can be additionally signaled to additionally determine whether to use 6-parameter affine prediction.

[0077] An example of the syntax signaled at the SPS level is as shown in Table 1 below.

[0078]

Table 1-1

[0079]

Table 1-2

[0080] In one embodiment, in low level coding syntax, as shown in Table 2 below, if the merge flag of the current coding block is 1, when the affine flag of the SPS is 1, based on the conditions of the current block (e.g., block size, block shape, etc.), a flag (e.g., the merge affine flag) can be signaled to indicate whether affine merge or normal merge is applicable to the current block. The merge affine flag can be represented as, for example, merge_affine_flag. In an example, when the value of the affine flag signaled at the SPS level is 0 and the value of the merge flag signaled at the coding unit level is 1, it can be determined that normal merge is applicable to the current block without signaling additional syntax elements.

[0081] An example of the syntax signaled at the coding unit level is as shown in Table 2 below.

[0082] [Table 2-1]

[0083] [Table 2-2]

[0084] [Table 2-3]

[0085] [Table 2-4]

[0086] On one hand, when the high-level syntax design in Table 1 and the low-level syntax design in Table 2 are applied, if the ATMVP is used as an affine merge candidate, design problems, logical problems, conceptual problems, etc. may occur. In one example, when the value of the affine flag signaled at the SPS level is 0 and the value of the sub-block TMVP flag signaled at the SPS level is 1, even though it is signaled to use the ATMVP in the SPS, the ATMVP candidate may not be used as any candidate. In addition to the design problems and logical problems as described above, conceptual problems may also exist. The ATMVP is a motion prediction method based on sub-blocks (in one example, SubPu). For the purpose of distinguishing between the motion prediction candidates of the non-sub-block base (in one example, non-SubPu base) and the motion prediction candidates of the sub-block base in the normal merge, by using it as a candidate for the affine merge mode that performs prediction on the sub-block base, for the merge of the current block, there is a purpose to distinguish whether it is a sub-block merge or a non-sub-block merge. However, despite such a purpose, the low-level syntax design according to Table 2 controls the sub-block ATMVP depending on whether the affine merge is used or not.

[0087] To complement the above-described design problems, logical problems, and conceptual problems, in one embodiment, a high-level and / or low-level syntax design based on at least one of Tables 3 to 11 below can be provided.

[0088] In one embodiment, a flag for controlling the motion prediction of the sub-block base can be signaled at the SPS level. The flag for controlling the motion prediction of the sub-block base can be represented as, for example, sps_subpumvp_enabled_flag, and can be used to determine whether the motion prediction of the sub-block base can be turned on / off. When the value of the sps_subpumvp_enabled_flag is 1, the affine_enabled_flag and the sbtmvp_enabled_flag can be signaled as shown in Table 3 below.

[0089] [Table 3]

[0090] When using the SPS-level syntax design of Table 3, the availability of affine prediction and ATMVP can be represented as shown in Table 4 below. In Table 4 below, 1 indicates that the method is available, and 0 indicates that the method is not available.

[0091] [Table 4]

[0092] In one embodiment, a high-level syntax design for jointly controlling the availability of affine prediction and ATMVP based on the sps_subpumvp_enabled_flag can be provided. When following this embodiment, in one example, if the value of the sps_subpumvp_enabled_flag is 1, it can be determined that both affine prediction and ATMVP are available. The high-level syntax design according to this embodiment can be as shown in Table 5 below.

[0093] [Table 5]

[0094] In one embodiment, based on the sps_subpumvp_enabled_flag included in the high-level syntax according to Table 5 above, both the availability of affine prediction and ATMVP are controlled. However, in order to finely control the availability of ATMVP for each slice unit, a method of using the slice_subpumvp_enabled_flag in the slice header syntax may be provided. The syntax at the slice header level according to this embodiment can be, for example, as shown in Table 6 below.

[0095]

Table 6

[0096] In one embodiment, when the affine prediction method is not used and the sps_sbtmvp_enabled_flag is 1, although the merge_affine_flag is signaled, the affine candidate is not configured as a candidate, and a method of configuring only ATMVP as a candidate may be provided. An example of the low-level syntax for representing this embodiment can be as shown in Table 7 below.

[0097]

Table 7

[0098] In Table 7 above, when the value of sps_affine_enabled_flag is 1 or the value of sps_sbtmvp_enabled_flag is 1, it can be determined that the merge_affine_flag, which represents the applicability of the merge affine mode, is decoded.

[0099] In one example, when the value of sps_affine_enabled_flag is 1 or the value of sps_sbtmvp_enabled_flag is 1, it can be determined to decode the merge sub-block flag merge_subblock_flag indicating whether the merge sub-block mode can be applied. In the merge sub-block mode, merge candidates can be determined based on sub-block units.

[0100] In Table 7 above, when the width (cbWidth) and height (cbHeight) of the current block are each 8 or more and the value of sps_affine_enabled_flag is 1 or the value of sps_sbtmvp_enabled_flag is 1, it can be determined to decode the merge affine flag merge_affine_flag.

[0101] In one example, when the maximum number of merge candidates of the sub-block of the current block is greater than 0, it can be determined to decode the already determined merge mode flag.

[0102] In one example, when the value of the affine flag is 1 or the value of the sub-block TMVP flag is 1, the maximum number of merge candidates of the sub-block of the current block can be greater than 0.

[0103] In one example, whether to decode the already determined merge mode flag or not can be determined based on whether if(MaxNumSubblockMergeCand>0&&cbWidth>=8&&cbHeight>=8) is satisfied. MaxNumSubblockMergeCand represents the maximum number of merge candidates of the sub-block, cbWidth represents the width of the current block, and cbHeight represents the height of the current block.

[0104] In Table 7 above, when the value of sps_affine_enabled_flag is 0 and the value of sps_sbtmvp_enabled_flag is 1, merge_affine_idx is not signaled and can be inferred as 0. When following the embodiment of Table 7, the availability of affine prediction and ATMVP can be represented as shown in Table 8 below.

[0105]

Table 8

[0106] In one embodiment, when the affine prediction method is not used and the value of sps_sbtmvp_enabled_flag is 1, a method of controlling so that ATMVP is used as a normal merge candidate may be provided. When following this embodiment, the availability of affine prediction and ATMVP can be represented as shown in Table 9 below.

[0107]

Table 9

[0108] In one embodiment, a method of designing the high-level syntax to signal sps_sbtmvp_enabled_flag only when the value of affine_enabled_flag is 1 may be provided. This can be considered in view of the structure of the low-level coding tool designed so that ATMVP is used as an affine merge candidate and cannot be used when the value of sps_affine_enabled_flag is 0. An exemplification of the high-level syntax according to this embodiment is as shown in Table 10 below.

[0109]

Table 10

[0110] When using the SPS level syntax design of Table 10 according to Table 10, the availability of affine prediction and ATMVP can be represented as shown in Table 11 below.

[0111]

Table 11

[0112] FIG. 4 is a flowchart showing the operation of an encoding device according to an embodiment, and FIG. 5 is a block diagram showing the configuration of an encoding device according to an embodiment.

[0113] The encoding device according to FIGS. 4 and 5 can perform corresponding operations with the decoding device according to FIGS. 6 and 7. Therefore, the operations of the decoding device described later in FIGS. 6 and 7 can be similarly applied to the encoding device according to FIGS. 4 and 5.

[0114] Each step disclosed in FIG. 4 can be performed by the encoding device 200 disclosed in FIG. 2. More specifically, S400 and S410 can be performed by the prediction unit 220 disclosed in FIG. 2, and S420 can be performed by the entropy encoding unit 240 disclosed in FIG. 2. Further, the operations according to S400 to S420 are based on a part of the content described above in FIG. 3. Therefore, specific content overlapping with the content described above in FIGS. 2 and 3 is omitted or simplified in the description.

[0115] As shown in FIG. 5, an encoding device according to an embodiment can include a prediction unit 220 and an entropy encoding unit 240. However, in some cases, not all of the components shown in FIG. 5 are essential components of the encoding device, and the encoding device can be realized with more or fewer components than those shown in FIG. 5.

[0116] In an encoding apparatus according to an embodiment, the prediction unit 220 and the entropy encoding unit 240 may be realized by separate chips (chips), or at least two or more components may be realized via one chip.

[0117] An encoding apparatus according to an embodiment can determine whether affine prediction can be applied to a current block and whether a temporal motion vector predictor based on sub-blocks of the current block can be used (S400). More specifically, the prediction unit 220 of the encoding apparatus can determine whether affine prediction can be applied to the current block and whether a temporal motion vector predictor based on sub-blocks of the current block can be used.

[0118] An encoding apparatus according to an embodiment can determine whether to encode a determined merge mode flag indicating whether to apply a predetermined merge mode to the current block based on the determination as to whether affine prediction can be applied to the current block and whether a temporal motion vector predictor based on sub-blocks of the current block can be used (S410). More specifically, the prediction unit 220 of the encoding apparatus can determine whether to encode a determined merge mode flag indicating whether to apply a predetermined merge mode to the current block based on the determination as to whether affine prediction can be applied to the current block and whether a temporal motion vector predictor based on sub-blocks of the current block can be used.

[0119] In one example, the predetermined merge mode can be a merge affine mode or a merge sub-block mode, and the determined merge mode flag can be a merge affine flag or a merge sub-block flag. The merge affine flag can be represented as merge_affine_flag, and the merge sub-block flag can be represented as merge_subblock_flag.

[0120] An encoding device according to an embodiment can apply the affine prediction to the current block based on a determination as to whether to encode the already determined merge mode flag, an affine flag indicating whether the time motion vector predictor based on the sub-block of the current block can be used, and a sub-block TMVP flag indicating whether the already determined merge mode flag can be encoded (S420). More specifically, an entropy encoding unit 240 of the encoding device can, based on the determination as to whether to encode the already determined merge mode flag, an affine flag indicating whether the affine prediction can be applied to the current block, a sub-block TMVP flag indicating whether the time motion vector predictor based on the sub-block of the current block can be used, and the already determined merge mode flag can be encoded.

[0121] In one embodiment, when the value of the affine flag is 1 or the value of the sub-block TMVP flag is 1, it can be determined that the already determined merge mode flag is to be encoded.

[0122] In one embodiment, when the width and height of the current block are each 8 or more and a first condition that the value of the affine flag is 1 is satisfied, or a second condition that the value of the sub-block TMVP flag is 1 is satisfied, it can be determined that the already determined merge mode flag is to be encoded.

[0123] In one embodiment, whether to encode the already determined merge mode flag can be determined based on the following mathematical formula 1.

[0124]

Equation

[0125] In the formula 1, sps_affine_enabled_flag represents the affine flag, cbWidth represents the width of the current block, cbHeight represents the height of the current block, and sps_sbtmvp_enabled_flag can represent the sub-block TMVP flag.

[0126] In one embodiment, the determined merge mode flag can be a merge affine flag indicating whether an affine merge mode is applied to the current block or a merge sub-block flag indicating whether a merge mode is applied in units of sub-blocks of the current block.

[0127] In one embodiment, when the maximum number of merge candidates of the sub-blocks of the current block is greater than 0, it can be determined that the determined merge mode flag is encoded.

[0128] In one embodiment, it can be characterized in that when the value of the affine flag is 1 or the value of the sub-block TMVP flag is 1, the maximum number of merge candidates of the sub-blocks of the current block is greater than 0.

[0129] In one embodiment, whether to encode the determined merge mode flag can be determined based on the following formula 2.

[0130]

Equation

[0131] In the formula 2, MaxNumSubblockMergeCand represents the maximum number of merge candidates of the sub-blocks, cbWidth represents the width of the current block, and cbHeight represents the height of the current block.

[0132] According to the encoding apparatus and the operation method of the encoding apparatus in FIGS. 4 and 5, the encoding apparatus determines whether affine prediction can be applied to the current block and whether a temporal motion vector predictor based on sub-blocks of the current block can be used (S400), and based on the determination regarding whether affine prediction can be applied to the current block and whether the temporal motion vector predictor based on the sub-blocks of the current block can be used, determines whether to encode a pre-determined merge mode flag indicating whether to apply the pre-determined merge mode to the current block (S410), and based on the determination regarding whether to encode the pre-determined merge mode flag, an affine flag indicating whether affine prediction can be applied to the current block, a sub-block TMVP flag indicating whether the temporal motion vector predictor based on the sub-blocks of the current block can be used, and the pre-determined merge mode flag are encoded (S420), and it can be characterized in that when the value of the affine flag is 1 or the value of the sub-block TMVP flag is 1, it is determined that the pre-determined merge mode flag is to be encoded. That is, based on the affine flag and the sub-block TMVP flag, by determining whether to decode a pre-determined merge mode flag indicating whether to apply the pre-determined merge mode to the current block, the image coding efficiency can be improved.

[0133] FIG. 6 is a flowchart showing the operation of a decoding apparatus according to an embodiment, and FIG. 7 is a block diagram showing the configuration of a decoding apparatus according to an embodiment.

[0134] Each step disclosed in FIG. 6 can be performed by the decoding device 300 disclosed in FIG. 3. More specifically, S600 and S610 can be performed by the entropy decoding unit 310 disclosed in FIG. 3, S620 can be performed by the prediction unit 330 disclosed in FIG. 3, and S630 can be performed by the addition unit 340 disclosed in FIG. 3. Further, the operations by S600 to S630 are based on a part of the content described above in FIG. 3. Therefore, specific content overlapping with the content described above in FIG. 3 is omitted or simplified in the description.

[0135] As shown in FIG. 7, a decoding device according to an embodiment may include an entropy decoding unit 310, a prediction unit 330, and an addition unit 340. However, in some cases, not all of the components shown in FIG. 7 may be essential components of the decoding device, and the decoding device may be implemented with more or fewer components than those shown in FIG. 7.

[0136] In a decoding device according to an embodiment, the entropy decoding unit 310, the prediction unit 330, and the addition unit 340 may each be implemented on a separate chip, or at least two or more components may be implemented via one chip.

[0137] A decoding device according to an embodiment can decode an affine flag indicating whether affine prediction can be applied to a current block and a sub-block TMVP flag indicating whether a temporal motion vector predictor based on a sub-block of the current block can be used, based on a bitstream (S600). More specifically, the entropy decoding unit 310 of the decoding device can decode an affine flag indicating whether affine prediction can be applied to a current block and a sub-block TMVP flag indicating whether a temporal motion vector predictor based on a sub-block of the current block can be used, based on a bitstream.

[0138] In one example, the affine flag can be represented as sps_affine_enabled_flag, and the sub-block TMVP flag can be represented as sps_sbtmvp_enabled_flag. The sub-block TMVP flag can also be referred to as the sub-PU TMVP flag in some cases.

[0139] In one example, the affine flag and the sub-block TMVP flag can be signaled at the SPS level.

[0140] A decoding device according to an embodiment can determine whether to decode a determined merge mode flag indicating whether to apply a predetermined merge mode to the current block based on the decoded affine flag and the decoded sub-block TMVP flag (S610). More specifically, the entropy decoding unit 310 of the decoding device can determine whether to decode a determined merge mode flag indicating whether to apply a predetermined merge mode to the current block based on the decoded affine flag and the decoded sub-block TMVP flag.

[0141] In one example, the predetermined merge mode can be a merge affine mode or a merge sub-block mode, and the determined merge mode flag can be a merge affine flag or a merge sub-block flag. The merge affine flag can be represented as merge_affine_flag, and the merge sub-block flag can be represented as merge_subblock_flag.

[0142] A decoding device according to an embodiment can derive a predicted sample for the current block based on the determination of whether to decode the determined merge mode flag (S620). More specifically, the prediction unit 330 of the decoding device can derive a predicted sample for the current block based on the determination of whether to decode the determined merge mode flag.

[0143] The decoding device according to one embodiment can derive the prediction mode applied to the current block based on the determination as to whether or not to decode the already determined merge mode flag, and can derive a prediction sample for the current block based on the derived prediction mode.

[0144] The decoding device according to one embodiment can generate a restored sample for the current block based on the prediction sample for the current block (S630). More specifically, the adding unit 340 of the decoding device can generate a restored sample for the current block based on the prediction sample for the current block.

[0145] In one embodiment, when the value of the affine flag is 1 or the value of the sub-block TMVP flag is 1, it can be determined to decode the already determined merge mode flag.

[0146] In one example, when the value of sps_affine_enabled_flag is 1 or the value of sps_sbtmvp_enabled_flag is 1, it can be determined to decode the already determined merge mode flag.

[0147] In another example, when the value of sps_affine_enabled_flag is 1 or the value of sps_sbtmvp_enabled_flag is 1, it can be determined to decode the merge affine flag merge_affine_flag.

[0148] In yet another example, when the value of sps_affine_enabled_flag is 1 or the value of sps_sbtmvp_enabled_flag is 1, it can be determined to decode the merge sub-block flag merge_subblock_flag.

[0149] In one embodiment, when the width and height of the current block are each 8 or more and satisfy a first condition that the value of the affine flag is 1, or when a second condition that the value of the sub-block TMVP flag is 1 is satisfied, it can be determined to decode the already determined merge mode flag.

[0150] In one embodiment, whether to decode the already determined merge mode flag can be determined based on the following Equation 3.

[0151]

Equation

[0152] In Equation 3, sps_affine_enabled_flag represents the affine flag, cbWidth represents the width of the current block, cbHeight represents the height of the current block, and sps_sbtmvp_enabled_flag can represent the sub-block TMVP flag.

[0153] In one embodiment, when the maximum number of merge candidates of the sub-blocks of the current block is greater than 0, it can be determined to decode the already determined merge mode flag.

[0154] In one embodiment, when the value of the affine flag is 1 or the value of the sub-block TMVP flag is 1, the maximum number of merge candidates of the sub-blocks of the current block can be greater than 0.

[0155] In one embodiment, whether to decode the already determined merge mode flag can be determined based on the following Equation 4.

[0156]

Equation

[0157] In the above formula 4, MaxNumSubblockMergeCand represents the maximum number of merge candidates of the sub-block, cbWidth can represent the width of the current block, and cbHeight can represent the height of the current block.

[0158] According to the decoding device and the operation method of the decoding device disclosed in FIGS. 6 and 7, the decoding device can decode an affine flag indicating whether affine prediction can be applied to a current block based on a bit stream and a sub-block TMVP flag indicating whether a temporal motion vector predictor based on sub-blocks of the current block can be used (S600), and based on the decoded affine flag and the decoded sub-block TMVP flag, determine whether to decode a predetermined merge mode flag indicating whether to apply the predetermined merge mode to the current block (S610), based on the determination regarding whether to decode the predetermined merge mode flag, derive prediction samples for the current block (S620), generate restored samples for the current block based on the prediction samples for the current block (S630), and it can be characterized in that when the value of the affine flag is 1 or the value of the sub-block TMVP flag is 1, it is determined to decode the predetermined merge mode flag. That is, by determining whether to decode a predetermined merge mode flag indicating whether to apply a predetermined merge mode to a current block based on the affine flag and the sub-block TMVP flag, the image coding efficiency can be improved.

[0159] In the above-described embodiments, the method has been described based on a sequence diagram as a series of steps or blocks. However, the present disclosure is not limited to the order of the steps, and a certain step can occur in an order different from that of the steps described above or simultaneously with different steps. Also, those skilled in the art will understand that the steps shown in the sequence diagram are not exclusive, and other steps may be included, or one or more steps of the sequence diagram may be deleted without affecting the scope of the present disclosure.

[0160] The method according to the present disclosure described above can be realized in software form, and the encoding device and / or decoding device according to the present disclosure can be included in devices that perform image processing such as TVs, computers, smartphones, set-top boxes, and display devices.

[0161] When embodiments and the like in the present disclosure are realized by software, the above-described method can be realized by modules (processes, functions, etc.) that perform the above-described functions. The modules can be stored in a memory and executed by a processor. The memory can be inside or outside the processor and can be connected to the processor by various well-known means. The processor can include an ASIC (application-specific integrated circuit), other chip sets, logic circuits, and / or data processing devices. The memory can include a ROM (read-only memory), a RAM (random access memory), a flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments and the like described in the present disclosure can be realized and performed on a processor, a microprocessor, a controller, or a chip. For example, the functional units illustrated in each drawing can be realized and performed on a computer, a processor, a microprocessor, a controller, or a chip. In this case, information for realization (for example, information on instructions) or an algorithm can be stored in a digital storage medium.

[0162] In addition, the decoding device and the encoding device to which the present disclosure is applied can be included in a multimedia broadcast transceiver, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video conferencing device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camcorder, an on-demand video (VoD) service providing device, an OTT video (Over the top video) device, an Internet streaming service providing device, a three-dimensional (3D) video device, a VR (virtual reality) device, an AR (augmented reality) device, a videophone video device, a transportation means terminal (e.g., a vehicle (including an autonomous driving vehicle) terminal, an airplane terminal, a ship terminal, etc.), and a medical video device, etc., and can be used to process a video signal or a data signal. For example, in an OTT video (Over the top video) device, it can include a game console, a Blu-ray player, an Internet-connected TV, a home theater system, a smartphone, a tablet PC, a DVR (Digital Video Recorder), etc.

[0163] In addition, the processing method to which the present disclosure is applicable can be produced in the form of a program executed by a computer and can be stored in a computer-readable recording medium. Multimedia data having the data structure according to the present disclosure can also be stored in a computer-readable recording medium. The computer-readable recording medium includes all kinds of storage devices and distributed storage devices in which data readable by a computer is stored. The computer-readable recording medium can include, for example, Blu-ray Disc (BD), Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disk, and optical data storage devices. Further, the computer-readable recording medium includes a medium realized in the form of a carrier wave (for example, transmission via the Internet). Also, a bit stream generated by an encoding method can be stored in a computer-readable recording medium or transmitted via a wired or wireless communication network.

[0164] In addition, embodiments of the present disclosure can be realized by a computer program product with program code, and the program code can be executed by a computer according to embodiments of the present disclosure. The program code can be stored on a carrier readable by a computer.

[0165] FIG. 8 shows an example of a content streaming system to which the disclosure of this document can be applied.

[0166] As shown in FIG. 8, the content streaming system to which the present disclosure is applicable can generally include an encoding server, a streaming server, a web server, a media repository, a user device, and a multimedia input device.

[0167] The encoding server compresses the content input from a multimedia input device such as a smartphone, camera, camcorder, etc. into digital data to generate a bitstream, and serves to transmit this to the streaming server. As another example, when a multimedia input device such as a smartphone, camera, camcorder, etc. directly generates a bitstream, the encoding server can be omitted.

[0168] The bitstream can be generated by an encoding method or a bitstream generation method to which the present disclosure is applied, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.

[0169] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server serves as a medium for informing the user of what services are available. If the user requests a desired service from the web server, the web server transmits this to the streaming server, and the streaming server transmits multimedia data to the user. At this time, the content streaming system can include another control server, and in this case, the control server serves to control commands / responses between each device in the content streaming system.

[0170] The streaming server can receive content from a media repository and / or an encoding server. For example, when it comes to receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.

[0171] In the example of the user device, there can be a mobile phone, a smart phone, a laptop computer, a digital broadcast terminal, a PDA (personal digital assistants), a PMP (portable multimedia player), a navigation device, a slate PC, a tablet PC, an ultrabook, a wearable device (for example, a smartwatch, a smart glass, an HMD (head mounted display)), a digital TV, a desktop computer, a digital signage, etc.

[0172] Each server in the content streaming system can be operated as a distributed server, and in this case, the data received by each server can be distributedly processed.

Claims

1. In an image decoding method performed by a decoding device, receiving image information including affine enable flag information and sub-block temporal motion vector prediction enable flag information; receiving specific flag information related to whether a sub-block based specific merge mode is applied to a current block; determining whether to receive a specific merge index for the sub-block based specific merge mode based on the specific flag information, the affine enable flag information, and the sub-block temporal motion vector prediction enable flag information; deriving a predicted sample for the current block by applying inter prediction to the current block based on the specific flag information; generating a reconstructed sample based on the predicted sample, wherein the step of receiving the specific flag information includes determining whether to receive the specific flag information based on at least one of the affine enable flag information and the sub-block temporal motion vector prediction enable flag information; and receiving the specific flag information based on a result of determining whether to receive the specific flag information; wherein it is determined that the specific merge index is not received based on a case where the value of the specific flag information is equal to 1, the value of the affine enable flag information is equal to 0, and the sub-block temporal motion vector prediction enable flag information is equal to 1; the sub-block temporal motion vector prediction enable flag information is configured in a sequence parameter set, an image decoding method.

2. An image encoding method performed by an encoding device, deriving affine enable flag information and sub-block temporal motion vector prediction enable flag information; deriving a predicted sample for the current block by applying inter prediction to the current block; deriving specific flag information related to whether a sub-block based specific merge mode is applied to the current block, Determining whether to signal a specific merge index for the sub-block based specific merge mode based on the specific flag information, the affine enable flag information, and the sub-block temporal motion vector prediction enable flag information; Encoding image information including at least one of the affine enable flag information, the sub-block temporal motion vector prediction enable flag information, the specific flag information, and the specific merge index; The step of deriving the specific flag information includes: Determining whether to signal the specific flag information based on at least one of the affine enable flag information and the sub-block temporal motion vector prediction enable flag information; Deriving the specific flag information based on the result of determining whether to signal the specific flag information; Based on the case where the value of the specific flag information is equal to 1, the value of the affine enable flag information is equal to 0, and the sub-block temporal motion vector prediction enable flag information is equal to 1, it is determined that the specific merge index is not signaled; The sub-block temporal motion vector prediction enable flag information is an image encoding method configured in a sequence parameter set.

3. A method for transmitting data for an image, comprising: Obtaining a bitstream, which is a step of: The bitstream is generated by performing: deriving affine enable flag information and sub-block temporal motion vector prediction enable flag information; deriving prediction samples for the current block by applying inter prediction to the current block; deriving specific flag information related to whether a sub-block based specific merge mode is applied to the current block; determining whether to signal a specific merge index for the sub-block based specific merge mode based on the specific flag information, the affine enable flag information, and the sub-block temporal motion vector prediction enable flag information; and encoding image information including at least one of the affine enable flag information, the sub-block temporal motion vector prediction enable flag information, the specific flag information, and the specific merge index to generate the bitstream, step; including the step of transmitting the data including the bitstream; Deriving the specific flag information includes: determining whether to signal the specific flag information based on at least one of the affine enable flag information and the sub-block temporal motion vector prediction enable flag information; deriving the specific flag information based on the result of determining whether to signal the specific flag information; when the value of the specific flag information is equal to 1, the value of the affine enable flag information is equal to 0, and the sub-block temporal motion vector prediction enable flag information is equal to 1, it is determined that the specific merge index is not signaled; The sub-block temporal motion vector prediction enable flag information is configured in a sequence parameter set, method.

Citation Information

Patent Citations

  • Affine motion prediction for video coding

    US20170332095A1