A method for encoding / decoding video, a method for transmitting a bitstream, and a recording medium storing a bitstream.

The video encoding/decoding method enhances efficiency by applying geometric partitioning modes with compensation, addressing the high-resolution video transmission and storage challenges through improved prediction accuracy and residual signal signaling.

JP2026510382APending Publication Date: 2026-04-02LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-13
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

The increasing demand for high-resolution, high-quality video has led to higher transmission and storage costs due to the increased amount of information, necessitating highly efficient video compression technology.

Method used

A video encoding/decoding method that applies geometric partitioning modes with compensation, deriving compensation information based on division direction and offset, and signaling whether compensation applies, along with a method for transmitting and storing the generated bitstream.

Benefits of technology

Improves encoding/decoding efficiency, prediction accuracy, and residual signal signaling by compensating for reference blocks, enabling efficient storage and transmission of high-resolution video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026510382000001_ABST
    Figure 2026510382000001_ABST
Patent Text Reader

Abstract

A video encoding / decoding method, a bitstream transmission method, and a computer-readable recording medium for storing a bitstream are provided. The video decoding method according to this disclosure is a video decoding method performed by a video decoding device and may include the steps of: acquiring at least one motion information for a current block to which a Geometry Partitioning Mode (GPM) is applied; inducing compensation information related to compensation of the reference block based on the peripheral region of at least one reference block indicated by the motion information and the peripheral region of the current block; and generating a predicted sample of the current block based on the compensation information.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This disclosure relates to a video encoding / decoding method, a bitstream transmission method, and a recording medium storing a bitstream, and more particularly to a method for compensating a reference block used in a geometric partitioning mode. [Background technology]

[0002] In recent years, demand for high-resolution, high-quality video, such as HD (High Definition) and UHD (Ultra High Definition) video, has been increasing in various fields. The higher the resolution and quality of video data, the greater the amount of information or bits transmitted compared to existing video data. This increase in the amount of information or bits transmitted leads to increased transmission and storage costs.

[0003] Therefore, highly efficient video compression technology is desired to effectively transmit, store, and play back high-resolution, high-quality video information. [Overview of the project] [Problems that the invention aims to solve]

[0004] The purpose of this disclosure is to provide a video encoding / decoding method and apparatus with improved encoding / decoding efficiency.

[0005] Furthermore, this disclosure aims to provide a method for performing geometric partitioning modes by applying compensation.

[0006] Furthermore, this disclosure aims to provide a method for inducing compensation information used for compensation purposes.

[0007] Furthermore, this disclosure aims to provide a method for deriving compensation information based on the division direction and / or division offset of a geometric division mode.

[0008] Furthermore, this disclosure aims to provide a method for signaling information regarding whether or not compensation applies.

[0009] Furthermore, this disclosure aims to provide a non-temporary computer-readable recording medium for storing a bitstream generated by the video encoding method relating to this disclosure.

[0010] Furthermore, this disclosure aims to provide a non-temporary computer-readable recording medium that stores a bitstream received and decoded by the video decoding device relating to this disclosure and used for video restoration.

[0011] Furthermore, this disclosure aims to provide a method for transmitting a bitstream generated by the video encoding method relating to this disclosure.

[0012] The technical challenges addressed in this disclosure are not limited to those mentioned above, and other technical challenges not mentioned above will be clearly understood by those with ordinary skill in the art to which this disclosure pertains from the following description. [Means for solving the problem]

[0013] A video decoding method according to one aspect of the present disclosure is a video decoding method performed by a video decoding device, which includes (composes; constructs; sets up; includes; incorporates; contains; has; has; includes; has; has; has; has; has; has; has; has; has; has; has; has; has; has; has; has; has; has; has; has; has a Geometry Partitioning Mode (GPM) applied

[0014] Other aspects of the present disclosure include a video encoding method performed by a video encoding device, which includes the steps of: determining at least one reference block for predicting a current block to which a Geometry Partitioning Mode (GPM) is applied; deriving compensation information related to the compensation of the reference block based on the peripheral region of the reference block and the peripheral region of the current block; and generating a prediction sample of the current block based on the compensation information.

[0015] A computer-readable recording medium relating to yet another aspect of this disclosure may store a bitstream generated by a video encoding method or apparatus of this disclosure.

[0016] A transmission method relating to yet another aspect of the present disclosure may transmit a bitstream generated by the video encoding method or apparatus of the present disclosure.

[0017] The features briefly summarized above are merely illustrative examples of the detailed description of the disclosure described below and do not limit the scope of the disclosure. [Effects of the Invention]

[0018] According to this disclosure, it is possible to provide a video encoding / decoding method and apparatus with improved encoding / decoding efficiency.

[0019] Furthermore, according to this disclosure, the accuracy of predictions can be improved by compensating for reference blocks.

[0020] Furthermore, according to this disclosure, the efficiency of residual signal signaling can be improved by increasing the prediction accuracy through compensation.

[0021] Furthermore, this disclosure makes it possible to provide a non-temporary computer-readable recording medium for storing a bitstream generated by the video encoding method relating to this disclosure.

[0022] Also, according to the present disclosure, a non-transitory computer-readable recording medium for storing a bitstream received and decoded by the video decoding apparatus according to the present disclosure and used for video restoration can be provided.

[0023] Also, according to the present disclosure, a method for transmitting a bitstream generated by a video encoding method can be provided.

[0024] The effects obtained by the present disclosure are not limited to the effects mentioned above, and other effects not mentioned will be clearly understood by those having ordinary knowledge in the technical field to which the present disclosure pertains from the following description.

Brief Description of Drawings

[0025] [Figure 1] It is a schematic diagram showing a video coding system to which an embodiment according to the present disclosure can be applied. [Figure 2] It is a schematic diagram showing a video encoding apparatus to which an embodiment according to the present disclosure can be applied. [Figure 3] It is a schematic diagram showing a video decoding apparatus to which an embodiment according to the present disclosure can be applied. [Figure 4] It is a diagram schematically showing an inter prediction unit of a video encoding apparatus. [Figure 5] It is a flowchart showing a method for encoding a video based on inter prediction. [Figure 6] It is a diagram schematically showing an inter prediction unit of a video decoding apparatus. [Figure 7] It is a flowchart showing a method for decoding a video based on inter prediction. [Figure 8] It is a flowchart showing an inter prediction method. [Figure 9] It is a diagram showing a division example of a geometric division mode. [Figure 10] It is a diagram for explaining a candidate list of geometric division modes. [Figure 11]This figure illustrates an example of a method for generating blending weights for geometric partitioning modes. [Figure 12] This diagram illustrates an example of the surrounding regions of a current block and a reference block. [Figure 13] This diagram illustrates an example of the division direction for geometric division modes. [Figure 14] This diagram illustrates examples of subblocks based on the division direction in geometric division modes. [Figure 15-18] This diagram illustrates an example of a reference region based on the division direction and / or division offset. [Figure 19] This is a flowchart showing a video encoding method according to one embodiment of the present disclosure. [Figure 20] This is a flowchart showing a video decoding method according to one embodiment of the present disclosure. [Figure 21] This is a flowchart showing a video encoding / decoding method according to another embodiment of the present disclosure. [Figure 22] This is a flowchart showing a video encoding / decoding method according to yet another embodiment of the present disclosure. [Figure 23] This is a flowchart showing a video encoding / decoding method according to yet another embodiment of the present disclosure. [Figure 24] This is a flowchart showing a video encoding / decoding method according to yet another embodiment of the present disclosure. [Figure 25] This flowchart shows a video encoding method according to yet another embodiment of the present disclosure. [Figure 26] This flowchart shows a video decoding method according to yet another embodiment of the present disclosure. [Figure 27] This figure illustrates a content streaming system to which the embodiments of this disclosure can be applied. [Modes for carrying out the invention]

[0026] Hereafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings, so that they can be easily implemented by a person with ordinary skill in the art to which the present disclosure pertains. However, the present disclosure may be embodied in various other forms and is not limited to the embodiments described herein.

[0027] In describing embodiments of this disclosure, if a specific description of a known configuration or function is deemed to obscure the gist of this disclosure, such detailed description will be omitted. In the figures, parts unrelated to the description of this disclosure will be omitted, and similar parts will be denoted by similar reference numerals.

[0028] In this disclosure, when one component is described as being “linked,” “joined,” or “connected” to another component, this may include not only direct linkages but also indirect linkages where other components exist in between. Furthermore, when one component is described as “containing” or “having” another component, this means, unless otherwise specified, that it may contain further other components rather than excluding them.

[0029] In this disclosure, terms such as "first," "second," etc., are used solely to distinguish one component from another, and do not limit the order or importance of the components unless otherwise specified. Therefore, within the scope of this disclosure, a first component in one embodiment may be referred to as a second component in another embodiment, and similarly, a second component in one embodiment may be referred to as a first component in another embodiment.

[0030] In this disclosure, components are distinguished from each other solely to clearly describe their respective characteristics, and this does not necessarily mean that these components are separate. That is, multiple components may be integrated to constitute a single hardware or software unit, or a single component may be distributed to constitute multiple hardware or software units. Therefore, such integrated or distributed embodiments are also included in the scope of this disclosure, even without specific mention.

[0031] In this disclosure, the components described in various embodiments are not necessarily essential components, and some may be optional components. Therefore, embodiments consisting of a subset of the components described in one embodiment are also included in the scope of this disclosure. Furthermore, embodiments that further include other components in addition to the components described in various embodiments are also included in the scope of this disclosure.

[0032] This disclosure relates to the encoding and decoding of video, and unless otherwise defined herein, the terms used herein may have their ordinary meanings in the art to which this disclosure pertains.

[0033] In this disclosure, "picture" generally refers to a unit representing a single video image for a specific time period, and "slice / tile" is an encoding unit that constitutes a part of a picture. A single picture may consist of one or more slices / tiles. A slice / tile may also contain one or more CTUs (coding tree units).

[0034] In this disclosure, “pixel” or “pel” can mean the smallest unit that constitutes a picture (or video). The term “sample” may also be used as a counterpart to pixel. A sample may generally represent a pixel or a pixel value, or it may represent only the pixel / pixel value of the luma component, or only the pixel / pixel value of the chroma component.

[0035] In this disclosure, “unit” may represent a basic unit of image processing. A unit may include at least one of a specific region of a picture and information associated with that region. A unit may, as it may be, be replaced by terms such as “sample array,” “block,” or “area.” In general, an MxN block may include a sample (or sample array) or a set (or array) of transform coefficients consisting of M columns and N rows.

[0036] In this disclosure, “current block” can mean one of the following: “current coding block,” “current coding unit,” “block to encode,” “block to decode,” or “block to process.” When prediction is performed, “current block” can mean “current prediction block” or “block to predict.” When transformation (inverse transformation) / quantization (inverse quantization) is performed, “current block” can mean “current transformation block” or “block to transform.” When filtering is performed, “current block” can mean “block to filter.”

[0037] In this disclosure, "current block" may mean a block containing both a luma component block and a chroma component block, or "the luma block of the current block," unless otherwise explicitly stated as a chroma block. The luma component block of the current block may be expressed with an explicit mention of a luma component block, such as "luma block" or "current luma block." Similarly, the chroma component block of the current block may be expressed with an explicit mention of a chroma component block, such as "chroma block" or "current chroma block."

[0038] In this disclosure, " / " and "," may be interpreted as "and / or." For example, "A / B" and "A, B" may be interpreted as "A and / or B." Also, "A / B / C" and "A, B, C" may mean "at least one of A, B and / or C."

[0039] In this disclosure, “or” may be interpreted as “and / or.” For example, “A or B” may mean 1) “A” only, 2) “B” only, or 3) “A and B.” Alternatively, in this disclosure, “or” may mean “additionally or alternatively.”

[0040] Overview of the video coding system

[0041] Figure 1 is a schematic diagram showing a video coding system to which the embodiments of this disclosure can be applied.

[0042] A video coding system according to one embodiment may include an encoding device 10 and a decoding device 20. The encoding device 10 can transmit encoded video and / or image information or data to the decoding device 20 in file or streaming form via a digital storage medium or network.

[0043] An encoding device 10 according to one embodiment may include a video source generation unit 11, an encoding unit 12, and a transmission unit 13. A decoding device 20 according to one embodiment may include a receiving unit 21, a decoding unit 22, and a rendering unit 23. The encoding unit 12 may be called a video / image encoding unit, and the decoding unit 22 may be called a video / image decoding unit. The transmission unit 13 may be included in the encoding unit 12. The receiving unit 21 may be included in the decoding unit 22. The rendering unit 23 may include a display unit, and the display unit may be composed of a separate device or external component.

[0044] The video source generation unit 11 can acquire video / images through video / image capture, synthesis, or generation processes. The video source generation unit 11 may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, or a video / image archive containing previously captured video / images. The video / image generation device may include, for example, a computer, a tablet, and a smartphone, and can generate video / images (electronically). For example, virtual video / images may be generated by a computer, in which case the video / image capture process may be replaced by a process in which related data is generated.

[0045] The encoding unit 12 can encode the input video / image data. The encoding unit 12 can perform a series of procedures such as prediction, transformation, and quantization for compression and encoding efficiency. The encoding unit 12 can output the encoded data (encoded video / image information) in the form of a bitstream.

[0046] The transmitting unit 13 can acquire encoded video / image information or data output in bitstream form and transmit it in file or streaming form to the receiving unit 21 of the decoding device 20 or other external object via a digital storage medium or network. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray (registered trademark: same hereinafter), HDD, SSD, etc. The transmitting unit 13 may include elements for generating media files in a predetermined file format and may include elements for transmission via a broadcast / communication network. The transmitting unit 13 may be provided as a transmission device separate from the encoding unit 12, in which case the transmission device may include at least one processor that acquires encoded video / image information or data output in bitstream form and a transmitting unit that transmits it in file or streaming form. The receiving unit 21 can extract / receive the bitstream from the storage medium or network and transmit it to the decoding unit 22.

[0047] The decoding unit 22 can decode the video / image by performing a series of procedures such as inverse quantization, inverse transform, and prediction, which correspond to the operation of the encoding unit 12.

[0048] The rendering unit 23 can render the decoded video / image. The rendered video / image may be displayed through the display unit.

[0049] Overview of video encoding equipment

[0050] Figure 2 is a schematic diagram showing a video encoding device to which the embodiments of this disclosure can be applied.

[0051] As shown in Figure 2, the video encoding device 100 may include a video splitting unit 110, a subtraction unit 115, a conversion unit 120, a quantization unit 130, an inverse quantization unit 140, an inverse conversion unit 150, an addition unit 155, a filtering unit 160, a memory 170, an inter-prediction unit 180, an intra-prediction unit 185, and an entropy encoding unit 190. The inter-prediction unit 180 and the intra-prediction unit 185 may be collectively called the "prediction unit". The conversion unit 120, the quantization unit 130, the inverse quantization unit 140, and the inverse conversion unit 150 may be included in the residual processing unit. The residual processing unit may further include a subtraction unit 115.

[0052] Depending on the embodiment, all or at least some of the multiple components constituting the video encoding device 100 may be embodied as a single hardware component (e.g., an encoder or a processor). Furthermore, the memory 170 may include a DPB (decoded picture buffer) and may be embodied by a digital storage medium.

[0053] The video splitting unit 110 can split the input video (or picture, frame) input to the video encoding device 100 into one or more processing units. For example, the processing units may be called coding units (CUs). Coding units can be obtained by recursively splitting a coding tree unit (CTU) or the largest coding unit (LCU) using a QT / BT / TT (Quad-tree / binary-tree / ternary-tree) structure. For example, one coding unit may be split into multiple coding units of deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary-tree structure. For the splitting of coding units, a quad-tree structure may be applied first, followed by a binary-tree structure and / or a ternary-tree structure. The coding procedure according to this disclosure may be performed based on the final coding unit that is not further split. The maximum coding unit may be used directly as the final coding unit, or a lower-depth coding unit obtained by dividing the maximum coding unit may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and / or restoration, which will be described later. As another example, the processing unit of the coding procedure may be a prediction unit (PU) or a transformation unit (TU). The prediction unit and the transformation unit may be divided or partitioned from the final coding unit, respectively. The prediction unit may be a unit of sample prediction, and the transformation unit may be a unit that derives transformation coefficients and / or a unit that derives a residual signal from transformation coefficients.

[0054] The prediction unit (inter-prediction unit 180 or intra-prediction unit 185) can make predictions for the block to be processed (current block) and generate a predicted block that includes prediction samples for the current block. The prediction unit can determine whether intra-prediction or inter-prediction is applied to the current block or on a CU basis. The prediction unit can generate various information regarding the prediction of the current block and transmit it to the entropy encoding unit 190. The prediction information may be encoded by the entropy encoding unit 190 and output in the form of a bitstream.

[0055] The intra-prediction unit 185 can predict the current block by referring to a sample in the current picture. The referenced sample may be located in the vicinity of the current block or at a distance from it, depending on the intra-prediction mode and / or intra-prediction method. The intra-prediction mode may include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes may include, for example, a DC mode and a planar mode. The directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes, depending on the accuracy of the prediction direction. However, this is an example, and more or fewer directional prediction modes may be used depending on the settings. The intra-prediction unit 185 can also determine the prediction mode to be applied to the current block using the prediction modes applied to the surrounding blocks.

[0056] The interprediction unit 180 can derive a predicted block relative to the current block based on a reference block (reference sample array) identified by motion vectors on the reference picture. In this case, in order to reduce the amount of motion information transmitted in interprediction mode, motion information can be predicted in units of blocks, subblocks, or samples based on the correlation of motion information between the surrounding blocks and the current block. The motion information may include motion vectors and reference picture indices. The motion information may further include interprediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of interprediction, the surrounding blocks may include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. The reference picture containing the reference block and the reference picture containing the temporal neighboring block may be the same or different from each other. The temporal neighboring block may be called a collocated reference block, colCU, etc. The reference picture containing the temporal neighboring block may be called a collocated picture (colPic). For example, the interpretation unit 180 can construct a motion information candidate list based on surrounding blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Interpretation may be performed based on various prediction modes; for example, in skip mode and merge mode, the interpretation unit 180 can use the motion information of surrounding blocks as the motion information of the current block. In skip mode, unlike merge mode, the residual signal does not need to be transmitted.In motion vector prediction (MVP) mode, the motion vectors of surrounding blocks are used as motion vector predictors, and the motion vector of the current block can be signaled by encoding the motion vector difference and an indicator for the motion vector predictor. The motion vector difference represents the difference between the motion vector of the current block and the motion vector predictor.

[0057] The prediction unit can generate a prediction signal based on various prediction methods and / or prediction techniques described later. For example, the prediction unit may apply intra-prediction or inter-prediction to predict the current block, or it may apply intra-prediction and inter-prediction simultaneously. A prediction method that applies intra-prediction and inter-prediction simultaneously to predict the current block may be called CIIP (combined inter and intra prediction). The prediction unit can also perform intra-block copy (IBC) to predict the current block. Intra-block copy may be used, for example, for coding content images / videos such as games, as in SCC (screen content coding). IBC is a method of predicting the current block using a reference block that has already been restored in the current picture at a predetermined distance from the current block. When IBC is applied, the position of the reference block in the current picture may be encoded as a vector (block vector) corresponding to the predetermined distance. IBC basically performs prediction within the current picture, but it may be performed similarly to inter-prediction in that it derives the reference block within the current picture. In other words, IBC can use at least one of the interpretation methods described in this disclosure.

[0058] The predicted signal generated by the prediction unit may be used to generate a restored signal or a residual signal. The subtraction unit 115 can generate a residual signal (residual block, residual sample array) by subtracting the predicted signal output from the prediction unit (predicted block, predicted sample array) from the input video signal (original block, original sample array). The generated residual signal may be transmitted to the conversion unit 120.

[0059] The transformation unit 120 can generate transformation coefficients by applying a transformation method to the residual signal. For example, the transformation method may include at least one of the following: DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), KLT (Karhunen-Loeve Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT refers to the transformation obtained from a graph when the relationship information between pixels is represented by this graph. CNT refers to the transformation obtained by generating a prediction signal using all previously reconstructed pixels and obtaining a transformation based on it. The transformation process may be applied to pixel blocks of the same size and square shape, or to blocks of a variable size instead of square shape.

[0060] The quantization unit 130 can quantize the conversion coefficients and transmit them to the entropy encoding unit 190. The entropy encoding unit 190 can encode the quantized signal (information about the quantized conversion coefficients) and output it as a bitstream. The information about the quantized conversion coefficients may be called residual information. The quantization unit 130 can rearrange the block-shaped quantized conversion coefficients into a one-dimensional vector form based on the coefficient scan order, and can also generate information about the quantized conversion coefficients based on the one-dimensional vector form of the quantized conversion coefficients.

[0061] The entropy encoding unit 190 can perform various encoding methods, such as exponential Golomb, CAVLC (context-adaptive variable length coding), and CABAC (context-adaptive binary arithmetic coding). In addition to the quantized conversion coefficients, the entropy encoding unit 190 can also encode information necessary for video / image restoration (e.g., the values ​​of syntax elements) together or separately. The encoded information (e.g., encoded video / image information) may be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer) units. The video / image information may further include information about various parameter sets, such as an adaptation parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS). The video / image information may also further include general constraint information. The signaling information, transmitted information, and / or syntax elements referred to in this disclosure may be encoded by the encoding procedure described above and included in the bitstream.

[0062] The bitstream may be transmitted over a network or stored on a digital storage medium. Here, the network may include broadcasting networks and / or communication networks, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitting unit (not shown) for transmitting the signal output from the entropy encoding unit 190 and / or a storage unit (not shown) for storing it may be provided as an internal / external element of the video encoding device 100, or the transmitting unit may be provided as a component of the entropy encoding unit 190.

[0063] The quantized conversion coefficients output from the quantization unit 130 may be used to generate a resistive signal. For example, by applying inverse quantization and inverse transformation to the quantized conversion coefficients in the inverse quantization unit 140 and the inverse transformation unit 150, a resistive signal (residual block or resistive sample) can be reconstructed.

[0064] The adder 155 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the predicted signal output from the inter-prediction unit 180 or the intra-prediction unit 185. When there is no residual for the block to be processed, such as when skip mode is applied, the predicted block may be used as the reconstructed block. The adder 155 may be called the reconstruction unit or the reconstructed block generation unit. The generated reconstructed signal may be used for intra-prediction of the next block to be processed in the current picture, or, as described later, may be used for inter-prediction of the next picture after filtering.

[0065] The filtering unit 160 can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit 160 can apply various filtering methods to the restored picture to generate a modified restored picture, and store the modified restored picture in the memory 170, specifically in the DPB of the memory 170. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, and bilateral filter. The filtering unit 160 can generate various filtering information and transmit it to the entropy encoding unit 190, as will be described later in the description of each filtering method. The filtering information may be encoded by the entropy encoding unit 190 and output in the form of a bitstream.

[0066] The corrected restored picture transmitted to memory 170 may be used as a reference picture in the interpretation unit 180. This allows the video encoding device 100 to avoid prediction mismatches between the video encoding device 100 and the video decoding device when interpretation is applied, and also improves encoding efficiency.

[0067] The DPB in memory 170 can store the corrected restored picture for use as a reference picture in the inter-prediction unit 180. Memory 170 can store motion information of blocks from which motion information in the current picture has been derived (or encoded) and / or motion information of blocks in the picture that have already been restored. The stored motion information may be transmitted to the inter-prediction unit 180 for use as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. Memory 170 can store restored samples of restored blocks in the current picture and transmit them to the intra-prediction unit 185.

[0068] Overview of the video decoding device

[0069] Figure 3 is a schematic diagram showing an image decoding device to which the embodiments of this disclosure can be applied.

[0070] As shown in Figure 3, the video decoding device 200 may include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transformation unit 230, an addition unit 235, a filtering unit 240, a memory 250, an inter-prediction unit 260, and an intra-prediction unit 265. The inter-prediction unit 260 and the intra-prediction unit 265 can be collectively referred to as the "prediction unit". The inverse quantization unit 220 and the inverse transformation unit 230 may be included in the residual processing unit.

[0071] All or at least some of the multiple components constituting the video decoding device 200 may be embodied as a single hardware component (e.g., a decoder or processor) depending on the embodiment. Furthermore, the memory 170 may include a DPB and may be embodied by a digital storage medium.

[0072] A video decoding device 200 that receives a bitstream containing video / image information can restore the image by performing a process corresponding to the process performed by the video encoding device 100 in Figure 2. For example, the video decoding device 200 can perform decoding using the processing unit applied in the video encoding device. Therefore, the decoding processing unit may be, for example, a coding unit. The coding unit may be a coding tree unit, or it may be obtained by dividing the largest coding unit. The restored video signal decoded and output by the video decoding device 200 may then be played back by a playback device (not shown).

[0073] The video decoding device 200 can receive the signal output from the video encoding device shown in Figure 2 in the form of a bitstream. The received signal may be decoded by the entropy decoding unit 210. For example, the entropy decoding unit 210 can parse the bitstream to derive information necessary for video restoration (or picture restoration) (e.g., video / image information). The video / image information may further include information about various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / image information may also further include general constraint information. The video decoding device may further utilize the parameter set information and / or the general constraint information to decode the video. The signaling information, received information, and / or syntax elements referred to in this disclosure may be obtained from the bitstream by decoding through the decoding procedure. For example, the entropy decoding unit 210 can decode information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the values ​​of syntax elements necessary for image restoration and the quantized values ​​of conversion coefficients related to the residual. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element in the bitstream, determines a context model using the syntax element information to be decoded and the decoding information of the surrounding blocks and the blocks to be decoded, or symbol / bin information decoded in a previous stage, predicts the probability of bin occurrence based on the determined context model, performs arithmetic decoding of the bins, and generates symbols corresponding to the values ​​of each syntax element.In this case, the CABAC entropy decoding method can update the context model using the decoded symbol / bin information for the context model of the next symbol / bin after determining the context model. Information related to prediction from the information decoded by the entropy decoding unit 210 is provided to the prediction unit (inter-prediction unit 260 and intra-prediction unit 265), and residual values ​​that have been entropy decoded by the entropy decoding unit 210, i.e., quantized conversion coefficients and related parameter information, may be input to the inverse quantization unit 220. In addition, information related to filtering from the information decoded by the entropy decoding unit 210 may be provided to the filtering unit 240. On the other hand, a receiving unit (not shown) that receives signals output from the video encoding device may be further provided as an internal / external element of the video decoding device 200, or the receiving unit may be provided as a component of the entropy decoding unit 210.

[0074] On the other hand, the video decoding device according to this disclosure may be called a video / image / picture decoding device. The video decoding device may include an information decoder (video / image / picture information decoder) and / or a sample decoder (video / image / picture sample decoder). The information decoder may include an entropy decoding unit 210, and the sample decoder may include at least one of an inverse quantization unit 220, an inverse transformation unit 230, an addition unit 235, a filtering unit 240, a memory 250, an inter-prediction unit 260, and an intra-prediction unit 265.

[0075] The inverse quantization unit 220 can inverse quantize the quantized transformation coefficients and output the transformation coefficients. The inverse quantization unit 220 can rearrange the quantized transformation coefficients in a two-dimensional block form. In this case, the rearrangement may be performed based on the coefficient scan order performed by the video encoding device. The inverse quantization unit 220 can perform inverse quantization on the quantized transformation coefficients using quantization parameters (e.g., quantization step size information) and obtain the transformation coefficients.

[0076] The inverse conversion unit 230 can inversely convert the conversion coefficients to obtain residual signals (residual blocks, residual sample arrays).

[0077] The prediction unit can make predictions for the current block and generate a predicted block containing prediction samples for the current block. Based on the prediction information output from the entropy decoding unit 210, the prediction unit can determine whether intra-prediction or inter-prediction is applied to the current block and can determine a specific intra / inter-prediction mode (prediction method).

[0078] As mentioned in the description of the prediction unit of the video coding device 100, the prediction unit can generate prediction signals based on various prediction methods (techniques) described later.

[0079] The intra-prediction unit 265 can predict the current block by referring to the samples in the current picture. The description of the intra-prediction unit 185 may also apply to the intra-prediction unit 265.

[0080] The interprediction unit 260 can derive a predicted block relative to the current block based on a reference block (reference sample array) identified by motion vectors on the reference picture. In this case, in order to reduce the amount of motion information transmitted in interprediction mode, motion information can be predicted in block, subblock, or sample units based on the correlation of motion information between the surrounding block and the current block. The motion information may include motion vectors and reference picture indices. The motion information may further include interprediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of interprediction, the surrounding block may include spatially neighboring blocks present in the current picture and temporally neighboring blocks present in the reference picture. For example, the interprediction unit 260 can construct a motion information candidate list based on the surrounding blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Interprediction may be performed based on various prediction modes (methods), and the prediction information may include information indicating the mode (method) of interprediction for the current block.

[0081] The adder 235 can generate a restored signal (restored picture, restored block, restored sample array) by adding the acquired residual signal to the predicted signal (predicted block, predicted sample array) output from the prediction unit (including the inter-prediction unit 260 and / or intra-prediction unit 265). When there is no residual for the block to be processed, such as when skip mode is applied, the predicted block may be used as the restored block. The description of the adder 155 may also apply to the adder 235. The adder 235 may be called the restore unit or the restored block generation unit. The generated restored signal may be used for intra-prediction of the next block to be processed in the current picture, or, as described later, may be used for inter-prediction of the next picture after filtering.

[0082] The filtering unit 240 can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit 240 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be stored in the memory 250, specifically in the DPB of the memory 250. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, and bilateral filter.

[0083] The restored picture stored (modified) in the DPB of memory 250 may be used as a reference picture in the inter-prediction unit 260. Memory 250 can store motion information of blocks from which motion information in the current picture has been derived (or decoded) and / or motion information of blocks in the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit 260 for use as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. Memory 250 can store restored samples of restored blocks in the current picture and transmit them to the intra-prediction unit 265.

[0084] In this specification, the embodiments described for the filtering unit 160, inter-prediction unit 180, and intra-prediction unit 185 of the video encoding device 100 may be applied identically or in a corresponding manner to the filtering unit 240, inter-prediction unit 260, and intra-prediction unit 265 of the video decoding device 200, respectively.

[0085] Interpretation General

[0086] The prediction units of the video encoding device 100 and the video decoding device 200 can perform interpretation on a block-by-block basis to derive predicted samples. Interpretation may be a prediction derived by a method that depends on data elements of pictures other than the current picture (e.g., sample values ​​or motion information). When interpretation is applied to the current block, a predicted block (predicted sample array) for the current block may be derived based on a reference block (reference sample array) identified by motion vectors on the reference picture indicated by the reference picture index. In this case, in order to reduce the amount of motion information transmitted in interpretation mode, the motion information of the current block may be predicted on a block, subblock, or sample basis based on the correlation of motion information between the surrounding block and the current block. The motion information may include motion vectors and the reference picture index. The motion information may further include interpretation type (L0 prediction, L1 prediction, Bi prediction, etc.) information. When interpretation is applied, the surrounding block may include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. The reference picture containing the reference block and the reference picture containing the temporally surrounding block may be identical or different. The temporally surrounding block may also be called a collocated reference block (colCU), and the reference picture containing the temporally surrounding block may be called a collocated picture (colPic). For example, a list of motion information candidates may be constructed based on the surrounding blocks of the current block, and a flag or index information may be signaled indicating which candidate is selected (used) to derive the motion vector and / or reference picture index of the current block. Interpretation may be performed based on various prediction modes; for example, in skip mode and merge mode, the motion information of the current block may be identical to the motion information of the selected surrounding block.In skip mode, unlike merge mode, the resistive signal does not need to be transmitted. In motion vector prediction (MVP) mode, the motion vector of the selected surrounding block is used as the motion vector predictor, and the motion vector difference may be signaled. In this case, the motion vector of the current block can be derived using the sum of the motion vector predictor and the motion vector difference.

[0087] The motion information may include L0 motion information and / or L1 motion information according to the interpretation type (L0 prediction, L1 prediction, Bi prediction, etc.). A motion vector in the L0 direction may be called an L0 motion vector or MVL0, and a motion vector in the L1 direction may be called an L1 motion vector or MVL1. A prediction based on the L0 motion vector may be called an L0 prediction, a prediction based on the L1 motion vector may be called an L1 prediction, and a prediction based on both the L0 motion vector and the L1 motion vector may be called a bi (Bi) prediction. Here, the L0 motion vector can represent a motion vector associated with the reference picture list L0 (L0), and the L1 motion vector can represent a motion vector associated with the reference picture list L1 (L1). The reference picture list L0 may include pictures earlier in the output order than the current picture as reference pictures, and the reference picture list L1 may include pictures later in the output order than the current picture. The earlier pictures may be called forward (reference) pictures, and the later pictures may be called backward (reference) pictures. The aforementioned reference picture list L0 may further include pictures that are later in the output order than the current picture as reference pictures. In this case, the earlier picture may be indexed first in the reference picture list L0, and the later picture may be indexed next. The aforementioned reference picture list L1 may further include pictures that are earlier in the output order than the current picture as reference pictures. In this case, the later picture may be indexed first in the reference picture list L1, and the earlier picture may be indexed next. Here, the output order may correspond to the POC (picture order count) order.

[0088] Figure 4 is a schematic diagram showing the interpretation unit 180 of the video encoding device 100, and Figure 5 is a flowchart showing a method for encoding video based on interpretation.

[0089] The video encoding device 100 can perform inter prediction for the current block (S510). The video encoding device 100 can derive the inter prediction mode and motion information of the current block and generate prediction samples for the current block. Here, the procedures of determining the inter prediction mode, deriving motion information, and generating prediction samples may be performed simultaneously, or any one of the procedures may be performed before the others. For example, the inter prediction unit 180 of the video encoding device 100 may include a prediction mode determination unit 181, a motion information derivation unit 182, and a prediction sample derivation unit 183, wherein the prediction mode determination unit 181 determines the prediction mode for the current block, the motion information derivation unit 182 derives the motion information of the current block, and the prediction sample derivation unit 183 derives prediction samples for the current block. For example, the interpretation unit 180 of the video encoding device 100 can use motion estimation to search for blocks similar to the current block within a certain area (search area) of the reference picture, and derive a reference block whose difference from the current block is minimal or below a certain standard. Based on this, it can derive a reference picture index indicating the reference picture in which the reference block is located, and derive a motion vector based on the positional difference between the reference block and the current block. The video encoding device 100 can determine which of the various prediction modes is applied to the current block. The video encoding device 100 can compare the RD costs for the various prediction modes and determine the optimal prediction mode for the current block.

[0090] For example, when skip mode or merge mode is applied to the current block, the video encoding device 100 can configure a merge candidate list, as described later, and derive a reference block from among the reference blocks indicated by the merge candidates included in the merge candidate list whose difference from the current block is the minimum or below a certain standard. In this case, a merge candidate associated with the derived reference block is selected, merge index information indicating the selected merge candidate is generated, and may be signaled to the decoding device. The motion information of the current block can be derived using the motion information of the selected merge candidate.

[0091] As another example, when the (A)MVP mode is applied to the current block, the video encoding device 100 can configure the (A)MVP candidate list described later, and use the motion vector of an MVP candidate selected from the MVP (motion vector predictor) candidates included in the (A)MVP candidate list as the MVP of the current block. In this case, for example, the motion vector indicating the reference block derived by the motion estimation described above may be used as the motion vector of the current block, and the MVP candidate having the motion vector with the smallest difference from the motion vector of the current block may become the selected MVP candidate. The MVD (motion vector difference), which is the difference obtained by subtracting the MVP from the motion vector of the current block, may be derived. In this case, information regarding the MVD may be signaled to the video decoding device 200. Also, when the (A)MVP mode is applied, the value of the reference picture index may be configured as reference picture index information and separately signaled to the video decoding device 200.

[0092] The video encoding device 100 can derive a residual sample based on the predicted sample (S520). The video encoding device 100 can derive the residual sample by comparing the original sample of the current block with the predicted sample.

[0093] The video encoding device 100 can encode video information including prediction information and residual information (S530). The video encoding device 100 can output the encoded video information in the form of a bitstream. The prediction information is information related to the prediction procedure and may include prediction mode information (e.g., skip flag, merge flag, or mode index) and motion information. The motion information may include candidate selection information (e.g., merge index, mvp flag, or mvp index) which is information for deriving a motion vector. The motion information may also include the above-mentioned MVD information and / or reference picture index information. The motion information may also include information indicating whether L0 prediction, L1 prediction, or bi prediction is applied. The residual information is information about the residual sample. The residual information may include information about the quantized conversion coefficients for the residual sample.

[0094] The output bitstream may be stored on a (digital) storage medium and transmitted to a decoding device, or it may be transmitted to a video decoding device 200 via a network.

[0095] On the other hand, as described above, the video encoding device 100 can generate a restored picture (including restored samples and restored blocks) based on the reference sample and the residual sample. This is because the video encoding device 100 can derive the same prediction results as those performed by the video decoding device 200, thereby improving coding efficiency. Therefore, the video encoding device 100 can store the restored picture (or restored sample, restored block) in memory and use it as a reference picture for interpretation. As described above, in-loop filtering procedures and the like may be further applied to the restored picture.

[0096] Figure 6 is a schematic diagram of the interpretation unit 260 of the video decoding device 200, and Figure 7 is a flowchart showing a method for decoding video based on interpretation.

[0097] The video decoding device 200 can perform operations corresponding to those performed by the video encoding device 100. The video decoding device 200 can make predictions for the current block based on the received prediction information and derive prediction samples.

[0098] Specifically, the video decoding device 200 can determine a prediction mode for the current block based on the received prediction information (S710). Based on the prediction mode information in the prediction information, the video decoding device 200 can determine which inter-prediction mode is applied to the current block.

[0099] For example, based on the merge flag, it can be determined whether the merge mode is applied to the current block or whether the (A)MVP mode is determined. Alternatively, based on the mode index, one of several inter-prediction mode candidates can be selected. The inter-prediction mode candidates may include skip mode, merge mode, and / or (A)MVP mode, or may include several inter-prediction modes as described later.

[0100] The video decoding device 200 can derive motion information for the current block based on the determined interprediction mode (S720). For example, when a skip mode or merge mode is applied to the current block, the video decoding device 200 can configure a merge candidate list, which will be described later, and select one merge candidate from among the merge candidates included in the merge candidate list. The selection may be made based on the selection information (merge index) described above. The motion information for the current block can be derived using the motion information for the selected merge candidate. The motion information for the selected merge candidate may be used as the motion information for the current block.

[0101] As another example, when the (A)MVP mode is applied to the current block, the video decoding device 200 can configure the (A)MVP candidate list described later, and use the motion vector of the mvp candidate selected from the mvp (motion vector predictor) candidates included in the (A)MVP candidate list for the mvp of the current block. The selection may be made based on the selection information (mvp flag or mvp index) described above. In this case, the MVD of the current block may be derived based on the information regarding the MVD, and the motion vector of the current block may be derived based on the mvp of the current block and the MVD. In addition, the reference picture index of the current block may be derived based on the reference picture index information. The picture indicated by the reference picture index in the reference picture list for the current block may be derived as the reference picture referenced for interpretation of the current block.

[0102] On the other hand, as will be described later, the movement information of the current block may be derived without constructing a candidate list, in which case the movement information of the current block may be derived by the disclosure procedure in the prediction mode described later. In this case, the construction of the candidate list as described above may be omitted.

[0103] The video decoding device 200 can generate predicted samples for the current block based on the motion information of the current block (S730). In this case, the reference picture is derived based on the reference picture index of the current block, and the predicted samples for the current block can be derived using the reference block samples indicated on the reference picture by the motion vector of the current block. In this case, as will be described later, a further predicted sample filtering procedure may be performed on all or part of the predicted samples for the current block.

[0104] For example, the interpretation unit 260 of the video decoding device 200 may include a prediction mode determination unit 261, a motion information derivation unit 262, and a prediction sample derivation unit 263. The prediction mode determination unit 181 determines the prediction mode for the current block based on the received prediction mode information, the motion information derivation unit 182 derives motion information (motion vectors and / or reference picture indexes, etc.) for the current block based on the received motion information, and the prediction sample derivation unit 183 derives prediction samples for the current block.

[0105] The video decoding device 200 can generate a residual sample for the current block based on the received residual information (S740). The video decoding device 200 can generate a restored sample for the current block based on the predicted sample and the residual sample, and generate a restored picture based on this (S750). As described above, in-loop filtering procedures and the like may be further applied to the restored picture thereafter.

[0106] Referring to Figure 8, as described above, the interpretation procedure may include an interpretation mode determination step (S810), a motion information derivation step based on the determined prediction mode (S820), and a prediction execution (prediction sample generation) step (S830) based on the derived motion information. The interpretation procedure may be performed by the video encoding device 100 and the video decoding device 200 as described above.

[0107] Interpretation mode determination

[0108] Various interpretation modes may be used to predict the current block in the picture. For example, various modes may be used, such as merge mode, skip mode, MVP (motion vector prediction) mode, affine mode, subblock merge mode, and MMVD (merge with MVD) mode. DMVR (Decoder side motion vector refinement) mode, AMVR (adaptive motion vector resolution) mode, BCW (Bi-prediction with CU-level weight), and BDOF (Bi-directional optical flow) may be used as secondary modes or alternatives. The affine mode may also be called the affine motion prediction mode. The MVP mode may also be called the AMVP (advanced motion vector prediction) mode. Motion information candidates derived by some modes and / or some modes in this document may be included as one of the motion information related candidates in other modes. For example, the HMVP candidate may be added as a merge candidate in the merge / skip mode, or as an mvp candidate in the MVP mode.

[0109] Prediction mode information indicating the current block's inter-prediction mode may be signaled from the video encoding device 100 to the video decoding device 200. The prediction mode information may be included in the bitstream and received by the video decoding device 200. The prediction mode information may include index information indicating one of a plurality of candidate modes. Alternatively, the inter-prediction mode may be indicated by hierarchical signaling of flag information. In this case, the prediction mode information may include one or more flags. For example, a skip flag may be signaled to indicate whether or not the skip mode is applied, and if the skip mode is not applied, a merge flag may be signaled to indicate whether or not the merge mode is applied, and if the merge mode is not applied, it may be indicated that the MVP mode is applied, or additional flags for further distinctions may be signaled. Affine modes may be signaled as independent modes, or as modes dependent on the merge mode or MVP mode, etc. For example, affine modes may include affine merge mode and affine MVP mode.

[0110] Derivation of motion information

[0111] Interpretation may be performed using motion information of the current block. The video encoding device 100 can derive optimal motion information for the current block by a motion estimation procedure. For example, the video encoding device 100 can use the original block in the original picture for the current block to search for a highly correlated similar reference block in fractional pixel units within a defined search range in the reference picture, thereby deriving motion information. Block similarity can be derived based on the difference in phase-based sample values. For example, block similarity may be calculated based on the SAD between the current block (or the template of the current block) and the reference block (or the template of the reference block). In this case, motion information may be derived based on the reference block with the smallest SAD in the search area. The derived motion information may be signaled to the video decoding device 200 by various methods based on the interpretation mode.

[0112] Predictive sample generation

[0113] Based on the motion information derived by the prediction mode, a predicted block may be derived for the current block. The predicted block may include predicted samples (predicted sample array) of the current block. If the motion vector of the current block is in fractional sample units, an interpolation procedure may be performed so that predicted samples of the current block are derived based on reference samples in fractional sample units within the reference picture. When affine interpretation is applied to the current block, predicted samples may be generated based on sample / subblock unit MV. When biprediction is applied, predicted samples derived by a (phase-weighted) sum or weighted average of predicted samples derived based on L0 prediction (i.e., prediction using reference pictures in reference picture list L0 and MVL0) and predicted samples derived based on L1 prediction (i.e., prediction using reference pictures in reference picture list L1 and MVL1) may be used as predicted samples for the current block. When biprediction is applied, if the reference picture used for L0 prediction and the reference picture used for L1 prediction are located in different temporal directions relative to the current picture (i.e., it is biprediction but corresponds to bidirectional prediction), this can be called true biprediction.

[0114] As mentioned above, reconstructed samples and reconstructed pictures may be generated based on the derived predicted samples, and then procedures such as in-loop filtering may be performed.

[0115] Geometric Partitioning Mode (GPM)

[0116] GPM may be used to support inter-prediction. GPM may be signaled as one type of merge mode using a CU level flag, along with other merge modes such as regular merge mode, MMVD (merge mode with MVD) mode, CIIP mode, and subblock merge mode. GPM is used for each possible CU size other than 8×64 and 64×8, w×h=2 m ×2n In this case, a total of 64 partitions can be supported for m,n∈{36}.

[0117] When GPM is used, the CU may be divided into two parts by a geometrically positioned straight line (dividing line) as shown in Figure 9. The position of the dividing line may be mathematically derived from the angle and offset parameter of a particular division. Each part of the geometric division within the CU may be interpreted using its own motion, and only unidirectional predictions may be allowed for each division. That is, each part may have one motion vector and a reference index. The unidirectional prediction restriction is applied, and as with existing bidirectional predictions, only two motion-compensated predictions may be required for each CU.

[0118] When GPM is applied to a CU, a geometric partition index (GPM index) indicating the partition mode (angle and offset) of the geometric partition, and two merge indices for the partition may be signaled. The maximum number of GPM candidate sizes is explicitly signaled in the SPS, which can show syntax binary evolution for the GPM merge index. After each part of the geometric partition is predicted, sample values ​​at the geometric partition edges may be adjusted by blending processing with adaptive weights. This is the prediction signal for the overall CU, and as with other prediction modes, transformation and quantization processes may be applied to the overall CU. Finally, the motion field of the CU predicted using GPM may be preserved.

[0119] When GPM is applied, the syntax in Table 1 is signaled, and based on the signaled GPM index (merge_gpm_partition_idx), the angle and distance are derived using Table 2, and one block may be divided into two arbitrary regions based on the angle and distance. Prediction blocks are generated for each region using different motion prediction information, and predictions may be made by blending the generated prediction blocks and using them as the prediction blocks for that block.

[0120] [Table 1]

[0121] [Table 2]

[0122] Uni-prediction candidate list construction

[0123] In GPM, the unidirectional candidate list may be directly derived from the merge candidate list constructed by the extended merge prediction process. n may be represented as the index of the unidirectional motion in the geometric unidirectional candidate list. The LX motion vector of the nth extended merge candidate has the same X as the parity of n and may be used as the nth unidirectional motion vector for the GPM. Such a motion vector is represented as "x" in Figure 10. If such an LX motion vector does not exist for the nth extended merge candidate, the L(1-X) motion vector of the same candidate may be used instead as the unidirectional motion vector for the GPM.

[0124] Blending using geometrically divided edges

[0125] After each part of the geometric division is predicted using its own motion, blending may be applied to the two predicted signals to induce samples around the geometric division edges. Blending weights for each position in the CU may be derived based on the distance between the individual position and the division edge.

[0126] The distance from position (x,y) to the dividing edge may be derived based on equation 1.

[0127]

number

[0128] In Equation 1, i and j are indices for the angle and offset of the geometric partition, and may be determined by the signaled geometric partition index (merge_gpm_partition_idx). ρ x,j and ρ y,j This may be determined by the angle index i.

[0129] The weights for each geometric partition may be derived based on Equation 2.

[0130]

number

[0131] PartIdx may be determined by the angle index i. Weighted value w o An example of this is shown in Figure 11.

[0132] Saving motion fields to GPM

[0133] Mv1 for the first part of the geometric partition, Mv2 for the second part of the geometric partition, and Mv, which is a combination of Mv1 and Mv2, may be stored in the CU motion field coded in GPM.

[0134] The "save motion vector type" for each individual position within the motion field may be derived based on Equation 3.

[0135]

number

[0136] In equation 3, motionIdx is identical to d(4x+2,4y+2) and may be recalculated using the first equation in equation 1. partIdx may be determined by the angle index i.

[0137] If the value of sType is the same as 0 or 1, Mv0 or Mv1 may be stored in the corresponding motion field. Otherwise, if the value of sType is the same as 2, an Mv combined from Mv1 and Mv2 may be stored. The combined Mv may be generated using the following process:

[0138] - If Mv1 and Mv2 are derived from different reference picture lists (one from L0 and the other from L1), Mv1 and Mv2 simply combine to form a bidirectional motion vector.

[0139] - Instead, if Mv1 and Mv2 are derived from the same reference picture list, only the unidirectional motion Mv2 will be saved.

[0140] Examples

[0141] This disclosure proposes a method for improving prediction accuracy by deriving compensation information (illumination compensation parameters) for a reference block using GPM and applying it to the reference block.

[0142] The method of this disclosure can be summarized as follows:

[0143] 1) When the current block is partitioned and predicted using GPM, compensation information may be derived using the reconstructed peripheral region around the current block (the peripheral region of the current block) and the reconstructed peripheral region around the reference block (the peripheral region of the reference block), and the reference block may be compensated (corrected) using the derived compensation information to generate a predicted sample of the current block.

[0144] 2) When the current block is partitioned and predicted by the GPM, the peripheral region of the current block and the peripheral region of the reference block may be determined by the partitioning direction (partitioning angle) and / or partitioning offset (distance) of the GPM.

[0145] 3) When a block is currently divided and predicted by the GPM, information regarding whether or not to induce or apply compensation information may be signaled in various ways.

[0146] Example 1

[0147] Example 1 relates to a method for deriving compensation information using the peripheral regions of the current block and the reference block when the current block is partitioned and predicted by the GPM.

[0148] Figure 12 shows an example of the peripheral regions of the current block and the peripheral regions of the reference block. Referring to Figure 12, the current block C in the current video is a GPM-encoded block and may be divided into two parts P0' and P1'. Prediction samples for the current block may be generated using these two parts. In this disclosure, the parts may be referred to as the "first block" and the "second block," respectively. For example, P0' may be called the "first block" and P1' may be called the "second block." As another example, P0' may be called the "second block" and P1' may be called the "first block." In Figure 12, the region represented by the pattern may be the peripheral region of the current block or the peripheral region of the reference block.

[0149] Information regarding the GPM division direction may be determined by the bitstream or guided by the video decoder 200. If the motion vectors for P0' and P1' are mv0 and mv1, respectively, the two reference blocks P0 and P1 may be obtained from the reference video using their respective motion vectors. Figure 12 shows an example of obtaining two reference blocks from one reference video, where each reference block may be obtained from a different list (L0 or L1) and / or a different reference video. The two reference blocks may be blended by a weighted average based on the geometric division boundary.

[0150] Whether to use reference block correction using compensation information according to the present disclosure may be determined by a bitstream or may be induced by the video decoder 200. When compensation is applied, the compensation information may be induced using the peripheral area of the current block and the peripheral area of the reference block. The compensation information is information for compensating the luminance difference between the current block and the reference block, and may be defined as in Equation 4.

[0151] [Number]

[0152] In Equation 4, Cur N (x, y) represents the peripheral area of the current block, and Ref N (x, y) represents the peripheral area of the reference block. α and β may represent compensation information. For example, α represents scale information applied to the luminance value of the peripheral area of the current block or the luminance value of the peripheral area of the reference block, and β represents offset information applied to the luminance value to which the scale information is applied among the luminance values of the peripheral area of the current block or the luminance value of the peripheral area of the reference block. The compensation information may be induced as α and β that minimize the difference between Cur N (x, y) and Ref N (x, y). The difference between Cur N (x, y) and Ref N (x, y) may be the sum of absolute differences (SAD) or the sum of square differences (SSD). The compensation information may be induced based on various areas such as a picture, a tile, a slice, a coding tree unit (CTU), a coding unit, and the like.

[0153] The compensation information may be induced for P0’ and P1’ respectively. That is, the compensation information α0 and β0 for P0’ may be induced based on the peripheral area of the current block and the peripheral area of P0. Also, the compensation information α1 and β1 for P1’ may be induced based on the peripheral area of the current block and the peripheral area of P1.

[0154] The current block's prediction sample (predicted block or final predicted block) may be derived based on equation 5.

[0155]

number

[0156] JPEG2026510382000009.jpg39164

[0157] According to the Example 1 described above, by obtaining compensation information for the lower blocks of the current block and applying it to the reference block, prediction performance can be improved, and thus coding efficiency can be improved.

[0158] Example 2

[0159] Example 2 relates to a method for determining the peripheral regions of the current block and the reference block based on the division direction and / or division offset of the GPM. Below, Example 2 will be described focusing on the method for determining the peripheral region of the current block. The peripheral region of the reference block may be determined in the same way as the peripheral region of the current block. Alternatively, the peripheral region of the reference block may be determined at a position corresponding to the peripheral region of the current block.

[0160] Figure 13 shows an example of block division using GPM. The division direction of the GPM shown in Figure 13 is illustrative, and the current block may be divided at other angles (other directions) not shown in Figure 13. Also, the lines (division lines) indicating each division direction may be translated horizontally or vertically, thereby dividing the current block into parts with unequal proportions.

[0161] Figure 14 shows various examples of P0' and P1' depending on the division direction when the current block is divided. Referring to Figure 14, the surrounding area of ​​the current block may change depending on the division direction, the position of P0', and the position of P1'.

[0162] Figure 15(a) shows the peripheral regions of P0' and P1' when the division direction is 0. Referring to Figure 15(a), P0' is adjacent to the peripheral region of the current block in the left and upward directions, so compensation information for P0' may be derived using both the peripheral region in the left and upward directions. In contrast, P1' is adjacent to the peripheral region of the current block only in the upward direction, so compensation information for P1' may be derived using only the peripheral region in the upward direction.

[0163] Figure 15(b) shows the peripheral regions of P0' and P1' when the division direction is number 4. Referring to Figure 15(b), since P0' is adjacent to the peripheral region of the current block only in the leftward direction, compensation information for P0' can be derived using only the peripheral region in the leftward direction. In contrast, since P1' is adjacent to the peripheral region of the current block only in the upward direction, compensation information for P1' can be derived using only the peripheral region in the upward direction.

[0164] Thus, the compensation information for P0' and P1' can be derived using only the surrounding regions adjacent to each of them. Therefore, the compensation information for P0' and P1' can be derived more accurately, which can improve prediction performance and reduce the difference signal, thereby improving coding efficiency.

[0165] Table 3 shows the position of P0' relative to the surrounding region and the position of P1' relative to the surrounding region, depending on the division direction.

[0166] [Table 3]

[0167] Referring to Table 3, division directions 16 to 30 may be defined as the same direction as division directions 0 to 14, with only the division offset differing. In other words, the peripheral region used for inducing compensation information when the division direction is 16 to 30 may be defined as the same as the peripheral region used for inducing compensation information when the division direction is 0 to 14.

[0168] When the division directions are 12th and 28th, P1' does not need to be adjacent to the surrounding area of ​​the current block. Therefore, in such cases, compensation information for P1' does not need to be induced. That is, when the division directions are 12th and 28th, only compensation information for P0' may be induced and applied.

[0169] Depending on the embodiment, compensation information may be derived using only the pixels adjacent to P0' and P1', respectively. The pixels adjacent to P0' and P1' may vary depending on the division direction and the division offset (the displacement value that moves the division line horizontally or vertically).

[0170] Figure 16(a) shows the pixel-level peripheral regions for P0' and P1' when the division direction is 0. Referring to Figure 16(a), P0' is adjacent only to all pixels on the left and some pixels on the upper side, so compensation information for P0' may be derived using all pixels on the left and some pixels on the upper side. In contrast, P1' is adjacent only to some pixels on the upper side, so compensation information for P1' may be derived using only some pixels on the upper side.

[0171] Figure 16(b) shows the pixel-level peripheral regions for P0' and P1' when the division direction is 5. Referring to Figure 16(b), P0' is adjacent to only a portion of the pixels on the left side, so the compensation information for P0' can be derived using only a portion of the pixels on the left side. In contrast, P1' is adjacent to only a portion of the pixels on the left side and all the pixels on the upper side, so the compensation information for P1' can be derived using a portion of the pixels on the left side and all the pixels on the upper side.

[0172] Thus, the compensation information for P0' and P1' can be derived using only the surrounding regions adjacent to each of them. Therefore, the compensation information for P0' and P1' can be derived more accurately, which can improve prediction performance and reduce the difference signal, thereby improving coding efficiency.

[0173] The above describes how to define the peripheral region when the division offset value is 0. When the division offset value is not 0, P0' and P1' may be divided in an unequal ratio, and in this case, the respective peripheral regions of P0' and P1' may vary depending on the division offset value.

[0174] Figure 17(a) shows the peripheral regions of P0' and P1' when the division direction is 4 and the division offset value is 2. Referring to Figure 17(a), unlike the case where the division direction is 4 and the division offset value is 0, P0' is adjacent to the periphery in both the left and upper directions, so the compensation information for P0' may be derived using both the peripheral region in the left direction and the peripheral region in the upper direction. In contrast, P1' is adjacent to the periphery only in the upper direction, so the compensation information for P1' may be derived using only the peripheral region in the upper direction. That is, the compensation information for P1' may be derived in the same way as when the division offset value is 0.

[0175] Figure 17(b) shows the pixel-level peripheral regions for P0' and P1' when the division direction is 4 and the division offset value is 2. Referring to Figure 17(b), P0' is adjacent to all pixels on the left and some pixels on the upper side, so compensation information for P0' may be derived using all pixels on the left and some pixels on the upper side. In contrast, P1' is adjacent only to some pixels on the upper side, so compensation information for P1' may be derived using only some pixels on the upper side.

[0176] Thus, the compensation information for P0' and P1' can be adaptively induced by selecting surrounding regions or pixels using encoded information such as the division offset value or block size, which can improve the accuracy of the prediction.

[0177] On the other hand, when using only adjacent peripheral regions or adjacent pixels to derive compensation information for P0' and P1', there may be insufficient pixels to derive the compensation information. To solve this problem, the present invention proposes a method to derive compensation information by expanding a portion of the region or increasing the number of pixels when the peripheral region of P0' and P1' is not a portion of the surrounding region or is a portion of the surrounding pixels.

[0178] For example, if the size of the regions adjacent to P0' and P1', respectively, is located only on either the left or the top side compared to the reference region used to derive compensation information when GPM mode is not applied, the number of pixels is reduced by half, making it possible to perform the peripheral region expansion proposed in this application.

[0179] Figure 18(a) shows an example of peripheral region expansion when the division direction is number 4. Referring to Figure 18(a), P0' is adjacent to the periphery only in the left direction (a0), and P1' is adjacent to the periphery only above (a1). In such cases, the peripheral region of P0' may be expanded by a factor of 2 in the left direction from the existing left-side region a0 adjacent to P0' to the non-adjacent region na0 not adjacent to P0', and the peripheral region of P1' may be expanded by a factor of 2 in the upward direction from the existing upper-side region a1 adjacent to P1' to the non-adjacent region na1 not adjacent to P1'. That is, the non-adjacent region may be located in the same direction as the adjacent region.

[0180] Another example of peripheral region expansion is when either P0' or P1' is adjacent to only a portion of pixels in the leftward or upward direction, in which case the peripheral region may be expanded in a predetermined direction.

[0181] Figure 18(b) shows an example of peripheral region expansion when the division direction is 4 and the division offset value is 2. Referring to Figure 18(b), P0' is adjacent to only some pixels a0 in the upward direction, and P1' is also adjacent to only some pixels a1 in the upward direction. In such a case, the peripheral region of P0' may be expanded to include pixels na0 that are not adjacent. That is, the peripheral region of P0' may be expanded to the right to include some pixels na1 that are adjacent to P1' in the upward direction. Similarly, the peripheral region of P1' may be expanded to include pixels na1 that are not adjacent. That is, the peripheral region of P1' may be expanded to the left to include some pixels na0 that are adjacent to P0' in the upward direction.

[0182] Figure 18(c) shows an example of peripheral region expansion when the division direction is 5. Referring to Figure 18(c), P0' is adjacent to only some pixels a0 in the leftward direction, and P1' is also adjacent to only some pixels a1 in the leftward direction. In such a case, the peripheral region of P0' may be expanded to include pixels na0 that are not adjacent. That is, the peripheral region of P0' may be expanded upward to include some pixels na1 that are adjacent to P1' in the leftward direction. Similarly, the peripheral region of P1' may be expanded to include pixels na1 that are not adjacent. That is, the peripheral region of P1' may be expanded downward to include some pixels na0 that are adjacent to P0' in the leftward direction.

[0183] The above explanation focused on the case where the extended peripheral region and the existing peripheral region are adjacent to each other. However, the extended peripheral region and the existing peripheral region do not necessarily have to be adjacent to each other. Similarly, when peripheral region extension is applied, a larger number of pixels in the extended peripheral region may be used for compensation information guidance compared to the pixels in the existing peripheral region.

[0184] Example 3

[0185] Example 3 relates to a method for signaling and acquiring application information, which is information regarding whether or not compensation is applied. The compensation according to this disclosure may manifest in various forms depending on the input video, and for this reason, the compensation according to this disclosure may be performed adaptively to the input video. That is, whether or not compensation is applied to a specific unit may be determined in header information such as a higher-level syntax. In addition, whether or not compensation is applied may be signaled at lower units such as coding units or CTU units.

[0186] Table 4 shows an example of a method for transmitting whether or not compensation is used (allowed) in SPS. [Table 4]

[0187] The semantics of the sps_illumination_compensation_enabled_flag syntax element may be defined as follows:

[0188] A value of 1 for sps_illumination_compensation_enabled_flag indicates that ic_flag exists in the CLVS intercoding unit syntax, while a value of 0 for sps_illumination_compensation_enabled_flag indicates that ic_flag does not exist in the CLVS intercoding unit syntax. If sps_illumination_compensation_enabled_flag does not exist, its value may be inferred to be the same as 0.

[0189] If it is determined that compensation is to be used at the higher level syntax, the merge_data obtained by joining Tables 5 and 6 may signal a syntax indicating whether or not compensation is applied on a per-encoded-unit basis. The GPM of this disclosure may be applied when the values ​​of regular_merge_flag and ciip_flag are both 0, in which case the compensation application flag (application information) of this disclosure may be signaled.

[0190] [Table 5]

[0191] [Table 6]

[0192] A value of 1 for regular_merge_flag indicates that the regular merge mode is currently used to generate the inter-prediction parameters for the coding unit. In other words, regular_merge_flag indicates whether the merge mode (regular merge mode) is currently applied to the block.

[0193] A value of 1 for mmvd_merge_flag indicates that MMVD is currently used to generate interprediction parameters for the coding unit. In other words, mmvd_merge_flag indicates whether or not MMVD is currently applied to the block.

[0194] The mmvd_cand_flag indicates that the first (0) or second (1) candidate in the merge candidate list will be used with the MVD derived from mmvd_distance_idx and mmvd_direction_idx.

[0195] mmvd_distance_idx represents the index used to derive MmvdDistance, and mmvd_direction_idx represents the index used to derive MmvdSign.

[0196] The merge_subblock_flag indicates whether the subblock merge mode is currently used to generate subblock-based interprediction parameters for the coding unit. In other words, merge_subblock_flag indicates whether the subblock merge mode (or affine merge mode) is currently applied to the block.

[0197] `merge_subblock_idx` represents the merge candidate index in the subblock-based merge candidate list.

[0198] ciip_flag indicates whether combined interpicture merging and intrapicture prediction are currently applied to the coding block.

[0199] merge_triangle_idx0 represents the first merge candidate index in the triangle-based motion compensation candidate list, and merge_triangle_idx1 represents the second merge candidate index in the triangle-based motion compensation candidate list.

[0200] `merge_idx` represents the merge candidate index in the merge candidate list.

[0201] `merge_gpm_partition_idx` indicates the partition shape of the GPM. If `merge_gpm_partition_idx` does not exist, its value can be inferred to be the same as 0.

[0202] merge_gpm_idx0 indicates the first merge candidate index in the geometric partition-based motion compensation candidate list, and merge_gpm_idx1 indicates the second merge candidate index in the geometric partition-based motion compensation candidate list. If merge_gpm_idx0 does not exist, its value may be inferred to be the same as 0, and if merge_gpm_idx1 does not exist, its value may be inferred to be the same as 0.

[0203] The semantics of the `ic_flag` syntax element, which is application information, may be defined as follows:

[0204] The `ic_flag` can indicate whether illumination compensation is currently being used within the coding unit. A value of `ic_flag` of 1 indicates that illumination compensation is currently applied to the coding unit, while a value of `ic_flag` of 0 indicates that illumination compensation is not currently applied to the coding unit. If `ic_flag` is not present, its value may be inferred to be equivalent to 0.

[0205] As shown in the combination of Tables 5 and 6, when one ic_flag is signaled, the compensation proposed in this application may be applied to both P0' and P1' separated by the GPM.

[0206] In another embodiment, the presence or absence of compensation may be signaled separately for each of the P0' and P1' segments separated by the GPM. In such a case, the merge_data may signal the application information, such as the join of Tables 7 and 8.

[0207] [Table 7]

[0208] [Table 8]

[0209] The semantics of the part0_ic_flag and part1_ic_flag syntax elements may be defined as follows:

[0210] The part0_ic_flag can indicate whether illumination compensation is currently used for partition 0 within the coding unit. A value of 1 for part0_ic_flag indicates that illumination compensation is currently applied to partition 0 of the coding unit, while a value of 0 indicates that illumination compensation is not currently applied to partition 0 of the coding unit. If part0_ic_flag does not exist, its value may be inferred to be the same as 0. Partition 0 may be the aforementioned P0'. The part0_ic_flag can be called the "first application information".

[0211] part1_ic_flag can indicate whether illumination compensation is currently used for partition 1 within the coding unit. A value of 1 for part1_ic_flag indicates that illumination compensation is currently applied to partition 1 of the coding unit, while a value of 0 for part1_ic_flag indicates that illumination compensation is not currently applied to partition 1 of the coding unit. If part1_ic_flag does not exist, its value may be inferred to be the same as 0. Partition 1 may be P1' as described above. part1_ic_flag can be called "second application information".

[0212] In another embodiment, compensation may be applied to only one of the blocks P0' and P1' partitioned by the GPM. In such a case, the merge_data may signal application information such as the join of Tables 9 and 10.

[0213] [Table 9]

[0214] [Table 10]

[0215] The semantics of the ic_flag and part0_ic_flag syntax elements may be defined as follows:

[0216] `ic_flag` can indicate whether illumination compensation is currently used within the coding unit. A value of `ic_flag` of 1 indicates that illumination compensation is currently applied to the coding unit, while a value of `ic_flag` of 0 indicates that illumination compensation is not currently applied to the coding unit. If `ic_flag` is not present, its value may be inferred to be equivalent to 0. `ic_flag` can be referred to as "third application information".

[0217] part0_ic_flag can indicate whether illumination compensation currently applies to partition 0 or partition 1 within the geometric block. A value of 1 for part0_ic_flag indicates that illumination compensation currently applies to partition 0 within the geometric block, and a value of 0 for part0_ic_flag indicates that illumination compensation currently applies to partition 1 within the geometric block. If part0_ic_flag does not exist, its value may be inferred to be the same as 0. part0_ic_flag can be called the "fourth application information".

[0218] In another embodiment, if the division offset value is not zero, one of the blocks P0' and P1' is divided larger than the other. In this case, illumination compensation may be restricted to apply only to the block of P0' and P1' that has the larger size. Therefore, in the combination of Tables 9 and 10, the fourth application information (part0_ic_flag) is not signaled, thereby reducing the number of bits required for signaling the fourth application information and improving encoding efficiency.

[0219] Video encoding method and video decoding method

[0220] The following describes the video encoding method and video decoding method for GPM proposed in this disclosure. The video encoding method described below may be performed by a video encoding device 100, and the video decoding method may be performed by a video decoding device 200.

[0221] Figure 19 shows a video encoding method according to one embodiment of the present disclosure, and Figure 20 shows a video decoding method according to one embodiment of the present disclosure.

[0222] Referring to Figure 19, at least one reference block may be determined for the prediction of the current block to which the GPM applies (S1910). The reference block may be P0 and / or P1 as described herein.

[0223] Compensation information related to the compensation of the reference block may be derived based on the surrounding area of ​​the reference block and the surrounding area of ​​the current block (S1920). As described above, the compensation information is information for compensating for the luminance difference between the current block and the reference block, and may include scale information applied to the luminance value of the surrounding area of ​​the current block or the luminance value of the surrounding area of ​​the reference block and / or offset information applied to the luminance value of the surrounding area of ​​the current block or the luminance value of the surrounding area of ​​the reference block to which the scale information has been applied.

[0224] Predicted samples of the current block (predicted block or final predicted block) may be generated based on compensation information (S1930). Information related to whether or not GPM is applied, motion information indicating a reference block, and information related to whether or not compensation is applied (application information) may be encoded into a bitstream and signaled from the video encoding device 100 to the video decoding device 200.

[0225] Referring to Figure 20, at least one piece of motion information may be obtained for the current block to which GPM is applied (S2010). At least one reference block for predicting the current block may be determined based on the motion information.

[0226] Compensation information related to the compensation of the reference block may be derived based on the surrounding area of ​​the reference block and the surrounding area of ​​the current block (S2020). As described above, the compensation information is information for compensating for the luminance difference between the current block and the reference block, and may include scale information applied to the luminance value of the surrounding area of ​​the current block or the luminance value of the surrounding area of ​​the reference block and / or offset information applied to the luminance value of the surrounding area of ​​the current block or the luminance value of the surrounding area of ​​the reference block to which the scale information has been applied.

[0227] Prediction samples of the current block (prediction block or final prediction block) may be generated based on compensation information (S2030).

[0228] The surrounding regions of the current block and the reference block may be determined based on the GPM partitioning direction. Furthermore, the surrounding regions of the current block and the reference block may also be determined based on the GPM partitioning offset.

[0229] On the other hand, the surrounding area of ​​the current block may include an adjacent area that is adjacent to the upper or left side of the current block. In such a case, under predetermined conditions, a non-adjacent area that is not adjacent to the first block may be included in the surrounding area to induce compensation information, and a non-adjacent area that is not adjacent to the second block may be included in the surrounding area to induce compensation information. That is, if predetermined conditions are met, the surrounding area may be expanded to include the non-adjacent area.

[0230] Figures 21 to 23 show various examples of video encoding and decoding methods for determining the surrounding region.

[0231] Referring to Figure 21, it may be determined whether the adjacent area is located only above or to the left of the current block (S2110). If it is determined that the adjacent area is located only above or to the left of the current block, the surrounding area may be determined including the non-adjacent area located in the same direction as the adjacent area (S2120). Conversely, if it is not determined that the adjacent area is located only above or to the left of the current block, the surrounding area may be determined without including the non-adjacent area (S2130).

[0232] Referring to Figure 22, it may be determined whether the region adjacent to the upper side of the current block includes the region adjacent only to the upper side of the first block (e.g., P0') and the region adjacent only to the upper side of the second block (e.g., P1') (S2210). If it is determined that the region adjacent to the upper side of the current block includes the region adjacent only to the upper side of the first block and the region adjacent only to the upper side of the second block, the surrounding region (of the first block) may be determined including the region adjacent to the upper side of the second block (S2220). Conversely, if it is determined that the region adjacent to the upper side of the current block includes only the region adjacent only to the upper side of the first block (i.e., it does not include the region adjacent only to the upper side of the second block), the upper surrounding region (of the first block) may be determined without including the region adjacent to the upper side of the second block (i.e., only the region adjacent only to the upper side of the first block) (S2230).

[0233] Referring to Figure 23, it may be determined whether the region adjacent to the left of the current block includes the region adjacent only to the left of the first block (e.g., P0') and the region adjacent only to the left of the second block (e.g., P1') (S2310). If it is determined that the region adjacent to the left of the current block includes the region adjacent only to the left of the first block and the region adjacent only to the left of the second block, the surrounding region (of the first block) may be determined including the region adjacent to the left of the second block (S2320). Conversely, if it is determined that the region adjacent to the left of the current block includes only the region adjacent only to the left of the first block (i.e., it does not include the region adjacent only to the left of the second block), the left surrounding region (of the first block) may be determined without including the region adjacent to the left of the second block (i.e., only the region adjacent only to the left of the first block) (S2330).

[0234] Figure 24 shows a video encoding method and a video decoding method that determine whether compensation is performed or whether compensation information is induced depending on the GPM division direction.

[0235] Referring to Figure 24, it can be determined whether the first block (e.g., P0') and the second block (e.g., P1') are triangular in shape when divided diagonally upwards to the right (division direction 12 or 28, division offset value 0) (S2410). If the first and second blocks are triangular in shape when divided diagonally upwards to the right, the second block does not need to have a peripheral area above and to the left of the current block. Therefore, if it is determined that the first and second blocks are triangular in shape when divided diagonally upwards to the right, compensation information for the second block is not induced, and illumination compensation for the second block is not performed (S2420). Conversely, if the first and second blocks are not triangular in shape when divided diagonally upwards to the right, both the first and second blocks may have a peripheral area above or to the left of the current block. Therefore, compensation information for both the first and second blocks may be induced, and illumination compensation for the second block may be performed (S2430).

[0236] Figure 25 shows a video encoding method for encoding application information, and Figure 26 shows a video decoding method for acquiring application information.

[0237] Referring to Figure 25, it is determined whether compensation is permissible (S2510), and if compensation is permissible, it may be determined whether compensation is applied (S2520). If compensation is applied, the value of sps_illumination_compensation_enabled_flag, which indicates whether compensation is permissible, may be encoded to 1, and the value of ic_flag, which is application information indicating whether compensation is applied, may be encoded to 1 (S2530). Conversely, if compensation is not permissible, the value of sps_illumination_compensation_enabled_flag is encoded to 0, and the application information does not need to be encoded (S2540). Also, if compensation is permissible but not applied to the current block, the value of sps_illumination_compensation_enabled_flag may be encoded to 1, and the value of ic_flag may be encoded to 0 (S2550).

[0238] Referring to Figure 26, sps_illumination_compensation_enabled_flag is obtained from the bitstream (S2610), and based on the value of the obtained sps_illumination_compensation_enabled_flag, it may be determined whether compensation is permitted or not (S2620). If compensation is permitted, application information (ic_flag) is obtained from the bitstream (S2630), and based on the obtained ic_flag, it may be determined whether compensation is applied (i.e., whether compensation information is induced or not) (S2640). If it is determined that compensation is applied, compensation information may be induced (S2650). Conversely, if it is determined that compensation is not applied or that compensation is not permitted, compensation information does not need to be induced (S2660).

[0239] Depending on the embodiment, the application information may include first application information (part0_ic_flag) indicating whether compensation applies to the first block, and second application information (part1_ic_flag) indicating whether compensation applies to the second block. That is, whether compensation applies is determined for the first block and the second block respectively, and the application information may also be signaled for the first block and the second block respectively.

[0240] Depending on the embodiment, the application information may include a third application information (ic_flag in Table 9) indicating whether compensation applies to the current block, and a fourth application information (part0_ic_flag in Table 9) indicating which of the first and second blocks is subject to compensation. In this case, the fourth application information may be signaled when the third application information indicates that compensation applies to the current block.

[0241] In this embodiment, the fourth application information does not need to be signaled if the size of the first block is larger than the size of the second block. That is, compensation may be limited to apply only to the block with the larger size among the first and second blocks, and in such a case, if the size of the first block is larger than the size of the second block, compensation will apply only to the first block. Therefore, in such a case, the fourth application information does not need to be included in the application information.

[0242] Figure 27 illustrates a content streaming system to which the embodiments of this disclosure can be applied.

[0243] As shown in Figure 27, a content streaming system to which an embodiment of the present disclosure is applied may broadly include an encoding server, a streaming server, a web server, media storage, user equipment, and multimedia input devices.

[0244] The encoding server is responsible for compressing content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream, and transmitting this bitstream to the streaming server. As another example, if a multimedia input device such as a smartphone, camera, or camcorder directly generates the bitstream, the encoding server may be omitted.

[0245] The bitstream may be generated by a video encoding method and / or video encoding apparatus to which an embodiment of the present disclosure is applied, and the streaming server may temporarily store the bitstream in the process of transmitting or receiving the bitstream.

[0246] The streaming server transmits multimedia data to user devices based on user requests via a web server, and the web server can act as an intermediary to inform users of available services. When a user requests a desired service from the web server, the web server transmits it to the streaming server, and the streaming server can transmit multimedia data to the user. In this case, the content streaming system may include a separate control server, in which case the control server can play a role in controlling commands and responses between the devices within the content streaming system.

[0247] The streaming server can receive content from media storage and / or encoding servers. For example, when receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.

[0248] Examples of the user device may include a mobile phone, a smart phone, a laptop computer, a digital broadcast terminal, a PDA (personal digital assistants), a PMP (portable multimedia player), a navigation device, a slate PC, a tablet PC, an ultrabook, a wearable device (e.g., a smartwatch, smart glass, a head mounted display (HMD)), a digital TV, a desktop computer, a digital signage, and the like.

[0249] Each server in the content streaming system may be operated as a distributed server, and in this case, the data received by each server may be distributedly processed.

[0250] The scope of the present disclosure includes software or machine-executable instructions (e.g., an operation system, an application, firmware, a program, etc.) that enable operations according to the methods of various embodiments to be executed on a device or a computer, and a non-transitory computer-readable medium on which such software or instructions are stored and executable on the device or the computer.

[0251] 〔Industrial Applicability〕 The embodiments according to the present disclosure can be used for encoding / decoding video.

[0252] 〔Claims at the Time of International Application〕 〔Claim 1〕 A video decoding method performed by a video decoding apparatus, obtaining at least one motion information for a current block to which a Geometry Partitioning Mode (GPM) is applied; A step of deriving compensation information related to the compensation of the reference block based on the peripheral region of at least one reference block indicated by the motion information and the peripheral region of the current block, A video decoding method comprising the step of generating a predicted sample of the current block based on the compensation information. [Claim 2] The video decoding method according to claim 1, wherein the compensation information includes scale information applied to the luminance value of the peripheral region of the current block or the luminance value of the peripheral region of the reference block. [Claim 3] The video decoding method according to claim 2, further comprising offset information applied to the luminance value from among the luminance values ​​of the surrounding area of ​​the current block or the luminance values ​​of the surrounding area of ​​the reference block to which the scale information is applied. [Claim 4] The video decoding method according to claim 1, wherein the peripheral region of the current block is determined based on the division direction of the GPM. [Claim 5] The video decoding method according to claim 4, wherein the peripheral region of the current block is determined based on the division offset of the GPM. [Claim 6] The video decoding method according to claim 1, wherein the surrounding area of ​​the current block includes an adjacent area adjacent to the upper or left side of the current block. [Claim 7] Based on the fact that the adjacent region is located only above or to the left of the current block, the surrounding region of the current block further includes non-adjacent regions that are not adjacent to the current block. The video decoding method according to claim 6, wherein the non-adjacent region is located above or to the left of the current block, in the same direction as the adjacent region. [Claim 8] The aforementioned current block includes a first block and a second block divided by the GPM, Based on the fact that the region adjacent to the upper side of the current block includes the region adjacent only to the upper side of the first block and the region adjacent only to the upper side of the second block, the peripheral region of the first block further includes the region adjacent only to the upper side of the second block. The video decoding method according to claim 6, wherein the region adjacent to the left of the current block includes the region adjacent only to the left of the first block and the region adjacent only to the left of the second block, and the peripheral region of the first block further includes the region adjacent only to the left of the second block. [Claim 9] The aforementioned current block includes a first block and a second block divided by the GPM, The video decoding method according to claim 1, wherein no compensation information for the second block is induced based on the fact that the first block and the second block are triangular in shape, divided in a diagonal direction sloping upwards to the right. [Claim 10] The video decoding method according to claim 1, wherein the step of inducing the compensation information is performed on the basis that applicable information obtained from the bitstream indicates that compensation is applied. [Claim 11] The aforementioned current block includes a first block and a second block divided by the GPM, The video decoding method according to claim 10, wherein the application information includes first application information indicating whether or not compensation is applied to the first block, and second application information indicating whether or not compensation is applied to the second block. [Claim 12] The aforementioned current block includes a first block and a second block divided by the GPM, The aforementioned application information includes a third application information indicating whether or not compensation applies to the current block, and a fourth application information indicating which block of the first block or the second block to which compensation applies. The video decoding method according to claim 10, wherein the fourth application information is obtained based on the third application information indicating that compensation is applied to the current block. [Claim 13] The video decoding method according to claim 12, wherein the application information does not include the fourth application information, based on the size of the first block being larger than the size of the second block. [Claim 14] A video encoding method performed by a video encoding device, A step of determining at least one reference block for predicting the current block to which the Geometry Partitioning Mode (GPM) is applied, A step of deriving compensation information related to the compensation of the reference block based on the surrounding area of ​​the reference block and the surrounding area of ​​the current block, A video encoding method comprising the step of generating a predicted sample of the current block based on the compensation information. [Claim 15] A method for transmitting a bitstream generated by a video encoding method, The aforementioned video encoding method is A step of determining at least one reference block for predicting the current block to which the Geometry Partitioning Mode (GPM) is applied, A step of deriving compensation information related to the compensation of the reference block based on the surrounding area of ​​the reference block and the surrounding area of ​​the current block, A method comprising the step of generating a predicted sample of the current block based on the compensation information.

Claims

1. A video decoding method performed by a video decoding device, A step of obtaining at least one piece of motion information for the current block to which the Geometry Partitioning Mode (GPM) is applied, A step of deriving compensation information related to the compensation of the reference block based on the peripheral region of at least one reference block indicated by the motion information and the peripheral region of the current block, A video decoding method comprising the step of generating a predicted sample of the current block based on the compensation information.

2. The video decoding method according to claim 1, wherein the compensation information includes scale information applied to the luminance value of the peripheral region of the current block or the luminance value of the peripheral region of the reference block.

3. The video decoding method according to claim 2, further comprising offset information applied to the luminance value from among the luminance values ​​of the surrounding area of ​​the current block or the luminance values ​​of the surrounding area of ​​the reference block to which the scale information is applied.

4. The video decoding method according to claim 1, wherein the peripheral region of the current block is determined based on the division direction of the GPM.

5. The video decoding method according to claim 4, wherein the peripheral region of the current block is determined based on the division offset of the GPM.

6. The video decoding method according to claim 1, wherein the surrounding area of ​​the current block includes an adjacent area adjacent to the upper or left side of the current block.

7. Based on the fact that the adjacent region is located only above or to the left of the current block, the surrounding region of the current block further includes non-adjacent regions that are not adjacent to the current block. The video decoding method according to claim 6, wherein the non-adjacent region is located above or to the left of the current block, in the same direction as the adjacent region.

8. The aforementioned current block includes a first block and a second block divided by the GPM, Based on the fact that the region adjacent to the upper side of the current block includes the region adjacent only to the upper side of the first block and the region adjacent only to the upper side of the second block, the peripheral region of the first block further includes the region adjacent only to the upper side of the second block. The video decoding method according to claim 6, wherein the region adjacent to the left side of the current block includes the region adjacent only to the left side of the first block and the region adjacent only to the left side of the second block, and the peripheral region of the first block further includes the region adjacent only to the left side of the second block.

9. The aforementioned current block includes a first block and a second block divided by the GPM, The video decoding method according to claim 1, wherein no compensation information for the second block is induced based on the fact that the first block and the second block are triangular in shape, divided in a diagonal direction sloping upwards to the right.

10. The video decoding method according to claim 1, wherein the step of inducing the compensation information is performed on the basis that applicable information obtained from the bitstream indicates that compensation is applied.

11. The aforementioned current block includes a first block and a second block divided by the GPM, The video decoding method according to claim 10, wherein the application information includes first application information indicating whether or not compensation is applied to the first block, and second application information indicating whether or not compensation is applied to the second block.

12. The aforementioned current block includes a first block and a second block divided by the GPM, The application information includes a third application information indicating whether or not compensation applies to the current block, and a fourth application information indicating the block to which compensation applies among the first block or the second block. The video decoding method according to claim 10, wherein the fourth application information is obtained based on the third application information indicating that compensation is applied to the current block.

13. The video decoding method according to claim 12, wherein the application information does not include the fourth application information, based on the fact that the size of the first block is larger than the size of the second block.

14. A video encoding method performed by a video encoding device, A step of determining at least one reference block for predicting the current block to which the Geometry Partitioning Mode (GPM) is applied, A step of deriving compensation information related to the compensation of the reference block based on the surrounding area of ​​the reference block and the surrounding area of ​​the current block, A video encoding method comprising the step of generating a predicted sample of the current block based on the compensation information.

15. A method for transmitting a bitstream generated by a video encoding method, The aforementioned video encoding method is A step of determining at least one reference block for predicting the current block to which the Geometry Partitioning Mode (GPM) is applied, A step of deriving compensation information related to the compensation of the reference block based on the surrounding area of ​​the reference block and the surrounding area of ​​the current block, A method comprising the step of generating a predicted sample of the current block based on the compensation information.