Image coding / decoding method, apparatus, and method for transmitting a bitstream that performs weighted prediction.

The image coding/decoding method enhances encoding/decoding efficiency by employing weighted prediction and bitstream management, addressing the challenge of high-resolution image data transmission and storage costs.

JP2026063366APending Publication Date: 2026-04-10LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
LG ELECTRONICS INC
Filing Date
2026-01-26
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

The increasing demand for high-resolution, high-quality images leads to a significant increase in the amount of information transmitted, necessitating highly efficient image compression technology to reduce transmission and storage costs.

Method used

An image coding/decoding method and apparatus that performs weighted prediction, including Bi-prediction with CU-level Weight (BCW) and PROF, with mechanisms for determining default or explicit weighted prediction based on flags and indices, and a method for transmitting and storing generated bitstreams.

Benefits of technology

Improves encoding/decoding efficiency and enables effective transmission and storage of high-resolution, high-quality images by optimizing prediction methods and bitstream handling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026063366000001_ABST
    Figure 2026063366000001_ABST
Patent Text Reader

Abstract

An image encoding / decoding method and apparatus are provided. [Solution] The image decoding method according to the present disclosure is an image decoding method performed by an image decoding device, and includes the steps of: a first flag indicating whether or not to perform weighted prediction on the current block; and inducing a weight index (BcwIdx) of BCW (Bi-prediction with CU-level Weight) on the current block; determining whether to perform a default weighted prediction or an explicit weighted prediction on the current block based on the first flag and the BcwIdx; and generating a predicted block for the current block by performing the determined method.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This disclosure relates to an image coding / decoding method, apparatus, and method for transmitting a bitstream, and more particularly to an image coding / decoding method, apparatus, and method for transmitting a bitstream generated by the image coding method / apparatus of this disclosure that performs weighted prediction considering PROF (Prediction Refinement with Optical Flow). [Background technology]

[0002] Recently, the demand for high-resolution, high-quality images, such as HD (High Definition) and UHD (Ultra High Definition) images, has been increasing in various fields. As image data becomes higher resolution and higher quality, the amount of information or bits transmitted increases relatively compared to conventional image data. This increase in the amount of information or bits transmitted leads to increased transmission and storage costs.

[0003] This necessitates highly efficient image compression technology to effectively transmit, store, and reproduce high-resolution, high-quality image information. [Overview of the Initiative] [Problems that the invention aims to solve]

[0004] The purpose of this disclosure is to provide an image coding / decoding method and apparatus with improved coding / decoding efficiency.

[0005] Furthermore, this disclosure aims to provide an image coding / decoding method and apparatus for weighted prediction.

[0006] Furthermore, this disclosure aims to provide an image coding / decoding method and apparatus that performs weighted prediction or BCW (Bi-prediction with CU-level Weight) taking PROF into consideration.

[0007] Furthermore, this disclosure aims to provide a method for transmitting a bitstream generated by an image encoding method or apparatus according to this disclosure.

[0008] Furthermore, this disclosure aims to provide a recording medium that stores a bitstream generated by the image encoding method or apparatus according to this disclosure.

[0009] Furthermore, this disclosure aims to provide a recording medium that stores a bitstream received by the image decoding device provided herein, decoded, and used for image restoration.

[0010] The technical problems that this disclosure seeks to solve are not limited to those described above, and other technical problems not mentioned above will be clearly understood by a person with ordinary skill in the art to which this disclosure pertains from the following description. [Means for solving the problem]

[0011] An image decoding method according to one aspect of the present disclosure may include: a first flag indicating whether or not to perform weighted prediction on the current block; a step of inducing a weight index (BcwIdx) of BCW (Bi-prediction with CU-level Weight) on the current block; a step of determining whether to perform default weighted prediction or explicit weighted prediction on the current block based on the first flag and the BcwIdx; and a step of generating a predicted block for the current block by performing the determined method.

[0012] In the image decoding method according to this disclosure, the first flag can be determined differently based on the slice type of the current slice to which the current block belongs.

[0013] In the image decoding method according to the present disclosure, when the slice type of the current slice is a P slice, the first flag is derived from the value of pps_weighted_pred_flag signaled in the PPS, and when the slice type of the current slice is a B slice, the first flag can be derived from the value of pps_weighted_bipred_flag signaled in the PPS.

[0014] In the image decoding method according to the present disclosure, the BcwIdx is derived based on the syntax element bcw_idx signaled via the bitstream, and when the bcw_idx does not exist in the bitstream, the BcwIdx can be derived to be 0.

[0015] In the image decoding method according to the present disclosure, the bcw_idx can be parsed from the bitstream based on the weighted prediction flag of the reference picture of the current block.

[0016] In the image decoding method according to the present disclosure, the bcw_idx can be parsed from the bitstream when all of the weighted prediction flags of the reference pictures of the current block are 0.

[0017] In the image decoding method according to the present disclosure, when the first flag is 0 or the BcwIdx is not 0, default weighted prediction can be performed on the current block.

[0018] In the image decoding method according to the present disclosure, when the first flag is 1 and the BcwIdx is 0, explicit weighted prediction can be performed on the current block.

[0019] In the image decoding method according to the present disclosure, the default weighted prediction can perform BCW or average sum based on the BcwIdx.

[0020] In the image decoding method according to the present disclosure, when the BcwIdx is 0, an average sum is performed on the current block, and when the BcwIdx is not 0, BCW can be performed on the current block.

[0021] In the image decoding method according to the present disclosure, the explicit weighted prediction can be performed based on a weighting parameter (weight and offset) for a reference picture of the current block.

[0022] In the image decoding method according to the present disclosure, the weighting parameter can be explicitly signaled via a bitstream.

[0023] An image decoding apparatus according to another aspect of the present disclosure includes a memory and at least one processor, and the at least one processor induces a first flag indicating whether to perform weighted prediction on a current block and a weight index (BcwIdx) of BCW (Bi-prediction with CU-level Weight) on the current block, and determines whether to perform default weighted prediction or explicit weighted prediction on the current block based on the first flag and the BcwIdx, and can generate a predicted block for the current block by performing the determined method.

[0024] An image coding method according to another aspect of the present disclosure may include the steps of: a first flag indicating whether or not to perform a weighted prediction on a current block; determining a weight index (BcwIdx) for BCW (Bi-prediction with CU-level Weight) on the current block; determining whether to perform a default weighted prediction or an explicit weighted prediction on the current block based on the first flag and the BcwIdx; and generating a predicted block for the current block by performing the determined method.

[0025] A computer-readable recording medium according to another aspect of the present disclosure can store a bitstream generated by an image encoding method or image encoding apparatus of the present disclosure.

[0026] The features described above, which are a brief summary of this disclosure, are merely illustrative examples of the detailed description of this disclosure described below and do not limit the scope of this disclosure. [Effects of the Invention]

[0027] According to this disclosure, an image encoding / decoding method and apparatus with improved encoding / decoding efficiency can be provided.

[0028] Furthermore, this disclosure provides an image coding / decoding method and apparatus for performing weighted prediction.

[0029] Furthermore, according to this disclosure, an image coding / decoding method and apparatus can be provided that performs weighted prediction or BCW considering PROF.

[0030] Furthermore, this disclosure provides a method for transmitting a bitstream generated by an image encoding method or apparatus according to this disclosure.

[0031] Furthermore, according to this disclosure, a recording medium storing a bitstream generated by the image encoding method or apparatus according to this disclosure can be provided.

[0032] Furthermore, according to this disclosure, a recording medium can be provided that stores a bitstream that is received by the image decoding device according to this disclosure, decoded, and used for image restoration.

[0033] The effects obtained from this disclosure are not limited to those described above, and other effects not mentioned above will be clearly understood by a person with ordinary skill in the art to which this disclosure pertains from the following description. [Brief explanation of the drawing]

[0034] [Figure 1] This figure schematically illustrates a video coding system to which the embodiments described herein can be applied. [Figure 2] This figure schematically shows an image encoding device to which the embodiments of this disclosure can be applied. [Figure 3] This figure schematically shows an image decoding apparatus to which the embodiments of this disclosure can be applied. [Figure 4] This is a flowchart showing a video / image encoding method based on interpretation. [Figure 5] This figure illustrates the configuration of the interpretation unit 180 according to this disclosure. [Figure 6] This is a flowchart showing a video / image decoding method based on interpretation. [Figure 7] This figure illustrates the configuration of the interpretation unit 260 according to this disclosure. [Figure 8] This diagram illustrates surrounding blocks that can be used as candidates for spatial merge. [Figure 9] This diagram schematically illustrates a method for constructing a merge candidate list according to an example of this disclosure. [Figure 10] This diagram illustrates candidate pairs for redundancy checks performed on spatial candidates. [Figure 11]This diagram illustrates how to scale the motion vector of a time candidate. [Figure 12] This diagram illustrates the position that guides the time candidates. [Figure 13] This figure schematically illustrates a method for constructing a motion vector predictor candidate list according to an example of this disclosure. [Figure 14] This is a diagram illustrating the parameter model of affine modes. [Figure 15] This is a diagram illustrating how to generate a list of affine merge candidates. [Figure 16] This diagram illustrates CPMV induced from surrounding blocks. [Figure 17] This diagram illustrates the surrounding blocks used to guide candidate combinational affine merges. [Figure 18] This diagram illustrates how to generate a list of Affine MVP candidates. [Figure 19] This is a diagram illustrating the peripheral blocks of the subblock-based TMVP mode. [Figure 20] This diagram illustrates how to induce a motion vector field according to a subblock-based TMVP mode. [Figure 21] This figure shows the extended CU for performing BDOF. [Figure 22] This diagram shows the relationship between Δv(i,j), v(i,j), and the subblock motion vector. [Figure 23] This flowchart shows an example of how PROF, BCW, WP, and / or average sum are performed according to this disclosure. [Figure 24] This flowchart shows other examples of performing PROF, BCW, WP, and / or average sum according to this disclosure. [Figure 25] Table 7 shows a flowchart illustrating an example of performing BCW or WP using the method described. [Figure 26] Table 8 shows a flowchart illustrating an example of performing BCW or WP using the method described. [Figure 27] Table 9 shows a flowchart illustrating an example of performing BCW or WP using the method described. [Figure 28] Table 10 shows a flowchart illustrating an example of performing BCW or WP using the method described. [Figure 29] This figure illustrates a content streaming system to which the embodiments of this disclosure can be applied. [Modes for carrying out the invention]

[0035] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the attached drawings, so that they can be easily implemented by a person with ordinary skill in the art to which the present disclosure pertains. However, the present disclosure can be implemented in a variety of different forms and is not limited to the embodiments described herein.

[0036] In describing embodiments of this disclosure, if it is determined that a specific description of a known configuration or function would obscure the gist of this disclosure, such detailed description will be omitted. In the drawings, parts unrelated to the description of this disclosure will be omitted, and similar parts will be denoted by the same reference numerals.

[0037] In this disclosure, when one component is described as being “connected,” “joined,” or “linked” to another component, this can include not only direct connections but also indirect connections where another component exists between them. Furthermore, when one component is described as “containing” or “having” another component, this means, unless otherwise stated to the contrary, that it may include another component rather than excluding it.

[0038] In this disclosure, terms such as "first," "second," etc., are used solely for the purpose of distinguishing one component from another, and do not limit the order or importance of the components unless otherwise specified. Therefore, within the scope of this disclosure, a first component in one embodiment may be called a second component in another embodiment, and similarly, a second component in one embodiment may be called a first component in another embodiment.

[0039] In this disclosure, components that are distinguished from each other are used to clearly describe their respective characteristics and do not necessarily mean that the components are separate. In other words, multiple components may be integrated to constitute a single hardware or software unit, or a single component may be distributed to constitute multiple hardware or software units. Therefore, such integrated or distributed embodiments are also included in the scope of this disclosure, without needing to be specifically mentioned.

[0040] In this disclosure, the components described in various embodiments are not necessarily essential components, and some may be optional components. Therefore, embodiments consisting of a subset of the components described in one embodiment are also included in the scope of this disclosure. Furthermore, embodiments that include additional components in addition to the components described in various embodiments are also included in the scope of this disclosure.

[0041] This disclosure relates to the encoding and decoding of images, and the terms used in this disclosure may have their ordinary meanings in the art to which this disclosure pertains, unless otherwise defined herein.

[0042] In this disclosure, "picture" generally means a unit representing any one image within a specific time period, and "slice / tile" is an encoding unit that constitutes part of a picture, and a single picture can consist of one or more slices / tiles. Furthermore, a slice / tile may contain one or more CTUs (coding tree units).

[0043] In this disclosure, “pixel” or “pel” may mean the smallest unit that constitutes a picture (or image). The term “sample” may also be used as a counterpart to pixel. A sample may generally represent a pixel or a pixel value, or it may represent only the pixel / pixel value of the luma component, or only the pixel / pixel value of the chroma component.

[0044] In this disclosure, “unit” can refer to a basic unit of image processing. A unit may include at least one of a specific region of a picture and information associated with that region. A unit may be used interchangeably with terms such as “sample array,” “block,” or “area,” as it may be used. Generally, an M×N block may include a set (or array) of samples (or sample arrays) or transform coefficients consisting of M columns and N rows.

[0045] In this disclosure, “current block” can mean any one of the following: “current coding block,” “current coding unit,” “block to encode,” “block to decode,” or “block to process.” If prediction is performed, “current block” can mean “current prediction block” or “block to predict.” If transformation (inverse transformation) / quantization (inverse quantization) is performed, “current block” can mean “current transformation block” or “block to transform.” If filtering is performed, “current block” can mean “block to filter.”

[0046] In this disclosure, " / " and "," may be interpreted as "and / or." For example, "A / B" and "A, B" may be interpreted as "A and / or B." Also, "A / B / C" and "A, B, C" may mean "at least one of A, B and / or C."

[0047] In this disclosure, “or” may be interpreted as “and / or.” For example, “A or B” may mean 1) “A” only, 2) “B” only, or 3) “A and B.” Alternatively, in this disclosure, “or” may mean “additionally or alternatively.”

[0048] Overview of the video coding system

[0049] Figure 1 shows the video coding system according to this disclosure.

[0050] A video coding system according to one embodiment may include an encoding device 10 and a decoding device 20. The encoding device 10 can transmit encoded video and / or image information or data to the decoding device 20 via a digital storage medium or network in file or streaming format.

[0051] An encoding device 10 according to one embodiment may include a video source generation unit 11, an encoding unit 12, and a transmission unit 13. A decoding device 20 according to one embodiment may include a receiving unit 21, a decoding unit 22, and a rendering unit 23. The encoding unit 12 may be called a video / image encoding unit, and the decoding unit 22 may be called a video / image decoding unit. The transmission unit 13 may be included in the encoding unit 12. The receiving unit 21 may be included in the decoding unit 22. The rendering unit 23 may also include a display unit, which may be configured as a separate device or external component.

[0052] The video source generation unit 11 can acquire video / images through processes such as video / image capture, synthesis, or generation. The video source generation unit 11 may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, or a video / image archive containing previously captured video / images. The video / image generation device may include, for example, a computer, tablet, and smartphone, and may generate video / images (electronically). For example, virtual video / images may be generated via a computer, in which case the video / image capture process may be replaced by a process in which the relevant data is generated.

[0053] The encoding unit 12 can encode the input video / image. The encoding unit 12 can perform a series of steps such as prediction, transformation, and quantization for compression and encoding efficiency. The encoding unit 12 can output the encoded data (encoded video / image information) in bitstream format.

[0054] The transmission unit 13 can transmit encoded video / image information or data, output in bitstream format, to the receiving unit 21 of the decoding device 20 via a digital storage medium or network in file or streaming format. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray®, HDD, and SSD. The transmission unit 13 may include elements for generating media files via a predetermined file format and may include elements for transmission via a broadcast / communication network. The receiving unit 21 can extract / receive the bitstream from the storage medium or network and transmit it to the decoding unit 22.

[0055] The decoding unit 22 can decode the video / image by performing a series of steps such as inverse quantization, inverse transform, and prediction, corresponding to the operation of the encoding unit 12.

[0056] The rendering unit 23 can render the decoded video / image. The rendered video / image can be displayed via the display unit.

[0057] Overview of Image Encoding Devices

[0058] Figure 2 is a schematic diagram showing an image encoding device to which the embodiments of this disclosure can be applied.

[0059] As shown in Figure 2, the image coding device 100 may include an image splitting unit 110, a subtraction unit 115, a transformation unit 120, a quantization unit 130, an inverse quantization unit 140, an inverse transformation unit 150, an addition unit 155, a filtering unit 160, a memory 170, an inter-prediction unit 180, an intra-prediction unit 185, and an entropy coding unit 190. The inter-prediction unit 180 and the intra-prediction unit 185 can together be called the "prediction unit". The transformation unit 120, the quantization unit 130, the inverse quantization unit 140, and the inverse transformation unit 150 may be included in a residual processing unit. The residual processing unit may further include a subtraction unit 115.

[0060] All or at least some of the multiple components constituting the image encoding device 100 can be implemented by a single hardware component (e.g., an encoder or processor) depending on the embodiment. Furthermore, the memory 170 may include a DPB (decoded picture buffer) and can be implemented by a digital storage medium.

[0061] The image splitting unit 110 can split an input image (or picture, frame) input to the image encoding device 100 into one or more processing units. For example, the processing units may be called coding units (CUs). Coding units can be obtained by recursively splitting a coding tree unit (CTU) or the largest coding unit (LCU) using a QT / BT / TT (Quad-tree / binary-tree / ternary-tree) structure. For example, a single coding unit can be split into multiple coding units of deeper depth based on a quad-tree structure, a binary-tree structure and / or a ternary-tree structure. For the splitting of coding units, a quad-tree structure may be applied first, followed by a binary-tree structure and / or a ternary-tree structure. Based on the final coding unit that cannot be further split, the coding procedure according to this disclosure can be performed. The largest coding unit can be used as the final coding unit, or a lower-depth coding unit obtained by dividing the largest coding unit can be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and / or restoration, as described later. As another example, the processing units of the coding procedure may be prediction units (PU) or transformation units (TU). The prediction unit and the transformation unit may be divided or partitioned from the final coding unit, respectively. The prediction unit may be a unit of sample prediction, and the transformation unit may be a unit that derives transformation coefficients and / or a unit that derives a residual signal from transformation coefficients.

[0062] The prediction unit (inter-prediction unit 180 or intra-prediction unit 185) can make predictions for the block to be processed (current block) and generate a predicted block that includes prediction samples for the current block. The prediction unit can determine whether intra-prediction or inter-prediction is applied to the current block or on a CU basis. The prediction unit can generate various information regarding the prediction of the current block and transmit it to the entropy coding unit 190. The prediction information can be encoded by the entropy coding unit 190 and output in bitstream format.

[0063] The intra-prediction unit 185 can predict the current block by referring to a sample in the current picture. The referenced sample may be located in the vicinity (neighbor) or at a distance from the current block, according to the intra-prediction mode and / or intra-prediction technique. The intra-prediction mode may include multiple non-directional modes and multiple directional modes. The non-directional modes may include, for example, a DC mode and a Planar mode. The directional modes may include, for example, 33 or 65 directional prediction modes, depending on the degree of fineness of the prediction direction. However, this is merely an example, and more or fewer directional prediction modes may be used depending on the settings. The intra-prediction unit 185 may also determine the prediction mode to be applied to the current block using the prediction modes applied to the surrounding blocks.

[0064] The interprediction unit 180 can derive a predicted block relative to the current block based on a reference block (reference sample array) identified by motion vectors on the reference picture. In this case, in order to reduce the amount of motion information transmitted in interprediction mode, motion information can be predicted in units of blocks, subblocks, or samples based on the correlation of motion information between the surrounding blocks and the current block. The motion information may include motion vectors and reference picture indices. The motion information may further include interprediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of interprediction, the surrounding blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture containing the reference block and the reference picture containing the temporal neighboring block may be the same or different from each other. The temporal neighboring block may be called a collocated reference block, collocated CU (colCU), etc. The reference picture containing the temporal neighboring block may be called a collocated picture (colPic). For example, the interpretation unit 180 can construct a motion information candidate list based on surrounding blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Interpretation can be performed based on various prediction modes; for example, in skip mode and merge mode, the interpretation unit 180 can use the motion information of surrounding blocks as the motion information of the current block. In skip mode, unlike merge mode, the residual signal may not be transmitted.In motion vector prediction (MVP) mode, the motion vector of the surrounding block is used as the motion vector predictor, and the motion vector of the current block can be signaled by encoding the motion vector difference and an indicator for the motion vector predictor. The motion vector difference can represent the difference between the motion vector of the current block and the motion vector predictor.

[0065] The prediction unit can generate a prediction signal based on various prediction methods and / or techniques described later. For example, the prediction unit can apply intra-prediction or inter-prediction to predict the current block, and can also apply intra-prediction and inter-prediction simultaneously. A prediction method that applies intra-prediction and inter-prediction simultaneously to predict the current block can be called CIIP (combined inter and intra prediction). The prediction unit can also perform intra-block copy (IBC) to predict the current block. Intra-block copy can be used for content image / video coding such as in games, for example, in SCC (screen content coding). IBC is a method of predicting the current block using a reference block that has already been restored in the current picture at a predetermined distance from the current block. When IBC is applied, the position of the reference block in the current picture can be encoded as a vector (block vector) corresponding to the predetermined distance.

[0066] The predicted signal generated by the prediction unit can be used to generate a reconstructed signal or a residual signal. The subtraction unit 115 can generate a residual signal (residual block, residual sample array) by subtracting the predicted signal output from the prediction unit (predicted block, predicted sample array) from the input image signal (original block, original sample array). The generated residual signal can be transmitted to the conversion unit 120.

[0067] The transformation unit 120 can generate transformation coefficients by applying transformation techniques to the residual signal. For example, the transformation techniques may include at least one of the following: DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), KLT (Karhunen-Loeve Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT refers to a transformation obtained from a graph, where the relationship information between pixels is represented by a graph. CNT refers to a transformation obtained by generating a prediction signal using all previously reconstructed pixels. The transformation process can be applied to pixel blocks of the same size and square shape, or to non-square, variable-sized blocks.

[0068] The quantization unit 130 can quantize the conversion coefficients and transmit them to the entropy coding unit 190. The entropy coding unit 190 can encode the quantized signal (information about the quantized conversion coefficients) and output it in bitstream format. The information about the quantized conversion coefficients can be called residual information. The quantization unit 130 can rearrange the block-form quantized conversion coefficients into a one-dimensional vector format based on the coefficient scan order, and can also generate information about the quantized conversion coefficients based on the one-dimensional vector format of the quantized conversion coefficients.

[0069] The entropy coding unit 190 can perform various coding methods, such as exponential Golomb, CAVLC (context-adaptive variable length coding), and CABAC (context-adaptive binary arithmetic coding). In addition to the quantized conversion coefficients, the entropy coding unit 190 can also encode information necessary for video / image restoration (e.g., the values ​​of syntax elements) together or separately. The encoded information (e.g., encoded video / image information) can be transmitted or stored in bitstream format in units of NAL (network abstraction layer) units. The video / image information may further include information about various parameter sets, such as adaptive parameter sets (APS), picture parameter sets (PPS), sequence parameter sets (SPS), or video parameter sets (VPS). The video / image information may also further include general constraint information. The signaling information, transmitted information and / or syntax elements referred to in this disclosure may be encoded via the encoding procedure described above and included in the bitstream.

[0070] The bitstream can be transmitted over a network or stored on a digital storage medium. Here, the network may include broadcast networks and / or communication networks, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. A transmission unit (not shown) for transmitting the signal output from the entropy encoding unit 190 and / or a storage unit (not shown) for storing it may be provided as an internal / external element of the image encoding device 100, or the transmission unit may be provided as a component of the entropy encoding unit 190.

[0071] The quantized conversion coefficients output from the quantization unit 130 can be used to generate a residual signal. For example, by applying inverse quantization and inverse transformation to the quantized conversion coefficients via the inverse quantization unit 140 and the inverse transformation unit 150, a residual signal (residual block or residual sample) can be reconstructed.

[0072] The adder 155 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter-prediction unit 180 or the intra-prediction unit 185. If there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as the reconstructed block. The adder 155 may be called the reconstruction unit or the reconstructed block generation unit. The generated reconstructed signal can be used for intra-prediction of the next block to be processed in the current picture, or, as described later, for inter-prediction of the next picture after filtering.

[0073] On the other hand, as will be discussed later, LMCS (luma mapping with chroma scaling) can also be applied during the picture encoding process.

[0074] The filtering unit 160 can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit 160 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be stored in the memory 170, specifically in the DPB of the memory 170. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, and bilateral filter. The filtering unit 160 can generate various filtering-related information, as will be described later in the explanation of each filtering method, and transmit it to the entropy coding unit 190. The filtering-related information can be encoded by the entropy coding unit 190 and output in bitstream format.

[0075] The corrected restored picture transmitted to memory 170 can be used as a reference picture in the interpretation unit 180. When interpretation is applied via this, the image encoding device 100 can avoid prediction mismatches between the image encoding device 100 and the image decoding device, and can also improve encoding efficiency.

[0076] The DPB in memory 170 can store the modified restored picture for use as a reference picture in the inter-prediction unit 180. Memory 170 can store motion information of blocks from which motion information in the current picture has been derived (or encoded) and / or motion information of blocks in the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit 180 for use as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. Memory 170 can store restored samples of restored blocks in the current picture and transmit them to the intra-prediction unit 185.

[0077] Overview of the image decoding device

[0078] Figure 3 is a schematic diagram showing an image decoding apparatus to which the embodiments of this disclosure can be applied.

[0079] As shown in Figure 3, the image decoding device 200 can be configured to include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an additive unit 235, a filtering unit 240, a memory 250, an inter-prediction unit 260, and an intra-prediction unit 265. The inter-prediction unit 260 and the intra-prediction unit 265 can together be called the "prediction unit". The inverse quantization unit 220 and the inverse transform unit 230 can be included in the residual processing unit.

[0080] All or at least some of the multiple components constituting the image decoding device 200 can be implemented by a single hardware component (e.g., a decoder or processor) depending on the embodiment. Furthermore, the memory 170 may include a DPB and can be implemented by a digital storage medium.

[0081] An image decoding device 200, upon receiving a bitstream containing video / image information, can restore the image by executing a process corresponding to the process performed in the image encoding device 100 in Figure 1. For example, the image decoding device 200 can perform decoding using the processing unit applied in the image encoding device. Therefore, the decoding processing unit can be, for example, a coding unit. The coding unit can be obtained by dividing a coding tree unit or a maximum coding unit. The restored image signal decoded and output via the image decoding device 200 can then be reproduced via a playback device (not shown).

[0082] The image decoding device 200 can receive the signal output from the image encoding device 1 in bitstream format. The received signal can be decoded via the entropy decoding unit 210. For example, the entropy decoding unit 210 can parse the bitstream to derive information necessary for image restoration (or picture restoration) (e.g., video / image information). The video / image information may further include information about various parameter sets, such as adaptive parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS). The video / image information may also further include general constraint information. The image decoding device may further use the parameter set information and / or the general constraint information to decode the image. The signaling information, received information, and / or syntax elements referred to in this disclosure can be obtained from the bitstream by decoding via the decoding procedure. For example, the entropy decoding unit 210 can decode information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the values ​​of syntax elements necessary for image reconstruction and the quantized values ​​of conversion coefficients related to the residual. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element from the bitstream, determines a context model using the syntax element information to be decoded, the decoding information of the surrounding blocks and the blocks to be decoded, or the symbol / bin information decoded in a previous step, predicts the probability of bin occurrence based on the determined context model, and performs arithmetic decoding of the bins to generate symbols corresponding to the values ​​of each syntax element. At this time, after determining the context model, the CABAC entropy decoding method can update the context model using the decoded symbol / bin information for the context model of the next symbol / bin.Of the information decoded by the entropy decoding unit 210, information related to prediction is provided to the prediction unit (inter-prediction unit 260 and intra-prediction unit 265), and the residual values ​​that have undergone entropy decoding in the entropy decoding unit 210, i.e., quantized conversion coefficients and related parameter information, can be input to the inverse quantization unit 220. In addition, of the information decoded by the entropy decoding unit 210, information related to filtering can be provided to the filtering unit 240. On the other hand, a receiving unit (not shown) that receives signals output from the image coding device may be further provided as an internal / external element of the image decoding device 200, or the receiving unit may be provided as a component of the entropy decoding unit 210.

[0083] On the other hand, the image decoding device according to this disclosure may be called a video / image / picture decoding device. The image decoding device may also include an information decoder (video / image / picture information decoder) and / or a sample decoder (video / image / picture sample decoder). The information decoder may include an entropy decoding unit 210, and the sample decoder may include at least one of an inverse quantization unit 220, an inverse transform unit 230, an adder unit 235, a filtering unit 240, a memory 250, an inter-prediction unit 260, and an intra-prediction unit 265.

[0084] The inverse quantization unit 220 can inverse quantize the quantized transformation coefficients and output the transformation coefficients. The inverse quantization unit 220 can rearrange the quantized transformation coefficients in a two-dimensional block format. In this case, the rearrangement can be performed based on the coefficient scan order performed by the image encoding device. The inverse quantization unit 220 can perform inverse quantization on the quantized transformation coefficients using quantization parameters (e.g., quantization step size information) to obtain the transformation coefficients.

[0085] The inverse conversion unit 230 can inversely convert the conversion coefficients to obtain residual signals (residual blocks, residual sample arrays).

[0086] The prediction unit can make predictions for the current block and generate a predicted block containing prediction samples for the current block. Based on the prediction information output from the entropy decoding unit 210, the prediction unit can determine whether intra-prediction or inter-prediction is applied to the current block and can determine a specific intra / inter-prediction mode (prediction technique).

[0087] As described in the explanation of the prediction unit of the image coding device 100, the prediction unit can generate prediction signals based on various prediction methods (techniques) described later.

[0088] The intra-prediction unit 265 can predict the current block by referring to the samples in the current picture. The description of the intra-prediction unit 185 can also be applied to the intra-prediction unit 265.

[0089] The interprediction unit 260 can derive a predicted block relative to the current block based on a reference block (reference sample array) identified by motion vectors on a reference picture. In this case, to reduce the amount of motion information transmitted in interprediction mode, motion information can be predicted in block, sub-block, or sample units based on the correlation of motion information between surrounding blocks and the current block. The motion information may include motion vectors and reference picture indices. The motion information may further include interprediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In interprediction, surrounding blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. For example, the interprediction unit 260 can construct a motion information candidate list based on surrounding blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Interprediction can be performed based on various prediction modes (techniques), and the prediction information may include information indicating the mode (technique) of interprediction for the current block.

[0090] The adder 235 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the acquired residual signal to the predicted signal (predicted block, predicted sample array) output from the prediction unit (including the inter-prediction unit 260 and / or intra-prediction unit 265). The description of the adder 155 can also be applied to the adder 235.

[0091] On the other hand, as will be discussed later, LMCS (luma mapping with chroma scaling) can also be applied during the picture decoding process.

[0092] The filtering unit 240 can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit 240 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be stored in the memory 250, specifically in the DPB of the memory 250. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, and bilateral filter.

[0093] The restored picture stored (modified) in the DPB of memory 250 can be used as a reference picture in the inter-prediction unit 260. Memory 250 can store motion information of blocks from which motion information in the current picture has been derived (or decoded) and / or motion information of blocks in the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit 260 for use as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. Memory 250 can store restored samples of restored blocks in the current picture and transmit them to the intra-prediction unit 265.

[0094] In this specification, the embodiments described for the filtering unit 160, inter-prediction unit 180, and intra-prediction unit 185 of the image coding device 100 can be applied similarly or in a corresponding manner to the filtering unit 240, inter-prediction unit 260, and intra-prediction unit 265 of the image decoding device 200, respectively.

[0095] Overview of Interpretation

[0096] Image encoding / decoding devices can perform interpretation on a block-by-block basis to derive predicted samples. Interpretation can refer to a prediction technique derived in a manner that is dependent on the data elements of pictures other than the current picture. When interpretation is applied to the current block, predicted blocks for the current block can be induced based on the reference block identified by the motion vector on the reference picture.

[0097] In this case, in order to reduce the amount of motion information transmitted in interprediction mode, motion information of the current block can be derived based on the correlation of motion information between the surrounding block and the current block, and motion information can be derived at the block, subblock, or sample level. In this case, the motion information may include motion vectors and reference picture indices. The motion information may further include interprediction type information. Here, interprediction type information can mean direction information of the interprediction. Interprediction type information can indicate that the current block is predicted using one of L0 prediction, L1 prediction, or Bi prediction.

[0098] When interpretation is applied to the current block, the surrounding blocks of the current block may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. In this case, the reference picture containing the reference block for the current block and the reference picture containing the temporal neighboring block may be the same or different. The temporal neighboring block may be called a collocated reference block or collocated coding unit (colCU). The reference picture containing the temporal neighboring block may be called a collocated picture (colPic).

[0099] On the other hand, a list of motion information candidates can be constructed based on the surrounding blocks of the current block, and in this case, flags or index information can be signaled to indicate which candidate is used to derive the motion vector and / or reference picture index of the current block.

[0100] Motion information may include L0 motion information and / or L1 motion information based on the interpretation type. A motion vector in the L0 direction can be defined as an L0 motion vector or MVL0, and a motion vector in the L1 direction can be defined as an L1 motion vector or MVL1. A prediction based on an L0 motion vector can be defined as an L0 prediction, a prediction based on an L1 motion vector can be defined as an L1 prediction, and a prediction based on both the L0 motion vector and the L1 motion vector can be defined as a biprediction. Here, an L0 motion vector can mean a motion vector associated with reference picture list L0, and an L1 motion vector can mean a motion vector associated with reference picture list L1.

[0101] The reference picture list L0 can include pictures earlier in the output order than the current picture as reference pictures, and the reference picture list L1 can include pictures later in the output order than the current picture. In this case, earlier pictures can be defined as forward (reference) pictures, and later pictures can be defined as reverse (reference) pictures. On the other hand, the reference picture list L0 can further include pictures later in the output order than the current picture. In this case, earlier pictures can be indexed first in the reference picture list L0, and later pictures can be indexed next. The reference picture list L1 can further include pictures earlier in the output order than the current picture. In this case, later pictures can be indexed first in the reference picture list L1, and earlier pictures can be indexed next. Here, the output order can correspond to the POC (picture order count) order.

[0102] Figure 4 is a flowchart illustrating a video / image coding method based on interpretation.

[0103] Figure 5 is a diagram illustrating the configuration of the interpretation unit 180 according to this disclosure.

[0104] The encoding method in Figure 4 can be performed by the image encoding device in Figure 2. Specifically, step S410 can be performed by the interprediction unit 180, and step S420 can be performed by the residual processing unit. Specifically, step S420 can be performed by the subtraction unit 115. Step S430 can be performed by the entropy encoding unit 190. The prediction information in step S430 is derived by the interprediction unit 180, and the residual information in step S430 can be derived by the residual processing unit. The residual information may include information regarding the quantized conversion coefficients for the residual sample. As described above, the residual sample is derived as a conversion coefficient via the conversion unit 120 of the image encoding device, and the conversion coefficient can be derived as a quantized conversion coefficient via the quantization unit 130. Information regarding the quantized conversion coefficients can be encoded by the entropy encoding unit 190 via the residual coding procedure.

[0105] The image coding device can perform inter prediction for the current block (S410). The image coding device can derive the inter prediction mode and motion information of the current block and generate prediction samples for the current block. Here, the procedures for determining the inter prediction mode, deriving motion information, and generating prediction samples may be performed simultaneously, or one of the procedures may be performed before the others. For example, as shown in Figure 5, the inter prediction unit 180 of the image coding device may include a prediction mode determination unit 181, a motion information derivation unit 182, and a prediction sample derivation unit 183. The prediction mode determination unit 181 can determine the prediction mode for the current block, the motion information derivation unit 182 can derive the motion information of the current block, and the prediction sample derivation unit 183 can derive prediction samples for the current block. For example, the inter prediction unit 180 of the image coding device can search for blocks similar to the current block within a certain area (search area) of the reference picture via motion estimation and derive reference blocks whose difference from the current block is the minimum or below a certain standard. Based on this, a reference picture index pointing to the reference picture in which the reference block is located can be derived, and a motion vector can be derived based on the positional difference between the reference block and the current block. The image encoding device can determine which of the various prediction modes is applied to the current block. The image encoding device can compare the rate-distortion (RD) cost for the various interprediction modes and determine the optimal prediction mode for the current block. However, the method by which the image encoding device determines the interprediction mode for the current block is not limited to the above example, and various methods can be used.

[0106] For example, the current inter-prediction mode for a block can be determined to be at least one of the following: merge mode, skip mode, MVP mode (Motion Vector Prediction mode), SMVD mode (Symmetric Motion Vector Difference), affine mode, subblock-based merge mode, AMVR mode (Adaptive Motion Vector Resolution mode), HMVP mode (History-based Motion Vector Predictor mode), pair-wise average merge mode, MMVD mode (Merge mode with Motion Vector Differences mode), DMVR mode (Decoder side Motion Vector Refinement mode), CIIP mode (Combined Inter and Intra Prediction mode), and GPM (Geometric Partitioning mode).

[0107] For example, when skip mode or merge mode is applied to the current block, the image encoding device can derive merge candidates from surrounding blocks of the current block and construct a merge candidate list using the deriveted merge candidates. The image encoding device can also derive a reference block from among the reference blocks pointed to by the merge candidates included in the merge candidate list whose difference from the current block is the minimum or below a certain standard. In this case, a merge candidate associated with the derived reference block is selected, and merge index information indicating the selected merge candidate is generated and signaled to the image decoding device. The motion information of the current block can be derived using the motion information of the selected merge candidate.

[0108] As another example, when MVP mode is applied to the current block, the image encoding device can derive motion vector predictor (MVP) candidates from the surrounding blocks of the current block and construct an MVP candidate list using the deriveted MVP candidates. The image encoding device can also use the motion vector of a selected MVP candidate from among the MVP candidates included in the MVP candidate list as the MVP of the current block. In this case, for example, the motion vector pointing to the reference block derived by the motion estimation described above can be used as the motion vector of the current block, and the MVP candidate with the smallest difference between its motion vector and the motion vector of the current block can become the selected MVP candidate. The MVD (motion vector difference), which is the difference obtained by subtracting the MVP from the motion vector of the current block, can be derived. In this case, index information indicating the selected MVP candidate and information regarding the MVD can be signaled to the image decoding device. Furthermore, when MVP mode is applied, the value of the reference picture index can be composed of reference picture index information and separately signaled to the image decoding device.

[0109] The image coding device can derive a residual sample based on the predicted sample (S420). The image coding device can derive the residual sample by comparing the original sample of the current block with the predicted sample. For example, the residual sample can be derived by subtracting the corresponding predicted sample from the original sample.

[0110] The image encoding device can encode image information including prediction information and residual information (S430). The image encoding device can output the encoded image information in bitstream format. The prediction information is information related to the prediction procedure and may include prediction mode information (e.g., skip flag, merge flag, or mode index) and motion information. Of the prediction mode information, the skip flag indicates whether or not the skip mode is applied to the current block, and the merge flag indicates whether or not the merge mode is applied to the current block. Alternatively, the prediction mode information may be information that indicates one of several prediction modes, such as the mode index. If the skip flag and merge flag are both 0, it can be determined that the MVP mode is applied to the current block. The motion information may include candidate selection information (e.g., merge index, mvp flag, or mvp index) which is information for deriving a motion vector. Of the candidate selection information, the merge index can be signaled when the merge mode is applied to the current block and may be information for selecting one of the merge candidates included in the merge candidate list. Of the candidate selection information, the MVP flag or MVP index can be signaled when MVP mode is applied to the current block, and can be information for selecting one of the MVP candidates included in the MVP candidate list. Specifically, the MVP flag can be signaled using the syntax elements mvp_l0_flag or mvp_l1_flag. The motion information may also include the MVD information and / or reference picture index information described above. The motion information may also include information indicating whether L0 prediction, L1 prediction, or bi (Bi) prediction is applied. The residual information is information about the residual sample.The residual information may include information regarding the quantized transformation coefficients for the residual sample.

[0111] The output bitstream can be stored on a (digital) storage medium and transmitted to an image decoding device, or it can be transmitted to an image decoding device via a network.

[0112] On the other hand, as mentioned above, the image coding device can generate a reconstructed picture (a picture including a reconstructed sample and a reconstructed block) based on the reference sample and the residual sample. This is because the image coding device can derive the same prediction results as the image decoding device, thereby improving coding efficiency. Therefore, the image coding device can store the reconstructed picture (or reconstructed sample, reconstructed block) in memory and use it as a picture for interpretation. As mentioned above, in-loop filtering procedures and the like can be further applied to the reconstructed picture.

[0113] Figure 6 is a flowchart showing a video / image decoding method based on interpretation.

[0114] Figure 7 is a diagram illustrating the configuration of the interpretation unit 260 according to this disclosure.

[0115] The image decoding device can perform operations corresponding to those performed by the image encoding device. The image decoding device can make predictions for the current block based on the received prediction information and derive prediction samples.

[0116] The decoding method in Figure 6 can be performed by the image decoding device in Figure 3. Steps S610 to S630 can be performed by the interpretation unit 260, and the prediction information in step S610 and the residual information in step S640 can be obtained from the bitstream by the entropy decoding unit 210. The residual processing unit of the image decoding device can derive a residual sample for the current block based on the residual information (S640). Specifically, the inverse quantization unit 220 of the residual processing unit derives a conversion coefficient by performing inverse quantization based on the quantized conversion coefficient derived from the residual information, and the inverse transformation unit 230 of the residual processing unit can derive a residual sample for the current block by performing an inverse transformation on the conversion coefficient. Step S650 can be performed by the addition unit 235 or the reconstruction unit.

[0117] Specifically, the image decoding device can determine the prediction mode for the current block based on the received prediction information (S610). Based on the prediction mode information in the prediction information, the image decoding device can determine which inter-prediction mode is applied to the current block.

[0118] For example, based on the skip flag, it can be determined whether the skip mode is applied to the current block. Alternatively, based on the merge flag, it can be determined whether the merge mode is applied to the current block or whether the MVP mode is determined. Or, based on the mode index, one of several candidate inter-prediction modes can be selected. The candidate inter-prediction modes may include the skip mode, merge mode, and / or the MVP mode, or may include various inter-prediction modes as described later.

[0119] The image decoding device can derive motion information for the current block based on the determined interprediction mode (S620). For example, if a skip mode or merge mode is applied to the current block, the image decoding device can configure a merge candidate list, which will be described later, and select one of the merge candidates included in the merge candidate list. This selection can be made based on the candidate selection information (merge index) described above. The motion information for the current block can be derived using the motion information for the selected merge candidate. For example, the motion information for the selected merge candidate can be used as the motion information for the current block.

[0120] As another example, when the MVP mode is applied to the current block, the image decoding device can configure an MVP candidate list and use the motion vector of an MVP candidate selected from the MVP candidates included in the MVP candidate list as the MVP of the current block. The selection can be made based on the candidate selection information (mvp flag or mvp index) described above. In this case, the MVD of the current block can be derived based on the information regarding the MVD, and the motion vector of the current block can be derived based on the MVP of the current block and the MVD. Furthermore, the reference picture index of the current block can be derived based on the reference picture index information. The picture pointed to by the reference picture index in the reference picture list for the current block can be derived as the reference picture referenced for interpretation of the current block.

[0121] The image decoding device can generate predicted samples for the current block based on the motion information of the current block (S630). In this case, the reference picture can be derived based on the reference picture index of the current block, and the predicted samples for the current block can be derived using the sample of the reference block pointed to by the motion vector of the current block on the reference picture. Depending on the case, a prediction sample filtering procedure can be further performed on all or some of the predicted samples of the current block.

[0122] For example, as shown in Figure 7, the interpretation unit 260 of the image decoding device may include a prediction mode determination unit 261, a motion information derivation unit 262, and a prediction sample derivation unit 263. The interpretation unit 260 of the image decoding device can determine a prediction mode for the current block based on prediction mode information received by the prediction mode determination unit 261, derive motion information (such as motion vectors and / or reference picture indices) for the current block based on motion information received by the motion information derivation unit 262, and derive prediction samples for the current block by the prediction sample derivation unit 263.

[0123] The image decoding device can generate a residual sample for the current block based on the received residual information (S640). The image decoding device can generate a reconstructed sample for the current block based on the predicted sample and the residual sample, and generate a reconstructed picture based on this (S650). As previously mentioned, further procedures such as in-loop filtering can be applied to the reconstructed picture thereafter.

[0124] As described above, the interpretation procedure may include an interpretation mode determination step, a motion information derivation step based on the determined prediction mode, and a prediction execution (prediction sample generation) step based on the derived motion information. The interpretation procedure can be performed by an image encoding device and an image decoding device, as described above.

[0125] The following describes in more detail the steps for deriving motion information using the prediction mode.

[0126] As mentioned above, interpretation can be performed using motion information of the current block. The image encoding device can derive optimal motion information for the current block through a motion estimation procedure. For example, the image encoding device can use the original block in the original picture to search for a highly correlated similar reference block in fractional pixel units within a defined search range in the reference picture, thereby deriving motion information. Block similarity can be calculated based on the sum of absolute differences (SAD) between the current block and the reference block. In this case, motion information can be derived based on the reference block with the smallest SAD within the search area. The derived motion information can be signaled to the image decoding device in various ways based on the interpretation mode.

[0127] When merge mode is applied to a current block, the movement information of the current block is not transmitted directly, but rather the movement information of the current block is guided using the movement information of surrounding blocks. Therefore, the movement information of the current predicted block can be instructed by transmitting flag information indicating that merge mode has been used and candidate selection information (e.g., merge index) indicating which surrounding blocks were used as merge candidates. In this disclosure, since the current block is the unit of prediction execution, the current block can be used in the same sense as the current predicted block, and the surrounding block can be used in the same sense as the surrounding predicted block.

[0128] The image encoding device can search for merge candidate blocks to be used to guide the motion information of the current block in order to perform a merge mode. For example, up to five merge candidate blocks may be used, but this is not limited to them. The maximum number of merge candidate blocks may be transmitted from the slice header or tile group header, but this is not limited to them. After finding the merge candidate blocks, the image encoding device can generate a merge candidate list and select the merge candidate block with the lowest RD cost from among them as the final merge candidate block.

[0129] This disclosure provides various embodiments of merge candidate blocks that constitute the merge candidate list. The merge candidate list may use, for example, five merge candidate blocks. For example, it may use four spatial merge candidates and one temporal merge candidate.

[0130] Figure 8 illustrates surrounding blocks used as candidates for spatial merge.

[0131] Figure 9 is a schematic diagram illustrating a method for constructing a merge candidate list according to an example of this disclosure.

[0132] The image encoding / decoding device can search for spatially surrounding blocks of the current block and insert the derived spatial merge candidates into the merge candidate list (S910). The spatially surrounding blocks may include, as shown in Figure 8, the left-lower corner surrounding block A0, the left-side surrounding block A1, the right-upper corner surrounding block B0, the upper-side surrounding block B1, and the left-upper corner surrounding block B2 of the current block. However, this is an example, and additional surrounding blocks such as the right-side surrounding block, the lower-side surrounding block, and the right-lower surrounding block can also be used as spatially surrounding blocks. The image encoding / decoding device can detect available blocks by searching for the spatially surrounding blocks based on priority and derive the movement information of the detected blocks as spatial merge candidates. For example, the image encoding / decoding device can construct a merge candidate list by searching the five blocks shown in Figure 8 in the order A1, B1, B0, A0, B2 and sequentially indexing the available candidates.

[0133] The image encoding / decoding device can search for time-peripheral blocks of the current block and insert the derived time merge candidates into the merge candidate list (S920). The time-peripheral blocks can be located on a reference picture which is a different picture from the current picture on which the current block is located. The reference picture on which the time-peripheral blocks are located can be called a collocated picture or a col picture. The time-peripheral blocks can be searched on the col picture in the order of the lower right corner block and the lower right center block of the co-located block relative to the current block. On the other hand, when motion data compression is applied to reduce memory load, specific motion information can be stored as representative motion information for each certain storage unit in the col picture. In this case, it is not necessary to store motion information for all blocks in the certain storage unit, thereby achieving the motion data compression effect. In this case, the certain storage unit can be predetermined, for example, a 16x16 sample unit or an 8x8 sample unit, or size information for the certain storage unit can be signaled from the image encoding device to the image decoding device. When the motion data compression described above is applied, the motion information of the time-period block can be replaced with representative motion information of the constant storage unit in which the time-period block is located. In other words, in this case, from an implementation standpoint, instead of a predicted block lock located at the coordinates of the time-period block, the time merge candidate can be derived based on the motion information of the predicted block that covers the position after being arithmetically shifted to the right by a certain value based on the coordinates of the time-period block (upper left sample position) and then arithmetically shifted to the left. For example, if the constant storage unit is 2 n ×2 nWhen it is in the sample unit, if the coordinates of the time neighboring block are (xTnb, yTnb), the motion information of the prediction block located at the corrected position ((xTnb>>n)<<n), (yTnb>>n)<<n)) can be used for the time merge candidate. Specifically, for example, when the fixed storage unit is 16×16 sample units, if the coordinates of the time neighboring block are (xTnb, yTnb), the motion information of the prediction block located at the corrected position ((xTnb>>4)<<4), (yTnb>>4)<<4)) can be used for the time merge candidate. Or, for example, when the fixed storage unit is 8×8 sample units, if the coordinates of the time neighboring block are (xTnb, yTnb), the motion information of the prediction block located at the corrected position ((xTnb>>3)<<3), (yTnb>>3)<<3)) can be used for the time merge candidate.

[0134] Referring to FIG. 9 again, the image encoding device / image decoding device can check whether the number of current merge candidates is smaller than the number of maximum merge candidates (S930). The number of the maximum merge candidates can be defined in advance or signaled from the image encoding device to the image decoding device. For example, the image encoding device can generate information regarding the number of the maximum merge candidates, encode it, and transmit it to the image decoding device in the form of a bit stream. When all the numbers of the maximum merge candidates are satisfied, the subsequent candidate addition process (S940) cannot be performed.

[0135] If the result of step S930 confirms that the number of current merge candidates is less than the number of maximum merge candidates, the image encoding / decoding device may induce additional merge candidates based on a predetermined scheme and then insert them into the merge candidate list (S940). The additional merge candidates may include, for example, at least one of history-based merge candidate(s), pair-wise average merge candidate(s), ATMVP, combined bi-predictive merge candidates (if the slice / tile group type of the current slice / tile group is type B), and / or zero vector merge candidates.

[0136] If the result of step S930 confirms that the number of current merge candidates is not less than the number of maximum merge candidates, the image encoding / decoding device can terminate the configuration of the merge candidate list. In this case, the image encoding device can select the optimal merge candidate from among the merge candidates constituting the merge candidate list based on the RD cost, and can signal candidate selection information (e.g., merge candidate index) pointing to the selected merge candidate to the image decoding device. The image decoding device can select the optimal merge candidate based on the merge candidate list and the candidate selection information.

[0137] As described above, the motion information of the selected merge candidate can be used as the motion information of the current block, and predicted samples of the current block can be derived based on the motion information of the current block. The image encoding device can derive the residual samples of the current block based on the predicted samples and signal the image decoding device to the residual information of the residual samples. As described above, the image decoding device can generate restored samples based on the residual samples derived based on the residual information and the predicted samples, and generate a restored picture based on these.

[0138] When skip mode is applied to a block, the motion information of the current block can be derived in the same way as when merge mode is applied previously. However, when skip mode is applied, the residual signal for that block is omitted. Therefore, the predicted sample can be immediately used as the restored sample. The skip mode can be applied, for example, when the value of cu_skip_flag is 1.

[0139] The following describes how to guide spatial candidates in merge mode and / or skip mode. Spatial candidates can be the spatial merge candidates mentioned above.

[0140] Spatial candidate deriving can be based on spatially adjacent blocks. For example, up to four spatial candidates can be derived from candidate blocks located at the positions shown in Figure 8. The order in which spatial candidates are derived may be A1→B1→B0→A0→B2. However, the order in which spatial candidates are derived is not limited to the above order; for example, it may be B1→A1→B0→A0→B2. The last position in terms of order (position B2 in the above example) can be considered if at least one of the preceding four positions (A1, B1, B0, and A0 in the above example) is unavailable. In this case, the block at a given position being unavailable may include cases where the block belongs to a different slice or tile than the current block, or where the block is an intra-predicted block. If a spatial candidate is derive from the first position in terms of order (A1 or B1 in the above example), redundancy checks can be performed on the spatial candidates at subsequent positions. For example, if the motion information of a subsequent spatial candidate is identical to the motion information of a spatial candidate already included in the merge candidate list, the encoding efficiency can be improved by not including the subsequent spatial candidate in the merge candidate list. The redundancy check performed on subsequent spatial candidates can be reduced by performing it only on some candidate pairs, rather than on all candidate pairs as possible, thereby reducing computational complexity.

[0141] Figure 10 illustrates candidate pairs for redundancy checks performed on spatial candidates.

[0142] In the example shown in Figure 10, a redundancy check for the space candidate at position B0 can be performed only for the space candidate at position A0. Similarly, a redundancy check for the space candidate at position B1 can be performed only for the space candidate at position B0. Furthermore, a redundancy check for the space candidate at position A1 can be performed only for the space candidate at position A0. Finally, a redundancy check for the space candidate at position B2 can be performed only for the space candidates at positions A0 and B0.

[0143] The example shown in Figure 10 is an example where the order in which spatial candidates are derived is A0 → B0 → B1 → A1 → B2. However, it is not limited to this, and even if the order in which spatial candidates are derived is changed, redundancy checks can be performed on only some candidate pairs, as in the example shown in Figure 10.

[0144] The following describes how to derive time candidates in merge mode and / or skip mode. Time candidates can represent the time merge candidates described above. The motion vectors of time candidates can also correspond to time candidates in MVP mode.

[0145] Only one time candidate can be included in the merge candidate list. During the process of guiding time candidates, the motion vector of the time candidate can be scaled. For example, this scaling can be performed based on a co-located CU (hereinafter referred to as "col block") belonging to a colocated reference picture (colPic) (hereinafter referred to as "col picture"). The list of reference pictures used to guide the col block can be explicitly signaled in the slice header.

[0146] Figure 11 is a diagram illustrating how to scale the motion vector of a time candidate.

[0147] In Figure 11, curr_CU and curr_pic represent the current block and current picture, while col_CU and col_pic represent the col block and col picture. col_ref represents the reference picture of the col block. Additionally, tb represents the distance between the reference picture of the current block and the current picture, and td represents the distance between the reference picture of the col block and the col picture. tb and td can be represented by values ​​corresponding to the difference in POC (Picture Order Count) between pictures. Scaling of the motion vector of the time candidate can be performed based on tb and td. The reference picture index of the time candidate can also be set to 0.

[0148] Figure 12 is a diagram illustrating the position that guides the time candidates.

[0149] In Figure 12, the thick solid line blocks represent the current block. Time candidates can be derived from the block in the col picture corresponding to position C0 (lower right) or C1 (center) in Figure 12. First, it is determined that position C0 is available. If position C0 is available, time candidates can be derived based on position C0. If position C0 is not available, time candidates can be derived based on position C1. For example, if the block in the col picture at position C0 is an intra-predicted block or is currently outside the CTU row, it can be determined that position C0 is not available.

[0150] As described above, when motion data compression is applied, the motion vectors of the col blocks can be saved for each predetermined unit block. In this case, the C0 or C1 position can be modified in order to induce the motion vector of the block covering the C0 or C1 position. For example, if the predetermined unit block is an 8x8 block and the C0 or C1 position is (xColCi, yColCi), the position for inducing the time candidate can be modified to ((xColCi>>3)<<3, (yColCi>>3)<<3).

[0151] The following describes how to derive history-based candidates in merge mode and / or skip mode. History-based candidates can be represented by history-based merge candidates.

[0152] History-based candidates can be added to the merge candidate list after spatial and temporal candidates have been added to the merge candidate list. For example, motion information of a previously encoded / decoded block can be stored in a table and used as a History-based candidate for the current block. The table can store multiple History-based candidates during the encoding / decoding process. The table can be initialized when a new CTU row begins. Initialization of the table can mean that all History-based candidates stored in the table are deleted and the table becomes empty. Whenever an inter-predicted block is found, the associated motion information can be added to the table as the last entry. In this case, the inter-predicted block does not have to be a block predicted based on a subblock. The motion information added to the table can be used as a new History-based candidate.

[0153] The History-based candidate table can have a predetermined size. For example, this size could be 5. In this case, the table can store up to 5 History-based candidates. When a new candidate is added to the table, a limited first-in-first-out (FIFO) rule can be applied, which first performs a redundancy check to see if the same candidate already exists in the table. If the same candidate already exists in the table, that candidate is removed from the table, and the position of all subsequent History-based candidates can be shifted forward.

[0154] History-based candidates can be used in the process of constructing the merge candidate list. At this time, the most recently included History-based candidates in the table are sequentially checked and can be included in the merge candidate list at positions after the time candidate. When a History-based candidate is included in the merge candidate list, a redundancy check can be performed with spatial or time candidates already included in the merge candidate list. If a History-based candidate and a spatial or time candidate already included in the merge candidate list overlap simultaneously, the History-based candidate may not be included in the merge candidate list. The computational complexity of the redundancy check can be reduced by simplifying it as follows.

[0155] The number of history-based candidates used to generate the merge candidate list can be set to (N<=4)?M:(8-N). Here, N represents the number of candidates already included in the merge candidate list, and M represents the number of available history-based candidates stored in the table. In other words, if the merge candidate list contains 4 or fewer candidates, the number of history-based candidates used to generate the merge candidate list is M, and if the merge candidate list contains N candidates (more than 4), the number of history-based candidates used to generate the merge candidate list can be set to (8-N).

[0156] The construction of the merge candidate list using history-based candidates can be terminated when the total number of available merge candidates reaches (maximum allowed number of merge candidates - 1).

[0157] The following describes how to derive pair-wise average candidates in merge mode and / or skip mode. Pair-wise average candidates can be represented by pair-wise average merge candidates or pair-wise candidates.

[0158] Pair-wise average candidates can be generated by taking already defined candidate pairs from the candidates included in the merge candidate list and averaging them. The already defined candidate pairs are {(0,1),(0,2),(1,2),(0,3),(1,3),(2,3)}, and the numbers that make up each candidate pair can be indices in the merge candidate list. In other words, the already defined candidate pair (0,1) represents the pair of candidate with index 0 and candidate with index 1 in the merge candidate list, and the pair-wise average candidate can be generated by averaging candidate with candidate with index 0 and candidate with index 1. The induction of pair-wise average candidates can be performed in the order of the aforementioned already defined candidate pairs. That is, after inducing a pair-wise average candidate for candidate pair (0,1), the pair-wise average candidate induction process can be performed for candidate pair (0,2) and candidate pair (1,2) in that order. The pair-wise average candidate induction process can be performed until the construction of the merge candidate list is complete. For example, the pair-wise average candidate induction process can be performed until the number of merge candidates included in the merge candidate list reaches the maximum number of merge candidates.

[0159] Pair-wise average candidates can be calculated individually for each reference picture list. If two motion vectors are available for a reference picture list (L0 list or L1 list), the average of these two motion vectors can be calculated. In this case, the average can be performed even if the two motion vectors point to different reference pictures. If only one motion vector is available for a reference picture list, the available motion vector can be used as the motion vector for the pair-wise average candidate. If neither of the two motion vectors is available for a reference picture list, that reference picture list can be determined to be invalid.

[0160] If the merge candidate list has not been completed after the pair-wise average candidate has been added, zero vectors can be added to the merge candidate list until the maximum number of merge candidates is reached.

[0161] When MVP mode is applied to the current block, a motion vector predictor (MVP) candidate list can be generated using the motion vectors of the restored spatial surrounding blocks (e.g., the surrounding blocks shown in Figure 8) and / or the motion vectors corresponding to the temporal surrounding blocks (or Col blocks). In other words, the motion vectors of the restored spatial surrounding blocks and / or the motion vectors corresponding to the temporal surrounding blocks can be used as motion vector predictor candidates for the current block. When dual prediction is applied, an MVP candidate list for L0 motion information derivation and an MVP candidate list for L1 motion information derivation are generated and available separately. The prediction information (or information related to prediction) for the current block may include candidate selection information (e.g., an MVP flag or MVP index) that indicates the optimal motion vector predictor candidate selected from among the motion vector predictor candidates included in the MVP candidate list. In this case, the prediction unit can use the candidate selection information to select the motion vector predictor for the current block from among the motion vector predictor candidates included in the MVP candidate list. The prediction unit of the image encoding device can calculate the motion vector difference (MVD) between the motion vector of the current block and the motion vector predictor, and can encode this and output it in bitstream format. In other words, the MVD can be calculated by subtracting the motion vector predictor from the motion vector of the current block. The prediction unit of the image decoding device can obtain the motion vector difference contained in the prediction information and derive the motion vector of the current block by adding the motion vector difference and the motion vector predictor. The prediction unit of the image decoding device can obtain or derive a reference picture index that indicates a reference picture from the prediction information.

[0162] Figure 13 is a schematic diagram illustrating a method for constructing a motion vector predictor candidate list according to an example of this disclosure.

[0163] First, the system searches for spatial candidate blocks in the current block and inserts any available candidate blocks into the MVP candidate list (S1010). Then, it is determined whether there are fewer than two MVP candidates in the MVP candidate list (S1020). If there are two, the construction of the MVP candidate list can be completed.

[0164] In step S1020, if there are fewer than two available spatial candidate blocks, the system can search for a time candidate block for the current block and insert any available candidate blocks into the MVP candidate list (S1030). If no time candidate blocks are available, the system can complete the construction of the MVP candidate list by inserting a zero motion vector into the MVP candidate list (S1040).

[0165] On the other hand, when MVP mode is applied, the reference picture index can be explicitly signaled. In this case, the picture index for L0 prediction (refidxL0) and the reference picture index for L1 prediction (refidxL1) can be separately signaled. For example, when MVP mode is applied and bi-indicator prediction (BI prediction) is applied, both the information regarding refidxL0 and the information regarding refidxL1 can be signaled.

[0166] As mentioned above, when MVP mode is applied, information about the MVD derived from the image encoding device can be signaled to the image decoding device. The information about the MVD may include, for example, information indicating the x and y components for the MVD absolute value and sign. In this case, information indicating whether the MVD absolute value is greater than 0, whether it is greater than 1, and the rest of the MVD can be signaled stepwise. For example, information indicating whether the MVD absolute value is greater than 1 can be signaled only if the value of the flag information indicating whether the MVD absolute value is greater than 0 is 1.

[0167] Overview of Affine Mode

[0168] The following describes in detail the affine mode, which is an example of an interpretation mode. Conventional video encoding / decoding systems currently use only one motion vector to represent the motion information of a block. However, this method has the problem that it can only represent the optimal motion information at the block level, and cannot represent the optimal motion information at the pixel level. To solve this problem, the affine mode was proposed, which defines the motion information of a block at the pixel level. According to the affine mode, the motion vectors for each pixel / or subblock of a block can be determined using two to four motion vectors currently associated with the block.

[0169] In contrast to conventional methods where motion information is represented using the translation (or displacement) of pixel values, affine modes can represent pixel-specific motion information using at least one of the following: translation, scaling, rotation, or shearing. Among these, affine modes in which pixel-specific motion information is represented using displacement, shearing, or rotation can be defined as similarity or simplified affine modes. In the following explanation, "affine mode" can refer to similarity or simplified affine modes.

[0170] Motion information in affine mode can be represented using two or more CPMVs (Control Point Motion Vectors). The motion vector of a specific pixel position in the current block can be induced using a CPMV. In this case, the set of pixel-specific and / or sub-block-specific motion vectors in the current block can be defined as an Affine Motion Vector Field (Affine MVF).

[0171] Figure 14 is a diagram illustrating the parameter model of affine modes.

[0172] Currently, when an affine mode is applied to a block, an affine MVF can be induced using either a 4-parameter model or a 6-parameter model. In this case, the 4-parameter model refers to a model type in which two CPMVs are used, and the 6-parameter model refers to a model type in which three CPMVs are used. Figures 14(a) and 14(b) illustrate the CPMVs used in the 4-parameter model and the 6-parameter model, respectively.

[0173] Currently, if we define the block position as (x,y), the motion vector due to the pixel position can be derived according to equation 1 or 2 below. For example, the motion vector for a 4-parameter model can be derived according to equation 1, and the motion vector for a 6-parameter model can be derived according to equation 2.

[0174]

number

[0175]

number

[0176] In equations 1 and 2, mv0 = {mv_0x, mv_0y} is the CPMV at the upper left corner of the current block, v1 = {mv_1x, mv_1y} is the CPMV at the upper right corner of the current block, and mv2 = {mv_2} is the CPMV at the lower left position of the current block. Here, W and H correspond to the width and height of the current block, respectively, and mv = {mv_x, mv_y} can represent the motion vector of the pixel position {x, y}.

[0177] During the encoding / decoding process, the affine MVF can be determined on a pixel-by-pixel and / or predefined subblock basis. When the affine MVF is determined on a pixel-by-pixel basis, a motion vector can be derived based on each pixel value. On the other hand, when the affine MVF is determined on a subblock basis, a motion vector for a subblock can be derived based on the central pixel value of that subblock. The central pixel value can represent a virtual pixel located at the center of the subblock, or the lower-right pixel of four pixels located in the center. Alternatively, the central pixel value can be a specific pixel within the subblock that represents that subblock. In this disclosure, the case where the affine MVF is determined on a 4x4 subblock basis is described. However, this is for the sake of explanation, and the size of the subblocks can be varied.

[0178] In other words, when affine prediction is available, the motion models currently applicable to a block can include three: the Translational motion model, the 4-parameter affine motion model, and the 6-parameter affine motion model. Here, the Translational motion model can represent a model in which conventional block-unit motion vectors are used, the 4-parameter affine motion model can represent a model in which two CPMVs are used, and the 6-parameter affine motion model can represent a model in which three CPMVs are used. Affine modes can be further subdivided into detailed modes depending on the method of encoding / decoding motion information. For example, affine modes can be subdivided into affine MVP modes and affine merge modes.

[0179] When affine merge mode is applied to the current block, the CPMV can be derived from the surrounding blocks of the current block that have been encoded / decoded in affine mode. If at least one of the surrounding blocks of the current block has been encoded / decoded in affine mode, then affine merge mode can be applied to the current block. That is, when affine merge mode is applied to the current block, the CPMV of the current block can be derived using the CPMV of the surrounding blocks. For example, the CPMV of the surrounding blocks can be determined as the CPMV of the current block, or the CPMV of the current block can be derived based on the CPMV of the surrounding blocks. When the CPMV of the current block is derived based on the CPMV of the surrounding blocks, at least one of the encoding parameters of the current block or the surrounding blocks can be used. For example, the CPMV of the surrounding blocks can be modified based on the size of the surrounding blocks and the size of the current block, etc., and used as the CPMV of the current block.

[0180] On the other hand, in the case of an affine merge where the MV is derived on a subblock basis, this can be called a subblock merge mode, which can be indicated by a merge_subblock_flag with a first value (e.g., "1"). In this case, the affine merging candidate list described later can also be called a subblock merging candidate list. In this case, the subblock merging candidate list may further include candidates derived by SbTMVP described later. In this case, the candidates derived by sbTMVP can be used as the candidate at index 0 of the subblock merging candidate list. In other words, the candidates derived by sbTMVP may be located earlier in the subblock merging candidate list than the inherited affine candidates and constructed affine candidates described later.

[0181] As an example, an affine mode flag can be defined to indicate whether or not affine mode can be applied to the current block. This can be signaled at at least one level above the current block, such as sequence, picture, slice, tile, tile group, or brick. For example, the affine mode flag could be named sps_affine_enabled_flag.

[0182] When affine merge mode is applied, an affine merge candidate list can be constructed for CPMV induction of the current block. This affine merge candidate list may include at least one of the following: inherited affine merge candidates, combined affine merge candidates, and zero merge candidates. Inherited affine merge candidates can represent candidates that would be induced using the CPMVs of the surrounding blocks if those blocks were encoded / decoded in affine mode. Combined affine merge candidates can represent candidates whose CPMVs are induced based on the motion vectors of the surrounding blocks of each Control Point (CP). Zero merge candidates, on the other hand, can represent candidates consisting of CPMVs of size 0. In the following description, CP can represent a specific location on the block used to induce the CPMV. For example, CPs can be the vertex positions of the block.

[0183] Figure 15 illustrates how to generate a list of affine merge candidates.

[0184] Referring to the flowchart in Figure 15, affine merge candidates can be added to the affine merge candidate list in the following order: inheritance affine merge candidates (S1210), combination affine merge candidates (S1220), and zero merge candidates (S1230). Zero merge candidates can be added if, despite all inheritance affine merge candidates and combination affine merge candidates being added to the affine merge candidate list, the number of candidates in the candidate list does not reach the maximum number of candidates. In this case, zero merge candidates can be added until the number of candidates in the affine merge candidate list reaches the maximum number of candidates.

[0185] Figure 16 is a diagram illustrating CPMV induced from surrounding blocks.

[0186] As an example, up to two inheritance affine merge candidates can be derived, and each candidate can be derived based on at least one of the left peripheral block and the upper peripheral block. The peripheral blocks for deriving inheritance affine merge candidates are explained with reference to Figure 8. An inheritance affine merge candidate derived based on the left peripheral block can be derived based on at least one of A0 and A1, and an inheritance affine merge candidate derived based on the upper peripheral block can be derived based on at least one of B0, B1 and B2. In this case, the scan order of each peripheral block may be, but is not limited to, A0 to A1 and B0 to B1, B2. For each of the left and upper sides, an inheritance affine merge candidate can be derived based on the first peripheral block available in the scan order. In this case, redundancy checks may not be performed between the candidates derived from the left peripheral block and the upper peripheral block.

[0187] As an example, as shown in Figure 16, when the left peripheral block A is encoded / decoded in affine mode, at least one of the motion vectors v2, v3, and v4 corresponding to the CP of peripheral block A can be induced. When peripheral block A is encoded / decoded via a 4-parameter affine model, the inherited affine merge candidate can be induced using v2 and v3. On the other hand, when peripheral block A is encoded / decoded via a 6-parameter affine model, the inherited affine merge candidate can be induced using v2, v3, and v4.

[0188] Figure 17 is a diagram illustrating the surrounding blocks used to induce candidate combination affine merges.

[0189] A combinational affine candidate can mean a candidate from which the CPMV is derived using a combination of general motion information of surrounding blocks. Motion information for each CP can be derived using the spatial or temporal surrounding blocks of the current block. In the following description, CPMVk can mean the motion vector representing the kth CP. As an example, referring to Figure 17, CPMV1 can be determined as the first available motion vector among the motion vectors of B2, B3, and A2, in which case the scan order may be B2, B3, A2. CPMV2 can be determined as the first available motion vector among the motion vectors of B1 and B0, in which case the scan order may be B1, B0. CPMV3 can be determined as the first available motion vector among the motion vectors of A1 and A0, in which case the scan order may be A1, A0. If TMVP can be applied to the current block, CPMV4 can be determined as the motion vector of the temporal surrounding block T.

[0190] After four motion vectors are derived for each CP, a combinational affine merge candidate can be derived based on these. The combinational affine merge candidate can consist of at least two motion vectors selected from the four motion vectors derived for each CP. For example, a combinational affine merge candidate can consist of at least one set of motion vectors in the following order: {CPMV1,CPMV2,CPMV3}, {CPMV1,CPMV2,CPMV4}, {CPMV1,CPMV3,CPMV4}, {CPMV2,CPMV3,CPMV4}, {CPMV1,CPMV2}, and {CPMV1,CPMV3}. A combinational affine candidate consisting of three motion vectors may be a candidate for a 6-parameter affine model. Conversely, a combinational affine candidate consisting of two motion vectors may be a candidate for a 4-parameter affine model. To avoid the scaling process of motion vectors, if the reference picture indices of the CPs are different, the associated CPMV combinations can be ignored and not used in the derivation of the combinational affine candidate.

[0191] When affine MVP mode is applied to the current block, the image encoder can induce two or more CPMV predictors and CPMVs for the current block and, based on these, derive CPMV differences. At this time, the CPMV differences can be signaled from the encoder to the decoder. The image decoder can induce CPMV predictors for the current block, recover the signaled CPMV differences, and then derive the CPMV of the current block based on the CPMV predictors and CPMV differences.

[0192] On the other hand, affine MVP mode can be applied to the current block only if affine merge mode or subblock-based TMVP is not currently applied to the block. Affine MVP mode can also be referred to as affine CP MVP mode.

[0193] If an affine MVP is currently applied to a block, an affine MVP candidate list can be constructed to induce a CPMV on the current block. Here, the affine MVP candidate list may include at least one of the following: inheritance affine MVP candidates, combination affine MVP candidates, translation affine MVP candidates, and zero MVP candidates.

[0194] In this context, inherited affine MVP candidates can refer to candidates that are induced based on the CPMV of the surrounding blocks when the surrounding blocks of the current block are encoded / decoded in affine mode. Combinatorial affine MVP candidates can refer to candidates that are induced by generating a CPMV combination based on the motion vector of the CP surrounding blocks. Zero MVP candidates can refer to candidates consisting of CPMVs with a value of 0. The induction methods and characteristics of inherited affine MVP candidates and combinational affine MVP candidates are the same as those described above for inherited affine candidates and combinational affine candidates, so their explanation is omitted.

[0195] If the maximum number of candidates in the affine MVP candidate list is 2, then combination affine MVP candidates, translation affine MVP candidates, and zero MVP candidates can be added if the current number of candidates is less than 2. In particular, translation affine MVP candidates can be derived in the following order:

[0196] For example, if the number of candidates in the affine MVP candidate list is less than two, and the combined affine MVP candidate CPMV0 is valid, then CPMV0 can be used as an affine MVP candidate. That is, an affine MVP candidate whose motion vectors CP0, CP1, and CP2 are all CPMV0 can be added to the affine MVP candidate list.

[0197] Next, if the number of candidates in the affine MVP candidate list is less than two, and the combined affine MVP candidate CPMV1 is valid, then CPMV1 can be used as an affine MVP candidate. That is, an affine MVP candidate whose motion vectors CP0, CP1, and CP2 are all CPMV1 can be added to the affine MVP candidate list.

[0198] Next, if the number of candidates in the affine MVP candidate list is less than two, and the combined affine MVP candidate CPMV2 is valid, then CPMV2 can be used as an affine MVP candidate. That is, an affine MVP candidate whose motion vectors CP0, CP1, and CP2 are all CPMV2 can be added to the affine MVP candidate list.

[0199] Despite the conditions mentioned above, if the number of candidates in the Affine MVP candidate list is less than two, the current block's TMVP (temporal motion vector predictor) can be added to the Affine MVP candidate list.

[0200] Despite the addition of a translated affine MVP candidate, if the number of candidates in the affine MVP candidate list is less than two, a zero MVP candidate can be added to the affine MVP candidate list.

[0201] Figure 18 illustrates how to generate an affine MVP candidate list.

[0202] Referring to the flowchart in Figure 18, candidates can be added to the affine MVP candidate list in the following order: inherited affine MVP candidates (S1610), combined affine MVP candidates (S1620), translated affine MVP candidates (S1630), and zero MVP candidates (S1640). As mentioned above, steps S1620 to S1640 can be performed depending on whether the number of candidates included in the affine MVP candidate list at each step is less than 2.

[0203] The scan order for inherited affine MVP candidates may be the same as that for inherited affine merge candidates. However, for inherited affine MVP candidates, only surrounding blocks that reference the same reference picture as the current block's reference picture can be considered. When adding inherited affine MVP candidates to the affine MVP candidate list, redundancy checks may be disabled.

[0204] To guide the combination affine MVP candidates, only the spatial surrounding blocks shown in Figure 17 can be considered. Furthermore, the scan order of the combination affine MVP candidates may be the same as the scan order of the combination affine merge candidates. Additionally, to guide the combination affine MVP candidates, the reference picture indices of the surrounding blocks are checked, and the first surrounding block that is intercoded and references the same reference picture as the current block is available in the scan order.

[0205] Overview of Subblock-based TMVP (SbTMVP) mode

[0206] The following describes in detail the subblock-based TMVP mode, which is an example of an interpretation mode. According to the subblock-based TMVP mode, a motion vector field (MVF) is induced for the current block, so motion vectors can be induced on a subblock basis.

[0207] Unlike the conventional TMVP mode, which is performed on a coding unit basis, the subblock-based TMVP mode allows encoding / decoding of motion vectors to be performed on a subcoding unit basis. Furthermore, while the conventional TMVP mode derives time-motion vectors from collocated blocks, the subblock-based TMVP mode allows motion vector fields to be derived from reference blocks indicated by motion vectors derive from surrounding blocks of the current block. Hereafter, the motion vectors derive from surrounding blocks can be referred to as the motion shift or representative motion vector of the current block.

[0208] Figure 19 is a diagram illustrating the peripheral blocks of the subblock-based TMVP mode.

[0209] When a subblock-based TMVP mode is applied to the current block, the surrounding block for determining the motion shift can be determined. For example, a scan of surrounding blocks for determining the motion shift can be performed in the order of blocks A1, B1, B0, and A0 in Figure 19. As another example, the surrounding block for determining the motion shift can be restricted to a specific surrounding block of the current block. For example, the surrounding block for determining the motion shift can always be determined to be block A1. If the surrounding block has a motion vector that references a col picture, that motion vector can be determined as the motion shift. The motion vector determined as the motion shift can also be called the time motion vector. On the other hand, if the above motion vector cannot be derived from the surrounding block, the motion shift can be set to (0,0).

[0210] Figure 20 is a diagram illustrating how to induce a motion vector field according to a subblock-based TMVP mode.

[0211] Next, the reference block on the collated picture indicated by the motion shift can be determined. For example, by adding the motion shift to the coordinates of the current block, subblock-based motion information (motion vector, reference picture index) can be obtained from the col picture. In the example shown in Figure 20, we assume that the motion shift is the motion vector of block A1. By applying the motion shift to the current block, the subblocks in the col picture corresponding to each subblock that makes up the current block (col subblocks) can be identified. Then, using the motion information of the corresponding subblocks (col subblocks) in the col picture, the motion information of each subblock in the current block can be derived. For example, the motion information of a corresponding subblock can be obtained from the central position of that corresponding subblock. In this case, the central position may be the position of the lower right sample among the four samples located in the center of the corresponding subblock. If the motion information of a specific subblock in the col block corresponding to the current block is not available, the motion information of the central subblock of the col block can be determined as the motion information of that subblock. Once the motion information of the corresponding subblock is derived, it can be switched to the motion vector and reference picture index of the current subblock, similar to the TMVP process described above. In other words, when subblock-based motion vectors are induced, the motion vectors can be scaled by taking into account the POC of the reference picture of the reference block.

[0212] As described above, subblock-based TMVP candidates for the current block can be derived using the motion vector field or motion information of the current block, which is derived based on the subblock.

[0213] Hereinafter, a list of merge candidates composed of subblocks will be defined as a subblock merge candidate list. The affine merge candidates and subblock-based TMVP candidates mentioned above can be merged to form a subblock merge candidate list.

[0214] On the other hand, a subblock-based TMVP mode flag can be defined to indicate whether or not subblock-based TMVP mode can be applied to the current block. This can be signaled at at least one level above the current block, such as sequence, picture, slice, tile, tile group, or brick. For example, the subblock-based TMVP mode flag can be named sps_sbtmvp_enabled_flag. If subblock-based TMVP mode is applicable to the current block, subblock-based TMVP candidates can be added first to the subblock-unit merge candidate list. Subsequently, affine merge candidates can be added to the subblock-unit merge candidate list. On the other hand, the maximum number of candidates that can be included in the subblock-unit merge candidate list can be signaled. For example, the maximum number of candidates that can be included in the subblock-unit merge candidate list may be 5.

[0215] The size of the subblock used to guide the subblock merge candidate list can be signaled or already set to M×N. For example, M×N could be 8×8. Therefore, affine mode or subblock-based TMVP mode can only be applied to the current block if the current block size is 8×8 or larger.

[0216] An embodiment of the prediction execution method of this disclosure will be described below. The following prediction execution method can be performed in step S410 of Figure 4 or step S630 of Figure 6.

[0217] Based on motion information derived according to the prediction mode, a predicted block can be generated for the current block. The predicted block (predicted block) may include predicted samples (predicted sample array) of the current block. If the motion vector of the current block points to fractional sample units, an interpolation procedure can be performed, thereby deriving predicted samples of the current block based on reference samples in fractional sample units within the reference picture. When affine interpretation is applied to the current block, predicted samples can be generated based on sample / subblock unit MV. When bi-prediction is applied, predicted samples derived by a (phase-weighted) sum or weighted average of predicted samples derived based on L0 prediction (i.e., prediction using reference pictures in reference picture list L0 and MVL0) and predicted samples derived based on L1 prediction (i.e., prediction using reference pictures in reference picture list L1 and MVL1) can be used as predicted samples of the current block. When biprediction is applied, if the reference picture used for L0 prediction and the reference picture used for L1 prediction are located in different time directions relative to the current picture (i.e., it is biprediction but corresponds to bidirectional prediction), this can be called true biprediction.

[0218] In the image decoding device, restored samples and restored pictures can be generated based on the derived predicted samples, and then procedures such as in-loop filtering can be performed. In the image encoding device, residual samples can be derived based on the derived predicted samples, and image information including the predicted information and residual information can be encoded.

[0219] Bi-prediction with CU-level weights (BCW)

[0220] As described above, when dual prediction is applied to a block, the predicted samples can be derived based on a weighted average. Traditionally, the dual prediction signal (i.e., the dual prediction sample) could be derived via a simple average of the L0 prediction signal (L0 prediction sample) and the L1 prediction signal (L1 prediction sample). That is, the dual prediction sample was derived by the average of the L0 prediction sample based on the L0 reference picture and MVL0 and the L1 prediction sample based on the L1 reference picture and MVL1. However, according to this disclosure, when dual prediction is applied, the dual prediction signal (dual prediction sample) can be derived via a weighted average of the L0 prediction signal and the L1 prediction signal as follows.

[0221]

number

[0222] In the above formula 3, P bi-pred The graph shows the dual prediction signal (dual prediction block) derived by weighted averaging, where P0 and P1 represent the L0 prediction sample (L0 prediction block) and L1 prediction sample (L1 prediction block), respectively. (8-w) and w represent the weights applied to P0 and P1, respectively.

[0223] In generating a biprediction signal using weighted averaging, five weights are acceptable. For example, the weight w can be selected from {-2, 3, 4, 5, 10}. For each bipredicted CU, the weight w can be determined in one of two ways. The first of these two methods is that, if the CU is not currently in merge mode (non-merge CU), the weight index can be signaled along with the motion vector difference. For example, the bitstream may contain information about the weight index after information about the motion vector difference. The second of these two methods is that, if the CU is currently in merge mode (merge CU), the weight index can be derived from surrounding blocks based on a merge candidate index (merge index).

[0224] The generation of a weighted average biprediction signal can be restricted to apply only to CUs with a size containing 256 or more samples (luma component samples). That is, weighted average biprediction can only be performed for CUs where the product of the width and height of the current block is 256 or greater. Furthermore, the weight w may be one of the five weights as described above, or one of a different number of weights may be used. For example, depending on the characteristics of the current image, five weights may be used for low-delay pictures and three weights for non-low-delay pictures. In this case, the three weights could be {3, 4, 5}.

[0225] The image encoding device can determine the weight index without significantly increasing complexity by applying a fast search algorithm. In this case, the fast search algorithm can be summarized as follows. Below, unequal weight can mean that the weights applied to P0 and P1 are not equal. Equal weight can mean that the weights applied to P0 and P1 are equal.

[0226] -If AMVR mode, which adaptively changes the resolution of the motion vectors, is also applied, then if the picture is currently a low-delay picture, only unequal weighting can be conditionally checked for the 1-pel motion vector resolution and the 4-pel motion vector resolution, respectively.

[0227] -If an affine mode is applied and selected as the optimal mode for the block, the image encoding device can perform affine ME (motion estimation) for each of the unequal weights.

[0228] - If the two reference pictures used for biprediction are identical, only unequal weights can be conditionally checked.

[0229] - Unequal weighting may be ignored if certain conditions are met. These conditions may be based on the point-of-concept (POC) distance between the current picture and the reference picture, the quantization parameter (QP), the temporal level, etc.

[0230] The BCW weight index can be encoded using one context-coded bin followed by one or more bypass-coded bins. The first context-coded bin indicates whether equal weights are used or not. If unequal weights are used, additional bins can be bypass-coded and signaled. These additional bins can signal which weights are used.

[0231] Weighted prediction (WP) is a tool for efficiently encoding images that include fading. According to weighted prediction, weighting parameters (weights and offsets) can be signaled to each reference picture contained in reference picture lists L0 and L1, respectively. Then, when motion compensation is performed, the weights and offsets can be applied to the corresponding reference images. Weighted prediction and BCW can be used for different types of images. To avoid interaction between weighted prediction and BCW, the BCW weight index may not be signaled for CUs using weighted prediction. In this case, the weight can be inferred as 4; that is, equal weighting can be applied.

[0232] For CUs with merge mode applied, the weight index can be inferred from surrounding blocks based on the merge candidate index. This applies to both normal merge mode and inherited affine merge mode.

[0233] In combined affine merge mode, affine motion information can be constructed based on motion information of up to three blocks. In this case, the following steps can be taken to derive the BCW weight index for CUs using combined affine merge mode.

[0234] (1) First, the range of BCW weight index {0,1,2,3,4} can be divided into three groups {0}, {1,2,3} and {4}. If the BCW weight index of all CPs is derived from the same group, the BCW weight index can be derived by the following step (2). Otherwise, the BCW weight index can be set to 2.

[0235] (2) If at least two CPs have the same BCW weight index, the same BCW weight index can be assigned as the weight index of the combination affine merge candidate. Otherwise, the weight index of the combination affine merge candidate can be set to 2.

[0236] Bi-directional optical flow (BDOF)

[0237] According to this disclosure, BDOF can be used to refine (improve) a bi-prediction signal. BDOF is used to calculate improved motion information and generate prediction samples when bi-prediction is applied to a current block (e.g., CU). Therefore, the process of applying BDOF to calculate improved motion information may be included in the motion information derivation step described above.

[0238] For example, BDOF can be applied at the 4x4 subblock level. That is, BDOF can be performed on a 4x4 subblock basis within the current block.

[0239] BODF can be applied to CUs that meet the following conditions, for example:

[0240] 1) When the height of the CU is not 4 and the size of the CU is not 4x8

[0241] 2) If the CU is not in affine mode or ATMVP merge mode

[0242] 3) When CU is encoded in true biprediction mode, that is, when one of the two reference pictures has a time order that precedes the current picture and the other has a time order that follows the current picture.

[0243] Furthermore, BDOF can be applied only to the luma component. However, it is not limited to this; BDOF may be applied only to the chromatic component, or to both the luma and chromatic components.

[0244] The BDOF mode is based on the concept of optical flow, that is, it assumes that the object's movement is smooth. When BDOF is applied, an improved motion vector (v) is applied to each 4x4 subblock. x ,v y The improved motion vector can be calculated by minimizing the difference between the L0 predicted sample and the L1 predicted sample. The improved motion vector can be used to adjust the bipredicted sample values ​​within a 4x4 subblock.

[0245] The following will explain the BDOF process in more detail.

[0246] First, the horizontal gradient of the two prediction signals. JPEG2026063366000005.jpg15140 and vertical gradient The image JPEG2026063366000006.jpg14149 can be calculated. In this case, k can be 0 or 1. The gradient can be calculated by directly calculating the difference between two adjacent samples, as shown in equation 4 below.

[0247] [Number]

[0248] In the above formula (4), I (k) (i, j) means the sample value of the coordinate (i, j) of the predicted signal in the list k (k = 0, 1). For example, I (0) (i, j) means the sample value at the position (i, j) in the L0 prediction block, and I (1) (i, j) can mean the sample value at the position (i, j) in the L1 prediction block.

[0249] In the above formula (4), the difference between the two samples is right-shifted by only 4. However, it is not limited to this, and the amount of right shift (shift1) can be determined based on the bit depth of the luma component. For example, when the bit depth of the luma component is bitDepth, shift1 can be determined as max(6, bitDepth - 6). Or, it can also be determined as a simply fixed value of 6. In the above formula (4), after calculating the difference between the two samples first for gradient calculation, a right shift operation is applied to the difference value. However, it is not limited to this, and the gradient can also be calculated by applying a right shift operation to the values of the two samples and then calculating the difference between the right-shifted values.

[0250] After the gradient is calculated as described above, the auto-correlation and cross-correlation S1, S2, S3, S5, and S6 between the gradients can be calculated as follows.

[0251] [Number]

[0252] Using the auto-correlation and cross-correlation between the gradients described above, an improved motion vector (v x , v y ) can be derived as follows.

[0253] [Number]

[0254] Based on the induced and improved motion vectors and gradients, the following adjustments can be made for each sample within the 4×4 sub-block.

[0255] [Number]

[0256] Finally, by adjusting the dual-prediction samples of the CU as follows, the prediction samples (pred BDOF ) of the CU to which BDOF is applied can be calculated.

[0257] [Number]

[0258] In the above equation, n a , n b and n S2 can be 3, 6, and 12 respectively. These values can be selected so that the multiplier in the BDOF process does not exceed 15 bits and the bit-width of the intermediate parameters can be maintained within 32 bits.

[0259] To derive the gradient values, the prediction samples I (k) (i,j) within the list k (k = 0, 1) existing outside the current CU can be generated. FIG. 21 is a diagram showing the CU expanded for performing BDOF.

[0260] As shown in Figure 21, extended rows / columns around the CU boundary can be used to perform BDOF. To control the computational complexity for generating prediction samples outside the boundary, prediction samples within the extended region (white area in Figure 21) can be generated using a bilinear filter, while prediction samples within the CU (gray area in Figure 21) can be generated using a normal 8-tap motion compensation interpolation filter. The sample values ​​at the extended location can only be used for gradient calculations. If sample values ​​and / or gradient values ​​located outside the CU boundary are needed to perform the remaining steps of the BDOF process, the nearest adjacent sample values ​​and / or gradient values ​​can be used for padding (iteration).

[0261] If the width and / or height of a CU is greater than 16 luma samples, the CU can be divided into subblocks, each with a width and / or height of 16 luma samples. The boundaries of each subblock can be treated identically to the CU boundaries described above during the BDOF process. The maximum unit size on which the BDOF process is performed can be limited to 16 × 16.

[0262] If BCW is currently available for a block, for example, if the BCW weight index indicates unequal weighting, BDOF may not be applied. Similarly, if WP is currently available for a block, for example, if luma_weight_lx_flag for at least one of the two reference pictures is 1, BDOF may not be applied. In this case, luma_weight_lx_flag may indicate whether the weighting factors of WP for the luma component of the lx prediction (x is 0 or 1) exist in the bitstream. Alternatively, it may indicate whether WP is applied to the luma component of the lx prediction. If CU is encoded in SMVD mode, BDOF may not be applied.

[0263] Prediction refinement with optical flow(PROF)

[0264] The following describes a method for improving sub-block-based affine motion compensation predicted blocks by applying optical flow. Predicted samples generated by sub-block-based affine motion compensation can be improved based on differences induced by the optical flow equation. Such improvement of predicted samples can be referred to in this disclosure as prediction refinement with optical flow (PROF). PROF can achieve pixel-level granularity interpretation without increasing memory access bandwidth.

[0265] The parameters of the affine motion model can be used to derive the motion vector for each pixel in the CU. However, since pixel-based affine motion compensation prediction results in high complexity and increased memory access bandwidth, subblock-based affine motion compensation prediction can be performed. When subblock-based affine motion compensation prediction is performed, the CU is divided into 4x4 subblocks, and a motion vector can be determined for each subblock. In this case, the motion vector for each subblock can be derived from the CPMV of the CU. Subblock-based affine motion compensation has a trade-off relationship between coding efficiency and complexity and memory access bandwidth. Because the motion vector is derived on a subblock basis, complexity and memory access bandwidth are reduced, but the prediction accuracy is lower.

[0266] Therefore, by applying optical flow to subblock-based affine motion compensation prediction, improved granularity motion compensation can be achieved.

[0267] As described above, after subblock-based affine motion compensation is performed, the Luma predicted sample can be improved by adding the difference induced by the optical flow equation. More specifically, PROF can be performed in the following four steps.

[0268] Step 1) Subblock-based affine motion compensation is performed to generate the predicted subblock I(i,j).

[0269] Step 2) Predicted spatial gradients of the subblocks x (i,j) and g y (i,j) is calculated at each sample position. A 3-tap filter can be used, and the filter coefficients can be [-1,0,1]. For example, the spatial gradient can be calculated as follows:

[0270]

number

[0271] To calculate the gradient, the predicted subblock can be extended by one pixel on each side. In this case, to reduce memory bandwidth and complexity, the pixels of the extended boundary can be copied from the nearest integer pixel in the reference picture. Therefore, additional interpolation for the padding area can be omitted.

[0272] Step 3) The luma prediction refinement (ΔI(i,j)) can be calculated using the optical flow equation. For example, the following formula can be used:

[0273]

number

[0274] In the above formula, Δv(i,j) means the difference between the pixel motion vector (pixel MV, v(i,j)) calculated at the sample position (i,j) and the sub-block motion vector (sub-block MV) of the sub-block to which the sample (i,j) belongs.

[0275] FIG. 22 is a diagram showing the relationship between Δv(i,j), v(i,j) and the sub-block motion vector.

[0276] In the example shown in FIG. 22, for example, the difference between the motion vector v(i,j) of the upper left sample position of the current sub-block and the motion vector v SB of the current sub-block can be represented by a thick dashed arrow, and the vector indicated by the thick dashed arrow can correspond to Δv(i,j).

[0277] The affine model parameters and the pixel positions from the center of the sub-block are not changed. Therefore, Δv(i,j) is calculated only for the first sub-block and can be reused for different sub-blocks within the same CU. When the horizontal offset and the vertical offset from the pixel position to the center of the sub-block are x and y respectively, Δv(x,y) can be derived as follows.

[0278] [Number]

[0279] In the above, (v 0x , v 0y ), (v 1x , v 1y ) and (v 2x , v 2y ) correspond to the upper left CPMV, the upper right CPMV and the lower left CPMV, and w and h mean the width and height of the CU.

[0280] Step 4) Finally, the final predicted block I'(i,j) can be generated based on the calculated improvement in the Luma prediction ΔI(i,j) and the predicted subblock I(i,j). For example, the final predicted block I' can be generated as follows:

[0281]

number

[0282] As mentioned above, BDOF can be applied during the interpretation process to improve the reference sample during the motion compensation process, thereby enhancing image compression performance. BDOF can be performed in general modes; that is, it is not performed in affine mode, GPM mode, CIIP mode, etc.

[0283] For blocks encoded in affine mode, PROF can be performed in a manner similar to BDOF. As mentioned above, image compression performance can be improved by improving the reference samples within each 4x4 subblock via PROF.

[0284] Since both PROF and BDOF utilize the characteristics of optical flow, the application of PROF can be determined according to the same conditions as the application conditions for BDOF. Furthermore, this disclosure provides a variety of embodiments for WP and BCW.

[0285] In this disclosure, setting or inducing any information (e.g., a flag) to true may mean that the information is inducing a first value (e.g., "1"). Furthermore, setting any information to true may indicate that the process indicated by that information (e.g., BDOF, PROF, WP, etc.) will be performed. Conversely, setting or inducing any information (e.g., a flag) to false may mean that the information is inducing a second value (e.g., "0"). Furthermore, setting any information to false may indicate that the process indicated by that information will not be performed.

[0286] BDOF can be performed during the motion compensation process when bdofFlag is set to true and various conditions are met, as follows:

[0287] [Table 1]

[0288] The execution conditions for BDOF listed in Table 1 above can be described as shown in Table 2 below.

[0289] [Table 2]

[0290] However, the conditions for BDOF to be performed are not limited to the examples in Tables 1 and 2 above, and some of these conditions can be omitted. Furthermore, other conditions may also be considered.

[0291] If BDOF is not applied to the current block according to the above conditions, for example, if the prediction mode of the current block is affine mode, PROF can be applied in the same way as BDOF. For example, if the prediction mode of the current block is affine mode, it is determined whether or not to apply PROF (cbProfFlagLX), and if cbProfFlagLX is true, PROF can be performed.

[0292] BDOF determines the sample offset using optical flow characteristics. Therefore, BDOF is not performed when the brightness values ​​between reference pictures differ, i.e., when BCW or WP (weighted prediction) is applied. However, PROF can be performed without considering whether BCW or WP is applied, even though it also uses optical flow characteristics to induce the sample offset.

[0293] According to one embodiment of this disclosure, in order to harmonize BDOF and PROF from a design perspective, PROF may not be applied to blocks to which BCW or WP is applied. For example, when BcwIdx is not 0, or luma_weight_l0_flag[refIdxL0] is 1, or luma_weight_l1_flag[refIdxL1] is 1, the information cbProfFlagLX indicating whether or not to apply PROF may be set to false. A non-zero BcwIdx may mean that BCW is applied to the current block, and a luma_weight_lX_flag[refIdxLX](X=0 or 1) being 1 may mean that WP is applied to the current block. In this disclosure, a BcwIdx of 0 may mean that equal weights are applied, i.e., that bidirectional prediction blocks are generated with the average sum of the L0 prediction block and the L1 prediction block. Therefore, when setting cbProfFlagLX, by adding the above condition, it is possible to control whether PROF is applied if BCW or WP is currently applied to the block.

[0294] The table below shows an example of how to configure cbProfFlagLX according to this disclosure, with the underlined portion indicating the added conditions.

[0295] [Table 3]

[0296] The table below shows other examples of how to set cbProfFlagLX according to this disclosure, with the underlined portion indicating the additional conditions.

[0297] [Table 4]

[0298] As described above, it is possible to determine whether or not PROF is applied to the current block. For example, cbProfFlagLX(X=0 or 1) can indicate whether or not PROF is applied to the L0 prediction direction or the L1 prediction direction, and according to the method in Table 3, cbProfFlagLX can be determined based on at least one of bcwIdx, luma_wighted_l0_flag, and / or luma_wighted_l1_flag. As another example, according to the method in Table 4, cbProfFlagLX can be determined based on at least one of bcwIdx, slice_type, pps_weighted_pred_flag, and / or pps_weighted_bipred_flag. The slice_type indicates the slice type of the current slice to which the picture currently belongs, pps_weighted_pred_flag is a PPS (Picture Parameter Set) parameter indicating whether WP is applied to the P slice that references the PPS, and pps_weighted_bipred_flag is a PPS (Picture Parameter Set) parameter indicating whether WP is applied to the B slice that references the PPS.

[0299] According to other embodiments of this disclosure, PROF can be applied when BCW or WP (explicit weighted prediction) is performed.

[0300] Generally, the BCW weight index (bcw_idx) is signaled only when WP is unavailable. Therefore, the weight coefficients of bcw_idx and WP are not signaled simultaneously. The table below shows an example of the syntax structure for signaling bcw_idx.

[0301] [Table 5]

[0302] According to Table 5 above, the condition for signaling bcw_idx is to check whether all of the weight prediction flags (e.g., luma_weight_l0_flag, luma_weight_l1_flag, chroma_weight_l0_flag, chroma_weight_l1_flag) of the reference picture pointed to by the reference picture index (ref_idx_l0, ref_idx_l1) of the CU are 0. Therefore, according to Table 5, even if the WP application flag transmitted from the PPS (e.g., pps_weighted_pred_flag and / or pps_weighted_bipre_flag) is TRUE, bcw_idx can be signaled if the weight prediction flag of a particular reference picture index is 0. In other words, according to the example in Table 5, bcw_idx can be signaled even though WP is applied.

[0303] Figure 23 is a flowchart illustrating an example of how PROF, BCW, WP, and / or average sum are performed according to this disclosure.

[0304] Referring to Figure 23, first, a cbProfFlag (e.g., cbProfFlagLX) can be derived to indicate whether or not to apply PROF to the current block (S2310). The cbProfFlag can be derived based on the various methods described herein.

[0305] Subsequently, in step S2320, it is checked whether cbProfFlag is TRUE. If it is TRUE, a PROF can be performed on the current block in step S2330. The PROF can be performed in the manner described above. As a result of the PROF, improved prediction samples for the current block can be obtained. If cbProfFlag is False in step S2320, step S2330 can be skipped.

[0306] Next, in step S2340, a weightedPredFlag is induced to indicate whether or not to apply weighted prediction (WP) to the current block, and it can be checked whether its value is TRUE. The method for inducing weightedPredFlag will be described later. If weightedPredFlag is TRUE, it is decided to apply weighted prediction to the current block, and weighted prediction can be performed on the current block (S2350). Weighted prediction for the current block can be performed based on the weighting parameters (weight and offset) for the reference picture of the current block. As mentioned above, the weighting parameters for the reference picture can be explicitly signaled via a bitstream.

[0307] If weightedPredFlag is False, it is decided not to apply weighted prediction to the current block, and it can be checked whether bcwIdx is 0 or not (S2360). bcwIdx can be induced differently depending on the prediction mode of the current block. For example, if the prediction mode of the current block is skip mode or merge mode, bcwIdx for the current block can be induced to bcwIdx for the merge candidate indicated by the merge candidate index of the current block. If the prediction mode of the current block is not merge mode, for example, MVP mode, bcwIdx for the current block can be reconstructed by parsing the syntax element bcw_idx which is signaled via the bitstream. If bcw_idx is not signaled via the bitstream, the bcwIdx value can be estimated to be 0. If bcwIdx is 0, it can be indicated that BCW is not applied to the current block. As mentioned above, bcwIdx being 0 means that equal weights are applied, which can be interpreted as meaning that bidirectional prediction blocks are generated using the average sum of the L0 and L1 prediction blocks.

[0308] In step S2360, if bcwIdx is not 0, it is determined that BCW should be applied to the current block, and BCW can be performed on the current block based on the weight indicated by bcwIdx (S2370). In step S2360, if bcwIdx is 0, it is determined that BCW should not be applied to the current block, and an average sum can be performed on the current block (S2380).

[0309] Table 6 shows an example of how the present disclosure induces weightedPredFlag and thereby enables WP or BCW.

[0310] [Table 6]

[0311] According to the method in Table 6, weightedPredFlag can be derived based on the slice type of the current slice to which the current block belongs and the WP application flags signaled via PPS (e.g., pps_weighted_pred_flag, pps_weighted_bipred_flag). Specifically, if the slice type of the current block is a P slice, weightedPredflag can be determined to the value of pps_weighted_pred_flag. Also, if the slice type of the current block is a B slice, weightedPredflag can be determined to the value of pps_weighted_bipred_flag.

[0312] According to the method in Table 6, if the weightedPredFlag determined as described above is false, a default weighted prediction can be performed. If bcwIdx is not 0, or if BCW or bcwIdx is 0, an average sum can be performed. Furthermore, if weightedPredFlag is true, an explicit weighted prediction can be performed based on the signaled weighting parameters.

[0313] According to the method in Table 6, weightedPredFlag is determined solely by the slice type of the current slice and PPS information, not by the weighted prediction flag for each reference picture (e.g., luma_weight_l0_flag, luma_weight_l1_flag, chroma_weight_l0_flag, chroma_weight_l1_flag), whereas bcw_idx can be signaled based on the weighted prediction flag for each reference picture (e.g., luma_weight_l0_flag, luma_weight_l1_flag, chroma_weight_l0_flag, chroma_weight_l1_flag), as shown in Table 5.

[0314] Therefore, even if weightedPredFlag is determined to be TRUE according to the method in Table 6, bcw_idx can still be signaled. In this case, even if bcw_idx is not 0 (default), WP (explicit weighted prediction) will always be performed.

[0315] The embodiments described with reference to Figure 23 and Tables 5 and 6 include the problem that WP is performed even when bcw_idx is not the default, and other embodiments of the present disclosure that solve such problems are described below.

[0316] Figure 24 is a flowchart showing other examples of performing PROF, BCW, WP and / or average sum according to this disclosure.

[0317] Referring to Figure 24, first, a cbProfFlag can be derived indicating whether or not to apply PROF to the current block (S2410). The cbProfFlag can be derived based on the various methods described herein.

[0318] Subsequently, in step S2420, it is checked whether cbProfFlag is TRUE. If it is TRUE, a PROF can be performed on the current block in step S2430. The PROF can be performed in the manner described above. As a result of the PROF, improved prediction samples for the current block can be obtained. If cbProfFlag is False in step S2420, step S2430 can be skipped.

[0319] Subsequently, in step S2440, it can be checked whether bcwIdx is 0 or not. bcwIdx can be derived differently depending on the prediction mode of the current block. For example, if the prediction mode of the current block is skip mode or merge mode, bcwIdx for the current block can be derived to bcwIdx for the merge candidate indicated by the merge candidate index of the current block. If the prediction mode of the current block is not merge mode, for example, MVP mode, bcwIdx for the current block can be reconstructed by parsing the syntax element bcw_idx which is signaled via the bitstream. If bcw_idx is not signaled via the bitstream, the bcwIdx value can be inferred to be 0. In step S2440, if bcwIdx is not 0, it is determined that BCW should be applied to the current block, and BCW can be performed on the current block based on the weight indicated by bcwIdx (S2450).

[0320] In step S2440, if bcwIdx is 0, it is determined that BCW is not applied to the current block. Then, in step S2460, it can be checked whether weightedPredFlag, which indicates whether or not weighted prediction (WP) is applied to the current block, is TRUE.

[0321] In step S2460, if weightedPredFlag is TRUE, it is decided to apply weighted prediction to the current block, and weighted prediction can be performed on the current block (S2470). Weighted prediction for the current block can be performed based on weighting parameters (weights and offsets) for the reference picture of the current block. As mentioned above, the weighting parameters for the reference picture can be explicitly signaled via the bitstream.

[0322] In step S2460, if weightedPredFlag is False, it is decided not to apply weighted prediction to the current block and an average sum can be performed on the current block (S2480).

[0323] According to the embodiment described with reference to Figure 24, if both explicit weighted prediction (WP) and block weighting (BCW) are applicable to a given block, BCW can be applied preferentially.

[0324] Table 7 shows other examples of how the present disclosure induces weightedPredFlag and thereby performs WP or BCW.

[0325] [Table 7]

[0326] The methods for inducing weightedPredFlag are the same as those in Table 6 and Table 7, so a detailed explanation is omitted. According to the method in Table 7, if the weightedPredFlag determined as described above is false, or if bcwIdx is not 0, a default weighted prediction is performed, and BCW or average sum can be calculated according to the bcwIdx value. Also, if weightedPredFlag is true and bcwIdx is 0, an explicit weighted prediction can be performed based on the signaled weighting parameters.

[0327] According to the method in Table 7, if weightedPredFlag is true and bcwIdx is not 0, that is, if both explicit weighted prediction (WP) and BCW are applicable to the current block, BCW can be applied preferentially.

[0328] Figure 25 is a flowchart showing an example of performing BCW or WP using the method in Table 7.

[0329] First, the slice type of the current slice to which the current block belongs can be determined (S2510). If the slice type is a P slice, weightedPredFlag can be redirected to pps_weighted_pred_flag (S2520). If the slice type is a B slice, weightedPredFlag can be redirected to pps_weighted_bipred_flag (S2530).

[0330] Subsequently, in step S2540, it can be determined whether weightedPredFlag is 0 or whether BcwIdx is not 0. bcwIdx can be derived differently depending on the prediction mode of the current block. For example, if the prediction mode of the current block is skip mode or merge mode, bcwIdx for the current block can be derived to bcwIdx for the merge candidate indicated by the merge candidate index of the current block. If the prediction mode of the current block is not merge mode, for example MVP mode, bcwIdx for the current block can be reconstructed by parsing the syntax element bcw_idx which is signaled via the bitstream. If bcw_idx is not signaled via the bitstream, the bcwIdx value can be inferred to be 0.

[0331] If weightedPredFlag is 0 or BcwIdx is not 0, a default weighted prediction can be performed (S2550). In this case, if BcwIdx is 0, the average sum described in step S2380 or step S2480 can be performed. If BcwIdx is not 0, the BCW described in step S2370 or step S2450 can be performed.

[0332] If weightedPredFlag is not 0 and BcwIdx is 0, explicit weighted prediction can be performed (S2560). In this case, the WP described in step S2350 or step S2470 can be performed.

[0333] Table 8 shows another example of how the present disclosure induces weightedPredFlag and thereby performs WP or BCW.

[0334] [Table 8]

[0335] According to the method in Table 8, weightedPredFlag can be induced by further considering bcwIdx. Specifically, if bcwIdx for the current block is not 0, weightedPredFlag can be induced to false. If bcwIdx for the current block is 0, weightedPredFlag can be induced according to the method in Table 6 based on the slice type of the current slice to which the current block belongs and the WP application flags signaled via PPS (e.g., pps_weighted_pred_flag, pps_weighted_bipred_flag). According to the method in Table 8, if the weightedPredFlag determined as described above is false (secondary value, e.g., 0), a default weighted prediction is performed, and BCW or average sum can be calculated depending on the bcwIdx value. Also, if weightedPredFlag is true (first value, e.g., 1), an explicit weighted prediction can be performed based on the signaled weighting parameters.

[0336] According to the method in Table 8, if bcwIdx is not 0, weightedPredFlag can be set to false, allowing BCW to be applied preferentially when both explicit weighted prediction (WP) and BCW are applicable to the current block.

[0337] Figure 26 is a flowchart showing an example of performing BCW or WP using the method in Table 8.

[0338] First, we can determine whether BcwIdx is non-zero (S2610). bcwIdx can be induced differently depending on the prediction mode of the current block. For example, if the prediction mode of the current block is skip mode or merge mode, bcwIdx for the current block can be induced to the bcwIdx for the merge candidate indicated by the current block's merge candidate index. If the prediction mode of the current block is not merge mode, for example, MVP mode, bcwIdx for the current block can be reconstructed by parsing the syntax element bcw_idx, which is signaled via the bitstream. If bcw_idx is not signaled via the bitstream, the bcwIdx value can be inferred to be 0.

[0339] If BcwIdx is not 0, weightedPredFlag can be guided to 0 (S2620). Then, after the decision in step S2660, a default weighted prediction can be performed (S2670). Alternatively, steps S2620 and S2660 can be skipped, and step S2670 can be performed immediately. Step S2670 is performed in the same way as step S2550, so a detailed explanation is omitted.

[0340] If BcwIdx is 0, the slice type of the current slice to which the current block belongs can be determined (S2630). If the slice type is a P slice, weightedPredFlag can be redirected to pps_weighted_pred_flag (S2640). If the slice type is a B slice, weightedPredFlag can be redirected to pps_weighted_bipred_flag (S2650).

[0341] Subsequently, in step S2660, it can be determined whether weightedPredFlag is 0 or not.

[0342] If weightedPredFlag is 0, a default weighted prediction can be performed (S2670). If weightedPredFlag is not 0, an explicit weighted prediction can be performed (S2680). Steps S2670 and S2680 are performed in the same way as steps S2550 and S2560, respectively, so a detailed explanation is omitted.

[0343] Table 9 shows another example of how the present disclosure induces weightedPredFlag and thereby performs WP or BCW.

[0344] [Table 9]

[0345] According to the method in Table 9, weightedPredFlag can be induced by considering the slice type of the current slice to which the current block belongs and the weight prediction flags (e.g., luma_weight_l0_flag, luma_weight_l1_flag, chroma_weight_l0_flag, chroma_weight_l1_flag) of the reference picture pointed to by the current block's reference picture indices (ref_idx_l0, ref_idx_l1). Specifically, if the current slice to which the current block belongs is a P slice and all the weight prediction flags for the L0 direction (e.g., luma_weight_l0_flag, chroma_weight_l0_flag) are 0, weightedPredFlag can be induced to 0. Furthermore, if the current slice to which the current block belongs is a P slice, and at least one of the weighted prediction flags for the L0 direction (e.g., luma_weight_l0_flag, chroma_weight_l0_flag) is not 0, then weightedPredFlag can be induced to 1.

[0346] If the current slice to which the current block belongs is a B slice, and all weight prediction flags for the L0 and L1 directions (e.g., luma_weight_l0_flag, luma_weight_l1_flag, chroma_weight_l0_flag, chroma_weight_l1_flag) are 0, then weightedPredFlag can be induced to 0. Also, if the current slice to which the current block belongs is a B slice, and at least one of the weight prediction flags for the L0 and L1 directions (e.g., luma_weight_l0_flag, luma_weight_l1_flag, chroma_weight_l0_flag, chroma_weight_l1_flag) is not 0, then weightedPredFlag can be induced to 1.

[0347] According to the method in Table 9, if the weightedPredFlag determined as described above is false (second value, e.g., 0), a default weighted prediction is performed, and BCW or average sum can be calculated according to the bcwIdx value. If weightedPredFlag is true (first value, e.g., 1), an explicit weighted prediction can be performed based on the signaled weighting parameters.

[0348] According to the method in Table 9, weightedPredFlag is induced based on the conditions under which bcw_idx is signaled. That is, by inducing weightedPredFlag to false when bcw_idx is signaled, if explicit weighted prediction (WP) and BCW are all applicable to the current block, BCW can be applied preferentially.

[0349] Figure 27 is a flowchart showing an example of performing BCW or WP using the method in Table 9.

[0350] First, the slice type of the current slice to which the current block belongs can be determined (S2710). If the slice type is a P slice, weightedPredFlag can be derived based on luma_weight_l0_flag and / or chroma_weight_l0_flag, as explained with reference to Table 9 (S2720). If the slice type is a B slice, weightedPredFlag can be derived based on luma_weight_l0_flag, chroma_weight_l0_flag, luma_weight_l1_flag and / or chroma_weight_l1_flag (S2730).

[0351] Subsequently, in step S2740, it can be determined whether weightedPredFlag is 0 or not.

[0352] If weightedPredFlag is 0, a default weighted prediction can be performed (S2750). If weightedPredFlag is not 0, an explicit weighted prediction can be performed (S2760). Steps S2750 and S2760 are performed in the same way as steps S2550 and S2560, respectively, so a detailed explanation is omitted.

[0353] Table 10 shows another example of how the present disclosure induces weightedPredFlag and thereby performs WP or BCW.

[0354] [Table 10]

[0355] The methods for inducing weightedPredFlag are the same for the methods in Table 9 and Table 10, so a detailed explanation is omitted. According to the method in Table 10, if the weightedPredFlag determined as described above is false (secondary value, e.g., 0) or if bcwIdx is not 0, a default weighted prediction is performed, and BCW or average sum can be calculated according to the bcwIdx value. Also, if weightedPredFlag is true (first value, e.g., 1) and bcwIdx is 0, an explicit weighted prediction can be performed based on the signaled weighting parameters.

[0356] According to the method in Table 10, weightedPredFlag is induced based on the conditions under which bcw_idx is signaled. That is, by inducing weightedPredFlag to false when bcw_idx is signaled, if both explicit weighted prediction (WP) and BCW are applicable to the current block, BCW can be applied preferentially.

[0357] Furthermore, according to the method in Table 10, if weightedPredFlag is true and bcwIdx is not 0, that is, if both explicit weighted prediction (WP) and BCW are applicable to the current block, BCW can be applied preferentially.

[0358] Figure 28 is a flowchart showing an example of performing BCW or WP using the method in Table 10.

[0359] Steps S2810 to S2830 in Figure 28 are the same as steps S2710 to S2730 in Figure 27, so a detailed explanation will be omitted.

[0360] Referring to Figure 28, in step S2840, it is determined whether weightedPredFlag is 0 or BcwIdx is not 0, and based on the determination result, either the default weighted prediction in step S2850 or the explicit weighted prediction in step S2860 can be performed. Steps S2840 to S2860 in Figure 28 are the same as steps S2540 to S2560 in Figure 25, so a detailed explanation is omitted.

[0361] As described above, the embodiments described with reference to Figure 23 and Tables 5 and 6 include the problem that WP is performed even when bcw_idx is not the default, and other embodiments of the present disclosure that solve such problems are described below.

[0362] PROF also applies when BCW or WP (explicit weighted prediction) is performed, and generally, bcw_idx is parsed from the bitstream only if WP is not available, so BCW and WP should not exist simultaneously. However, as shown in the syntax structure in Table 5, in order to parse bcw_idx from the bitstream, only the weighted prediction flags (e.g., luma_weight_l0_flag, luma_weight_l1_flag, chroma_weight_l0_flag, chroma_weight_l1_flag) of the reference picture pointed to by the reference picture index of the current block are checked. Also, as shown in Table 6, weightedPredFlag is induced based on the slice type of the current slice to which the current block belongs and the WP application flags (e.g., pps_weighted_pred_flag, pps_weighted_bipred_flag) signaled via PPS. Therefore, even if the WP application flag is true, bcw_idx is parsed if the weighted prediction flag for a particular reference picture is 0. Ultimately, there may be cases where WP and BCW are applied simultaneously.

[0363] Table 11 shows a modified syntax structure for parsing bcw_idx by another example in this disclosure.

[0364] [Table 11]

[0365] The syntax structure in Table 11 modifies the parsing condition of bcw_idx by considering the induction conditions of weightedPredFlag in Table 6. Therefore, if bcw_idx is parsed, weightedPredFlag is induced to false, and only if bcw_idx is not parsed is weightedPredFlag induced to true, thereby eliminating the case where WP and BCW are applied simultaneously. For example, the embodiment described with reference to Figure 23 and Tables 5 and 6 can be solved by replacing Table 5 with Table 11.

[0366] The exemplary methods in this disclosure are presented as a series of actions for clarity of explanation, but this is not intended to restrict the order in which the steps are performed, and each step may be performed simultaneously or in a different order, if necessary. To implement the methods according to this disclosure, the exemplary steps may be further expanded to include other steps, or some steps may be expanded to include the remaining steps, or some steps may be expanded to include additional steps.

[0367] In this disclosure, an image encoding device or image decoding device that performs a predetermined operation (step) may perform an operation (step) to confirm the conditions or status of the execution of said operation (step). For example, if it is stated that a predetermined operation is performed when a predetermined condition is satisfied, the image encoding device or image decoding device may perform an operation to confirm whether or not the predetermined condition is satisfied, and then perform the predetermined operation.

[0368] The various embodiments of this disclosure are not intended to list all possible combinations, but rather to illustrate representative aspects of this disclosure. The matters described in the various embodiments may be applied independently or in combination of two or more.

[0369] Furthermore, various embodiments of this disclosure can be implemented by hardware, firmware, software, or a combination thereof. In the case of hardware implementation, it can be implemented by one or more ASICs (Application Specific Integrated Circuits), DSPs (Digital Signal Processors), DSPDs (Digital Signal Processing Devices), PLDs (Programmable Logic Devices), FPGAs (Field Programmable Gate Arrays), general processors, controllers, microcontrollers, microprocessors, etc.

[0370] Furthermore, the image decoding and image encoding devices to which the embodiments of this disclosure are applied can be included in multimedia broadcasting transceivers, mobile communication terminals, home cinema video equipment, digital cinema video equipment, surveillance cameras, video conferencing equipment, real-time communication equipment such as video communications, mobile streaming equipment, storage media, camcorders, video-on-demand (VoD) service providers, over-the-top (OTT) video equipment, internet streaming service providers, 3D video equipment, image-phone video equipment, and medical video equipment, and can be used to process video signals or data signals. For example, over-the-top (OTT) video equipment can include game consoles, Blu-ray players, internet-connected TVs, home theater systems, smartphones, tablet PCs, and digital video recorders (DVRs).

[0371] Figure 29 illustrates a content streaming system to which the embodiments of this disclosure can be applied.

[0372] As shown in Figure 29, a content streaming system to which an embodiment of the present disclosure is applied may broadly include an encoding server, a streaming server, a web server, media storage, user equipment, and multimedia input devices.

[0373] The encoding server is responsible for compressing content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream, and transmitting this bitstream to the streaming server. In other cases, if a multimedia input device such as a smartphone, camera, or video camera directly generates the bitstream, the encoding server can be omitted.

[0374] The bitstream can be generated by an image encoding method and / or image encoding apparatus to which an embodiment of the present disclosure is applied, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.

[0375] The streaming server transmits multimedia data to the user's device based on the user's request via a web server, and the web server can act as an intermediary to inform the user of available services. When a user requests a desired service from the web server, the web server transmits this to the streaming server, and the streaming server can transmit multimedia data to the user. In this case, the content streaming system may include a separate control server, in which case the control server can play a role in controlling the commands and responses between the devices within the content streaming system.

[0376] The streaming server can receive content from media storage and / or encoding servers. For example, when receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.

[0377] Examples of user devices include mobile phones, smartphones, laptop computers, digital broadcasting terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices such as smartwatches, smart glasses, HMDs (head-mounted displays), digital TVs, desktop computers, and digital signage.

[0378] Each server within the aforementioned content streaming system can be operated as a distributed server, in which case the data received from each server can be processed in a distributed manner.

[0379] The scope of this disclosure includes software or machine-executable commands (e.g., operating systems, applications, firmware, programs, etc.) that enable the operation of various embodiments to be performed on a device or computer, and non-transitory computer-readable medium on which such software or commands etc. are stored and can be executed on a device or computer. [Industrial applicability]

[0380] The embodiments described herein can be used for encoding / decoding images.

Claims

1. An image decoding method performed by an image decoding device, The steps include receiving a bitstream, The steps include: deriving a first flag that indicates whether to apply BDOF (Bi-directional optical flow) to the current block of the picture based on the received bitstream; A step of performing a BDOF process on the current block based on the first flag indicating that the BDOF is applied to the current block, Based on the first flag indicating that the BDOF is not applied to the current block, Based on the received bitstream, the steps include: (i) checking a second flag indicating whether to perform weighted prediction on the current block, and (ii) checking a weight index BcwIdx for performing BCW (Bi-prediction with CU-level Weight) on the current block; A step of determining whether to perform a default weighted prediction or an explicit weighted prediction on the current block based on the second flag and the weight index BcwIdx, Based on whether the second flag is equal to 0, or whether the weight index BcwIdx is not equal to 0, the default weighted prediction is performed on the current block. Based on the fact that the second flag is equal to 1 and the weight index BcwIdx is equal to 0, the explicit weighted prediction is made to the current block. Based on the fact that the weight index BcwIdx is equal to 0, the default weighted prediction performs an average sum on the current block. Based on the fact that the weight index BcwIdx is not equal to 0, the default weighted prediction performs the BCW on the current block. An image decoding method comprising the steps of: performing the explicit weighted prediction based on weighting parameters (weight and offset) for the reference picture of the current block.

2. The image decoding method according to claim 1, wherein the second flag is determined to be different based on the slice type of the current slice to which the current block belongs.

3. Based on the fact that the slice type of the current slice is a P slice, the second flag is derived as the value of pps_weighted_pred_flag signaled in the PPS (picture parameter set), The image decoding method according to claim 2, wherein the second flag is derived as the value of pps_weighted_bipred_flag signaled by the PPS, based on the slice type of the current slice being a B slice.

4. The weight index BcwIdx is derived based on the syntax element bcw_idx signaled via the bitstream, The image decoding method according to claim 1, wherein the weight index BcwIdx is derived as 0 based on the fact that the syntax element bcw_idx does not exist in the bitstream.

5. The image decoding method according to claim 4, wherein the syntax element bcw_idx is parsed from the bitstream based on the weighted prediction flag of the reference picture of the current block.

6. The image decoding method according to claim 4, wherein the syntax element bcw_idx is parsed from the bitstream based on the fact that all weighted prediction flags of the reference picture of the current block are equal to 0.

7. The image decoding method according to claim 1, wherein the weighting parameters are explicitly signaled via the bitstream.

8. An image encoding method performed by an image encoding device, A step of deriving a first flag that indicates whether to apply BDOF (Bi-directional optical flow) to the current block of the picture, A step of performing a BDOF process on the current block based on the first flag indicating that the BDOF is applied to the current block, Based on the first flag indicating that the BDOF is not applied to the current block, (i) a second flag indicating whether to perform weighted prediction on the current block, and (ii) a step of checking the weight index BcwIdx for performing BCW (Bi-prediction with CU-level Weight) on the current block, A step of determining whether to perform a default weighted prediction or an explicit weighted prediction on the current block based on the second flag and the weight index BcwIdx, Based on whether the second flag is equal to 0, or whether the weight index BcwIdx is not equal to 0, the default weighted prediction is performed on the current block. The step is to perform the explicit weighted prediction on the current block based on the second flag being equal to 1 and the weight index BcwIdx being equal to 0, A step of encoding information related to the first flag, the second flag, and the weight index BcwIdx, Based on the fact that the weight index BcwIdx is equal to 0, the default weighted prediction performs an average sum on the current block. Based on the fact that the weight index BcwIdx is not equal to 0, the default weighted prediction performs the BCW on the current block. An image encoding method comprising the steps of: 1) performing the explicit weighted prediction based on weighting parameters (weights and offsets) for the reference picture of the current block.

9. A method for transmitting a bitstream, A step of generating a bitstream, The aforementioned bitstream is A step of deriving a first flag that indicates whether to apply BDOF (Bi-directional optical flow) to the current block of the picture, A step of performing a BDOF process on the current block based on the first flag indicating that the BDOF is applied to the current block, Based on the first flag indicating that the BDOF is not applied to the current block, (i) a second flag indicating whether to perform weighted prediction on the current block, and (ii) a step of checking the weight index BcwIdx for performing BCW (Bi-prediction with CU-level Weight) on the current block, A step of determining whether to perform a default weighted prediction or an explicit weighted prediction on the current block based on the second flag and the weight index BcwIdx, Based on whether the second flag is equal to 0, or whether the weight index BcwIdx is not equal to 0, the default weighted prediction is performed on the current block. The step is to perform the explicit weighted prediction on the current block based on the second flag being equal to 1 and the weight index BcwIdx being equal to 0, A step of encoding information related to the first flag, the second flag, and the weight index BcwIdx, Based on the fact that the weight index BcwIdx is equal to 0, the default weighted prediction performs an average sum on the current block. Based on the fact that the weight index BcwIdx is not equal to 0, the default weighted prediction performs the BCW on the current block. Based on the weighting parameters (weight and offset) for the reference picture of the current block, the explicit weighted prediction is performed in the following steps: A method comprising the step of transmitting the bitstream.