Image Encoding / Decoding Method and Apparatus for Executing PROF and Method for Transmitting Bitstream

By deducing the predicted samples and reference picture resampling conditions in image encoding/decoding, and applying optical flow prediction refinement, the problem of high-resolution image transmission and storage costs is solved, and a more efficient encoding/decoding method and equipment is realized.

CN114731428BActive Publication Date: 2025-07-04BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080078522.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-10-01
Filing Date
2020-09-10
Publication Date
2025-07-04
Estimated Expiration
2040-09-10

AI Technical Summary

Technical Problem

The prior art When transmitting and storing high-resolution and high-quality images, the increase in the amount of information leads to high transmission and storage costs, and it is necessary to improve image encoding/decoding efficiency.

Method used

By deriving the predicted samples of the current block, the reference picture resampling conditions are determined, and based on these conditions, whether to apply optical flow prediction refinement (PROF) is determined to improve encoding/decoding efficiency.

Benefits of technology

Improves image encoding/decoding efficiency, reduces transmission and storage costs, and supports flexible processing of current and reference picture sizes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114731428B_ABST
    Figure CN114731428B_ABST
Patent Text Reader

Abstract

An image encoding / decoding method and apparatus are provided. A method for an image decoding apparatus according to the present disclosure to decode an image may include the following steps: deriving a prediction sample of a current block based on motion information about the current block; deriving a reference picture resampling (RPR) condition for the current block; determining whether to apply a prediction refinement by optical flow (PROF) to the current block based on the RPR condition; and applying PROF to the current block to derive an improved prediction sample of the current block.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an image encoding / decoding method and apparatus, and a method of transmitting a bitstream. More specifically, the present disclosure relates to an image encoding / decoding method and apparatus for performing optical flow prediction refinement (PROF), and a method of transmitting a bitstream generated by the image encoding method / apparatus of the present disclosure. Background Art

[0002] Recently, the demand for high-resolution and high-quality images, such as high-definition (HD) images and ultra-high-definition (UHD) images, is increasing in various fields. As the resolution and quality of image data are improved, the amount of information or bits to be transmitted relatively increases compared to existing image data. The increase in the amount of information or bits to be transmitted results in an increase in transmission costs and storage costs.

[0003] Therefore, there is a need for efficient image compression techniques to effectively transmit, store, and reproduce information on high-resolution and high-quality images. Summary of the Invention

[0004] Technical Problem

[0005] An object of the present disclosure is to provide an image encoding / decoding method and apparatus having improved encoding / decoding efficiency.

[0006] Another object of the present disclosure is to provide an image encoding / decoding method and apparatus for performing PROF.

[0007] Another object of the present disclosure is to provide an image encoding / decoding method and apparatus for performing PROF in consideration of the size of a current picture and the size of a reference picture.

[0008] Another object of the present disclosure is to provide a method of transmitting a bitstream generated by an image encoding method or apparatus according to the present disclosure.

[0009] Another object of the present disclosure is to provide a recording medium storing a bitstream generated by an image encoding method or apparatus according to the present disclosure.

[0010] Another object of the present disclosure is to provide a recording medium storing a bitstream received, decoded, and used for reconstructing an image by an image decoding apparatus according to the present disclosure.

[0011] The technical problems solved by the present disclosure are not limited to the above technical problems, and other technical problems not described herein will be apparent to those skilled in the art from the following description.

[0012] Technical Solution

[0013] An image decoding method according to an aspect of the present disclosure may include the following steps: deriving a prediction sample of a current block based on motion information of the current block; deriving reference picture resampling (RPR) conditions of the current block; determining whether optical flow prediction refinement (PROF) is applied to the current block based on the RPR conditions; and deriving a refined prediction sample of the current block by applying PROF to the current block.

[0014] In the image decoding method of the present disclosure, the RPR conditions may be determined based on the size of a reference picture of the current block and the size of the current picture.

[0015] In the image decoding method of the present disclosure, based on the size of the reference picture of the current block being different from the size of the current picture, the RPR conditions may be derived as a first value, and based on the size of the reference picture of the current block being equal to the size of the current picture, the RPR conditions may be derived as a second value.

[0016] In the image decoding method of the present disclosure, based on the RPR conditions being the first value, it may be determined that PROF is not applied to the current block.

[0017] In the image decoding method of the present disclosure, it may be determined whether PROF is applied to the current block based on the size of the current block.

[0018] In the image decoding method of the present disclosure, based on the product of the width w and the height h of the current block being less than 128, it may be determined that PROF is not applied to the current block.

[0019] In the image decoding method of the present disclosure, information specifying whether the current block is in an affine merge mode is parsed from the bitstream based on the size of the current block.

[0020] In the image decoding method of the present disclosure, based on each of the width w and the height h of the current block being equal to or greater than 8 and w*h being equal to or greater than 128, information specifying whether the current block is in an affine merge mode may be parsed from the bitstream.

[0021] In the image decoding method of the present disclosure, information specifying whether the current block is in an affine MVP mode may be parsed from the bitstream based on the size of the current block.

[0022] In the image decoding method of the present disclosure, based on each of the width w and the height h of the current block being equal to or greater than 8 and w*h being equal to or greater than 128, information specifying whether the current block is in an affine MVP mode may be parsed from the bitstream.

[0023] In the image decoding method of the present disclosure, it may be determined whether PROF is applied to the current block based on whether BCW or WP is applied to the current block.

[0024] In the image decoding method of the present disclosure, based on BCW or WP being applied to the current block, it can be determined that PROF is not applied to the current block.

[0025] An image decoding device according to another aspect of the present disclosure may include a memory and at least one processor. The at least one processor may derive a prediction sample of the current block based on motion information of the current block; derive reference picture resampling (RPR) conditions of the current block; determine whether optical flow prediction refinement (PROF) is applied to the current block based on the RPR conditions; and derive a refined prediction sample of the current block by applying PROF to the current block.

[0026] An image encoding method according to another aspect of the present disclosure may include the following steps: deriving a prediction sample of the current block based on motion information of the current block; deriving reference picture resampling (RPR) conditions of the current block; determining whether optical flow prediction refinement (PROF) is applied to the current block based on the RPR conditions; and deriving a refined prediction sample of the current block by applying PROF to the current block.

[0027] A transmission method according to another aspect of the present disclosure may send a bitstream generated by the image encoding method and / or image encoding device of the present disclosure to an image decoding device.

[0028] Furthermore, a computer-readable recording medium according to another aspect of the present disclosure may store a bitstream generated by the image encoding device or image encoding method of the present disclosure.

[0029] The features of the above brief overview of the present disclosure are merely exemplary aspects of the following detailed description of the present disclosure and do not limit the scope of the present disclosure.

[0030] Advantageous Effects

[0031] According to the present disclosure, it is possible to provide an image encoding / decoding method and device having improved encoding / decoding efficiency.

[0032] Furthermore, according to the present disclosure, it is possible to provide an image encoding / decoding method and device for performing PROF.

[0033] Furthermore, according to the present disclosure, it is possible to provide an image encoding / decoding method and device for considering the size of the current picture and the size of the reference picture to perform PROF.

[0034] Furthermore, according to the present disclosure, it is possible to provide a method for sending a bitstream generated by the image encoding method or device according to the present disclosure.

[0035] Furthermore, according to the present disclosure, it is possible to provide a recording medium for storing a bitstream generated by the image encoding method or device according to the present disclosure.

[0036] In addition, according to the present disclosure, it is possible to provide a recording medium that stores a bitstream received, decoded, and used for reconstructing an image by an image decoding device according to the present disclosure.

[0037] Those skilled in the art will understand that the effects achievable by the present disclosure are not limited to what has been specifically described above, and other advantages of the present disclosure will be more clearly understood from the detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 is a view schematically illustrating a video coding system to which an embodiment of the present disclosure is applicable.

[0039] Figure 2 is a view schematically illustrating an image coding device to which an embodiment of the present disclosure is applicable.

[0040] Figure 3 is a view schematically illustrating an image decoding device to which an embodiment of the present disclosure is applicable.

[0041] Figure 4 is a flowchart illustrating a video / image coding method based on inter prediction.

[0042] Figure 5 is a view illustrating the configuration of an inter prediction unit 180 according to the present disclosure.

[0043] Figure 6 is a flowchart illustrating a video / image decoding method based on inter prediction.

[0044] Figure 7 is a view illustrating the configuration of an inter prediction unit 260 according to the present disclosure.

[0045] Figure 8 is a view illustrating a motion that can be expressed in an affine mode.

[0046] Figure 9 is a view illustrating a parameter model of an affine mode.

[0047] Figure 10 is a view illustrating a method of generating an affine merge candidate list.

[0048] Figure 11 is a view illustrating a CPMV derived from neighboring blocks.

[0049] Figure 12 is a view illustrating neighboring blocks for deriving inherited affine merge candidates.

[0050] Figure 13 is a view illustrating neighboring blocks for deriving constructed affine merge candidates.

[0051] Figure 14 A view exemplifying a method of generating an affine MVP candidate list.

[0052] Figure 15 A view exemplifying neighboring blocks of a sub-block based TMVP mode.

[0053] Figure 16 A view exemplifying a method of deriving a motion vector field according to a sub-block based TMVP mode.

[0054] Figure 17 A view exemplifying a CU extended to perform BDOF.

[0055] Figure 18 A view exemplifying the relationship between Δv(i,j), v(i,j), and the sub-block motion vector.

[0056] Figure 19 A view exemplifying an example of a process for determining whether to apply BDOF according to the present disclosure.

[0057] Figure 20 A view exemplifying an example of a process for determining whether to apply PROF according to the present disclosure.

[0058] Figure 21 A view exemplifying a signaling for specifying information on whether to apply a sub-block merge mode according to an example of the present disclosure.

[0059] Figure 22 A view exemplifying a signaling for specifying information on whether to apply an affine MVP mode according to an embodiment of the present disclosure.

[0060] Figure 23 A view exemplifying a process for determining whether to apply PROF according to another embodiment of the present disclosure.

[0061] Figure 24 A view exemplifying a signaling for specifying information on whether to apply a sub-block merge mode according to another embodiment of the present disclosure.

[0062] Figure 25 A view exemplifying a signaling for specifying information on whether to apply an affine MVP mode according to another embodiment of the present disclosure.

[0063] Figure 26 A view exemplifying a process for determining whether to apply PROF according to another embodiment of the present disclosure.

[0064] Figure 27 A view exemplifying a process for determining whether to apply PROF according to another embodiment of the present disclosure.

[0065] Figure 28It is a view illustrating a method of executing PROF according to the present disclosure.

[0066] Figure 29 It is a view illustrating a process of determining whether to apply PROF according to another embodiment of the present disclosure.

[0067] Figure 30 It is a view illustrating a process of determining whether to apply PROF according to another embodiment of the present disclosure.

[0068] Figure 31 It is a view showing a content stream system to which an embodiment of the present disclosure is applicable. Detailed Embodiment

[0069] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings to facilitate implementation by those skilled in the art. However, the present disclosure can be implemented in various different forms and is not limited to the embodiments described herein.

[0070] When describing the present disclosure, if a detailed description of a related known function or configuration makes the scope of the present disclosure unnecessarily ambiguous, its detailed description will be omitted. In the drawings, parts irrelevant to the description of the present disclosure are omitted, and similar reference numerals are assigned to similar parts.

[0071] In the present disclosure, when a component is "connected", "coupled" or "linked" to another component, it may include not only a direct connection relationship but also an indirect connection relationship with an intermediate component present. Additionally, when a component "includes" or "has" other components, unless otherwise stated, it means that other components may also be included, rather than excluding other components.

[0072] In the present disclosure, terms such as first, second, etc. are only used for the purpose of distinguishing one component from other components and do not limit the order or importance of the components, unless otherwise stated. Accordingly, within the scope of the present disclosure, the first component in one embodiment may be referred to as the second component in another embodiment, and similarly, the second component in one embodiment may be referred to as the first component in another embodiment.

[0073] In the present disclosure, the components distinguished from each other are intended to clearly describe each feature and do not mean that the components must be separated. That is, multiple components can be integrated and implemented in one hardware or software unit, or one component can be distributed and implemented in multiple hardware or software units. Therefore, even without specific description, these embodiments of component integration or distribution are included within the scope of the present disclosure.

[0074] In the present disclosure, the components described in each embodiment are not necessarily essential components, and some components may be optional components. Therefore, an embodiment composed of a subset of the components described in the embodiment is also included within the scope of the present disclosure. In addition, an embodiment that includes other components in addition to the components described in the various embodiments is included within the scope of the present disclosure.

[0075] The present disclosure relates to the encoding and decoding of images. Unless redefined in the present disclosure, the terms used in the present disclosure may have the general meanings commonly used in the technical field to which the present disclosure pertains.

[0076] In the present disclosure, a "picture" generally refers to a unit representing an image within a specific time period, while a slice / tile is an encoding unit that forms part of a picture, and a picture can be composed of one or more slices / tiles. In addition, a slice / tile can include one or more coding tree units (CTUs).

[0077] In the present disclosure, a "pixel" or "pel" can mean the smallest unit that constitutes a picture (or image). In addition, "sample" can be used as a term corresponding to a pixel. A sample generally can represent a pixel or the value of a pixel, or can represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component.

[0078] In the present disclosure, a "unit" can represent a basic unit of image processing. The unit can include at least one of a specific area of a picture and information related to the area. In some cases, the unit can be used interchangeably with terms such as "sample array", "block", or "region". Generally, an M×N block can include a set (or array) of samples (or sample arrays) or transform coefficients in M columns and N rows.

[0079] In the present disclosure, a "current block" can mean one of a "current encoding block", "current encoding unit", "encoding target block", "decoding target block", or "processing target block". When performing prediction, a "current block" can mean a "current prediction block" or a "prediction target block". When performing transform (inverse transform) / quantization (dequantization), a "current block" can mean a "current transform block" or a "transform target block". When performing filtering, a "current block" can mean a "filtering target block".

[0080] In the present disclosure, the term " / " or "," can be interpreted as indicating "and / or". For example, "A / B" and "A,B" can mean "A and / or B". In addition, "A / B / C" and "A,B,C" can mean at least one of A, B, and / or C.

[0081] In the present disclosure, the term "or" shall be construed to indicate "and / or". For example, the expression "A or B" may include 1) only "A", 2) only "B", or 3) both "A and B". In other words, in the present disclosure, "or" shall be construed to indicate "additionally or alternatively".

[0082] Overview of the video coding system

[0083] Figure 1 is a view schematically showing a video coding system according to the present disclosure.

[0084] A video coding system according to an embodiment may include an encoding device 10 and a decoding device 20. The encoding device 10 may deliver encoded video and / or image information or data in the form of a file or a stream to the decoding device 20 via a digital storage medium or a network.

[0085] The encoding device 10 according to an embodiment may include a video source generator 11, an encoding unit 12, and a transmitter 13. The decoding device 20 according to an embodiment may include a receiver 21, a decoding unit 22, and a renderer 23. The encoding unit 12 may be referred to as a video / image encoding unit, and the decoding unit 22 may be referred to as a video / image decoding unit. The transmitter 13 may be included in the encoding unit 12. The receiver 21 may be included in the decoding unit 22. The renderer 23 may include a display, and the display may be configured as a separate device or an external component.

[0086] The video source generator 11 may acquire video / images through a process of capturing, synthesizing, or generating video / images. The video source generator 11 may include a video / image capturing device and / or a video / image generating device. The video / image capturing device may include, for example, one or more cameras, a video / image archive including previously captured video / images, etc. The video / image generating device may include, for example, a computer, a tablet computer, and a smart phone, and may generate (electronically) video / images. For example, virtual video / images may be generated by a computer or the like. In this case, the video / image capturing process may be replaced by a process of generating relevant data.

[0087] The encoding unit 12 may encode the input video / images. For compression and encoding efficiency, the encoding unit 12 may perform a series of processes such as prediction, transformation, and quantization. The encoding unit 12 may output encoded data (encoded video / image information) in the form of a bitstream.

[0088] The transmitter 13 can transmit the encoded video / image information or data output in the form of a bitstream to the receiver 21 of the decoding device 20 in the form of a file or a stream via a digital storage medium or a network. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter 13 can include components for generating a media file in a predetermined file format and can include components for transmitting via a broadcast / communication network. The receiver 21 can extract / receive the bitstream from the storage medium or the network and transmit the bitstream to the decoding unit 22.

[0089] The decoding unit 22 can decode the video / image by performing a series of processes corresponding to the operations of the encoding unit 12, such as dequantization, inverse transformation, and prediction.

[0090] The renderer 23 can render the decoded video / image. The rendered video / image can be displayed via a display.

[0091] Overview of the image coding device

[0092] Figure 2 is a view schematically showing an image encoding device to which an embodiment of the present disclosure can be applied.

[0093] As Figure 2 shown, the image encoding device 100 can include an image splitter 110, a subtractor 115, a transformer 120, a quantizer 130, a dequantizer 140, an inverse transformer 150, an adder 155, a filter 160, a memory 170, an inter-frame prediction unit 180, an intra-frame prediction unit 185, and an entropy encoder 190. The inter-frame prediction unit 180 and the intra-frame prediction unit 185 can be collectively referred to as a "prediction unit". The transformer 120, the quantizer 130, the dequantizer 140, and the inverse transformer 150 can be included in a residual processor. The residual processor can also include the subtractor 115.

[0094] In some embodiments, all or at least some of the multiple components configuring the image encoding device 100 can be configured by one hardware component (e.g., an encoder or a processor). In addition, the memory 170 can include a decoded picture buffer (DPB) and can be configured by a digital storage medium.

[0095] The image splitter 110 may split an input image (or picture or frame) input to the image encoding device 100 into one or more processing units. For example, the processing unit may be referred to as a coding unit (CU). The coding unit may be obtained by recursively splitting a coding tree unit (CTU) or a largest coding unit (LCU) according to a quadtree binary tree ternary tree (QT / BT / TT) structure. For example, a coding unit may be split into multiple coding units of a deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. For the splitting of the coding unit, the quadtree structure may be applied first, and then the binary tree structure and / or the ternary tree structure may be applied. The encoding process according to the present disclosure may be performed based on the final coding unit that is no longer split. The largest coding unit may be used as the final coding unit, or the coding unit of a deeper depth obtained by splitting the largest coding unit may be used as the final coding unit. Here, the encoding process may include processes of prediction, transformation, and reconstruction to be described later. As another example, the processing unit of the encoding process may be a prediction unit (PU) or a transformation unit (TU). The prediction unit and the transformation unit may be divided or split from the final coding unit. The prediction unit may be a sample prediction unit, and the transformation unit may be a unit for deriving transformation coefficients and / or a unit for deriving a residual signal from the transformation coefficients.

[0096] The prediction unit (inter-frame prediction unit 180 or intra-frame prediction unit 185) may perform prediction on a block to be processed (current block) and generate a prediction block including prediction samples of the current block. The prediction unit may determine whether to apply intra-frame prediction or inter-frame prediction based on the current block or CU. The prediction unit may generate various information related to the prediction of the current block and transmit the generated information to the entropy encoder 190. The information about the prediction may be encoded in the entropy encoder 190 and output in the form of a bitstream.

[0097] The intra-frame prediction unit 185 may predict the current block by referring to samples in the current picture. According to the intra-frame prediction mode and / or intra-frame prediction technique, the reference samples may be located in the neighbors of the current block or may be placed separately. The intra-frame prediction mode may include multiple non-directional modes and multiple directional modes. The non-directional modes may include, for example, the DC mode and the planar mode. According to the detail level of the prediction direction, the directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes. However, this is only an example, and more or fewer directional prediction modes may be used according to the setting. The intra-frame prediction unit 185 may determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.

[0098] The inter - frame prediction unit 180 may derive a prediction block of a current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, in order to reduce the amount of motion information transmitted in the inter - frame prediction mode, the motion information may be predicted in units of blocks, sub - blocks, or samples based on the correlation of the motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may also include inter - frame prediction direction (L0 prediction, L1 prediction, bi - prediction, etc.) information. In the case of inter - frame prediction, the neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block may be referred to as a collocated reference block, a collocated CU (colCU), etc. The reference picture including the temporal neighboring block may be referred to as a collocated picture (colPic). For example, the inter - frame prediction unit 180 may configure a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or the reference picture index of the current block. The inter - frame prediction may be performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the inter - frame prediction unit 180 may use the motion information of neighboring blocks as the motion information of the current block. In the case of the skip mode, different from the merge mode, the residual signal may not be transmitted. In the case of the motion vector prediction (MVP) mode, the motion vector of a neighboring block may be used as a motion vector predictor, and the motion vector of the current block may be signaled by encoding the motion vector difference and an indicator of the motion vector predictor. The motion vector difference may mean the difference between the motion vector of the current block and the motion vector predictor.

[0099] The prediction unit may generate a prediction signal based on various prediction methods and prediction techniques described below. For example, the prediction unit may not only apply intra - frame prediction or inter - frame prediction, but also apply both intra - frame prediction and inter - frame prediction simultaneously to predict the current block. The prediction method of applying both intra - frame prediction and inter - frame prediction simultaneously to predict the current block may be referred to as combined intra - inter prediction (CIIP). In addition, the prediction unit may perform intra - block copy (IBC) to predict the current block. Intra - block copy may be used for content image / video coding of games, etc., for example, screen content coding (SCC). IBC is a method of predicting the current picture using a previously reconstructed reference block in the current picture at a position separated from the current block by a predetermined distance. When IBC is applied, the position of the reference block in the current picture may be encoded as a vector (block vector) corresponding to the predetermined distance.

[0100] The prediction signal generated by the prediction unit can be used to generate a reconstructed signal or a residual signal. The subtractor 115 can generate a residual signal (residual block or residual sample array) by subtracting the prediction signal (prediction block or prediction sample array) output from the prediction unit from the input image signal (original block or original sample array). The generated residual signal can be transmitted to the transformer 120.

[0101] The transformer 120 can generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique can include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen - Loève transform (KLT), a graph - based transform (GBT), or a conditional non - linear transform (CNT). Here, GBT refers to a transform obtained from a graph when the relationship information between pixels is represented by a graph. CNT refers to a transform obtained based on a prediction signal generated using all previously reconstructed pixels. In addition, the transform processing can be applied to square pixel blocks of the same size or can be applied to blocks of variable size rather than square.

[0102] The quantizer 130 can quantize the transform coefficients and transmit them to the entropy encoder 190. The entropy encoder 190 can encode the quantized signal (information about the quantized transform coefficients) and output a bitstream. The information about the quantized transform coefficients can be referred to as residual information. The quantizer 130 can rearrange the quantized transform coefficients in block form into a one - dimensional vector form based on the coefficient scan order and generate information about the quantized transform coefficients based on the quantized transform coefficients in one - dimensional vector form.

[0103] The entropy encoder 190 can perform various coding methods, such as exponential Golomb, context - adaptive variable - length coding (CAVLC), context - adaptive binary arithmetic coding (CABAC), etc. The entropy encoder 190 can encode information required for video / image reconstruction other than the quantized transform coefficients (e.g., values of syntax elements, etc.) together or separately. The encoded information (e.g., encoded video / image information) can be transmitted or stored in the form of a bitstream in units of network abstraction layer (NAL). The video / image information can also include information about various parameter sets, such as an adaptive parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information can also include general constraint information. The information signaled, transmitted, and / or syntax elements described in this disclosure can be encoded through the above - mentioned encoding process and included in the bitstream.

[0104] The bitstream can be transmitted over a network or stored in a digital storage medium. The network can include a broadcast network and / or a communication network, and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) for transmitting the signal output from the entropy encoder 190 and / or a storage unit (not shown) for storing the signal can be included as internal / external elements of the image encoding device 100. Alternatively, a transmitter can be provided as a component of the entropy encoder 190.

[0105] The quantized transform coefficients output from the quantizer 130 can be used to generate a residual signal. For example, the residual signal (residual block or residual samples) can be reconstructed by applying dequantization and inverse transformation to the quantized transform coefficients by the dequantizer 140 and the inverse transformer 150.

[0106] The adder 155 adds the reconstructed residual signal to the prediction signal output from the inter-frame prediction unit 180 or the intra-frame prediction unit 185 to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). If the block to be processed has no residual, for example, in the case of applying the skip mode, the predicted block can be used as the reconstructed block. The adder 155 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current picture and can be used for inter-frame prediction of the next picture by filtering as described below.

[0107] In addition, as described below, luminance mapping and chrominance scaling (LMCS) are applicable to picture encoding processing.

[0108] The filter 160 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 160 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture and store the modified reconstructed picture in the memory 170, specifically, in the DPB of the memory 170. Various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filtering, bilateral filtering, etc. The filter 160 can generate various information related to filtering and transmit the generated information to the entropy encoder 190, as described later in the description of each filtering method. The information related to filtering can be encoded by the entropy encoder 190 and output in the form of a bitstream.

[0109] The modified reconstructed picture transmitted to the memory 170 can be used as a reference picture in the inter-frame prediction unit 180. When inter-frame prediction is applied by the image encoding device 100, prediction mismatch between the image encoding device 100 and the image decoding device can be avoided and the encoding efficiency can be improved.

[0110] The DPB of the memory 170 may store the modified reconstructed picture to be used as a reference picture in the inter prediction unit 180. The memory 170 may store the motion information of the blocks from which the motion information in the current picture is derived (or encoded) and / or the motion information of the blocks that have been reconstructed in the picture. The stored motion information may be transmitted to the inter prediction unit 180 and used as the motion information of spatially neighboring blocks or temporally neighboring blocks. The memory 170 may store the reconstructed samples of the reconstructed blocks in the current picture and may transmit the reconstructed samples to the intra prediction unit 185.

[0111] Overview of the image decoding device

[0112] Figure 3 is a view schematically showing an image decoding device to which an embodiment of the present disclosure is applicable.

[0113] As Figure 3 shown, the image decoding device 200 may include an entropy decoder 210, a dequantizer 220, an inverse transformer 230, an adder 235, a filter 240, a memory 250, an inter prediction unit 260, and an intra prediction unit 265. The inter prediction unit 260 and the intra prediction unit 265 may be collectively referred to as a "prediction unit". The dequantizer 220 and the inverse transformer 230 may be included in a residual processor.

[0114] According to an embodiment, all or at least some of the multiple components configuring the image decoding device 200 may be configured by hardware components (e.g., a decoder or a processor). In addition, the memory 250 may include a decoded picture buffer (DPB) or may be configured by a digital storage medium.

[0115] The image decoding device 200 that has received a bitstream including video / image information may reconstruct an image by performing processing corresponding to the processing performed by Figure 2 the image encoding device 100. For example, the image decoding device 200 may perform decoding using the processing units applied in the image encoding device. Therefore, the decoding processing units may be, for example, encoding units. The encoding units may be obtained by splitting a coding tree unit or a maximum coding unit. The reconstructed image signal decoded and output by the image decoding device 200 may be reproduced by a reproduction device (not shown).

[0116] The image decoding device 200 may receive, in the form of a bitstream, from Figure 2The signal output by the image encoding device. The received signal can be decoded by the entropy decoder 210. For example, the entropy decoder 210 can parse the bitstream to derive the information (e.g., video / image information) required for image reconstruction (or picture reconstruction). The video / image information can also include information about various parameter sets, such as adaptive parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS). In addition, the video / image information can also include general constraint information. The image decoding device can also decode the picture based on the information about the parameter set and / or the general constraint information. The information and / or syntax elements signaled / received described in this disclosure can be decoded through the decoding process and obtained from the bitstream. For example, the entropy decoder 210 decodes the information in the bitstream based on an encoding method such as exponential Golomb coding, CAVLC, or CABAC, and outputs the values of the syntax elements required for image reconstruction and the quantization values of the transformed coefficients of the residuals. More specifically, the CABAC entropy decoding method can receive the bins corresponding to each syntax element in the bitstream, use the decoding target syntax element information, neighboring blocks, and the decoding information of the decoding target block or the information of the symbols / bins decoded in the previous stage to determine the context model, perform arithmetic decoding on the bins by predicting the occurrence probability of the bins according to the determined context model, and generate symbols corresponding to the values of each syntax element. In this case, after determining the context model, the CABAC entropy decoding method can update the context model by using the information of the decoded symbols / bins for the context model of the next symbol / bin. The information related to prediction in the information decoded by the entropy decoder 210 can be provided to the prediction units (inter-frame prediction unit 260 and intra-frame prediction unit 265), and the residual values, i.e., the quantized transform coefficients and related parameter information, for which entropy decoding is performed in the entropy decoder 210 can be input to the dequantizer 220. Additionally, the information about filtering among the information decoded by the entropy decoder 210 can be provided to the filter 240. Furthermore, the receiver (not shown) for receiving the signal output by the image encoding device can be further configured as an internal / external component of the image decoding device 200, or the receiver can be a component of the entropy decoder 210.

[0117] In addition, the image decoding device according to the present disclosure can be referred to as a video / image / picture decoding device. The image decoding device can be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder can include the entropy decoder 210. The sample decoder can include at least one of the dequantizer 220, inverse transformer 230, adder 235, filter 240, memory 250, inter-frame prediction unit 260, or intra-frame prediction unit 265.

[0118] The dequantizer 220 may dequantize the quantized transform coefficients and output the transform coefficients. The dequantizer 220 may rearrange the quantized transform coefficients in the form of two-dimensional blocks. In this case, the rearrangement may be performed based on the coefficient scan order executed in the image coding device. The dequantizer 220 may dequantize the quantized transform coefficients by using a quantization parameter (e.g., quantization step information) and obtain the transform coefficients.

[0119] The inverse transformer 230 may perform an inverse transform on the transform coefficients to obtain a residual signal (residual block, residual sample array).

[0120] The prediction unit may perform prediction on the current block and generate a prediction block including prediction samples of the current block. The prediction unit may determine whether to apply intra prediction or inter prediction to the current block based on the information about prediction output from the entropy decoder 210, and may determine a specific intra / inter prediction mode (prediction technique).

[0121] Similar to that described in the prediction unit of the image coding device 100, the prediction unit may generate a prediction signal based on various prediction methods (techniques) described later.

[0122] The intra prediction unit 265 may predict the current block by referring to samples in the current picture. The description of the intra prediction unit 185 is equally applicable to the intra prediction unit 265.

[0123] The inter prediction unit 260 may derive a prediction block of the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may also include information about an inter prediction direction (L0 prediction, L1 prediction, bi-prediction, etc.). In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. For example, the inter prediction unit 260 may configure a motion information candidate list based on neighboring blocks, and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. The inter prediction may be performed based on various prediction modes, and the information about prediction may include information indicating the inter prediction mode of the current block.

[0124] The adder 235 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the prediction unit (including the inter prediction unit 260 and / or the intra prediction unit 265). The description of the adder 155 is equally applicable to the adder 235.

[0125] In addition, as described below, Luminance Mapping and Chrominance Scaling (LMCS) is applicable to picture decoding processing.

[0126] Filter 240 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 240 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and store the modified reconstructed picture in memory 250, specifically, in the DPB of memory 250. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filtering, bilateral filtering, etc.

[0127] The (modified) reconstructed picture stored in the DPB of memory 250 can be used as a reference picture in inter prediction unit 260. Memory 250 can store the motion information of the blocks from which the motion information in the current picture is derived (or decoded) and / or the motion information of the blocks that have been reconstructed in the picture. The stored motion information can be transmitted to inter prediction unit 260 for use as the motion information of spatially adjacent blocks or temporally adjacent blocks. Memory 250 can store the reconstructed samples of the reconstructed blocks in the current picture and transmit the reconstructed samples to intra prediction unit 265.

[0128] In the present disclosure, the embodiments described in filter 160, inter prediction unit 180, and intra prediction unit 185 of image encoding device 100 can be equally or correspondingly applied to filter 240, inter prediction unit 260, and intra prediction unit 265 of image decoding device 200.

[0129] Overview of inter-frame prediction

[0130] The image encoding device / image decoding device can perform inter prediction in units of blocks to derive prediction samples. Inter prediction can mean deriving a prediction in a manner that depends on data elements of a picture other than the current picture. When inter prediction is applied to a current block, the prediction block of the current block can be derived based on a reference block specified by a motion vector on a reference picture.

[0131] In this case, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information of the current block can be derived based on the correlation of the motion information between adjacent blocks and the current block, and the motion information can be derived in units of blocks, sub - blocks, or samples. The motion information can include a motion vector and a reference picture index. The motion information can also include inter prediction type information. Here, the inter prediction type information can mean the direction information of inter prediction. The inter prediction type information can indicate that one of L0 prediction, L1 prediction, or bi - prediction is used to predict the current block.

[0132] When inter - frame prediction is applied to a current block, neighboring blocks of the current block may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in a reference picture. The reference picture including the reference block of the current block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block may be referred to as a collocated reference block or a collocated CU (colCU), and the reference picture including the temporal neighboring block may be referred to as a collocated picture (colPic).

[0133] In addition, a candidate list of motion information may be constructed based on neighboring blocks of the current block, and in this case, a flag or index information indicating which candidate is used may be signaled in order to derive the motion vector and / or reference picture index of the current block.

[0134] According to the inter - frame prediction type, the motion information may include L0 motion information and / or L1 motion information. The motion vector in the L0 direction may be defined as the L0 motion vector or MVL0, and the motion vector in the L1 direction may be defined as the L1 motion vector or MVL1. The prediction based on the L0 motion vector may be defined as L0 prediction, the prediction based on the L1 motion vector may be defined as L1 prediction, and the prediction based on both the L0 motion vector and the L1 motion vector may be defined as bi - prediction. Here, the L0 motion vector may mean the motion vector associated with the reference picture list L0, and the L1 motion vector may mean the motion vector associated with the reference picture list L1.

[0135] The reference picture list L0 may include pictures before the current picture in the output order as reference pictures, and the reference picture list L1 may include pictures after the current picture in the output order. The previous picture may be defined as a forward (reference) picture, and the subsequent picture may be defined as a backward (reference) picture. In addition, the reference picture list L0 may further include pictures after the current picture in the output order as reference pictures. In this case, within the reference picture list L0, the previous pictures may be indexed first, and then the subsequent pictures may be indexed. The reference picture list L1 may further include pictures before the current picture in the output order as reference pictures. In this case, within the reference picture list L1, the subsequent pictures may be indexed first, and then the previous pictures may be indexed. Here, the output order may correspond to the picture order count (POC) order.

[0136] Figure 4 is a flowchart illustrating a video / image encoding method based on inter - frame prediction.

[0137] Figure 5 is a view illustrating the configuration of the inter - frame predictor 180 according to the present disclosure.

[0138] Figure 4 The encoding method ofFigure 2 is performed by the image encoding device. Specifically, step S410 may be performed by the inter-frame predictor 180, and step S420 may be performed by the residual processor. Specifically, step S420 may be performed by the subtractor 115. Step S430 may be performed by the entropy encoder 190. The prediction information of step S630 may be derived by the inter-frame predictor 180, and the residual information of step S630 may be derived by the residual processor. The residual information is information about the residual samples. The residual information may include information about the quantization transform coefficients for the residual samples. As described above, the residual samples may be derived as transform coefficients by the transformer 120 of the image encoding device, and the transform coefficients may be derived as quantization transform coefficients by the quantizer 130. The information about the quantization transform coefficients may be encoded by the entropy encoder 190 through the residual encoding process.

[0139] The image encoding device may perform inter-frame prediction (S410) for the current block. The image encoding device may derive the inter-frame prediction mode and motion information of the current block and generate the prediction samples of the current block. Here, the inter-frame prediction mode determination, motion information derivation, and prediction sample generation processes may be performed simultaneously or any one of them may be performed before other processes. For example, as Figure 5 shown, the inter-frame prediction unit 180 of the image encoding device may include a prediction mode determination unit 181, a motion information derivation unit 182, and a prediction sample derivation unit 183. The prediction mode determination unit 181 may determine the prediction mode of the current block, the motion information derivation unit 182 may derive the motion information of the current block, and the prediction sample derivation unit 183 may derive the prediction samples of the current block. For example, the inter-frame prediction unit 180 of the image encoding device may search for a block similar to the current block within a predetermined area (search area) of the reference picture through motion estimation, and derive a reference block whose difference from the current block is equal to or less than a predetermined criterion or minimum value. Based on this, a reference picture index indicating the reference picture in which the reference block is located may be derived, and a motion vector may be derived based on the position difference between the reference block and the current block. The image encoding device may determine the mode applied to the current block among various inter-frame prediction modes. The image encoding device may compare the rate-distortion (RD) cost for various prediction modes and determine the best inter-frame prediction mode for the current block. However, the method for the image encoding device to determine the inter-frame prediction mode of the current block is not limited to the above example, and various methods may be used.

[0140] For example, the inter prediction mode of the current block may be determined as at least one of a merge mode, a merge skip mode, a motion vector prediction (MVP) mode, a symmetric motion vector difference (SMVD) mode, an affine mode, a sub-block based merge mode, an adaptive motion vector resolution (AMVR) mode, a history-based motion vector predictor (HMVP) mode, a pairwise average merge mode, a merge mode with motion vector difference (MMVD) mode, a decoder-side motion vector refinement (DMVR) mode, a combined inter and intra prediction (CIIP) mode, or a geometric partition mode (GPM).

[0141] For example, when the skip mode or the merge mode is applied to the current block, the image coding device may derive merge candidates from neighboring blocks of the current block and use the derived merge candidates to construct a merge candidate list. Additionally, the image coding device may derive, among reference blocks indicated by the merge candidates included in the merge candidate list, a reference block whose difference from the current block is equal to or less than a predetermined criterion or a minimum value. In this case, a merge candidate associated with the derived reference block may be selected, and merge index information indicating the selected merge candidate may be generated and signaled to the image decoding device. The motion information of the current block may be derived using the motion information of the selected merge candidate.

[0142] As another example, when the MVP mode is applied to the current block, the image coding device may derive motion vector predictor (MVP) candidates from neighboring blocks of the current block and use the derived MVP candidates to construct an MVP candidate list. Additionally, the image coding device may use the motion vector of an MVP candidate selected from among the MVP candidates included in the MVP candidate list as the MVP of the current block. In this case, for example, the motion vector indicating the reference block derived by the above-described motion estimation may be used as the motion vector of the current block, and the MVP candidate having the smallest difference from the motion vector of the current block among the MVP candidates may be the selected MVP candidate. A motion vector difference (MVD) that is the difference obtained by subtracting the MVP from the motion vector of the current block may be derived. In this case, the index information indicating the selected MVP candidate and information about the MVD may be signaled to the image decoding device. Additionally, when the MVP mode is applied, the value of the reference picture index may be constructed as reference picture index information and signaled separately to the image decoding device.

[0143] The image coding device may derive a residual sample based on a prediction sample (S420). The image coding device may derive the residual sample by comparing the original sample of the current block with the prediction sample. For example, the residual sample may be derived by subtracting the corresponding prediction sample from the original sample.

[0144] An image encoding device may encode image information including prediction information and residual information (S430). The image encoding device may output the encoded image information in the form of a bitstream. The prediction information may include prediction mode information (e.g., skip flag, merge flag, or mode index, etc.) and information about motion information as information related to the prediction process. Among the prediction mode information, the skip flag indicates whether the skip mode is applied to the current block, and the merge flag indicates whether the merge mode is applied to the current block. Alternatively, the prediction mode information may indicate one of multiple prediction modes, such as a mode index. When the skip flag and the merge flag are 0, it may be determined that the MVP mode is applied to the current block. The information about motion information may include candidate selection information (e.g., merge index, mvp flag, or mvp index) as information for deriving a motion vector. Among the candidate selection information, the merge index may be signaled when the merge mode is applied to the current block, and may be information for selecting one of the merge candidates included in the merge candidate list. Among the candidate selection information, the MVP flag or MVP index may be signaled when the MVP mode is applied to the current block, and may be information for selecting one of the MVP candidates in the MVP candidate list. Specifically, the MVP flag may be signaled using the syntax elements mvp_10_flag or mvp_11_flag. Additionally, the information about motion information may include information about the above MVD and / or reference picture index information. Additionally, the information about motion information may include information indicating whether L0 prediction, L1 prediction, or bi-prediction is applied. The residual information is information about residual samples. The residual information may include information about quantization transform coefficients for the residual samples.

[0145] The output bitstream may be stored in a (digital) storage medium and sent to the image decoding device or may be sent to the image decoding device via a network.

[0146] As described above, the image encoding device may generate a reconstructed picture (a picture including reconstructed samples and reconstructed blocks) based on reference samples and residual samples. This is for the image encoding device to derive the same prediction result as the prediction result performed by the image decoding device, thereby improving the encoding efficiency. Therefore, the image encoding device may store the reconstructed picture (or reconstructed samples and reconstructed blocks) in a memory and use it as a reference picture for inter-frame prediction. As described above, the in-loop filtering process also applies to the reconstructed picture.

[0147] Figure 6 is a flowchart illustrating a video / image decoding method based on inter-frame prediction.

[0148] Figure 7 is a view illustrating the configuration of the inter-frame prediction unit 260 according to the present disclosure.

[0149] The image decoding device may perform operations corresponding to those performed by the image encoding device. The image decoding device may perform prediction for a current block and derive a prediction sample based on received prediction information.

[0150] Figure 6 The decoding method may be performed by Figure 3 the image decoding device. Steps S610 to S630 may be performed by the inter prediction unit 260, and the prediction information of step S610 and the residual information of step S640 may be obtained by the entropy decoder 210 from the bitstream. The residual processor of the image decoding device may derive the residual sample of the current block based on the residual information (S640). Specifically, the dequantizer 220 of the residual processor may perform dequantization based on the quantization transform coefficients derived according to the residual information to derive the transform coefficients, and the inverse transformer 230 of the residual processor may perform an inverse transform on the transform coefficients to derive the residual sample of the current block. Step S650 may be performed by the adder 235 or the reconstructor.

[0151] Specifically, the image decoding device may determine the prediction mode of the current block based on the received prediction information (S610). The image decoding device may determine which inter prediction mode is applied to the current block based on the prediction mode information in the prediction information.

[0152] For example, it may be determined whether the skip mode is applied to the current block based on the skip flag. Additionally, it may be determined whether the merge mode or the MVP mode is applied to the current block based on the merge flag. Alternatively, one of various inter prediction mode candidates may be selected based on the mode index. The inter prediction mode candidates may include the skip mode, the merge mode, and / or the MVP mode or may include various inter prediction modes to be described below.

[0153] The image decoding device may derive the motion information of the current block based on the determined inter prediction mode (S620). For example, when the skip mode or the merge mode is applied to the current block, the image decoding device may construct a merge candidate list to be described below and select one of the merge candidates included in the merge candidate list. The selection may be performed based on the above candidate selection information (merge index). The motion information of the selected merge candidate may be used to derive the motion information of the current block. For example, the motion information of the selected merge candidate may be used as the motion information of the current block.

[0154] As another example, when the MVP mode is applied to the current block, the image decoding device may construct an MVP candidate list and use the motion vector of the MVP candidate selected from among the MVP candidates included in the MVP candidate list as the MVP of the current block. The selection may be performed based on the above-mentioned candidate selection information (mvp flag or mvp index). In this case, the MVD of the current block may be derived based on the information about the MVD, and the motion vector of the current block may be derived based on the MVP and MVD of the current block. Additionally, the reference picture index of the current block may be derived based on the reference picture index information. The picture indicated by the reference picture index in the reference picture list of the current block may be derived as the reference picture for inter-frame prediction of the current block.

[0155] The image decoding device may generate a prediction sample for the current block based on the motion information of the current block (S630). In this case, the reference picture may be derived based on the reference picture index of the current block, and the prediction sample of the current block may be derived using the samples of the reference block indicated by the motion vector of the current block on the reference picture. In some cases, a prediction sample filtering process may also be performed on all or some of the prediction samples of the current block.

[0156] For example, as Figure 7 shown, the inter-frame prediction unit 260 of the image decoding device may include a prediction mode determination unit 261, a motion information derivation unit 262, and a prediction sample derivation unit 263. In the inter-frame prediction unit 260 of the image decoding device, the prediction mode determination unit 261 may determine the prediction mode of the current block based on the received prediction mode information, the motion information derivation unit 262 may derive the motion information (motion vector and / or reference picture index, etc.) of the current block based on the received motion information, and the prediction sample derivation unit 263 may derive the prediction sample of the current block.

[0157] The image decoding device may generate a residual sample for the current block based on the received residual information (S640). The image decoding device may generate a reconstructed sample for the current block based on the prediction sample and the residual sample and generate a reconstructed picture based on this (S650). Thereafter, an in-loop filtering process is applied to the reconstructed picture as described above.

[0158] As described above, the inter-frame prediction process may include steps of determining an inter-frame prediction mode, deriving motion information according to the determined prediction mode, and performing prediction (generating a prediction sample) based on the derived motion information. As described above, the inter-frame prediction process may be performed by an image encoding device and an image decoding device.

[0159] Hereinafter, the step of deriving motion information according to the prediction mode will be described in more detail.

[0160] As described above, the motion information of the current block can be used to perform inter-frame prediction. The image encoding device can derive the optimal motion information of the current block through a motion estimation process. For example, the image encoding device can search for a similar reference block with high correlation in the reference picture using the original block in the original picture of the current block in fractional pixel units, and use it to derive the motion information. The similarity of the block can be calculated based on the sum of absolute differences (SAD) between the current block and the reference block. In this case, the motion information can be derived based on the reference block with the minimum SAD in the search area. The derived motion information can be signaled to the image decoding device according to various methods based on the inter-frame prediction mode.

[0161] When the merge mode is applied to the current block, the motion information of the current block is not directly sent, and the motion information of the neighboring block is used to derive the motion information of the current block. Therefore, the motion information of the current prediction block can be indicated by sending flag information indicating that the merge mode is used and candidate selection information (e.g., merge index) indicating which neighboring block is used as a merge candidate. In the present disclosure, since the current block is a prediction execution unit, the current block can be used in the same meaning as the current prediction block, and the neighboring block can be used in the same meaning as the neighboring prediction block.

[0162] The image encoding device can search for merge candidate blocks for deriving the motion information of the current block to perform the merge mode. For example, up to five merge candidate blocks can be used, but this is not limited thereto. The maximum number of merge candidate blocks can be sent in the slice header or the tile group header, but this is not limited thereto. After finding the merge candidate blocks, the image encoding device can generate a merge candidate list and select the merge candidate block with the minimum RD cost as the final merge candidate block.

[0163] The merge candidate list can use, for example, five merge candidate blocks. For example, four spatial merge candidates and one temporal merge candidate can be used.

[0164] Overview of the affine mode

[0165] Hereinafter, the affine mode, which is an example of the inter-frame prediction mode, will be described in detail. In a conventional video encoding / decoding system, only one motion vector is used to express the motion information of the current block (translation motion model). However, in the conventional method, the optimal motion information is only expressed in units of blocks, but the optimal motion information cannot be expressed in units of pixels. To solve this problem, an affine motion mode that defines the motion information of a block in units of pixels has been proposed. According to the affine mode, two to four motion vectors associated with the current block can be used to determine the motion vectors of each pixel and / or sub-block unit of the block.

[0166] Compared with the existing motion information expressed using the translation (or displacement) of pixel values, in the affine mode, the motion information of each pixel can be expressed using at least one of translation, scaling, rotation, or shear.

[0167] Figure 8 is a view illustrating the motion that can be expressed in the affine mode.

[0168] In Figure 8 Among the motions shown, the affine mode expressing the motion information of each pixel using displacement, scaling, or rotation can be a similarity or simplified affine mode. The affine mode in the following description can mean a similarity or simplified affine mode.

[0169] The motion information in the affine mode can be expressed using two or more control point motion vectors (CPMV). The CPMV can be used to derive the motion vector of a specific pixel position in the current block. In this case, the set of motion vectors of each pixel and / or sub - block in the current block can be defined as an affine motion vector field (affine MVF).

[0170] Figure 9 is a view illustrating the parametric model of the affine mode.

[0171] When the affine mode is applied to the current block, one of the 4 - parameter model and the 6 - parameter model can be used to derive the affine MVF. In this case, the 4 - parameter model can mean a model type using two CPMV, and the 6 - parameter model can mean a model type using three CPMV. Figure 9 of (a) and Figure 9 of (b) respectively show the CPMV used in the 4 - parameter model and the 6 - parameter model.

[0172] When the position of the current block is (x, y), the motion vector according to the pixel position can be derived according to Equation 1 or Equation 2 below. For example, the motion vector according to the 4 - parameter model can be derived according to Equation 1, and the motion vector according to the 6 - parameter model can be derived according to Equation 2.

[0173] [Equation 1]

[0174]

[0175] [Equation 2]

[0176]

[0177] In Equations 1 and 2, mv0 = {mv_0x, mv_0y} may be the CPMV at the upper left corner position of the current block, mv1 = {mv_1x, mv_1y} may be the CPMV at the upper right position of the current block, and mv2 = {mv_2x, mv_2y} may be the CPMV at the lower left position of the current block. In this case, W and H respectively correspond to the width and height of the current block, and mv = {mv_x, mv_y} may mean the motion vector of the pixel position {x, y}.

[0178] In the encoding / decoding process, the affine MVF may be determined in units of pixels and / or predefined sub-blocks. When determining the affine MVF in units of pixels, the motion vector may be derived based on each pixel value. In addition, when determining the affine MVF in units of sub-blocks, the motion vector of the corresponding block may be derived based on the central pixel value of the sub-block. The central pixel value may mean a virtual pixel existing at the center of the sub-block or the lower right pixel among the four pixels existing at the center. In addition, the central pixel value may be a specific pixel in the sub-block and may be a pixel representing the sub-block. In the present disclosure, the case of determining the affine MVF in units of 4×4 sub-blocks will be described. However, this is only for convenience of description, and the size of the sub-block may be changed differently.

[0179] That is, when affine prediction is available, the motion models applicable to the current block may include three models, namely, a translational motion model, a 4-parameter affine motion model, and a 6-parameter affine motion model. Here, the translational motion model may represent the model used by the existing block unit motion vector, the 4-parameter affine motion model may represent the model used by two CPMVs, and the 6-parameter affine motion model may represent the model used by three CPMVs. The affine mode may be divided into detailed modes according to the motion information encoding / decoding method. For example, the affine mode may be further divided into an affine MVP mode and an affine merge mode.

[0180] When the affine merge mode is applied to the current block, the CPMV may be derived from neighboring blocks of the current block encoded / decoded in the affine mode. When at least one neighboring block of the current block is encoded / decoded in the affine mode, the affine merge mode may be applied to the current block. That is, when the affine merge mode is applied to the current block, the CPMV of the current block may be derived using the CPMV of the neighboring blocks. For example, the CPMV of the neighboring block may be determined as the CPMV of the current block, or the CPMV of the current block may be derived based on the CPMV of the neighboring block. When deriving the CPMV of the current block based on the CPMV of the neighboring block, at least one encoding parameter of the current block or the neighboring block may be used. For example, the CPMV of the neighboring block may be modified based on the size of the neighboring block and the size of the current block and used as the CPMV of the current block.

[0181] In addition, the affine merge of deriving the MV in units of sub - blocks can be referred to as the sub - block merge mode, which can be specified by a merge_subblock_flag having a first value (e.g., 1). In this case, the affine merge candidate list described below can be referred to as the sub - block merge candidate list. In this case, the candidate derived as the following SbTMVP can be further included in the sub - block merge candidate list. In this case, the candidate derived as sbTMVP can be used as the candidate of index #0 of the sub - block merge candidate list. In other words, the candidate derived as sbTMVP can be located in front of the inherited affine candidate and the constructed affine candidate described below in the sub - block merge candidate list.

[0182] For example, an affine mode flag can be defined to specify whether the affine mode is applicable to the current block, which can be signaled at at least one higher level (e.g., sequence, picture, slice, tile, tile group, patch, etc.) of the current block. For example, the affine mode flag can be named sps_affine_enabled_flag.

[0183] When applying the affine merge mode, the affine merge candidate list can be configured to derive the CPMV of the current block. In this case, the affine merge candidate list can include at least one of an inherited affine merge candidate, a constructed affine merge candidate, or a zero merge candidate. When a neighboring block of the current block is encoded / decoded in the affine mode, the inherited affine merge candidate can mean a candidate derived using the CPMV of the neighboring block. The constructed affine merge candidate can mean a candidate for deriving each CPMV based on the motion vectors of the neighboring blocks of each control point (CP). In addition, the zero merge candidate can mean a candidate consisting of a CPMV of size 0. In the following description, CP can mean a specific position of a block for deriving the CPMV. For example, CP can be the respective vertex positions of the block.

[0184] Figure 10 is a view illustrating a method of generating an affine merge candidate list.

[0185] Refer to Figure 10 According to the flowchart of, the affine merge candidates can be added to the affine merge candidate list in the order of the inherited affine merge candidate (S1210), the constructed affine merge candidate (S1220), and the zero merge candidate (S1230). When even if all the inherited affine merge candidates and the constructed affine merge candidates are added to the affine merge candidate list, the number of candidates included in the candidate list still does not satisfy the maximum number of candidates, the zero merge candidate can be added. In this case, the zero merge candidate can be added until the number of candidates in the affine merge candidate list satisfies the maximum number of candidates.

[0186] Figure 11A view exemplifying a control point motion vector (CPMV) derived from neighboring blocks.

[0187] For example, up to two inherited affine merge candidates can be derived, each of which can be derived based on at least one of the left neighboring block and the upper neighboring block.

[0188] Figure 12 A view exemplifying neighboring blocks for deriving inherited affine merge candidates.

[0189] An inherited affine merge candidate derived based on the left neighboring block is derived based on at least one of neighboring blocks A0 or A1 of Figure 12 and an inherited affine merge candidate derived based on the upper neighboring block can be derived based on at least one of neighboring blocks B0, B1, or B2 of Figure 12 In this case, the scanning order of the neighboring blocks may be A0 to A1 and B0, B1, and B2, but is not limited thereto. For each of the left and upper, an inherited affine merge candidate can be derived based on the first available neighboring block in the scanning order. In this case, no redundancy check may be performed between candidates derived from the left neighboring block and the upper neighboring block.

[0190] For example, as shown in Figure 11 , when the left neighboring block A is encoded / decoded in an affine mode, at least one of the motion vectors v2, v3, and v4 corresponding to the CP of the neighboring block A can be derived. When the neighboring block A is encoded / decoded by a 4-parameter affine model, the inherited affine merge candidate can be derived using v2 and v3. In contrast, when the neighboring block A is encoded / decoded by a 6-parameter affine model, the inherited affine merge candidate can be derived using v2, v3, and v4.

[0191] Figure 13 A view exemplifying neighboring blocks for deriving constructed affine merge candidates.

[0192] A constructed affine candidate may mean a candidate having a CPMV derived by combining general motion information of neighboring blocks. The motion information of each CP can be derived using spatial neighboring blocks or temporal neighboring blocks of the current block. In the following description, CPMVk may mean a motion vector representing the k-th CP. For example, referring to Figure 13 , CPMV1 can be determined as the first available motion vector among the motion vectors of B2, B3, and A2, and in this case, the scanning order may be B2, B3, and A2. CPMV2 can be determined as the first available motion vector among the motion vectors of B1 and B0, and in this case, the scanning order may be B1 and B0. CPMV3 can be determined as one of the motion vectors of A1 and A0, and in this case, the scanning order may be A1 and A0. When TMVP is applicable to the current block, CPMV4 can be determined as the motion vector of the temporal neighboring block T.

[0193] After deriving the four motion vectors of each CP, an affine merge candidate can be derived based on this. The construction of the affine merge candidate can be configured by including at least two motion vectors selected from among the four motion vectors of each derived CP. For example, the construction of the affine merge candidate can consist of at least one of {CPMV1, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV4}, {CPMV1, CPMV3, CPMV4}, {CPMV2, CPMV3, CPMV4}, {CPMV1, CPMV2}, or {CPMV1, CPMV3} in this order. The construction affine candidate consisting of three motion vectors can be a candidate for the 6-parameter affine model. In contrast, the construction affine candidate consisting of two motion vectors can be a candidate for the 4-parameter affine model. To avoid the scaling process of the motion vectors, when the reference picture indices of the CPs are different from each other, the combination of the relevant CPMVs can be ignored and not used for deriving the construction affine candidate.

[0194] When the affine MVP mode is applied to the current block, the encoding / decoding device can derive two or more CPMV predictors and CPMVs of the current block and derive the CPMV difference based on them. In this case, the CPMV difference can be signaled from the encoding device to the decoding device. The image decoding device can derive the CPMV predictor of the current block, reconstruct the signaled CPMV difference, and then derive the CPMV of the current block based on the CPMV predictor and the CPMV difference.

[0195] In addition, when the affine merge mode or sub-block based TMVP is not applied to the current block (for example, the value of the affine merge flag or merge_subblock_flag is 0), the affine MVP mode can be applied to the current block. Alternatively, when the value of inter_affine_flag is 1, the affine MVP mode can be applied to the current block. In addition, the affine MVP mode can be expressed as an affine CP MVP mode. The affine mvp candidate list described below can be referred to as the control point motion vector predictor candidate list.

[0196] When the affine MVP mode is applied to the current block, the affine MVP candidate list can be configured to derive the CPMV of the current block. In this case, the affine MVP candidate list can include at least one of an inherited affine MVP candidate, a constructed affine MVP candidate, a translational motion affine MVP candidate, or a zero MVP candidate. For example, the affine MVP candidate list can include up to n (e.g., n = 2) candidates.

[0197] In this case, an inherited affine MVP candidate may mean a candidate derived based on the CPMV of neighboring blocks when neighboring blocks of the current block are encoded / decoded in affine mode. Constructing an affine MVP candidate may mean a candidate derived by generating a CPMV combination based on the motion vectors of CP neighboring blocks. A zero MVP candidate may mean a candidate consisting of CPMVs with a value of 0. The derivation method and the characteristics of the inherited affine MVP candidate and the constructed affine MVP candidate are the same as those of the inherited affine candidate and the constructed affine candidate described above, and thus their descriptions will be omitted.

[0198] When the maximum number of candidates in the affine MVP candidate list is 2, a constructed affine MVP candidate, a translational motion affine MVP candidate, and a zero MVP candidate may be added when the current number of candidates is less than 2. Specifically, the translational motion affine MVP candidate may be derived in the following order.

[0199] For example, when the number of candidates included in the affine MVP candidate list is less than 2 and CPMV0 of the constructed affine MVP candidate is valid, CPMV0 may be used as an affine MVP candidate. That is, an affine MVP candidate in which all the motion vectors of CP0, CP1, and CP2 are CPMV0 may be added to the affine MVP candidate list.

[0200] Next, when the number of candidates in the affine MVP candidate list is less than 2 and CPMV1 of the constructed affine MVP candidate is valid, CPMV1 may be used as an affine MVP candidate. That is, an affine MVP candidate in which all the motion vectors of CP0, CP1, and CP2 are CPMV1 may be added to the affine MVP candidate list.

[0201] Next, when the number of candidates in the affine MVP candidate list is less than 2 and CPMV2 of the constructed affine MVP candidate is valid, CPMV2 may be used as an affine MVP candidate. That is, an affine MVP candidate in which all the motion vectors of CP0, CP1, and CP2 are CPMV2 may be added to the affine MVP candidate list.

[0202] Regardless of the above conditions, when the number of candidates in the affine MVP candidate list is less than 2, the temporal motion vector predictor (TMVP) of the current block may be added to the affine MVP candidate list. Regardless of the above, when the number of candidates in the affine MVP candidate list is less than 2, a zero MVP candidate may be added to the affine MVP candidate list.

[0203] Figure 14 is a view illustrating a method of generating an affine MVP candidate list.

[0204] Refer to Figure 14The flowchart can add candidates to the affine MVP candidate list in the order of inheriting affine MVP candidates (S1610), constructing affine MVP candidates (S1620), translational motion affine MVP candidates (S1630), and zero MVP candidates (S1640). As described above, steps S1620 to S1640 can be executed according to whether the number of candidates included in the affine MVP candidate list in each step is less than 2.

[0205] The scanning order of inheriting affine MVP candidates can be equal to the scanning order of inheriting affine merge candidates. However, in the case of inheriting affine MVP candidates, only neighboring blocks referring to the same reference picture as the reference picture of the current block can be considered. When an inheriting affine MVP candidate is added to the affine MVP candidate list, a redundancy check may not be performed.

[0206] To derive a constructed affine MVP candidate, only Figure 13 the spatial neighboring blocks shown can be considered. Additionally, the scanning order of constructing affine MVP candidates can be equal to the scanning order of constructing affine merge candidates. Additionally, to derive a constructed affine MVP candidate, the reference picture index of neighboring blocks can be checked, and in the scanning order, the first neighboring block that is inter-frame coded and refers to the same reference picture as the reference picture of the current block can be used.

[0207] Overview of the sub-block based temporal motion vector prediction (SbTMVP) mode

[0208] Hereinafter, the sub-block based TMVP mode, which is an example of an inter-frame prediction mode, will be described in detail. According to the sub-block based TMVP mode, the motion vector field (MVF) of the current block can be derived and the motion vector can be derived in units of sub-blocks.

[0209] Different from the conventional TMVP mode that is performed in units of coding units, for a coding unit to which the sub-block based TMVP mode is applied, the motion vector can be encoded / decoded in units of sub-coding units. Additionally, according to the conventional TMVP mode, the temporal motion vector can be derived from the collocated block in the collocated picture, but in the sub-block based TMVP mode, the motion vector field can be derived from the reference block in the collocated picture specified by the motion vector derived from the neighboring blocks of the current block. Hereinafter, the motion vector derived from the neighboring blocks may be referred to as the motion shift or representative motion vector of the current block.

[0210] Figure 15 is a view exemplifying the neighboring blocks of the sub-block based TMVP mode.

[0211] When the sub-block based TMVP mode is applied to the current block, the neighboring blocks for determining the motion shift can be determined. For example, it can be in accordance with Figure 15The order of the blocks of A1, B1, B0, and A0 performs a scan on neighboring blocks used to determine the motion shift. As another example, the neighboring blocks used to determine the motion shift can be restricted to specific neighboring blocks of the current block. For example, the neighboring blocks used to determine the motion shift can always be determined as block A1. When the neighboring block has a motion vector of the reference col picture, the corresponding motion vector can be determined as the motion shift. The motion vector determined as the motion shift can be referred to as the temporal motion vector. In addition, when the above motion vector cannot be derived from the neighboring block, the motion shift can be set to (0,0).

[0212] Figure 16 is a view illustrating a method of deriving a motion vector field according to the sub-block-based TMVP mode.

[0213] Next, the reference block on the collocated picture specified by the motion shift can be determined. For example, the sub-block-based motion information (motion vector or reference picture index) can be obtained from the col picture by adding the motion shift to the coordinates of the current block. In Figure 16 the example shown, it is assumed that the motion shift is the motion vector of block A1. By applying the motion shift to the current block, the sub-blocks (col sub-blocks) corresponding to the respective sub-blocks configuring the current block in the col picture can be specified. Thereafter, using the motion information of the corresponding sub-blocks (col sub-blocks) in the col picture, the motion information of the respective sub-blocks of the current block can be derived. For example, the motion information of the corresponding sub-block can be obtained from the center position of the corresponding sub-block. In this case, the center position can be the position of the lower-right sample among the four samples located at the center of the corresponding sub-block. When the motion information of a specific sub-block of the col block corresponding to the current block is not available, the motion information of the center sub-block of the col block can be determined as the motion information of the corresponding sub-block. When deriving the motion vector of the corresponding sub-block, similar to the above TMVP process, the reference picture index and the motion vector of the current sub-block can be switched. That is, when deriving the sub-block-based motion vector, the POC of the reference picture of the reference block can be considered to perform the scaling of the motion vector.

[0214] As described above, the sub-block-based TMVP candidate of the current block can be derived using the motion vector field or motion information of the current block derived based on sub-blocks.

[0215] Hereinafter, a merge candidate list configured in units of sub-blocks is defined as a sub-block unit merge candidate list. The above affine merge candidate and the sub-block-based TMVP candidate can be merged to configure a sub-block unit merge candidate list.

[0216] In addition, a sub-block based TMVP mode flag can be defined to specify whether the sub-block based TMVP mode is applicable to the current block, and it can be signaled at at least one level among higher levels of the current block (e.g., sequence, picture, slice, tile, tile group, patch, etc.). For example, the sub-block based TMVP mode flag can be named sps_sbtmvp_enabled_flag. When the sub-block based TMVP mode is applicable to the current block, the sub-block based TMVP candidates can be first added to the sub-block unit merge candidate list, and then the affine merge candidates can be added to the sub-block unit merge candidate list. In addition, the maximum number of candidates that can be included in the sub-block unit merge candidate list can be signaled. For example, the maximum number of candidates that can be included in the sub-block unit merge candidate list can be 5.

[0217] The size of the sub-blocks used to derive the sub-block unit merge candidate list can be signaled or preset to M×N. For example, M×N can be 8×8. Therefore, the affine mode or the sub-block based TMVP mode is applicable to the current block only when the size of the current block is 8×8 or larger.

[0218] Hereinafter, embodiments of the prediction execution method of the present disclosure will be described. It can be performed in Figure 4 step S410 of Figure 6 or step S630 of

[0219] The prediction block of the current block can be generated based on the motion information derived according to the prediction mode. The prediction block (predicted block) can include the prediction samples (prediction sample array) of the current block. When the motion vector of the current block specifies partial sample units, an interpolation process can be performed, and thus, the prediction samples of the current block can be derived based on the reference samples in units of partial samples within the reference picture. When the affine inter prediction is applied to the current block, the prediction samples can be generated based on the sample / sub-block unit MV. When dual prediction is applied, the prediction samples derived by the weighted sum or weighted average (according to the phase) of the prediction samples derived based on L0 prediction (i.e., prediction using MVL0 and the reference picture within the reference picture list L0) and the prediction samples derived based on L1 prediction (i.e., prediction using MLV1 and the reference picture within the reference picture list L1) can be used as the prediction samples of the current block. When dual prediction is applied and the reference picture for L0 prediction and the reference picture for L1 prediction are in different temporal directions with respect to the current picture (i.e., if it corresponds to dual prediction and bi-directional prediction), this can be referred to as true dual prediction.

[0220] In an image decoding device, reconstructed samples and a reconstructed picture can be generated based on derived prediction samples, and then an in-loop filtering process can be performed. Additionally, in an image encoding device, residual samples can be derived based on the derived prediction samples, and encoding of image information including prediction information and residual information can be performed.

[0221] Bi-prediction with CU-level weights (BCW)

[0222] When dual prediction is applied to a current block as described above, prediction samples can be derived based on weighted averaging. Conventionally, a dual prediction signal (i.e., dual prediction samples) can be derived by simply averaging an L0 prediction signal (L0 prediction samples) and an L1 prediction signal (L1 prediction samples). That is, the dual prediction samples are derived by averaging an L0 prediction sample based on an L0 reference picture and MVL0 and an L1 prediction sample based on an L1 reference picture and MVL1. However, according to the present disclosure, when dual prediction is applied, the dual prediction signal (dual prediction samples) can be derived by weighted averaging of the L0 prediction signal and the L1 prediction signal as follows.

[0223] [Equation 3]

[0224] P bi-pred = ((8 - w) * P0 + w * P1 + 4) >> 3

[0225] In Equation 3 above, P bi-pred represents a dual prediction signal (dual prediction block) derived by weighted averaging, and P0 and P1 represent an L0 prediction sample (L0 prediction block) and an L1 prediction sample (L1 prediction block), respectively. Additionally, (8 - w) and w represent weights applied to P0 and P1, respectively.

[0226] When generating a dual prediction signal by weighted averaging, five weights can be allowed. For example, the weight w can be selected from {-2, 3, 4, 5, 10}. For each dual prediction CU, the weight w can be determined by one of two methods. As the first of these two methods, when the current CU is not in merge mode (non-merge CU), the weight index can be signaled together with the motion vector difference. For example, the bitstream can include information about the weight index after information about the motion vector difference. As the second of these two methods, when the current CU is in merge mode (merge CU), the weight index can be derived from neighboring blocks based on the merge candidate index (merge index).

[0227] Generating a dual prediction signal through weighted averaging can be restricted to be applied only to CUs having a size including 256 or more samples (luminance component samples). That is, dual prediction through weighted averaging can be performed only for a CU in which the product of the width and height of the current block is 256 or greater. Additionally, the weight w can be used as one of the five weights described above, and one of different numbers of weights can be used. For example, depending on the characteristics of the current image, five weights can be used for low-delay pictures, and three weights can be used for non-low-delay pictures. In this case, the three weights can be {3, 4, 5}.

[0228] By applying a fast search algorithm, an image encoding device can determine a weight index without significantly increasing complexity. In this case, the fast search algorithm can be summarized as follows. Hereinafter, unequal weights can mean that the weights applied to P0 and P1 are not equal. Additionally, equal weights can mean that the weights applied to P0 and P1 can be equal.

[0229] - In the case where the AMVR mode in which the resolution of the motion vector is adaptively changed is applied together, when the current picture is a low-delay picture, unequal weights can be conditionally checked only for each of the 1-pixel motion vector resolution and the 4-pixel motion vector resolution.

[0230] - In the case where the affine mode is applied together and the affine mode is selected as the best mode of the current block, the image encoding device can perform affine motion estimation (ME) for each unequal weight.

[0231] - When the two reference pictures used for dual prediction are equal, unequal weights can be conditionally checked only.

[0232] - When a predetermined condition is satisfied, unequal weights can be not checked. The predetermined picture can be based on the POC distance, quantization parameter (QP), temporal level, etc. between the current picture and the reference picture.

[0233] The weight index of BCW can be encoded using one context coding bin and one or more subsequent bypass coding bins. The first context coding bin specifies whether equal weights are used. When unequal weights are used, additional bins can be bypass-coded and signaled. The additional bins can be signaled to specify which weight is used.

[0234] Weighted prediction (WP) is a tool for efficiently encoding images including fade. According to weighted prediction, weighted parameters (weights and offsets) can be signaled for each reference picture included in each of the reference picture lists L0 and L1. Then, when motion compensation is performed, the weights and offsets can be applied to the corresponding reference pictures. Weighted prediction and BCW can be used for different types of images. To avoid interaction between weighted prediction and BCW, for a CU using weighted prediction, the BCW weight index may not be signaled. In this case, the weight can be inferred to be 4. That is, equal weights can be applied.

[0235] In the case of a CU applying the merge mode, the weight index can be inferred from neighboring blocks based on the merge candidate index. This can be applied to both the general merge mode and the inherited affine merge mode.

[0236] In the case of constructing the affine merge mode, the affine motion information can be configured based on the motion information of up to three blocks. The BCW weight index of a CU using the constructed affine merge mode can be set to the BCW weight index of the first CP in the combination. That is, BCW may not be applied to a CU encoded in the CIIP mode. For example, the BCW weight index of a CU encoded in the CIIP mode can be set to a value specifying equal weights.

[0237] Bidirectional optical flow (BDOF)

[0238] According to the present disclosure, BDOF can be used to refine the bi-prediction signal. BDOF generates prediction samples by calculating refined motion information when bi-prediction is applied to the current block (e.g., CU). Therefore, the process of calculating refined motion information by applying BDOF can be included in the above motion information derivation step.

[0239] For example, BDOF can be applied at the 4×4 sub-block level. That is, BDOF can be performed in units of 4×4 sub-blocks within the current block.

[0240] For example, BODF can be applied to a CU that satisfies at least one or all of the following conditions.

[0241] - The CU is encoded in the true bi-prediction mode, that is, one of the two reference pictures is before the current picture in the display order, and the other is after the current picture in the display order

[0242] - The CU is not in the affine mode or the ATMVP merge mode

[0243] - The CU has more than 64 luma samples

[0244] - The height and width of the CU are 8 or more luma samples

[0245] - The BCW weight index specifies equal weights, i.e., equal weights are applied to the L0 prediction samples and the L1 prediction samples.

[0246] - Weighted prediction (WP) is not applied to the current CU.

[0247] - The CIIP mode is used for the current CU.

[0248] In addition, BDOF can be applied only to the luminance component. However, the present disclosure is not limited thereto, and BDOF can be applied to the chrominance component or both the luminance component and the chrominance component.

[0249] The BDOF mode is based on the concept of optical flow. That is, it is assumed that the motion of the object is smooth. When BDOF is applied, for each 4×4 sub-block, the motion refinement (v x , v y ) can be calculated. The motion refinement can be calculated by minimizing the difference between the L0 prediction samples and the L1 prediction samples. The motion refinement can be used to adjust the dual prediction sample values within the 4×4 sub-block.

[0250] Hereinafter, the process of performing BDOF will be described in more detail.

[0251] First, the horizontal gradient and the vertical gradient of the two prediction signals can be calculated. In this case, k can be 0 or 1. The gradient can be calculated by directly calculating the difference between two adjacent samples. For example, the gradient can be calculated as follows.

[0252] [Equation 4]

[0253]

[0254] In Equation 4 above, I (k) (i, j) represents the sample value of the coordinate (i, j) of the prediction signal in the list k (k = 0, 1). For example, I (0) (i, j) can represent the sample value at the position (i, j) in the L0 prediction block, and I (1) (i, j) can represent the sample value at the position (i, j) in the L1 prediction block. In Equation 4 above, the first shift shift1 can be determined based on the bit depth of the luminance component. For example, when the bit depth of the luminance component is bitDepth, shift1 can be determined as max(6, bitDepth - 6).

[0255] As described above, after calculating the gradient, the autocorrelation and cross-correlation S1, S2, S3, S5, and S6 between the gradients can be calculated as follows.

[0256] [Equation 5]

[0257] S1 = ∑(i,j)∈Ω Abs(ψ x (i,j)), S3 = ∑ (i,j)∈Ω θ(i,j)·Sign(ψ x (i,j))

[0258] S2 = ∑ (i,j)∈Ω ψ x (i,j)·Sign(ψ y (i,j))

[0259] S5 = ∑ (i,j)∈Ω Abs(ψ y (i,j)) S6 = ∑ (i,j)∈Ω θ(i,j)·ψ y (i,j)

[0260] Wherein,

[0261]

[0262] θ(i,j) = (I (1) (i,j) >> n b ) - (I (0) (i,j) >> n b )

[0263] Where Ω is a 6×6 window around the 4×4 sub-block.

[0264] In Equation 5 above, n a and n b can be set to min(1, bitDepth - 11) and min(4, bitDepth - 8), respectively.

[0265] The motion refinement (v x , v y ) can be derived as follows using the autocorrelation and cross-correlation between the above gradients.

[0266] [Equation 6]

[0267]

[0268] Wherein, S_(2,s) = S_2 & (2^(n_(S_2)) - 1), th′ BIO = 2 13-BD . And is the floor function.

[0269] In Equation 6 above, n S2 can be 12. Based on the derived motion refinement and gradients, the following adjustment can be performed for each sample in the 4×4 sub-block.

[0270] [Equation 7]

[0271]

[0272] Finally, the predicted sample pred of the CU applying BDOF can be calculated by adjusting the dual predicted samples of the CU as follows BDOF .

[0273] [Equation 8]

[0274] pred BDOF (x,y) = (I (0) (x,y) + I (1) (x,y) + b(x,y) + o offset ) >> shift

[0275] In the above formula, n a , n b and n S2 can be 3, 6, and 12 respectively. These values can be selected such that the multiplier does not exceed 15 bits in BDOF processing and the bit width of the intermediate parameter is maintained within 32 bits.

[0276] To derive the gradient value, the predicted samples I (k) (i,j) existing outside the current CU in the list k (k = 0,1) can be generated. Figure 17 is a view of the CU exemplifying the extension to perform BDOF.

[0277] As Figure 17 shown, to perform BDOF, the rows / columns extending around the boundary of the CU can be used. To control the computational complexity of generating the predicted samples outside the boundary, the predicted samples in the extended region ( Figure 17 the white region in) can be generated using a bilinear filter, and the predicted samples in the CU ( Figure 17 the gray region in) can be generated using a normal 8-tap motion compensation interpolation filter. The sample values at the extended positions can be used only for gradient calculation. When the sample values and / or gradient values outside the CU boundary are required to perform the remaining steps of the BDOF process, the nearest neighbor sample values and / or gradient values can be filled (repeated) and used.

[0278] When the width and / or height of the CU is greater than 16 luma samples, the corresponding CU can be divided into sub-blocks with a width and / or height of 16 luma samples. The boundaries of the sub-blocks can be processed in the same way as the CU boundary described above. The maximum unit size for performing the BDOF process can be limited to 16×16.

[0279] For each sub-block, it can be determined whether to perform BDOF. That is, the BDOF processing for each sub-block can be skipped. For example, when the sum of absolute differences (SAD) value between the initial L0 prediction sample and the initial L1 prediction sample is less than a predetermined threshold, the BDOF processing may not be applied to the corresponding sub-block. In this case, when the width and height of the corresponding sub-block are W and H, the predetermined threshold can be set to (8 * W * (H >> 1)). Considering the complexity of additional SAD calculation, the SAD between the initial L0 prediction sample and the initial L1 prediction sample calculated in the DMVR processing can be reused.

[0280] When BCW is available for the current block, for example, when the BCW weight index specifies unequal weights, BDOF may not be applied. Similarly, when WP is available for the current block, for example, when the luma_weight_lx_flag of at least one of the two reference pictures is 1, BDOF may not be applied. In this case, the luma_weight_lx_flag may be information specifying the weighting factor of WP for the luminance component indicating whether there is an lx prediction (x is 0 or 1) in the bitstream or information specifying whether WP is applied to the luminance component of the lx prediction. When the CU is encoded in the symmetric MVD (SMVD) mode or the CIIP mode, BDOF may not be applied.

[0281] Optical flow prediction refinement (PROF)

[0282] Hereinafter, a method for refining a sub-block-based affine motion compensation prediction block by applying optical flow will be described. The prediction samples generated by performing sub-block-based affine motion compensation can be refined based on the difference derived from the optical flow equation. In the present disclosure, the refinement of these prediction samples may be referred to as optical flow prediction refinement (PROF). Through PROF, inter-frame prediction at the pixel-level granularity can be achieved without increasing the memory access bandwidth.

[0283] The parameters of the affine motion model can be used to derive the motion vectors of each pixel in the CU. However, since pixel-based affine motion compensation prediction results in high complexity and an increase in memory access bandwidth, sub-block-based affine motion compensation prediction can be performed. When performing sub-block-based affine motion compensation prediction, the CU can be divided into 4×4 sub-blocks, and motion vectors can be determined for each sub-block. In this case, the motion vectors of each sub-block can be derived from the CPMV of the CU. Sub-block-based affine motion compensation has a trade-off relationship between coding efficiency and complexity and memory access bandwidth. Since motion vectors are derived in units of sub-blocks, the complexity and memory access bandwidth are reduced, but the prediction accuracy is reduced.

[0284] Therefore, by applying optical flow to the sub-block-based affine motion compensation prediction, refined-granularity motion compensation can be achieved through refinement.

[0285] As described above, the luminance prediction sample can be refined by adding the difference derived from the optical flow equation after performing sub-block based affine motion compensation. More specifically, PROF can be performed in the following four steps.

[0286] Step 1) Generate a prediction sub-block I(i,j) by performing sub-block based affine motion compensation.

[0287] Step 2) Calculate the spatial gradients g x (i,j) and g y (i,j) of the prediction sub-block at each sample position. In this case, a 3-tap filter can be used, and the filter coefficients can be [-1, 0, 1]. For example, the spatial gradients can be calculated as follows.

[0288] [Equation 9]

[0289] g x (i,j) = I(i + 1, j) - I(i - 1, j)

[0290] g y (i,j) = I(i, j + 1) - I(i, j - 1)

[0291] To calculate the gradients, the prediction sub-block can be extended by one pixel on each side. In this case, to reduce the memory bandwidth and complexity, the pixels at the extended boundaries can be copied from the nearest integer pixels in the reference picture. Therefore, additional interpolation of the padding area can be skipped.

[0292] Step 3) Calculate the luminance prediction refinement (ΔI(i,j)) by the optical flow equation. For example, the following equation can be used.

[0293] [Equation 10]

[0294] ΔI(i,j) = g x (i,j) * Δv x (i,j) + g y (i,j) * Δv y (i,j)

[0295] In the above equation, Δv(i,j) represents the difference between the pixel motion vector (pixel MV, v(i,j)) calculated at the sample position (i,j) and the sub-block MV of the sub-block to which the sample (i,j) belongs.

[0296] Figure 18 is a view illustrating the relationship between Δv(i,j), v(i,j), and the sub-block motion vector.

[0297] In Figure 18In the example shown, for example, the difference between the motion vector v(i,j) at the upper left sample position of the current sub-block and the motion vector v of the current sub-block SB can be represented by a thick dashed arrow, and the vector represented by the thick dashed arrow can correspond to Δv(i,j).

[0298] The affine model parameters and pixel positions from the center of the sub-block do not change. Therefore, Δv(i,j) can be calculated only for the first sub-block and reused for other sub-blocks in the same CU. Assuming that the horizontal offset and vertical offset from the pixel position to the center of the sub-block are x and y respectively, Δv(x,y) can be derived as follows.

[0299] [Equation 11]

[0300]

[0301] For a 4-parameter affine model,

[0302]

[0303] For a 6-parameter affine model,

[0304]

[0305] In the above, (v 0x , v 0y ), (v 1x , v 1y ) and (v 2x , v 2y ) correspond to the upper left CPMV, upper right CPMV and lower left CPMV respectively, and w and h represent the width and height of the CU respectively.

[0306] Step 4) Finally, based on the calculated luminance prediction refinement ΔI(i,j) and the predicted sub-block I(i,j), the final predicted block I’(i,j) can be generated. For example, the final predicted block I’ can be generated as follows.

[0307] [Equation 12]

[0308] I′(i,j) = I(i,j) + ΔI(i,j)

[0309] Figure 19 is a view illustrating an example of a process for determining whether to apply BDOF according to the present disclosure.

[0310] Whether BDOF is applied to the current CU can be specified by the flag bdofFlag. A bdofFlag with a first value (“true” or “1”) can specify that BDOF is applied to the current CU. A bdofFlag with a second value (“false” or “0”) can specify that BDOF is not applied to the current CU. For example, it can be based onFigure 19 Derive bdofFlag from the various conditions shown below. As Figure 19 shown, bdofFlag includes conditions related to the size of the block (cbWidth, cbHeight). More specifically, when both the width cbWidth and the height cbHeight of the block are equal to or greater than 8 (luminance samples) and cbHeight * cbWidth is equal to or greater than 128 (luminance samples), bdofFlag can be set to a first value. In this case, cbHeight * cbWidth can specify the number of luminance samples included in the current CU. According to Figure 19 the example shown below, for a CU of size 8×8, bdofFlag can be set to a second value, and thus BDOF is not applied.

[0311] As described above, by applying BDOF in the inter - frame prediction process to refine the reference samples in the motion compensation process, the compression performance of the image can be increased. When the prediction mode of the current block is the normal mode (conventional merge mode or conventional AMVP mode), BDOF can be performed. That is, when the prediction mode of the current block is the affine mode, GPM mode, or CIIP mode, BDOF is not applied.

[0312] As a method similar to BDOF, PROF can be performed on the blocks encoded in the affine mode. As described above, by refining the reference samples in each 4×4 sub - block via PROF, the compression performance of the image can be increased.

[0313] The PROF according to the present disclosure may be performed according to a prediction direction. The prediction direction may include an L0 prediction direction and an L1 prediction direction. When the PROF is performed in the L0 prediction direction, the above PROF process may be applied to the L0 prediction sample, thereby generating a refined L0 prediction sample. When the PROF is performed in the L1 prediction direction, the above PROF process may be applied to the L1 prediction sample, thereby generating a refined L1 prediction sample. Accordingly, it may be derived whether to apply the PROF in both the L0 prediction direction and the L1 prediction direction. For example, a flag cbProfFlag for specifying whether to apply the PROF may include cbProfFlagL0 related to the L0 prediction direction and cbProfFlagL1 related to the L1 prediction direction. It may be determined whether the PROF is applied to a current block (CU) based on cbProfFlagL0 and / or cbProfFlagL1 for each of the L0 prediction direction and the L1 prediction direction. In the present disclosure, when cbProfFlagL0 and / or cbProfFlagL1 has a first value, it may mean that the PROF is performed in the corresponding prediction direction of the current CU. More specifically, when cbProfFlagL0 has the first value, the PROF may be performed in the L0 prediction direction of the current CU. Additionally, when cbProfFlagL1 has the first value, the PROF may be performed in the L1 prediction direction of the current CU. In the present disclosure, applying the PROF to the current CU may mean that cbProfFlagLX (X = 0 and / or 1) has the first value. In various embodiments of the present disclosure, various conditions for deriving cbProfFlagLX may be conditions for the corresponding prediction direction LX.

[0314] Figure 20 FIG. is an example view illustrating a process of determining whether to apply the PROF according to the present disclosure.

[0315] Whether the PROF is applied to the current CU may be specified by a flag cbProfFlagLX (X = 0 or 1). The cbProfFlag having a first value (“true” or “1”) may specify that the PROF is applied to the current CU. The cbProfFlag having a second value (“false” or “0”) may specify that the PROF is not applied to the current CU. For example, the cbProfFlag may be derived based on Figure 20 the various conditions shown. As Figure 20 shown, the cbProfFlag does not include conditions related to the size of the block (cbWidth, cbHeight).

[0316] Since the PROF is applicable to a block encoded in an affine mode (affine block), the size of the block to which the PROF is applied may be constrained by the block size conditions of the affine block. Accordingly, as described below, the block size conditions of the PROF and the BDOF are different from each other.

[0317] Figure 21 It is a view of signaling information that specifies whether to apply the sub-block merge mode according to an example of the present disclosure.

[0318] Whether the sub-block merge mode (affine merge mode) is applied to the current CU can be determined based on information signaled through the bitstream (e.g., Figure 21 the merge_subblock_flag). The merge_subblock_flag with the first value ("true" or "1") can specify that the sub-block merge mode is applied to the current CU. In this case, the index specifying one of the candidates included in the sub-block merge candidate list can be signaled (e.g., Figure 21 the merge_subblock_idx). When there is one candidate in the sub-block merge candidate list (when MaxNumSubblockMergeCand is 1), the index information for selecting the candidate is not signaled and can be determined as a fixed value of 0. As Figure 21 shown, the signaling conditions of the merge_subblock_flag include conditions related to the block size. Specifically, when both the width cbWidth and height cbHeight of the block are equal to or greater than 8, the merge_subblock_flag can be signaled. That is, the sub-block merge mode can be applied to blocks of size 8×8 or larger. Therefore, the PROF of the affine merge block can be applied to blocks of size 8×8 or larger.

[0319] Figure 22 It is a view of signaling information that specifies whether to apply the affine MVP mode according to an embodiment of the present disclosure.

[0320] Whether the affine MVP mode (inter-frame affine mode) is applied to the current CU can be determined based on information signaled through the bitstream (e.g., Figure 22 the inter_affine_flag). The inter_affine_flag with the first value ("true" or "1") can specify that the affine MVP mode is applied to the current CU. In this case, the index specifying one of the candidates included in the affine MVP candidate list can be signaled. As Figure 22As shown, the signaling conditions of inter_affine_flag include conditions related to the block size. Specifically, when both the width cbWidth and height cbHeight of the current block are equal to or greater than 16, inter_affine_flag can be signaled. That is, the affine MVP mode can be applied to blocks of size 16×16 or larger. Therefore, the PROF of the affine MVP block can be applied to blocks of size 16×16 or larger.

[0321] As referred to Figures 20 to 22 As described, since PROF does not include conditions related to the block size, the block sizes applicable to PROF can be constrained according to the block sizes applicable to the affine merge mode and the affine MVP mode. For example, the affine merge mode is applicable to blocks of size 8×8 or larger, and in this case, PROF can be applied to 8×8 blocks. However, the application conditions of BDOF include the condition that cbHeight*cbWidth is equal to or greater than 128 samples, and BDOF is not applied to 8×8 blocks. Therefore, the block sizes to which PROF is applied are different from the block sizes to which BDOF is applied.

[0322] The present disclosure provides various embodiments for matching the application conditions of PROF and BDOF. Specifically, the present disclosure provides various embodiments for matching the conditions related to the block sizes of PROF and BDOF. Additionally, the present disclosure provides various embodiments for matching the application conditions of PROF and BDOF by considering BCW or WP. Additionally, the present disclosure provides various embodiments including conditions related to the resolution of the current picture and the resolution of the reference picture as the application conditions of PROF.

[0323] Figure 23 is a view illustrating a process of determining whether to apply PROF according to another embodiment of the present disclosure.

[0324] Compared with Figure 20 the example of Figure 23 the embodiment of Figure 23 can additionally include conditions related to the block size as the application conditions of PROF. More specifically, as shown in the underlined part of

[0325] When cbHeight*cbWidth is less than 128 (luminance samples), cbProfFlag can be set to a second value ("false" or "0"). Figure 23 Therefore, according to Figure 23As in the embodiments, by adding conditions related to the block size to the application conditions of PROF, conditions related to the block sizes applicable to PROF and BDOF can be matched.

[0326] According to Figure 23 the embodiments, the conditions related to the block size of the affine MVP mode, the affine merge mode, PROF, and BDOF can be changed as shown in the following table.

[0327] [Table 1]

[0328]

[0329] In Table 1 above, w and h may respectively denote the width and height of the current block.

[0330] Figure 24 is a view of a signaling that indicates information specifying whether to apply a sub-block merge mode according to another embodiment of the present disclosure.

[0331] In Figure 21 the example, the conditions related to the block size among the signaling conditions of merge_subblock_flag include the condition that both cbWidth and cbHeight are equal to or greater than 8. According to Figure 24 the embodiments, the signaling conditions of merge_subblock_flag may additionally include the condition that cbWidth * cbHeight is equal to or greater than 128 (luminance samples). According to Figure 24 the embodiments, the affine merge mode is only applicable to blocks including 128 or more samples as blocks sized 8×8 or larger. That is, since the affine merge mode is not applied to 8×8 blocks, PROF may not be applied to 8×8 blocks.

[0332] According to Figure 24 the embodiments, the conditions related to the block size of the affine MVP mode, the affine merge mode, PROF, and BDOF can be changed as shown in the following table.

[0333] [Table 2]

[0334]

[0335] Figure 25 is a view of a signaling that indicates information specifying whether to apply the affine MVP mode according to another embodiment of the present disclosure.

[0336] In Figure 22 the example, the conditions related to the block size among the signaling conditions of inter_affine_flag include the condition that both cbWidth and cbHeight are equal to or greater than 16. According to Figure 25In an embodiment, the condition related to the block size among the signaling conditions of inter_affine_flag may be changed to the condition that both cbWidth and cbHeight are equal to or greater than 16 and cbWidth * cbHeight is equal to or greater than 128 (luminance samples). According to Figure 25 In an embodiment, the affine MVP mode is only applicable to a block including 128 or more samples as a block of size 8×8 or larger. That is, according to Figure 25 In an embodiment, the condition related to the block size of the affine MVP mode may match the condition related to the block size of BDOF. Therefore, according to Figure 25 In an embodiment, since the affine MVP mode is not applied to an 8×8 block, PROF may not be applied to an 8×8 block.

[0337] In addition, Figure 25 an embodiment of Figure 24 may be combined with an embodiment of

[0338] That is, both the block size condition of the affine MVP mode and the block size condition of the affine merge mode may match the block size condition of BDOF. Therefore, the block size condition of PROF applicable to the affine block may match the block size condition of BDOF. Figure 24 and Figure 25 In an embodiment, the conditions related to the block size of the affine MVP mode, the affine merge mode, PROF, and BDOF may be changed as shown in the following table.

[0339] [Table 3]

[0340]

[0341] Figure 26 is a view illustrating a process of determining whether to apply PROF according to another embodiment of the present disclosure.

[0342] In BDOF, the characteristics of the optical flow are used to determine the offset of the samples. Therefore, when the luminance values of the reference pictures are different, that is, when BCW or weighted prediction (WP) is applied, BDOF is not performed. However, although the characteristics of the optical flow are used to derive the offset of the samples, PROF may be performed regardless of whether BCW or WP is applied.

[0343] According to Figure 26In an embodiment, from a design perspective, in order to coordinate between BDOF and PROF, PROF may not be applied to blocks that do not apply BCW or WP. For example, when BcwIdx is not 0, or when luma_weight_lX_flag[refIdxLX] (where X is 0 or 1) is 1, cbProfFlagLX may be set to a second value ("false" or "0"). BcwIdx not being 0 may mean that BCW is applied to the current block, and luma_weight_lX_flag[refIdxLX] being 1 may mean that WP in the LX prediction direction is applied to the current block. In the present disclosure, BcwIdx being 0 may mean that equal weights are applied, that is, a bi-predicted block is generated by the average sum of the L0 predicted block and the L1 predicted block. Therefore, when deriving cbProfFlagLX, if BCW or WP is applied to the current block, control may be performed by adding the above conditions not to apply PROF.

[0344] Figure 27 is a view illustrating a process of determining whether to apply PROF according to another embodiment of the present disclosure.

[0345] According to Figure 27 In an embodiment, the PROF application condition may further include a condition related to the resolutions of the current picture and the reference picture. PROF is a refinement method for predicted samples that considers optical flow similar to BDOF. Optical flow is a technique that reflects motion offset when a moving object has the same pixel value and moves bidirectionally constantly. Therefore, when the resolutions of the current picture and the reference picture are different, it is necessary to limit not to perform PROF.

[0346] As Figure 27 shown, when the width pic_width_in_luma_samples of the reference picture is different from that of the current picture, or the height pic_height_in_luma_samples of the reference picture is different from that of the current picture, it may be controlled not to apply PROF to the current block by setting cbProfFlag to a second value ("false" or "0").

[0347] In this case, the reference picture may be the reference picture in the prediction direction of cbProfFlag. Specifically, when deriving cbProfFlagL0, the sizes of the L0 reference picture and the current picture may be considered. When the width or height of the L0 reference picture is different from the width or height of the current picture, cbProfFlagL0 may be set to a second value and PROF for the L0 predicted samples may not be performed. Additionally, when the width and height of the L0 reference picture are equal to the width or height of the current picture, cbProfFlagL0 may be set to a first value and PROF is applied to the L0 predicted samples, thereby generating refined L0 predicted samples.

[0348] Similarly, when deriving cbProfFlagL1, the size of the L1 reference picture and the size of the current picture can be considered. When the width or height of the L1 reference picture is different from the width or height of the current picture, cbProfFlagL1 can be set to a second value and PROF of the L1 prediction samples may not be performed. Additionally, when the width and height of the L1 reference picture are the same as the width and height of the current picture, cbProfFlagL1 can be set to a first value and PROF is applied to the L1 prediction samples, thereby generating refined L1 prediction samples.

[0349] Figure 27 The underlined condition can mean a reference picture resampling (RPR) condition. When the size of the reference picture and the size of the current picture are different from each other, the RPR condition can have a first value ("true" or "1"). The RPR condition with the first value can mean that resampling of the reference picture is required. Additionally, when the size of the reference picture and the size of the current picture are the same, the RPR condition can have a second value ("false" or "0"). The RPR condition with the second value can mean that resampling of the reference picture is not required. That is, when the RPR is the first value, PROF may not be applied.

[0350] Figure 28 is a view illustrating a method of performing PROF according to the present disclosure.

[0351] Figure 28 The method can be performed by the inter prediction unit 180 of the image encoding device or the inter prediction unit 260 of the image decoding device. More specifically, Figure 28 The method can be performed by the prediction sample derivation unit 183 in the inter prediction unit 180 of the image encoding device or the prediction sample derivation unit 263 in the inter prediction unit 260 of the image decoding device.

[0352] According to Figure 28 , the motion information of the current block can be determined (S2810). The motion information of the current block can be determined based on various methods described in the present disclosure. The image encoding device can determine the optimal motion information as the motion information of the current block by calculating the rate distortion (RD) cost based on various inter prediction modes and motion information. The image encoding device can encode the determined inter prediction mode and motion information in the bitstream. The image decoding device can determine (derive) the motion information of the current block by decoding the information signaled through the bitstream.

[0353] Based on the motion information of the current block determined in step S2810, the prediction samples (prediction block) of the current block can be derived (S2820). The prediction samples of the current block can be derived based on various methods described in the present disclosure.

[0354] In step S2830, the reference picture resampling (RPR) condition for the current block can be derived. For example, when the width or height of the reference picture of the current block is different from the width or height of the current picture, the RPR condition can be set to a first value ("true" or "1"). Additionally, when the width and height of the reference picture of the current block are respectively equal to the width or height of the current picture, the RPR condition can be set to a second value ("false" or "0").

[0355] Information cbProfFlag indicating whether a specified PROF is applied to the current block can be derived based on the RPR condition (S2840). For example, when the RPR condition is the first value, cbProfFlag can be set to the second value. That is, when the size of the current picture is different from the size of the reference picture, it can be determined that PROF is not applied. Additionally, when the RPR condition is the second value, cbProfFlag can be set to the first value. That is, when the size of the current picture is equal to the size of the reference picture, it can be determined that PROF is applied. Although step S2840 is described as deriving cbProfFlag based on the RPR condition, this is for convenience of description, and the conditions for deriving cbProfFlag are not limited to the RPR condition. That is, to derive cbProfFlag, other conditions described in this disclosure or other conditions not described in this disclosure can be considered in addition to the RPR condition.

[0356] Based on cbProfFlag derived in step S2840, it can be determined whether to execute PROF (S2850). When cbProfFlag is the first value ("true" or "1"), PROF can be executed for the predicted samples of the current block (S2860). When cbProfFlag is the second value ("false" or "0"), PROF can be skipped without being executed for the predicted samples of the current block.

[0357] The PROF process in step S2860 can be executed according to the PROF process described in this disclosure. More specifically, when PROF is applied to the current block, the differential motion vectors for each sample position in the current block can be derived, the gradients for each sample position in the current block can be derived, the PROF offset can be derived based on the differential motion vectors and gradients, and then the refined predicted samples for the current block can be derived based on the PROF offset.

[0358] The image encoding device can derive the residual samples (residual block) of the current block based on the refined predicted samples (predicted block) and encode the information about the residual samples in the bitstream. The image decoding device can reconstruct the current block based on the refined predicted samples (predicted block) and residual samples (residual block) obtained by decoding the bitstream.

[0359] In Figure 28In the example shown, the RPR condition in step S2830 is not limited to being executed after step S2820. For example, it is sufficient to derive the RPR condition before deriving cbProfFlag (S2840), and embodiments of the present disclosure may include various examples of deriving the RPR condition before executing step S2840.

[0360] Figure 29 is a view illustrating a process of determining whether to apply PROF according to another embodiment of the present disclosure.

[0361] Figure 29 The embodiment of Figure 26 The embodiment of Figure 27 is an example of an embodiment that combines the embodiments of

[0362] According to Figure 29 , when applying WP in the L0 direction (e.g., luma_weight_l0_flag == 1) or applying WP in the L1 direction (e.g., luma_weight_l1_flag == 1), cbProfFlag may be set to not apply PROF. Additionally, when the size of the reference picture in the L0 direction is different from the size of the current picture or the size of the reference picture in the L1 direction is different from the size of the current picture, cbProfFlag may be set to not apply PROF.

[0363] Figure 30 is a view illustrating a process of determining whether to apply PROF according to another embodiment of the present disclosure.

[0364] Figure 30 The embodiment of Figure 26 The embodiment of Figure 27 is another example of an embodiment that combines the embodiments of

[0365] According to Figure 30, when applying WP in the L0 direction (e.g., luma_weight_l0_flag == 1) or when the size of the reference picture in the L0 direction is different from the size of the current picture, cbProfFlagL0 can be set to a second value ("false" or "0") to not apply PROF in the L0 direction. Additionally, when applying WP in the L1 direction (e.g., luma_weight_l1_flag == 1) or when the size of the reference picture in the L1 direction is different from the size of the current picture, cbProfFlagL1 can be set to a second value ("false" or "0") to not apply PROF in the L1 direction.

[0366] The various embodiments described in this disclosure can be implemented alone or in combination with other embodiments. Alternatively, some of one embodiment can be added to another embodiment, or some of one embodiment can be replaced with some of another embodiment.

[0367] According to the various embodiments described in this disclosure, by matching some of the application conditions of PROF and the application conditions of BDOF, from a design perspective, coordination between PROF and BDOF can be expected, and the implementation complexity can be further reduced.

[0368] Although for clarity of description, the exemplary methods of the present disclosure above are shown as a series of operations, it is not intended to limit the order of execution of the steps, and these steps can be performed simultaneously or in a different order when necessary. To implement the method according to the present invention, the described steps can further include other steps, can include the remaining steps except some steps, or can include other additional steps except some steps.

[0369] In this disclosure, an image encoding device or an image decoding device that performs a predetermined operation (step) can perform an operation (step) of confirming the execution conditions or circumstances of the corresponding operation (step). For example, if it is described that a predetermined operation is performed when a predetermined condition is satisfied, the image encoding device or the image decoding device can perform the predetermined operation after determining whether the predetermined condition is satisfied.

[0370] The various embodiments of this disclosure are not a list of all possible combinations and are intended to describe representative aspects of this disclosure, and the matters described in the various embodiments can be applied independently or in combinations of two or more.

[0371] The various embodiments of the present disclosure can be implemented in hardware, firmware, software, or a combination thereof. In the case where the present disclosure is implemented by hardware, the present disclosure can be implemented by an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a general purpose processor, a controller, a microcontroller, a microprocessor, etc.

[0372] In addition, an image decoding device and an image encoding device to which the embodiments of the present disclosure are applied can be included in a multimedia broadcast transmission and reception device, a mobile communication terminal, a home theater video device, a digital cinema video device, a surveillance camera, a video chat device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camera, a video on demand (VoD) service providing device, an over the top video (OTT video) device, an Internet streaming service providing device, a three-dimensional (3D) video device, a video phone video device, a medical video device, etc., and can be used to process video signals or data signals. For example, an OTT video device can include a game console, a Blu-ray player, an Internet access TV, a home theater system, a smart phone, a tablet PC, a digital video recorder (DVR), etc.

[0373] Figure 31 is a view showing a content stream system to which the embodiments of the present disclosure can be applied.

[0374] As Figure 31 shown, a content stream system to which the embodiments of the present disclosure are applied can mainly include an encoding server, a streaming server, a network server, a media storage device, a user device, and a multimedia input device.

[0375] The encoding server compresses content input from a multimedia input device such as a smart phone, a camera, a video camera, etc. into digital data to generate a bitstream and sends the bitstream to the streaming server. As another example, when a multimedia input device such as a smart phone, a camera, a video camera, etc. directly generates a bitstream, the encoding server can be omitted.

[0376] The bitstream can be generated by an image encoding method or an image encoding device to which the embodiments of the present disclosure are applied, and the streaming server can temporarily store the bitstream during the process of sending or receiving the bitstream.

[0377] The streaming server sends multimedia data to the user device based on a request from the user via the web server, and the web server serves as a medium for informing the user of the service. When the user requests a desired service from the web server, the web server can deliver it to the streaming server, and the streaming server can send multimedia data to the user. In this case, the content streaming system may include a separate control server. In this case, the control server is used to control commands / responses between devices in the content streaming system.

[0378] The streaming server can receive content from a media storage device and / or an encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a predetermined time.

[0379] Examples of user devices may include mobile phones, smartphones, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head-mounted displays), digital TVs, desktop computers, digital signage, etc.

[0380] Each server in the content streaming system can operate as a distributed server, and in this case, the data received from each server can be distributed.

[0381] The scope of the present disclosure includes software or machine-executable commands (e.g., operating systems, applications, firmware, programs, etc.) for enabling the operation of methods according to various embodiments to be executed on a device or computer, and non-transitory computer-readable media having such software or commands stored thereon and executable on the device or computer.

[0382] Industrial Applicability

[0383] Embodiments of the present disclosure can be used to encode or decode images.

Claims

1. An image decoding method performed by an image decoding device, the image decoding method comprising the following steps: Deriving prediction samples of the current block based on motion information of the current block; Deriving reference picture resampling (RPR) conditions of the current block; Determining whether optical flow prediction refinement (PROF) is applied to the current block based on the RPR conditions; And Deriving refined prediction samples of the current block by applying PROF to the current block.

2. The image decoding method according to claim 1, wherein, Determining the RPR conditions based on the size of the reference picture of the current block and the size of the current picture.

3. The image decoding method according to claim 2, Among them, Based on the size of the reference picture of the current block being different from the size of the current picture, the RPR conditions are derived as a first value, and wherein, based on the size of the reference picture of the current block being equal to the size of the current picture, the RPR conditions are derived as a second value.

4. The image decoding method according to claim 3, wherein, Based on the RPR conditions being the first value, determining that PROF is not applied to the current block.

5. The image decoding method according to claim 1, wherein, Determining whether PROF is applied to the current block based on the size of the current block.

6. The image decoding method according to claim 5, wherein, Based on the product of the width w and the height h of the current block being less than 128, determining that PROF is not applied to the current block.

7. The image decoding method according to claim 1, wherein, Parsing information specifying whether the current block is in an affine merge mode from a bitstream based on the size of the current block.

8. The image decoding method according to claim 7, wherein, Based on each of the width w and the height h of the current block being equal to or greater than 8 and w*h being equal to or greater than 128, parsing information specifying whether the current block is in the affine merge mode from the bitstream.

9. The image decoding method according to claim 1, wherein, Parsing information specifying whether the current block is in an affine MVP mode from a bitstream based on the size of the current block.

10. The image decoding method according to claim 9, wherein, Based on each of the width w and the height h of the current block being equal to or greater than 8 and w*h being equal to or greater than 128, parsing information specifying whether the current block is in the affine MVP mode from the bitstream.

11. The image decoding method according to claim 1, wherein, Determining whether PROF is applied to the current block based on whether BCW or WP is applied to the current block.

12. The image decoding method according to claim 11, wherein, Based on BCW or WP being applied to the current block, determining that PROF is not applied to the current block.

13. An image encoding method performed by an image encoding device, the image encoding method comprising the following steps: Deriving prediction samples of the current block based on motion information of the current block; Deriving reference picture resampling (RPR) conditions of the current block; Determining whether optical flow prediction refinement (PROF) is applied to the current block based on the RPR conditions; And Deriving refined prediction samples of the current block by applying PROF to the current block.

14. A method of transmitting a bitstream generated by an image encoding method, the image encoding method comprising the following steps: Deriving prediction samples of the current block based on motion information of the current block; Deriving reference picture resampling (RPR) conditions of the current block; Determining whether optical flow prediction refinement (PROF) is applied to the current block based on the RPR conditions; And Deriving refined prediction samples of the current block by applying PROF to the current block.

15. A non-transitory computer-readable recording medium having stored thereon a computer program which, when executed by a processor, implements the image encoding method according to claim 13.

Citation Information

Patent Citations

  • Image encoding / decoding method and device for same

    CN108141595A