Image decoding method, image encoding method, and bit stream transmission method
Through weighted prediction technology and dual prediction (BCW) of CU-level weights, the efficiency of image encoding/decoding is improved, and the inefficiency problem in high-resolution and high-quality image transmission and storage is solved, thus achieving cost reduction.
Patent Information
- Application Number
- CN202510693682.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-08-22
- Filing Date
- 2020-08-20
- Publication Date
- 2025-07-11
AI Technical Summary
The prior art is inefficient in the transmission and storage of high resolution and high quality images, resulting in increased costs and requires improved image encoding/decoding methods to improve encoding/decoding efficiency.
Weighted prediction technology, including double prediction (BCW) of CU-level weights, generate prediction blocks of image blocks by deriving weighted prediction flags and weight indexes, and determine the default or explicit weighted prediction method.
Improves image encoding/decoding efficiency, reduces transmission and storage costs, and supports efficient transmission and storage of bitstreams of high-resolution, high-quality images.
Smart Images

Figure CN120302048A_ABST
Abstract
Description
[0001] This application is a divisional application of the patent application with the original application number 202080068734.7 (International Application No.: PCT / KR2020 / 011103, filing date: August 20, 2020, invention title: Image encoding / decoding method and apparatus for performing weighted prediction and method of transmitting a bitstream). Technical Field
[0002] The present disclosure relates to an image encoding / decoding method and apparatus and a method of transmitting a bitstream, and more particularly, to an image encoding / decoding method and apparatus for performing weighted prediction in consideration of optical flow prediction refinement (PROF) and a method of transmitting a bitstream generated by the image encoding method / apparatus of the present disclosure. Background Art
[0003] Recently, the demand for high-resolution and high-quality images, such as high-definition (HD) images and ultra-high-definition (UHD) images, is increasing in various fields. As the resolution and quality of image data are improved, the amount of information or bits to be transmitted increases relatively compared to existing image data. The increase in the amount of information or bits to be transmitted leads to an increase in transmission costs and storage costs.
[0004] Therefore, an efficient image compression technique is needed to effectively transmit, store, and reproduce information on high-resolution and high-quality images. Summary of the Invention
[0005] Technical Problem
[0006] An object of the present disclosure is to provide an image encoding / decoding method and apparatus having improved encoding / decoding efficiency.
[0007] Another object of the present disclosure is to provide an image encoding / decoding method and apparatus for performing weighted prediction.
[0008] Another object of the present disclosure is to provide an image encoding / decoding method and apparatus for performing weighted prediction or bi-prediction with CU-level weights (BCW) in consideration of PROF.
[0009] Another object of the present disclosure is to provide a method of transmitting a bitstream generated by an image encoding method or apparatus according to the present disclosure.
[0010] Another object of the present disclosure is to provide a recording medium storing a bitstream generated by an image encoding method or apparatus according to the present disclosure.
[0011] Another object of the present disclosure is to provide a recording medium storing a bitstream received, decoded, and used for reconstructing an image by an image decoding apparatus according to the present disclosure.
[0012] The technical problems solved by the present disclosure are not limited to the above technical problems, and those skilled in the art will be clear about other technical problems not described herein through the following description.
[0013] Technical solution
[0014] An image decoding method according to an aspect of the present disclosure may include: deriving a first flag specifying whether to perform weighted prediction on a current block and a weight index BcwIdx for bi-prediction with CU-level weights (BCW) of the current block; determining whether to perform default weighted prediction or explicit weighted prediction on the current block based on the first flag and BcwIdx; and generating a predicted block of the current block by performing the determined method.
[0015] In the image decoding method according to the present disclosure, the first flag may be determined differently based on the slice type of the current slice to which the current block belongs.
[0016] In the image decoding method according to the present disclosure, based on the slice type of the current slice being a P slice, the first flag may be derived as the value of pps_weighted_pred_flag signaled in the picture parameter set (PPS), and based on the slice type of the current slice being a B slice, the first flag may be derived as the value of pps_weighted_bipred_flag signaled in the PPS.
[0017] In the image decoding method according to the present disclosure, BcwIdx may be derived based on the syntax element bcw_idx signaled through the bitstream, and based on the absence of bcw_idx in the bitstream, BcwIdx may be derived as 0.
[0018] In the image decoding method according to the present disclosure, bcw_idx may be parsed from the bitstream based on the weighted prediction flag of the reference picture of the current block.
[0019] In the image decoding method according to the present disclosure, based on all weighted prediction flags of the reference picture of the current block being 0, bcw_idx may be parsed from the bitstream.
[0020] In the image decoding method according to the present disclosure, based on the first flag being 0 or BcwIdx not being 0, default weighted prediction may be performed on the current block.
[0021] In the image decoding method according to the present disclosure, based on the first flag being 1 or BcwIdx being 0, explicit weighted prediction may be performed on the current block.
[0022] In the image decoding method according to the present disclosure, default weighted prediction may perform BCW or an average sum based on BcwIdx.
[0023] In the image decoding method according to the present disclosure, based on BcwIdx being 0, an average sum can be performed on the current block, and based on BcwIdx not being 0, BCW can be performed on the current block.
[0024] In the image decoding method according to the present disclosure, explicit weighted prediction can be performed based on the weighted parameters (weight sum and offset) of the reference picture of the current block.
[0025] In the image decoding method according to the present disclosure, the weighted parameters can be explicitly signaled via a bitstream.
[0026] An image decoding apparatus according to another aspect of the present disclosure may include a memory and at least one processor. The at least one processor may derive a first flag specifying whether to perform weighted prediction on the current block and a weight index BcwIdx for bi-prediction (BCW) with CU-level weights for the current block; determine whether to perform default weighted prediction or explicit weighted prediction on the current block based on the first flag and BcwIdx; and generate a predicted block of the current block by performing the determined method.
[0027] An image encoding method according to another aspect of the present disclosure may include: determining a first flag specifying whether to perform weighted prediction on the current block and a weight index BcwIdx for bi-prediction (BCW) with CU-level weights for the current block; determining whether to perform default weighted prediction or explicit weighted prediction on the current block based on the first flag and BcwIdx; and generating a predicted block of the current block by performing the determined method.
[0028] Furthermore, a computer-readable recording medium according to another aspect of the present disclosure may store a bitstream generated by the image encoding apparatus or image encoding method of the present disclosure.
[0029] The features of the above brief overview of the present disclosure are merely exemplary aspects of the following detailed description of the present disclosure and do not limit the scope of the present disclosure.
[0030] Advantageous Effects
[0031] According to the present disclosure, it is possible to provide an image encoding / decoding method and apparatus having improved encoding / decoding efficiency.
[0032] Furthermore, according to the present disclosure, it is possible to provide an image encoding / decoding method and apparatus for performing weighted prediction.
[0033] Furthermore, according to the present disclosure, it is possible to provide an image encoding / decoding method and apparatus for considering PROF to perform weighted prediction or bi-prediction with CU-level weights (BCW).
[0034] In addition, according to the present disclosure, a method of transmitting a bitstream generated by an image encoding method or apparatus according to the present disclosure can be provided.
[0035] In addition, according to the present disclosure, a recording medium storing a bitstream generated by an image encoding method or apparatus according to the present disclosure can be provided.
[0036] In addition, according to the present disclosure, a recording medium storing a bitstream received, decoded, and used for reconstructing an image by an image decoding apparatus according to the present disclosure can be provided.
[0037] Those skilled in the art will understand that the effects that can be achieved by the present disclosure are not limited to those specifically described above, and other advantages of the present disclosure will be more clearly understood from the detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 is a view schematically illustrating a video encoding system to which an embodiment of the present disclosure is applicable.
[0039] Figure 2 is a view schematically illustrating an image encoding apparatus to which an embodiment of the present disclosure is applicable.
[0040] Figure 3 is a view schematically illustrating an image decoding apparatus to which an embodiment of the present disclosure is applicable.
[0041] Figure 4 is a flowchart illustrating a video / image encoding method based on inter prediction.
[0042] Figure 5 is a view illustrating the configuration of an inter prediction unit 180 according to the present disclosure.
[0043] Figure 6 is a flowchart illustrating a video / image decoding method based on inter prediction.
[0044] Figure 7 is a view illustrating the configuration of an inter prediction unit 260 according to the present disclosure.
[0045] Figure 8 is a view illustrating neighboring blocks that can be used as spatial merge candidates.
[0046] Figure 9 is a view schematically illustrating a method of constructing a merge candidate list according to an example of the present disclosure.
[0047] Figure 10 is a view illustrating candidate pairs for redundancy checking performed on spatial candidates.
[0048] Figure 11A view of a method for exemplifying scaling time candidates of motion vectors.
[0049] Figure 12 A view of a method for exemplifying deriving positions of temporal candidates.
[0050] Figure 13 A view schematically exemplifying a method for configuring a list of motion vector prediction candidates according to an example of the present disclosure.
[0051] Figure 14 A view of a parametric model of an affine mode.
[0052] Figure 15 A view of a method for exemplifying generating an affine merge candidate list.
[0053] Figure 16 A view of a CPMV derived from neighboring blocks.
[0054] Figure 17 A view of a neighboring block for deriving an affine merge candidate.
[0055] Figure 18 A view of a method for exemplifying generating an affine MVP candidate list.
[0056] Figure 19 A view of a neighboring block of a sub-block-based TMVP mode.
[0057] Figure 20 A view of a method for exemplifying deriving a motion vector field according to a sub-block-based TMVP mode.
[0058] Figure 21 A view of a CU extended to perform BDOF.
[0059] Figure 22 A view of the relationship between Δv(i,j), v(i,j), and a sub-block motion vector.
[0060] Figure 23 A flowchart exemplifying an example of performing PROF, BCW, WP, and / or average sum according to the present disclosure.
[0061] Figure 24 A flowchart exemplifying another example of performing PROF, BCW, WP, and / or average sum according to the present disclosure.
[0062] Figure 25 A flowchart exemplifying an example of performing BCW or WP according to the method of Table 7.
[0063] Figure 26 A flowchart exemplifying an example of performing BCW or WP according to the method of Table 8.
[0064] Figure 27 is a flowchart exemplifying an example of performing BCW or WP according to the method of Table 9.
[0065] Figure 28 is a flowchart exemplifying an example of performing BCW or WP according to the method of Table 10.
[0066] Figure 29 is a flowchart exemplifying an example of performing BDOF, PROF, BCW, WP, and / or average sum according to the present disclosure.
[0067] Figure 30 is a view showing a content stream system to which an embodiment of the present disclosure is applicable. Detailed Description of the Invention
[0068] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings so that those skilled in the art can easily implement them. However, the present disclosure can be implemented in various different forms and is not limited to the embodiments described herein.
[0069] When describing the present disclosure, if it is determined that a detailed description of related known functions or configurations makes the scope of the present disclosure unnecessarily ambiguous, the detailed description thereof will be omitted. In the drawings, parts irrelevant to the description of the present disclosure are omitted, and similar reference numerals are given to similar parts.
[0070] In the present disclosure, when a component is "connected", "coupled", or "linked" to another component, it may include not only a direct connection relationship but also an indirect connection relationship with an intermediate component present. Additionally, when a component "includes" or "has" other components, unless otherwise specified, it means that other components may also be included, rather than excluding other components.
[0071] In the present disclosure, terms such as first, second, etc. are used only for the purpose of distinguishing one component from other components and do not limit the order or importance of the components, unless otherwise specified. Accordingly, within the scope of the present disclosure, a first component in one embodiment may be referred to as a second component in another embodiment, and similarly, a second component in one embodiment may be referred to as a first component in another embodiment.
[0072] In the present disclosure, the components distinguished from each other are intended to clearly describe each feature and do not mean that the components must be separate. That is, multiple components may be integrated and implemented in one hardware or software unit, or one component may be distributed and implemented in multiple hardware or software units. Therefore, even without specific description, these embodiments of integration or distribution of components are included within the scope of the present disclosure.
[0073] In the present disclosure, the components described in each embodiment are not necessarily essential components, and some components may be optional components. Therefore, an embodiment composed of a subset of the components described in the embodiment is also included within the scope of the present disclosure. In addition, an embodiment including other components in addition to the components described in various embodiments is included within the scope of the present disclosure.
[0074] The present disclosure relates to the encoding and decoding of images. Unless redefined in the present disclosure, the terms used in the present disclosure may have the general meanings commonly used in the technical field to which the present disclosure pertains.
[0075] In the present disclosure, a "picture" generally refers to a unit representing an image within a specific time period, and a slice / tile is an encoding unit that forms a part of the picture. A picture can be composed of one or more slices / tiles. In addition, a slice / tile can include one or more coding tree units (CTUs).
[0076] In the present disclosure, a "pixel" or "pel" can mean the smallest unit that constitutes a picture (or image). In addition, "sample" can be used as a term corresponding to a pixel. A sample generally can represent a pixel or the value of a pixel, or can represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component.
[0077] In the present disclosure, a "unit" can represent a basic unit of image processing. The unit can include at least one of a specific area of a picture and information related to the area. In some cases, the unit can be used interchangeably with terms such as "sample array", "block", or "area". Generally, an M×N block can include a set (or array) of samples (or sample arrays) or transform coefficients having M columns and N rows.
[0078] In the present disclosure, a "current block" can mean one of a "current encoding block", a "current encoding unit", a "block to be encoded", a "block to be decoded", or a "block to be processed". When performing prediction, a "current block" can mean a "current prediction block" or a "block to be predicted". When performing transformation (inverse transformation) / quantization (dequantization), a "current block" can mean a "current transformation block" or a "block to be transformed". When performing filtering, a "current block" can mean a "block to be filtered".
[0079] In the present disclosure, the term " / " or "," can be interpreted as indicating "and / or". For example, "A / B" and "A, B" can mean "A and / or B". In addition, "A / B / C" and "A / B / C" can mean "at least one of A, B, and / or C".
[0080] In the present disclosure, the term "or" shall be construed to indicate "and / or". For example, the expression "A or B" may include 1) only "A", 2) only "B", or 3) both "A and B". In other words, in the present disclosure, "or" shall be construed to indicate "additionally or alternatively".
[0081] Overview of Video Encoding System
[0082] Figure 1 is a view schematically showing a video coding system according to the present disclosure.
[0083] A video coding system according to an embodiment may include an encoding device 10 and a decoding device 20. The encoding device 10 may deliver encoded video and / or image information or data in the form of a file or a stream to the decoding device 20 via a digital storage medium or a network.
[0084] The encoding device 10 according to an embodiment may include a video source generator 11, an encoding unit 12, and a transmitter 13. The decoding device 20 according to an embodiment may include a receiver 21, a decoding unit 22, and a renderer 23. The encoding unit 12 may be referred to as a video / image encoding unit, and the decoding unit 22 may be referred to as a video / image decoding unit. The transmitter 13 may be included in the encoding unit 12. The receiver 21 may be included in the decoding unit 22. The renderer 23 may include a display, and the display may be configured as a separate device or an external component.
[0085] The video source generator 11 may obtain video / images through a process of capturing, synthesizing, or generating video / images. The video source generator 11 may include a video / image capturing device and / or a video / image generating device. The video / image capturing device may include, for example, one or more cameras, a video / image archive including previously captured video / images, etc. The video / image generating device may include, for example, a computer, a tablet computer, and a smart phone, and may (electronically) generate video / images. For example, virtual video / images may be generated by a computer or the like. In this case, the video / image capturing process may be replaced by a process of generating related data.
[0086] The encoding unit 12 may encode the input video / images. For compression and encoding efficiency, the encoding unit 12 may perform a series of processes such as prediction, transformation, and quantization. The encoding unit 12 may output the encoded data (encoded video / image information) in the form of a bit stream.
[0087] The transmitter 13 can transmit the encoded video / image information or data output in the form of a bitstream to the receiver 21 of the decoding device 20 in the form of a file or a stream via a digital storage medium or a network. The digital storage medium can include various storage media, such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter 13 can include elements for generating a media file in a predetermined file format and can include elements for transmitting via a broadcast / communication network. The receiver 21 can extract / receive the bitstream from the storage medium or the network and transmit the bitstream to the decoding unit 22.
[0088] The decoding unit 22 can decode the video / image by performing a series of processes corresponding to the operations of the encoding unit 12, such as dequantization, inverse transformation, and prediction.
[0089] The renderer 23 can render the decoded video / image. The rendered video / image can be displayed via a display.
[0090] Overview of Image Encoding Device
[0091] Figure 2 is a view schematically showing an image encoding device to which an embodiment of the present disclosure can be applied.
[0092] As Figure 2 shown, the image encoding device 100 can include an image splitter 110, a subtractor 115, a transformer 120, a quantizer 130, a dequantizer 140, an inverse transformer 150, an adder 155, a filter 160, a memory 170, an inter-frame prediction unit 180, an intra-frame prediction unit 185, and an entropy encoder 190. The inter-frame prediction unit 180 and the intra-frame prediction unit 185 can be collectively referred to as a "prediction unit". The transformer 120, the quantizer 130, the dequantizer 140, and the inverse transformer 150 can be included in a residual processor. The residual processor can also include the subtractor 115.
[0093] In some embodiments, all or at least some of the multiple components configuring the image encoding device 100 can be configured by one hardware component (e.g., an encoder or a processor). In addition, the memory 170 can include a decoded picture buffer (DPB) and can be configured by a digital storage medium.
[0094] The image splitter 110 may split an input image (or picture or frame) input to the image encoding device 100 into one or more processing units. For example, the processing unit may be referred to as a coding unit (CU). Coding units may be obtained by recursively splitting a coding tree unit (CTU) or a largest coding unit (LCU) according to a quadtree binary tree ternary tree (QT / BT / TT) structure. For example, a coding unit may be split into multiple coding units of a deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. For the splitting of the coding unit, the quadtree structure may be applied first, and then the binary tree structure and / or the ternary tree structure may be applied. The encoding process according to the present disclosure may be performed based on the final coding unit that is no longer split. The largest coding unit may be used as the final coding unit, or the coding units of a deeper depth obtained by splitting the largest coding unit may be used as the final coding unit. Here, the encoding process may include processes of prediction, transformation, and reconstruction that will be described later. As another example, the processing unit of the encoding process may be a prediction unit (PU) or a transformation unit (TU). The prediction unit and the transformation unit may be divided or split from the final coding unit. The prediction unit may be a sample prediction unit, and the transformation unit may be a unit for deriving transformation coefficients and / or a unit for deriving a residual signal from the transformation coefficients.
[0095] The prediction unit (inter-frame prediction unit 180 or intra-frame prediction unit 185) may perform prediction on a block to be processed (current block) and generate a prediction block including prediction samples of the current block. The prediction unit may determine whether to apply intra-frame prediction or inter-frame prediction based on the current block or CU. The prediction unit may generate various information related to the prediction of the current block and transmit the generated information to the entropy encoder 190. The information about the prediction may be encoded in the entropy encoder 190 and output in the form of a bitstream.
[0096] The intra-frame prediction unit 185 may predict the current block by referring to samples in the current picture. According to the intra-frame prediction mode and / or intra-frame prediction technique, the reference samples may be located in the neighbors of the current block or may be placed separately. The intra-frame prediction mode may include multiple non-directional modes and multiple directional modes. The non-directional modes may include, for example, the DC mode and the planar mode. According to the degree of detail of the prediction direction, the directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes. However, this is only an example, and more or fewer directional prediction modes may be used according to the setting. The intra-frame prediction unit 185 may determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.
[0097] The inter - frame prediction unit 180 may derive a prediction block of a current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, in order to reduce the amount of motion information transmitted in the inter - frame prediction mode, the motion information may be predicted in units of blocks, sub - blocks, or samples based on the correlation of the motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may also include inter - frame prediction direction (L0 prediction, L1 prediction, bi - prediction, etc.) information. In the case of inter - frame prediction, neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block may be referred to as a collocated reference block, a collocated CU (colCU), etc. The reference picture including the temporal neighboring block may be referred to as a collocated picture (colPic). For example, the inter - frame prediction unit 180 may configure a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter - frame prediction may be performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the inter - frame prediction unit 180 may use the motion information of neighboring blocks as the motion information of the current block. In the case of the skip mode, different from the merge mode, the residual signal may not be transmitted. In the case of the motion vector prediction (MVP) mode, the motion vector of a neighboring block may be used as a motion vector predictor, and the motion vector of the current block may be signaled by encoding the motion vector difference and an indicator of the motion vector predictor. The motion vector difference may mean the difference between the motion vector of the current block and the motion vector predictor.
[0098] The prediction unit may generate a prediction signal based on various prediction methods and prediction techniques described below. For example, the prediction unit may not only apply intra - frame prediction or inter - frame prediction, but also apply both intra - frame prediction and inter - frame prediction simultaneously to predict the current block. The prediction method of applying both intra - frame prediction and inter - frame prediction simultaneously to predict the current block may be referred to as combined intra - inter prediction (CIIP). In addition, the prediction unit may perform intra - block copy (IBC) to predict the current block. Intra - block copy may be used for content image / video coding such as games, for example, screen content coding (SCC). IBC is a method of predicting the current picture using a previously reconstructed reference block in the current picture at a position separated from the current block by a predetermined distance. When IBC is applied, the position of the reference block in the current picture may be encoded as a vector (block vector) corresponding to the predetermined distance.
[0099] The prediction signal generated by the prediction unit can be used to generate a reconstructed signal or a residual signal. The subtractor 115 can generate a residual signal (residual block or residual sample array) by subtracting the prediction signal (prediction block or prediction sample array) output from the prediction unit from the input image signal (original block or original sample array). The generated residual signal can be transmitted to the transformer 120.
[0100] The transformer 120 can generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique can include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a karhunen-loève transform (KLT), a graph-based transform (GBT), or a conditional non-linear transform (CNT). Here, the GBT refers to a transform obtained from a graph when the relationship information between pixels is represented by a graph. The CNT refers to a transform obtained based on a prediction signal generated using all previously reconstructed pixels. In addition, the transform processing can be applied to square pixel blocks of the same size or can be applied to blocks of variable size rather than square.
[0101] The quantizer 130 can quantize the transform coefficients and transmit them to the entropy encoder 190. The entropy encoder 190 can encode the quantized signal (information about the quantized transform coefficients) and output a bitstream. The information about the quantized transform coefficients can be referred to as residual information. The quantizer 130 can rearrange the quantized transform coefficients in block form into a one-dimensional vector form based on the coefficient scan order and generate information about the quantized transform coefficients based on the quantized transform coefficients in one-dimensional vector form.
[0102] The entropy encoder 190 can perform various coding methods, such as exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. The entropy encoder 190 can encode the information required for video / image reconstruction other than the quantized transform coefficients (e.g., the values of syntax elements, etc.) together or separately. The encoded information (e.g., encoded video / image information) can be transmitted or stored in the form of a bitstream in units of network abstraction layer (NAL). The video / image information can also include information about various parameter sets, such as an adaptive parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information can also include general constraint information. The information signaled, transmitted, and / or syntax elements described in this disclosure can be encoded through the above coding process and included in the bitstream.
[0103] The bitstream can be transmitted over a network or stored in a digital storage medium. The network can include a broadcast network and / or a communication network, and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) for transmitting the signal output from the entropy encoder 190 and / or a storage unit (not shown) for storing the signal can be included as internal / external components of the image encoding device 100. Alternatively, a transmitter can be provided as a component of the entropy encoder 190.
[0104] The quantized transform coefficients output from the quantizer 130 can be used to generate a residual signal. For example, the residual signal (residual block or residual samples) can be reconstructed by applying dequantization and inverse transformation to the quantized transform coefficients by the dequantizer 140 and the inverse transform unit 150.
[0105] The adder 155 adds the reconstructed residual signal to the prediction signal output from the inter-frame prediction unit 180 or the intra-frame prediction unit 185 to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). If the block to be processed has no residual, for example, in the case of applying the skip mode, the predicted block can be used as the reconstructed block. The adder 155 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current picture and can be used for inter-frame prediction of the next picture by filtering as described below.
[0106] In addition, as described below, luminance mapping and chrominance scaling (LMCS) are applicable to the picture encoding process.
[0107] The filter 160 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 160 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture and store the modified reconstructed picture in the memory 170, specifically, in the DPB of the memory 170. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filtering, bilateral filtering, etc. The filter 160 can generate various information related to the filtering and transmit the generated information to the entropy encoder 190, as described later in the description of each filtering method. The information related to the filtering can be encoded by the entropy encoder 190 and output in the form of a bitstream.
[0108] The modified reconstructed picture transmitted to the memory 170 can be used as a reference picture in the inter-frame prediction unit 180. When inter-frame prediction is applied by the image encoding device 100, prediction mismatch between the image encoding device 100 and the image decoding device can be avoided and the encoding efficiency can be improved.
[0109] The DPB of memory 170 may store the modified reconstructed picture to be used as a reference picture in the inter prediction unit 180. Memory 170 may store the motion information of the blocks from which the motion information in the current picture is derived (or encoded) and / or the motion information of the blocks that have been reconstructed in the picture. The stored motion information may be transmitted to the inter prediction unit 180 and used as the motion information of spatially neighboring blocks or temporally neighboring blocks. Memory 170 may store the reconstructed samples of the reconstructed blocks in the current picture and may transmit the reconstructed samples to the intra prediction unit 185.
[0110] Overview of Image Decoding Device
[0111] Figure 3 is a view schematically showing an image decoding apparatus to which an embodiment of the present disclosure is applicable.
[0112] As Figure 3 shown, the image decoding apparatus 200 may include an entropy decoder 210, a dequantizer 220, an inverse transformer 230, an adder 235, a filter 240, a memory 250, an inter prediction unit 260, and an intra prediction unit 265. The inter prediction unit 260 and the intra prediction unit 265 may be collectively referred to as a "prediction unit". The dequantizer 220 and the inverse transformer 230 may be included in a residual processor.
[0113] According to an embodiment, all or at least some of the plurality of components configuring the image decoding apparatus 200 may be configured by hardware components (e.g., a decoder or a processor). In addition, the memory 250 may include a decoded picture buffer (DPB) or may be configured by a digital storage medium.
[0114] The image decoding apparatus 200 that has received a bitstream including video / image information may reconstruct an image by performing a process corresponding to the process performed by the Figure 2 image encoding apparatus 100. For example, the image decoding apparatus 200 may perform decoding using the processing units applied in the image encoding apparatus. Therefore, the decoding processing units may be, for example, encoding units. The encoding units may be obtained by dividing a coding tree unit or a largest coding unit. The reconstructed image signal decoded and output by the image decoding apparatus 200 may be reproduced by a reproduction apparatus (not shown).
[0115] The image decoding apparatus 200 may receive, in the form of a bitstream, from Figure 2The signal output by the image encoding device. The received signal can be decoded by the entropy decoder 210. For example, the entropy decoder 210 can parse the bitstream to derive the information (e.g., video / image information) required for image reconstruction (or picture reconstruction). The video / image information can also include information about various parameter sets, such as the Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). In addition, the video / image information can also include general constraint information. The image decoding device can also decode the picture based on the information about the parameter sets and / or the general constraint information. The information and / or syntax elements signaled / received described in this disclosure can be decoded through the decoding process and obtained from the bitstream. For example, the entropy decoder 210 decodes the information in the bitstream based on an encoding method such as Exponential Golomb coding, CAVLC, or CABAC, and outputs the values of the syntax elements required for image reconstruction and the quantization values of the transform coefficients of the residuals. More specifically, the CABAC entropy decoding method can receive the bins corresponding to each syntax element in the bitstream, determine the context model using the decoding target syntax element information, neighboring blocks, and the decoding information of the decoding target block or the information of the symbols / bins decoded in the previous stage, perform arithmetic decoding on the bins by predicting the occurrence probability of the bins according to the determined context model, and generate the symbols corresponding to the values of each syntax element. In this case, after determining the context model, the CABAC entropy decoding method can update the context model by using the information of the decoded symbols / bins for the context model of the next symbol / bin. The information related to prediction in the information decoded by the entropy decoder 210 can be provided to the prediction units (inter-frame prediction unit 260 and intra-frame prediction unit 265), and the residual values for which entropy decoding is performed in the entropy decoder 210, that is, the quantized transform coefficients and the related parameter information, can be input to the dequantizer 220. In addition, the information about filtering among the information decoded by the entropy decoder 210 can be provided to the filter 240. Furthermore, the receiver (not shown) for receiving the signal output by the image encoding device can be further configured as an internal / external element of the image decoding device 200, or the receiver can be a component of the entropy decoder 210.
[0116] In addition, the image decoding device according to this disclosure can be referred to as a video / image / picture decoding device. The image decoding device can be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder can include the entropy decoder 210. The sample decoder can include at least one of the dequantizer 220, inverse transform unit 230, adder 235, filter 240, memory 250, inter-frame prediction unit 260, or intra-frame prediction unit 265.
[0117] The dequantizer 220 may dequantize the quantized transform coefficients and output the transform coefficients. The dequantizer 220 may rearrange the quantized transform coefficients in the form of two-dimensional blocks. In this case, the rearrangement may be performed based on the coefficient scan order executed in the image coding device. The dequantizer 220 may dequantize the quantized transform coefficients by using quantization parameters (e.g., quantization step information) and obtain the transform coefficients.
[0118] The inverse transformer 230 may perform an inverse transform on the transform coefficients to obtain a residual signal (residual block, residual sample array).
[0119] The prediction unit may perform prediction on the current block and generate a prediction block including prediction samples of the current block. The prediction unit may determine whether to apply intra prediction or inter prediction to the current block based on the information about prediction output from the entropy decoder 210, and may determine a specific intra / inter prediction mode (prediction technique).
[0120] Similar to that described for the prediction unit in the image coding device 100, the prediction unit may generate a prediction signal based on various prediction methods (techniques) described later.
[0121] The intra prediction unit 265 may predict the current block by referring to samples in the current picture. The description of the intra prediction unit 185 is equally applicable to the intra prediction unit 265.
[0122] The inter prediction unit 260 may derive a prediction block of the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information about the inter prediction direction (L0 prediction, L1 prediction, bi-prediction, etc.). In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. For example, the inter prediction unit 260 may configure a motion information candidate list based on neighboring blocks, and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. The inter prediction may be performed based on various prediction modes, and the information about prediction may include information indicating the inter prediction mode of the current block.
[0123] The adder 235 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the prediction unit (including the inter prediction unit 260 and / or the intra prediction unit 265). The description of the adder 155 is equally applicable to the adder 235.
[0124] In addition, as described below, luminance mapping and chrominance scaling (LMCS) are applicable to the picture decoding process.
[0125] The filter 240 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 240 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and store the modified reconstructed picture in the memory 250, specifically, in the DPB of the memory 250. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filtering, bilateral filtering, etc.
[0126] The (modified) reconstructed picture stored in the DPB of the memory 250 can be used as a reference picture in the inter prediction unit 260. The memory 250 can store the motion information of the blocks from which the motion information in the current picture is derived (or decoded) and / or the motion information of the blocks that have been reconstructed in the picture. The stored motion information can be transmitted to the inter prediction unit 260 to be used as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 250 can store the reconstructed samples of the reconstructed blocks in the current picture and transmit the reconstructed samples to the intra prediction unit 265.
[0127] In the present disclosure, the embodiments described in the filter 160, the inter prediction unit 180, and the intra prediction unit 185 of the image encoding device 100 can be equally or correspondingly applied to the filter 240, the inter prediction unit 260, and the intra prediction unit 265 of the image decoding device 200.
[0128] Overview of Inter-Frame Prediction
[0129] The image encoding device / image decoding device can perform inter prediction in units of blocks to derive prediction samples. Inter prediction can mean deriving a prediction in a manner that depends on data elements of a picture other than the current picture. When inter prediction is applied to a current block, the prediction block of the current block can be derived based on a reference block specified by a motion vector on a reference picture.
[0130] In this case, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information of the current block can be derived based on the correlation of the motion information between adjacent blocks and the current block, and the motion information can be derived in units of blocks, sub-blocks, or samples. The motion information can include a motion vector and a reference picture index. The motion information can also include inter prediction type information. Here, the inter prediction type information can mean the direction information of the inter prediction. The inter prediction type information can indicate using one of L0 prediction, L1 prediction, or bi-prediction to predict the current block.
[0131] When performing inter-frame prediction on a current block, neighboring blocks of the current block may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in a reference picture. The reference picture including the reference block of the current block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block may be referred to as a collocated reference block or a collocated CU (colCU), and the reference picture including the temporal neighboring block may be referred to as a collocated picture (colPic).
[0132] In addition, a candidate list of motion information may be constructed based on neighboring blocks of the current block, and in this case, a flag or index information indicating which candidate to use may be signaled to derive the motion vector and / or reference picture index of the current block.
[0133] According to the inter-frame prediction type, the motion information may include L0 motion information and / or L1 motion information. The motion vector in the L0 direction may be defined as the L0 motion vector or MVL0, and the motion vector in the L1 direction may be defined as the L1 motion vector or MVL1. The prediction based on the L0 motion vector may be defined as L0 prediction, the prediction based on the L1 motion vector may be defined as L1 prediction, and the prediction based on both the L0 motion vector and the L1 motion vector may be defined as bi-prediction. Here, the L0 motion vector may mean a motion vector associated with the reference picture list L0, and the L1 motion vector may mean a motion vector associated with the reference picture list L1.
[0134] The reference picture list L0 may include pictures before the current picture in the output order as reference pictures, and the reference picture list L1 may include pictures after the current picture in the output order. The previous picture may be defined as a forward (reference) picture, and the subsequent picture may be defined as a backward (reference) picture. In addition, the reference picture list L0 may further include pictures after the current picture in the output order as reference pictures. In this case, within the reference picture list L0, the previous pictures may be indexed first, and then the subsequent pictures may be indexed. The reference picture list L1 may further include pictures before the current picture in the output order as reference pictures. In this case, within the reference picture list L1, the subsequent pictures may be indexed first, and then the previous pictures may be indexed. Here, the output order may correspond to the picture order count (POC) order.
[0135] Figure 4 is a flowchart exemplifying a video / image encoding method based on inter-frame prediction.
[0136] Figure 5 is a view exemplifying the configuration of the inter-frame prediction unit 180 according to the present disclosure.
[0137] Figure 6 The encoding method ofFigure 2 is performed by an image encoding device. Specifically, step S410 may be performed by the inter-frame prediction unit 180, and step S420 may be performed by the residual processor. Specifically, step S420 may be performed by the subtractor 115. Step S430 may be performed by the entropy encoder 190. The prediction information of step S630 may be derived by the inter-frame prediction unit 180, and the residual information of step S630 may be derived by the residual processor. The residual information is information about residual samples. The residual information may include information about the quantization transform coefficients for the residual samples. As described above, the residual samples may be derived as transform coefficients by the transformer 120 of the image encoding device, and the transform coefficients may be derived as quantization transform coefficients by the quantizer 130. The information about the quantization transform coefficients may be encoded by the entropy encoder 190 through a residual encoding process.
[0138] The image encoding device may perform inter-frame prediction (S410) for a current block. The image encoding device may derive an inter-frame prediction mode and motion information of the current block and generate a prediction sample of the current block. Here, the inter-frame prediction mode determination, motion information derivation, and prediction sample generation processes may be performed simultaneously or any one of them may be performed before other processes. For example, as Figure 5 shown, the inter-frame prediction unit 180 of the image encoding device may include a prediction mode determination unit 181, a motion information derivation unit 182, and a prediction sample derivation unit 183. The prediction mode determination unit 181 may determine the prediction mode of the current block, the motion information derivation unit 182 may derive the motion information of the current block, and the prediction sample derivation unit 183 may derive the prediction sample of the current block. For example, the inter-frame prediction unit 180 of the image encoding device may search for a block similar to the current block within a predetermined area (search area) of a reference picture through motion estimation, and derive a reference block whose difference from the current block is equal to or less than a predetermined criterion or minimum value. Based on this, a reference picture index indicating the reference picture in which the reference block is located may be derived, and a motion vector may be derived based on the position difference between the reference block and the current block. The image encoding device may determine a mode to be applied to the current block among various inter-frame prediction modes. The image encoding device may compare rate-distortion (RD) costs for various prediction modes and determine the best inter-frame prediction mode of the current block. However, the method for the image encoding device to determine the inter-frame prediction mode of the current block is not limited to the above example, and various methods may be used.
[0139] For example, the inter prediction mode of the current block can be determined as at least one of a merge mode, a merge skip mode, a motion vector prediction (MVP) mode, a symmetric motion vector difference (SMVD) mode, an affine mode, a sub-block based merge mode, an adaptive motion vector resolution (AMVR) mode, a history-based motion vector predictor (HMVP) mode, a pairwise average merge mode, a merge mode with motion vector difference (MMVD) mode, a decoder-side motion vector refinement (DMVR) mode, a combined inter and intra prediction (CIIP) mode, or a geometric partitioning mode (GPM).
[0140] For example, when a skip mode or a merge mode is applied to the current block, the image coding device may derive merge candidates from neighboring blocks of the current block and use the derived merge candidates to construct a merge candidate list. Additionally, the image coding device may derive, among the reference blocks indicated by the merge candidates included in the merge candidate list, a reference block whose difference from the current block is equal to or less than a predetermined criterion or minimum value. In this case, a merge candidate associated with the derived reference block may be selected, and merge index information indicating the selected merge candidate may be generated and signaled to the image decoding device. The motion information of the current block may be derived using the motion information of the selected merge candidate.
[0141] As another example, when the MVP mode is applied to the current block, the image coding device may derive motion vector predictor (MVP) candidates from neighboring blocks of the current block and use the derived MVP candidates to construct an MVP candidate list. Additionally, the image coding device may use the motion vector of an MVP candidate selected from among the MVP candidates included in the MVP candidate list as the MVP of the current block. In this case, for example, the motion vector indicating the reference block derived through the above motion estimation may be used as the motion vector of the current block, and the MVP candidate having the smallest difference from the motion vector of the current block among the MVP candidates may be the selected MVP candidate. A motion vector difference (MVD) that is the difference obtained by subtracting the MVP from the motion vector of the current block may be derived. In this case, the index information indicating the selected MVP candidate and information about the MVD may be signaled to the image decoding device. Additionally, when the MVP mode is applied, the value of the reference picture index may be constructed as reference picture index information and signaled separately to the image decoding device.
[0142] The image coding device may derive a residual sample based on a prediction sample (S420). The image coding device may derive the residual sample by comparing the original sample of the current block with the prediction sample. For example, the residual sample may be derived by subtracting the corresponding prediction sample from the original sample.
[0143] An image encoding device may encode image information including prediction information and residual information (S430). The image encoding device may output the encoded image information in the form of a bitstream. The prediction information may include prediction mode information (e.g., a skip flag, a merge flag, or a mode index, etc.) and information about motion information as information related to the prediction process. Among the prediction mode information, the skip flag indicates whether the skip mode is applied to the current block, and the merge flag indicates whether the merge mode is applied to the current block. Alternatively, the prediction mode information may indicate one of multiple prediction modes, such as a mode index. When the skip flag and the merge flag are 0, it may be determined that the MVP mode is applied to the current block. The information about motion information may include candidate selection information (e.g., a merge index, an mvp flag, or an mvp index) as information for deriving a motion vector. Among the candidate selection information, the merge index may be signaled when the merge mode is applied to the current block, and may be information for selecting one of the merge candidates included in the merge candidate list. Among the candidate selection information, the MVP flag or the MVP index may be signaled when the MVP mode is applied to the current block, and may be information for selecting one of the MVP candidates in the MVP candidate list. Specifically, the MVP flag may be signaled using the syntax elements mvp_10_flag or mvp_11_flag. Additionally, the information about motion information may include information about the above-mentioned MVD and / or reference picture index information. Additionally, the information about motion information may include information indicating whether L0 prediction, L1 prediction, or bi-prediction is applied. The residual information is information about residual samples. The residual information may include information about quantization transform coefficients for the residual samples.
[0144] The output bitstream may be stored in a (digital) storage medium and sent to the image decoding device or may be sent to the image decoding device via a network.
[0145] As described above, the image encoding device may generate a reconstructed picture (a picture including reconstructed samples and reconstructed blocks) based on reference samples and residual samples. This is for the image encoding device to derive the same prediction result as the prediction result performed by the image decoding device, thereby improving the encoding efficiency. Therefore, the image encoding device may store the reconstructed picture (or reconstructed samples and reconstructed blocks) in the memory and use it as a reference picture for inter prediction. As described above, the in-loop filtering process also applies to the reconstructed picture.
[0146] Figure 6 is a flowchart illustrating a video / image decoding method based on inter prediction.
[0147] Figure 7 is a view illustrating the configuration of the inter prediction unit 260 according to the present disclosure.
[0148] An image decoding device may perform operations corresponding to those performed by an image encoding device. The image decoding device may perform prediction for a current block and derive prediction samples based on received prediction information.
[0149] Figure 6 The decoding method may be performed by Figure 3 the image decoding device. Steps S610 to S630 may be performed by the inter prediction unit 260, and the prediction information in step S610 and the residual information in step S640 may be obtained by the entropy decoder 210 from the bitstream. The residual processor of the image decoding device may derive the residual samples of the current block based on the residual information (S640). Specifically, the dequantizer 220 of the residual processor may perform dequantization based on the quantization transform coefficients derived according to the residual information to derive transform coefficients, and the inverse transformer 230 of the residual processor may perform inverse transformation on the transform coefficients to derive the residual samples of the current block. Step S650 may be performed by the adder 235 or the reconstructor.
[0150] Specifically, the image decoding device may determine the prediction mode of the current block based on the received prediction information (S610). The image decoding device may determine which inter prediction mode is applied to the current block based on the prediction mode information in the prediction information.
[0151] For example, it may be determined whether the skip mode is applied to the current block based on the skip flag. Additionally, it may be determined whether the merge mode or the MVP mode is applied to the current block based on the merge flag. Alternatively, one of various inter prediction mode candidates may be selected based on the mode index. The inter prediction mode candidates may include the skip mode, the merge mode, and / or the MVP mode or may include various inter prediction modes to be described below.
[0152] The image decoding device may derive the motion information of the current block based on the determined inter prediction mode (S620). For example, when the skip mode or the merge mode is applied to the current block, the image decoding device may construct a merge candidate list to be described below and select one of the merge candidates included in the merge candidate list. The selection may be performed based on the above candidate selection information (merge index). The motion information of the selected merge candidate may be used to derive the motion information of the current block. For example, the motion information of the selected merge candidate may be used as the motion information of the current block.
[0153] As another example, when the MVP mode is applied to the current block, the image decoding device may construct an MVP candidate list and use the motion vector of the MVP candidate selected from among the MVP candidates included in the MVP candidate list as the MVP of the current block. The selection may be performed based on the above candidate selection information (mvp flag or mvp index). In this case, the MVD of the current block may be derived based on the information about the MVD, and the motion vector of the current block may be derived based on the MVP and MVD of the current block. Additionally, the reference picture index of the current block may be derived based on the reference picture index information. The picture indicated by the reference picture index in the reference picture list of the current block may be derived as the reference picture referred to for the inter-frame prediction of the current block.
[0154] The image decoding device may generate a prediction sample for the current block based on the motion information of the current block (S630). In this case, the reference picture may be derived based on the reference picture index of the current block, and the prediction sample of the current block may be derived using the samples of the reference block indicated by the motion vector of the current block on the reference picture. In some cases, a prediction sample filtering process may also be performed on all or some of the prediction samples of the current block.
[0155] For example, as Figure 7 shown, the inter-frame prediction unit 260 of the image decoding device may include a prediction mode determination unit 261, a motion information derivation unit 262, and a prediction sample derivation unit 263. In the inter-frame prediction unit 260 of the image decoding device, the prediction mode determination unit 261 may determine the prediction mode of the current block based on the received prediction mode information, the motion information derivation unit 262 may derive the motion information (motion vector and / or reference picture index, etc.) of the current block based on the received motion information, and the prediction sample derivation unit 263 may derive the prediction sample of the current block.
[0156] The image decoding device may generate a residual sample for the current block based on the received residual information (S640). The image decoding device may generate a reconstructed sample for the current block based on the prediction sample and the residual sample and generate a reconstructed picture based on this (S650). Thereafter, an in-loop filtering process is applied to the reconstructed picture as described above.
[0157] As described above, the inter-frame prediction process may include steps of determining an inter-frame prediction mode, deriving motion information according to the determined prediction mode, and performing prediction (generating a prediction sample) based on the derived motion information. As described above, the inter-frame prediction process may be performed by an image encoding device and an image decoding device.
[0158] Hereinafter, the step of deriving motion information according to the prediction mode will be described in more detail.
[0159] As described above, the motion information of the current block can be used to perform inter prediction. The image encoding device can derive the optimal motion information of the current block through a motion estimation process. For example, the image encoding device can search for a similar reference block with high correlation in the reference picture using the original block in the original picture of the current block in fractional pixel units, and use it to derive the motion information. The similarity of the block can be calculated based on the sum of absolute differences (SAD) between the current block and the reference block. In this case, the motion information can be derived based on the reference block with the minimum SAD in the search area. The derived motion information can be signaled to the image decoding device according to various methods based on the inter prediction mode.
[0160] When the merge mode is applied to the current block, the motion information of the current block is not directly sent, and the motion information of the neighboring block is used to derive the motion information of the current block. Therefore, the motion information of the current prediction block can be indicated by sending flag information indicating that the merge mode is used and candidate selection information (e.g., merge index) indicating which neighboring block is used as a merge candidate. In the present disclosure, since the current block is a prediction execution unit, the current block can be used in the same meaning as the current prediction block, and the neighboring block can be used in the same meaning as the neighboring prediction block.
[0161] The image encoding device can search for merge candidate blocks for deriving the motion information of the current block to perform the merge mode. For example, up to five merge candidate blocks can be used, but it is not limited thereto. The maximum number of merge candidate blocks can be sent in the slice header or the tile group header, but it is not limited thereto. After finding the merge candidate blocks, the image encoding device can generate a merge candidate list and select the merge candidate block with the minimum RD cost as the final merge candidate block.
[0162] The present disclosure provides various embodiments for configuring merge candidate blocks of the merge candidate list. The merge candidate list can use, for example, five merge candidate blocks. For example, four spatial merge candidates and one temporal merge candidate can be used.
[0163] Figure 8 is a view illustrating neighboring blocks that can be used as spatial merge candidates.
[0164] Figure 9 is a view schematically illustrating a method for constructing a merge candidate list according to an example of the present disclosure.
[0165] The image encoding / decoding device can insert the spatial merge candidates derived by searching the spatial neighboring blocks of the current block into the merge candidate list (S910). For example, as Figure 8As shown, the spatially neighboring blocks may include the bottom-left neighboring block A0, the left neighboring block A1, the top-right neighboring block B0, the top neighboring block B1, and the top-left neighboring block B2 of the current block. However, this is an example, and in addition to the above spatially neighboring blocks, additional neighboring blocks such as the right neighboring block, the bottom neighboring block, and the bottom-right neighboring block may be further used as spatially neighboring blocks. The image encoding / decoding device may detect available blocks by searching for spatially neighboring blocks based on priority, and derive the motion information of the detected blocks as spatial merge candidates. For example, the image encoding / decoding device may search for the five blocks shown in Figure 8 the order of A1, B1, B0, A0, and B2, and index the available candidates in sequence to construct a merge candidate list.
[0166] The image encoding / decoding device may insert the temporal merge candidates derived by searching for the temporal neighboring blocks of the current block into the merge candidate list (S920). The temporal neighboring blocks may be located on a reference picture different from the current picture in which the current block is located. The reference picture in which the temporal neighboring blocks are located may be referred to as the collocated picture or col picture. The temporal neighboring blocks may be searched in the order of the bottom-right neighboring block and the bottom-right center block of the collocated block of the current block on the col picture. In addition, when motion data compression is applied to reduce the memory load, specific motion information may be stored as the representative motion information of each predetermined storage unit of the col picture. In this case, it is not necessary to store the motion information of all the blocks in the predetermined storage unit, thereby obtaining a motion data compression effect. In this case, the predetermined storage unit may be predetermined as, for example, a 16×16 sample unit or an 8×8 sample unit, or the size information of the predetermined storage unit may be signaled from the image encoding device to the image decoding device. When motion data compression is applied, the motion information of the temporal neighboring blocks may be replaced with the representative motion information of the predetermined storage unit in which the temporal neighboring blocks are located. That is, in this case, from the perspective of implementation, the temporal merge candidates may be derived based on the motion information of the prediction block (instead of the prediction block located at the coordinates of the temporal neighboring block) that covers the arithmetic left-shift position after the arithmetic right-shift of the coordinates (top-left sample position) of the temporal neighboring block by a predetermined value. For example, when the predetermined storage unit is 2 n ×2 nWhen the coordinates of the sample unit and the temporally neighboring block are (xTnb, yTnb), the motion information of the predicted block located at the modified position ((xTnb >> n) << n), (yTnb >> n) << n)) can be used for temporal merge candidates. Specifically, for example, when the predetermined storage unit is a 16×16 sample unit and the coordinates of the temporally neighboring block are (xTnb, yTnb), the motion information of the predicted block located at the modified position ((xTnb >> 4) << 4), (yTnb >> 4) << 4)) can be used for temporal merge candidates. Alternatively, for example, when the predetermined storage unit is an 8×8 sample unit and the coordinates of the temporally neighboring block are (xTnb, yTnb), the motion information of the predicted block located at the modified position ((xTnb >> 3) << 3), (yTnb >> 3) << 3)) can be used for temporal merge candidates.
[0167] Referring again to Figure 9 , the image encoding / decoding device may check whether the current number of merge candidates is less than the maximum number of merge candidates (S930). The maximum number of merge candidates may be predefined or signaled from the image encoding device to the image decoding device. For example, the image encoding device may generate information about the maximum number of merge candidates, encode it, and send the encoded information to the image decoding device in the form of a bitstream. When the maximum number of merge candidates is satisfied, the subsequent candidate addition process S940 may not be performed.
[0168] When, as a result of the check in step S930, the current number of merge candidates is less than the maximum number of merge candidates, the image encoding / decoding device may derive additional merge candidates according to a predetermined method and then insert the additional merge candidates into the merge candidate list (S940). For example, the additional merge candidates may include at least one of a history-based merge candidate, a pairwise average merge candidate, an ATMVP, a combined bi-prediction merge candidate (when the slice / tile group type of the current slice / tile group is of type B), and / or a zero vector merge candidate.
[0169] When, as a result of the check in step S930, the current number of merge candidates is not less than the maximum number of merge candidates, the image encoding / decoding device may end the construction of the merge candidate list. In this case, the image encoding device may select the best merge candidate from among the merge candidates configuring the merge candidate list and signal candidate selection information (e.g., a merge candidate index or a merge index) indicating the selected merge candidate to the image decoding device. The image decoding device may select the best merge candidate based on the merge candidate list and the candidate selection information.
[0170] As described above, the motion information of the selected merge candidate can be used as the motion information of the current block, and the predicted samples of the current block can be derived based on the motion information of the current block. The image encoding device can derive the residual samples of the current block based on the predicted samples, and signal the residual information of the residual samples to the image decoding device. As described above, the image decoding device can generate reconstructed samples based on the residual samples derived according to the residual information and the predicted samples, and generate a reconstructed picture based on the same.
[0171] When the skip mode is applied to the current block, the same method as in the case of applying the merge mode can be used to derive the motion information of the current block. However, when the skip mode is applied, the residual signal of the corresponding block is omitted, and thus the predicted samples can be directly used as the reconstructed samples. For example, when the value of cu_skip_flag is 1, the above skip mode can be applied.
[0172] Hereinafter, a method of deriving spatial candidates in the merge mode and / or the skip mode will be described. The spatial candidates can represent the above-described spatial merge candidates.
[0173] The derivation of the spatial candidates can be performed based on spatially neighboring blocks. For example, up to four spatial candidates can be derived from the candidate blocks existing at the Figure 8 positions shown. The order of deriving the spatial candidates can be A1->B1->B0->A0->B2. However, the order of deriving the spatial candidates is not limited to the above order, and can be, for example, B1->A1->B0->A0->B2. When at least one of the current four positions (A1, B1, B0, and A0 in the above example) is unavailable, the last position in the order (position B2 in the above example) can be considered. In this case, the blocks at the unavailable predetermined positions can include the corresponding blocks belonging to a different slice or tile from the current block or the corresponding blocks as intra-predicted blocks. When deriving a spatial candidate from the first position in the order (A1 or B1 in the above example), a redundancy check can be performed on the spatial candidates at the subsequent positions. For example, when the motion information of a subsequent spatial candidate is the same as the motion information of a spatial candidate already included in the merge candidate list, the subsequent spatial candidate may not be included in the merge candidate list, thereby improving the encoding efficiency. The redundancy check performed on the subsequent spatial candidates can be performed on some candidate pairs rather than all possible candidate pairs, thereby reducing the computational complexity.
[0174] Figure 10 is a view illustrating the candidate pairs for the redundancy check performed on the spatial candidates.
[0175] In Figure 10In the example shown, a redundancy check on the spatial candidate at position B0 can be performed only on the spatial candidate at position A0. Additionally, a redundancy check on the spatial candidate at position B1 can be performed only on the spatial candidate at position B0. Additionally, a redundancy check on the spatial candidate at position A1 can be performed only on the spatial candidate at position A0. Finally, a redundancy check on the spatial candidate at position B2 can be performed only on the spatial candidates at positions A0 and B0.
[0176] In Figure 10 the example shown, the order of deriving spatial candidates is A0 -> B0 -> B1 -> A1 -> B2. However, the present disclosure is not limited thereto, and even if the order of deriving spatial candidates changes, as Figure 10 shown in the example, a redundancy check can be performed only on some candidate pairs.
[0177] Hereinafter, a method of deriving temporal candidates in the case of a merge mode and / or a skip mode will be described. The temporal candidates may represent the above-described temporal merge candidates. Additionally, the motion vectors of the temporal candidates may correspond to the temporal candidates in the MVP mode.
[0178] In the case of temporal candidates, only one candidate can be included in the merge candidate list. During the process of deriving temporal candidates, the motion vectors of the temporal candidates can be scaled. For example, the scaling can be performed based on a collocated block (CU) (hereinafter referred to as a "col block") belonging to a collocated reference picture (colPic) (hereinafter referred to as a "col picture"). The reference picture list used to derive the col block can be explicitly signaled in the slice header.
[0179] Figure 11 is a view illustrating a method of scaling the motion vectors of temporal candidates.
[0180] In Figure 11 it, curr_CU and curr_pic represent the current block and the current picture, respectively, and col_CU and col_pic represent the col block and the col picture, respectively. Additionally, curr_ref represents the reference picture of the current block, and col_ref represents the reference picture of the col block. Additionally, tb represents the distance between the reference picture of the current block and the current picture, and td represents the distance between the reference picture of the col block and the col picture. tb and td can represent values corresponding to the difference in POC (Picture Order Count) between pictures. The scaling of the motion vectors of the temporal candidates can be performed based on tb and td. Additionally, the reference picture index of the temporal candidate can be set to 0.
[0181] Figure 12 is a view illustrating the position of deriving temporal candidates.
[0182] In Figure 12The block with a thick solid line indicates the current block. Figure 12 The temporal candidate is derived from the block corresponding to position C0 (lower right position) or C1 (center position). First, it can be determined whether position C0 is available, and when position C0 is available, the temporal candidate can be derived based on position C0. When position C0 is not available, the temporal candidate can be derived based on position C1. For example, when the block at position C0 in the col picture is an intra-frame prediction block or is located outside the current CTU row, it can be determined that position C0 is not available.
[0183] As described above, when motion data compression is applied, the motion vector of the col block may be stored for each predetermined unit block. In this case, in order to derive the motion vector of the block covering position C0 or position C1, position C0 or position C1 may be modified. For example, when the predetermined unit block is an 8×8 block and position C0 or position C1 is (xColCi, yColCi), the position for deriving the temporal candidate may be modified to ((xColCi>>3)<<3, (yColCi>>3)<<3).
[0184] Hereinafter, a method of deriving a history-based candidate in case of a merge mode and / or a skip mode will be described.The history-based candidate may be represented by a history-based merge candidate.
[0185] After the spatial candidates and temporal candidates are added to the merge candidate list, history-based candidates may be added to the merge candidate list. For example, the motion information of a previously encoded / decoded block may be stored in a table and used as a history-based candidate for the current block. The table may store multiple history-based candidates during the encoding / decoding process. The table may be initialized at the beginning of a new CTU row. Initializing the table may mean clearing the corresponding table by deleting all history-based candidates stored in the table. Whenever there is an inter-prediction block, the relevant motion information may be added to the table as the last entry. In this case, the inter-prediction block may not be a block based on sub-block prediction. The motion information added to the table may be used as a new history-based candidate.
[0186] The table of history-based candidates may have a predetermined size. For example, the size may be 5. In this case, the table may store up to five history-based candidates. When a new candidate is added to the table, a limited first-in, first-out (FIFO) rule of redundant checking of whether there is an identical candidate in the table may be applied. If an identical candidate already exists in the table, the identical candidate may be deleted from the table and the positions of all subsequent history-based candidates may be moved forward.
[0187] Historical candidates can be used in the process of configuring the merge candidate list. In this case, the historical candidates that were most recently included in the table can be checked in sequence and placed after the temporal candidates in the merge candidate list. When a historical candidate is included in the merge candidate list, a redundancy check can be performed with the spatial or temporal candidates that are already included in the merge candidate list. If a spatial or temporal candidate already included in the merge candidate list overlaps with the historical candidate, the historical candidate may not be included in the merge candidate list. By simplifying the redundancy check as described below, the computational load can be reduced.
[0188] The number of historical candidates used to generate the merge candidate list can be set to (N <= 4)? M : (8 - N). In this case, N can represent the number of candidates already included in the merge candidate list, and M can represent the number of available historical candidates included in the table. That is, when the merge candidate list includes 4 or fewer candidates, the number of historical candidates used to generate the merge candidate list can be M, and when the merge candidate list includes N candidates greater than 4, the number of historical candidates used to generate the merge candidate list can be set to (8 - N).
[0189] When the total number of available merge candidates reaches (the maximum allowable number of merge candidates - 1), the configuration of the merge candidate list using historical candidates can end.
[0190] Hereinafter, a method for deriving paired average candidates in the case of the merge mode and / or the skip mode will be described. The paired average candidates can be represented by paired average merge candidates or paired candidates.
[0191] Paired average candidates can be generated by obtaining predefined candidate pairs from the candidates included in the merge candidate list and averaging them. The predefined candidate pairs can be {(0, 1), (0, 2), (1, 2), (0, 3), (1, 3), (2, 3)} and the numbers configuring each candidate pair can be the indices of the merge candidate list. That is, the predefined candidate pair (0, 1) can mean a pair of the index 0 candidate and the index 1 candidate in the merge candidate list, and the paired average candidate can be generated by averaging the index 0 candidate and the index 1 candidate. The derivation of the paired average candidates can be performed in the order of the predefined candidate pairs. That is, after deriving the paired average candidate for the candidate pair (0, 1), the process of deriving the paired average candidate can be performed in the order of the candidate pair (0, 2) and the candidate pair (1, 2). The paired average candidate derivation process can be performed until the configuration of the merge candidate list is completed. For example, the paired average candidate derivation process can be performed until the number of merge candidates included in the merge candidate list reaches the maximum number of merge candidates.
[0192] Pairwise average candidates can be calculated separately for each reference picture list. When two motion vectors are available for a reference picture list (L0 list or L1 list), the average of the two motion vectors can be calculated. In this case, even if the two motion vectors indicate different reference pictures, the average of the two motion vectors can be performed. If only one motion vector is available for a reference picture list, the available motion vector can be used as the motion vector of the pairwise average candidate. If no motion vector is available for a reference picture list, the reference picture list can be determined to be invalid.
[0193] When the configuration of the merge candidate list is not completed even after the pairwise average candidate is included in the merge candidate list, a zero vector can be added to the merge candidate list until the maximum merge candidate number is reached.
[0194] When the MVP mode is applied to the current block, the motion vectors of the reconstructed spatial neighboring blocks (e.g., Figure 8 the neighboring blocks shown) and / or the motion vectors corresponding to the temporal neighboring blocks (or Col blocks) can be used to generate a motion vector predictor (MVP) candidate list. That is, the motion vectors of the reconstructed spatial neighboring blocks and the motion vectors corresponding to the temporal neighboring blocks can be used as the motion vector predictor candidates of the current block. When bidirectional prediction is applied, an MVP candidate list for L0 motion information derivation and an MVP candidate list for L1 motion information derivation are generated and used separately. The prediction information (or information about prediction) of the current block can include candidate selection information (e.g., MVP flag or MVP index) indicating the best motion vector predictor candidate selected from among the motion vector predictor candidates included in the MVP candidate list. In this case, the prediction unit can use the candidate selection information to select the motion vector predictor of the current block from among the motion vector predictor candidates included in the MVP candidate list. The prediction unit of the image coding device can obtain and encode the motion vector difference (MVD) between the motion vector of the current block and the motion vector predictor and output the encoded MVD in the form of a bitstream. That is, the MVD can be obtained by subtracting the motion vector predictor from the motion vector of the current block. The prediction unit of the image decoding device can obtain the motion vector difference included in the information about prediction and derive the motion vector of the current block by adding the motion vector difference and the motion vector predictor. The prediction unit of the image decoding device can obtain or derive a reference picture index indicating the reference picture from the information about prediction.
[0195] Figure 13 is a view schematically illustrating a method for constructing a motion vector predictor candidate list according to an example of the present disclosure.
[0196] First, spatial candidate blocks of the current block can be searched and available candidate blocks can be inserted into the MVP candidate list (S1010). Thereafter, it is determined whether the number of MVP candidates included in the MVP candidate list is less than 2 (S1020), and when the number of MVP candidates is 2, the construction of the MVP candidate list can be completed.
[0197] In step S1020, when the number of available spatial candidate blocks is less than 2, temporal candidate blocks of the current block can be searched and available candidate blocks can be inserted into the MVP candidate list (S1030). When temporal candidate blocks are not available, a zero motion vector can be inserted into the MVP candidate list (S1040), thereby completing the construction of the MVP candidate list.
[0198] In addition, when the mvp mode is applied, the reference picture index can be signaled explicitly. In this case, the reference picture index refidxL0 for L0 prediction and the reference picture index refidxL1 for L1 prediction can be signaled differentially. For example, when the MVP mode is applied and dual prediction is applied, information about refidxL0 and information about refidxL1 can be signaled.
[0199] As described above, when the MVP mode is applied, information about the MVP derived by the image coding device can be signaled to the image decoding device. For example, information about the MVD can include the absolute value of the MVD and information indicating the x and y components for the positive and negative signs (sign). In this case, when the absolute value of the MVD is greater than 0, whether the absolute value of the MVD is greater than 1 and information indicating the MVD remainder can be signaled step by step. For example, information indicating whether the absolute value of the MVD is greater than 1 can be signaled only when the value of the flag information indicating whether the absolute value of the MVD is greater than 0 is 1.
[0200] Overview of Affine Mode
[0201] Hereinafter, the affine mode, which is an example of an inter prediction mode, will be described in detail. In a conventional video coding / decoding system, only one motion vector is used to express the motion information of the current block. However, in this method, there is a problem that the best motion information can only be expressed in units of blocks, but the best motion information cannot be expressed in units of pixels. To solve this problem, an affine mode that defines the motion information of a block in units of pixels has been proposed. According to the affine mode, two to four motion vectors associated with the current block can be used to determine the motion vectors of each pixel and / or sub-block unit of the block.
[0202] Compared with the existing motion information expressed by the translational motion (or displacement) of pixel values, in the affine mode, the motion information of each pixel can be expressed using at least one of translational motion, scaling, rotation, or shear. Among them, the affine mode that uses displacement, scaling, or rotation to express the motion information of each pixel can be a similarity or simplified affine mode. The affine mode in the following description can mean a similarity or simplified affine mode.
[0203] The motion information in the affine mode can be expressed using two or more control point motion vectors (CPMVs). The CPMVs can be used to derive the motion vector of a specific pixel position in the current block. In this case, the set of motion vectors of each pixel and / or sub-block in the current block can be defined as an affine motion vector field (affine MVF).
[0204] Figure 14 is a view illustrating the parametric model of the affine mode.
[0205] When the affine mode is applied to the current block, one of the 4-parameter model and the 6-parameter model can be used to derive the affine MVF. In this case, the 4-parameter model can mean a model type that uses two CPMVs, and the 6-parameter model can mean a model type that uses three CPMVs. Figure 14 of (a) and Figure 14 of (b) respectively show the CPMVs used in the 4-parameter model and the 6-parameter model.
[0206] When the position of the current block is (x, y), the motion vector according to the pixel position can be derived according to Equation 1 or Equation 2 below. For example, the motion vector according to the 4-parameter model can be derived according to Equation 1, and the motion vector according to the 6-parameter model can be derived according to Equation 2.
[0207] [Equation 1]
[0208]
[0209] [Equation 2]
[0210]
[0211] In Equation 1 and Equation 2, mv0 = {mv_0x, mv_0y} can be the CPMV at the upper left corner position of the current block, v1 = {mv_1x, mv_1y} can be the CPMV at the upper right position of the current block, and mv2 = {mv_2x, mv_2y} can be the CPMV at the lower left position of the current block. In this case, W and H respectively correspond to the width and height of the current block, and mv = {mv_x, mv_y} can mean the motion vector of the pixel position {x, y}.
[0212] In the encoding / decoding process, the affine MVF can be determined in units of pixels and / or predefined sub-blocks. When determining the affine MVF in units of pixels, the motion vector can be derived based on each pixel value. In addition, when determining the affine MVF in units of sub-blocks, the motion vector of the corresponding block can be derived based on the central pixel value of the sub-block. The central pixel value can mean a virtual pixel existing at the center of the sub-block or the lower-right pixel among the four pixels at the center. Additionally, the central pixel value can be a specific pixel in the sub-block and can be a pixel representing the sub-block. In the present disclosure, the case of determining the affine MVF in units of 4×4 sub-blocks will be described. However, this is only for convenience of description, and the size of the sub-block can be changed differently.
[0213] That is, when affine prediction is available, the motion models applicable to the current block can include three models, namely, the translational motion model, the 4-parameter affine motion model, and the 6-parameter affine motion model. Here, the translational motion model can represent the model used by the existing block unit motion vector, the 4-parameter affine motion model can represent the model used by two CPMVs, and the 6-parameter affine motion model can represent the model used by three CPMVs. The affine mode can be divided into detailed modes according to the motion information encoding / decoding method. For example, the affine mode can be further divided into the affine MVP mode and the affine merge mode.
[0214] When the affine merge mode is applied to the current block, the CPMV can be derived from neighboring blocks of the current block encoded / decoded in the affine mode. When at least one neighboring block of the current block is encoded / decoded in the affine mode, the affine merge mode can be applied to the current block. That is, when the affine merge mode is applied to the current block, the CPMV of the current block can be derived using the CPMV of the neighboring blocks. For example, the CPMV of the neighboring block can be determined as the CPMV of the current block, or the CPMV of the current block can be derived based on the CPMV of the neighboring block. When deriving the CPMV of the current block based on the CPMV of the neighboring block, at least one encoding parameter of the current block or the neighboring block can be used. For example, the CPMV of the neighboring block can be modified based on the size of the neighboring block and the size of the current block and used as the CPMV of the current block.
[0215] In addition, the affine merge of deriving an MV in units of sub-blocks may be referred to as a sub-block merge mode, which may be specified by a merge_subblock_flag having a first value (e.g., 1). In this case, the affine merge candidate list described below may be referred to as a sub-block merge candidate list. In this case, candidates derived as the following SbTMVP may be further included in the sub-block merge candidate list. In this case, candidates derived as sbTMVP may be used as candidates for index #0 of the sub-block merge candidate list. In other words, candidates derived as sbTMVP may be located in front of the inherited affine candidates and constructed affine candidates described below in the sub-block merge candidate list.
[0216] For example, an affine mode flag may be defined to specify whether the affine mode is applicable to the current block, which may be signaled at at least one higher level (e.g., sequence, picture, slice, tile, tile group, patch, etc.) of the current block. For example, the affine mode flag may be named sps_affine_enabled_flag.
[0217] When the affine merge mode is applied, the affine merge candidate list may be configured to derive the CPMV of the current block. In this case, the affine merge candidate list may include at least one of an inherited affine merge candidate, a constructed affine merge candidate, or a zero merge candidate. When neighboring blocks of the current block are encoded / decoded in the affine mode, the inherited affine merge candidate may mean a candidate derived using the CPMV of the neighboring block. The constructed affine merge candidate may mean a candidate for deriving each CPMV based on the motion vectors of neighboring blocks of each control point (CP). In addition, the zero merge candidate may mean a candidate composed of a CPMV of size 0. In the following description, CP may mean a specific position of a block for deriving a CPMV. For example, CP may be the respective vertex positions of the block.
[0218] Figure 15 is a view illustrating a method of generating an affine merge candidate list.
[0219] Referring to Figure 15 the flowchart of, the affine merge candidates may be added to the affine merge candidate list in the order of an inherited affine merge candidate (S1210), a constructed affine merge candidate (S1220), and a zero merge candidate (S1230). When even if all the inherited affine merge candidates and constructed affine merge candidates are added to the affine merge candidate list, the number of candidates included in the candidate list still does not satisfy the maximum number of candidates, zero merge candidates may be added. In this case, zero merge candidates may be added until the number of candidates in the affine merge candidate list satisfies the maximum number of candidates.
[0220] Figure 16It is a view illustrating a control point motion vector (CPMV) derived from neighboring blocks.
[0221] For example, up to two inherited affine merge candidates can be derived, each of which can be derived based on at least one of the left neighboring block and the upper neighboring block. Reference will be made to Figure 8 to describe the neighboring blocks for deriving the inherited affine merge mode. The inherited affine merge candidate derived based on the left neighboring block is derived based on at least one of A0 or A1, and the inherited affine merge candidate derived based on the upper neighboring block can be derived based on at least one of B0, B1, or B2. In this case, the scanning order of the neighboring blocks can be A0 to A1 and B0, B1, and B2, but is not limited thereto. For each of the left and upper, the inherited affine merge candidate can be derived based on the first available neighboring block in the scanning order. In this case, no redundancy check may be performed between the candidates derived from the left neighboring block and the upper neighboring block.
[0222] For example, as Figure 16 shown, when the left neighboring block A is encoded / decoded in an affine mode, at least one of the motion vectors v2, v3, and v4 corresponding to the CP of the neighboring block A can be derived. When the neighboring block A is encoded / decoded by a 4-parameter affine model, the inherited affine merge candidate can be derived using v2 and v3. In contrast, when the neighboring block A is encoded / decoded by a 6-parameter affine model, the inherited affine merge candidate can be derived using v2, v3, and v4.
[0223] Figure 17 It is a view illustrating the neighboring blocks for deriving the constructed affine merge candidate.
[0224] The constructed affine candidate can mean a candidate having a CPMV derived by combining the general motion information of the neighboring blocks. The motion information of each CP can be derived using the spatial neighboring blocks or the temporal neighboring blocks of the current block. In the following description, CPMVk can mean the motion vector representing the k-th CP. For example, referring to Figure 17 , CPMV1 can be determined as the first available motion vector among the motion vectors of B2, B3, and A2, and in this case, the scanning order can be B2, B3, and A2. CPMV2 can be determined as the first available motion vector among the motion vectors of B1 and B0, and in this case, the scanning order can be B1 and B0. CPMV3 can be determined as one of the motion vectors of A1 and A0, and in this case, the scanning order can be A1 and A0. When TMVP is applicable to the current block, CPMV4 can be determined as the motion vector of the temporal neighboring block T.
[0225] After deriving the four motion vectors of each CP, an affine merge candidate can be derived based on this. The construction of the affine merge candidate can be configured by including at least two motion vectors selected from the four motion vectors of each derived CP. For example, the construction of the affine merge candidate can consist of at least one of {CPMV1, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV4}, {CPMV1, CPMV3, CPMV4}, {CPMV2, CPMV3, CPMV4}, {CPMV1, CPMV2}, or {CPMV1, CPMV3} in this order. The affine candidate construction consisting of three motion vectors can be a candidate for the 6-parameter affine model. In contrast, the affine candidate construction consisting of two motion vectors can be a candidate for the 4-parameter affine model. To avoid the scaling process of the motion vectors, when the reference picture indices of the CPs are different from each other, the combination of the relevant CPMVs can be ignored and not used for deriving the affine merge candidate.
[0226] When the affine MVP mode is applied to the current block, the encoding / decoding device can derive two or more CPMV predictors and CPMVs of the current block and derive the CPMV difference based on them. In this case, the CPMV difference can be signaled from the encoding device to the decoding device. The image decoding device can derive the CPMV predictor of the current block, reconstruct the signaled CPMV difference, and then derive the CPMV of the current block based on the CPMV predictor and the CPMV difference.
[0227] In addition, the affine MVP mode can be applied to the current block only when the affine merge mode or the sub-block based TMVP is not applied to the current block. In addition, the affine MVP mode can be expressed as the affine CP MVP mode.
[0228] When the affine MVP is applied to the current block, the affine MVP candidate list can be configured to derive the CPMV of the current block. In this case, the affine MVP candidate list can include at least one of the inherited affine MVP candidate, the constructed affine MVP candidate, the translational motion affine MVP candidate, or the zero MVP candidate.
[0229] In this case, the inherited affine MVP candidate can mean a candidate derived based on the CPMV of the neighboring block when the neighboring block of the current block is encoded / decoded in the affine mode. The constructed affine MVP candidate can mean a candidate derived by generating a CPMV combination based on the motion vectors of the CP neighboring blocks. The zero MVP candidate can mean a candidate consisting of CPMVs with a value of 0. The derivation method and the characteristics of the inherited affine MVP candidate and the constructed affine MVP candidate are the same as those of the inherited affine candidate and the constructed affine candidate described above, so their descriptions will be omitted.
[0230] When the maximum number of candidates in the affine MVP candidate list is 2, when the current number of candidates is less than 2, affine MVP candidates, translational motion affine MVP candidates, and zero MVP candidates can be added. Specifically, the translational motion affine MVP candidates can be derived in the following order.
[0231] For example, when the number of candidates included in the affine MVP candidate list is less than 2 and CPMV0 of the constructed affine MVP candidate is valid, CPMV0 can be used as an affine MVP candidate. That is, the affine MVP candidates whose motion vectors of CP0, CP1, and CP2 are all CPMV0 can be added to the affine MVP candidate list.
[0232] Next, when the number of candidates in the affine MVP candidate list is less than 2 and CPMV1 of the constructed affine MVP candidate is valid, CPMV1 can be used as an affine MVP candidate. That is, the affine MVP candidates whose motion vectors of CP0, CP1, and CP2 are all CPMV1 can be added to the affine MVP candidate list.
[0233] Next, when the number of candidates in the affine MVP candidate list is less than 2 and CPMV2 of the constructed affine MVP candidate is valid, CPMV2 can be used as an affine MVP candidate. That is, the affine MVP candidates whose motion vectors of CP0, CP1, and CP2 are all CPMV2 can be added to the affine MVP candidate list.
[0234] Regardless of the above conditions, when the number of candidates in the affine MVP candidate list is less than 2, the temporal motion vector predictor (TMVP) of the current block can be added to the affine MVP candidate list.
[0235] Regardless of the addition of translational motion affine MVP candidates, when the number of candidates in the affine MVP candidate list is less than 2, zero MVP candidates can be added to the affine MVP candidate list.
[0236] Figure 18 is a view illustrating a method of generating an affine MVP candidate list.
[0237] Refer to Figure 18 According to the flowchart of, candidates can be added to the affine MVP candidate list in the order of inheriting affine MVP candidates (S1610), constructing affine MVP candidates (S1620), translational motion affine MVP candidates (S1630), and zero MVP candidates (S1640). As described above, steps S1620 to S1640 can be executed according to whether the number of candidates included in the affine MVP candidate list in each step is less than 2.
[0238] The scan order for inherited affine MVP candidates may be equal to the scan order for inherited affine merge candidates. However, in the case of inherited affine MVP candidates, only neighboring blocks that refer to the same reference picture as the reference picture of the current block may be considered. When an inherited affine MVP candidate is added to the affine MVP candidate list, a redundancy check may not be performed.
[0239] To derive and construct affine MVP candidates, only Figure 17 the spatial neighboring blocks shown may be considered. Additionally, the scan order for constructing affine MVP candidates may be equal to the scan order for constructing affine merge candidates. Additionally, to derive and construct affine MVP candidates, the reference picture index of neighboring blocks may be checked, and in the scan order, the first neighboring block that is inter-frame coded and refers to the same reference picture as the reference picture of the current block may be used.
[0240] Overview of Sub-Block Based TMVP (SbTMVP) Mode
[0241] Hereinafter, the sub-block-based TMVP mode, which is an example of an inter-frame prediction mode, will be described in detail. According to the sub-block-based TMVP mode, the motion vector field (MVF) of the current block may be derived and the motion vectors may be derived on a sub-block basis.
[0242] Different from the conventional TMVP mode that is performed on a coding unit basis, for a coding unit to which the sub-block-based TMVP mode is applied, the motion vectors may be coded / decoded on a sub-coding unit basis. Additionally, according to the conventional TMVP mode, the temporal motion vectors may be derived from neighboring blocks, but in the sub-block-based TMVP mode, the motion vector field may be derived from the reference blocks specified by the motion vectors derived from the neighboring blocks of the current block. Hereinafter, the motion vectors derived from neighboring blocks may be referred to as the motion shift or representative motion vectors of the current block.
[0243] Figure 19 is a view exemplifying the neighboring blocks of the sub-block-based TMVP mode.
[0244] When the sub-block-based TMVP mode is applied to the current block, the neighboring blocks for determining the motion shift may be determined. For example, the neighboring blocks for determining the motion shift may be scanned in the order of the blocks A1, B1, B0, and A0 Figure 19 . As another example, the neighboring blocks for determining the motion shift may be limited to specific neighboring blocks of the current block. For example, the neighboring block for determining the motion shift may always be determined as block A1. When the neighboring block has a motion vector referring to the col picture, the corresponding motion vector may be determined as the motion shift. The motion vector determined as the motion shift may be referred to as the temporal motion vector. Additionally, when the above motion vector cannot be derived from the neighboring blocks, the motion shift may be set to (0,0).
[0245] Figure 20 It is a view illustrating a method for deriving a motion vector field according to the sub-block-based TMVP mode.
[0246] Next, the reference block on the collocated picture specified by the motion shift can be determined. For example, sub-block-based motion information (motion vector or reference picture index) can be obtained from the col picture by adding the motion shift to the coordinates of the current block. In Figure 20 the example shown, it is assumed that the motion shift is the motion vector of the A1 block. By applying the motion shift to the current block, the sub-blocks (col sub-blocks) corresponding to the respective sub-blocks configuring the current block in the col picture can be specified. Thereafter, using the motion information of the corresponding sub-blocks (col sub-blocks) in the col picture, the motion information of the respective sub-blocks of the current block can be derived. For example, the motion information of the corresponding sub-block can be obtained from the center position of the corresponding sub-block. In this case, the center position can be the position of the bottom-right sample among the four samples located at the center of the corresponding sub-block. When the motion information of a specific sub-block of the col block corresponding to the current block is not available, the motion information of the center sub-block of the col block can be determined as the motion information of the corresponding sub-block. When deriving the motion vector of the corresponding sub-block, similar to the above-mentioned TMVP process, the reference picture index and the motion vector of the current sub-block can be switched. That is, when deriving the sub-block-based motion vector, the POC of the reference picture of the reference block can be considered to perform the scaling of the motion vector.
[0247] As described above, the sub-block-based TMVP candidates of the current block can be derived using the motion vector field or motion information of the current block derived based on sub-blocks.
[0248] Hereinafter, a merge candidate list configured in units of sub-blocks is defined as a sub-block unit merge candidate list. The above-mentioned affine merge candidates and sub-block-based TMVP candidates can be merged to configure a sub-block unit merge candidate list.
[0249] In addition, a sub-block-based TMVP mode flag specifying whether the sub-block-based TMVP mode is applicable to the current block can be defined, which can be signaled at at least one level among higher levels (e.g., sequence, picture, slice, tile, tile group, patch, etc.) of the current block. For example, the sub-block-based TMVP mode flag can be named sps_sbtmvp_enabled_flag. When the sub-block-based TMVP mode is applicable to the current block, the sub-block-based TMVP candidates can be first added to the sub-block unit merge candidate list, and then the affine merge candidates can be added to the sub-block unit merge candidate list. In addition, the maximum number of candidates that can be included in the sub-block unit merge candidate list can be signaled. For example, the maximum number of candidates that can be included in the sub-block unit merge candidate list can be 5.
[0250] The size of the sub - block for deriving the sub - block unit merge candidate list can be signaled or preset to M×N. For example, M×N can be 8×8. Thus, the affine mode or the sub - block - based TMVP mode is applicable to the current block only when the size of the current block is 8×8 or larger.
[0251] Hereinafter, embodiments of the prediction execution method of the present disclosure will be described. It can be performed in Figure 4 step S410 of Figure 6 or step S630 of
[0252] The prediction block of the current block can be generated based on the motion information derived according to the prediction mode. The prediction block (predicted block) can include the prediction samples (prediction sample array) of the current block. When the motion vector of the current block specifies partial sample units, an interpolation process can be performed, and thus, the prediction samples of the current block can be derived based on the reference samples in units of partial samples within the reference picture. When affine inter - prediction is applied to the current block, the prediction samples can be generated based on the sample / sub - block unit MV. When dual prediction is applied, the prediction samples derived from the weighted sum or weighted average (according to the phase) of the prediction samples derived based on L0 prediction (i.e., prediction using MVL0 and the reference picture within the reference picture list L0) and the prediction samples derived based on L1 prediction (i.e., prediction using MLV1 and the reference picture within the reference picture list L1) can be used as the prediction samples of the current block. When dual prediction is applied and the reference picture for L0 prediction and the reference picture for L1 prediction are in different temporal directions with respect to the current picture (i.e., if it corresponds to dual prediction and bidirectional prediction), this can be referred to as true dual prediction.
[0253] In the image decoding device, the reconstructed samples and the reconstructed picture can be generated based on the derived prediction samples, and then the in - loop filtering process can be performed. Additionally, in the image encoding device, the residual samples can be derived based on the derived prediction samples, and the encoding of the image information including the prediction information and the residual information can be performed.
[0254] Dual prediction with CU - level weight (BCW)
[0255] When dual prediction is applied to the current block as described above, a predicted sample can be derived based on a weighted average. Conventionally, a dual prediction signal (i.e., a dual prediction sample) can be derived by a simple average of an L0 prediction signal (L0 prediction sample) and an L1 prediction signal (L1 prediction sample). That is, the dual prediction sample is derived by averaging an L0 prediction sample based on an L0 reference picture and MVL0 and an L1 prediction sample based on an L1 reference picture and MVL1. However, according to the present disclosure, when dual prediction is applied, the dual prediction signal (dual prediction sample) can be derived by a weighted average of the L0 prediction signal and the L1 prediction signal as follows.
[0256] [Equation 3]
[0257] P bi-pred = ((8 - w) * P0 + w * P1 + 4) >> 3
[0258] In Equation 3 above, P bi-pred represents a dual prediction signal (dual prediction block) derived by a weighted average, and P0 and P1 represent an L0 prediction sample (L0 prediction block) and an L1 prediction sample (L1 prediction block), respectively. Additionally, (8 - w) and w represent weights applied to P0 and P1, respectively.
[0259] When generating a dual prediction signal by a weighted average, five weights can be allowed. For example, the weight w can be selected from {-2, 3, 4, 5, 10}. For each dual prediction CU, the weight w can be determined by one of two methods. As the first method of these two methods, when the current CU is not in merge mode (non-merge CU), the weight index can be signaled together with the motion vector difference. For example, the bitstream can include information about the weight index after the information about the motion vector difference. As the second method of these two methods, when the current CU is in merge mode (merge CU), the weight index can be derived from neighboring blocks based on the merge candidate index (merge index).
[0260] Generating a dual prediction signal by a weighted average can be restricted to be applied only to CUs having a size including 256 or more samples (luminance component samples). That is, dual prediction by a weighted average can be performed only for CUs where the product of the width and height of the current block is 256 or greater. Additionally, the weight w can be used as one of the five weights as described above, and one of different numbers of weights can be used. For example, according to the characteristics of the current image, five weights can be used for low-latency pictures, and three weights can be used for non-low-latency pictures. In this case, the three weights can be {3, 4, 5}.
[0261] By applying a fast search algorithm, an image encoding device can determine a weight index without significantly increasing the complexity. In this case, the fast search algorithm can be summarized as follows. Hereinafter, unequal weights may mean that the weights applied to P0 and P1 are not equal. Additionally, equal weights may mean that the weights applied to P0 and P1 may be equal.
[0262] - When applied together with the AMVR mode in which the resolution of the motion vector adaptively changes, when the current picture is a low-latency picture, unequal weights may be conditionally checked only for each of the 1-pixel motion vector resolution and the 4-pixel motion vector resolution.
[0263] - When applied together with the affine mode and the affine mode is selected as the best mode for the current block, the image encoding device may perform affine motion estimation (ME) for each unequal weight.
[0264] - When the two reference pictures for bi-prediction are equal, unequal weights may be conditionally checked only.
[0265] - When a predetermined condition is satisfied, unequal weights may not be checked. The predetermined picture may be based on the POC distance, quantization parameter (QP), temporal level, etc. between the current picture and the reference picture.
[0266] The weight index of BCW may be encoded using one context encoding bin and one or more subsequent bypass encoding bins. The first context encoding bin specifies whether equal weights are used. When unequal weights are used, additional bins may be bypass-encoded and signaled. The additional bins may be signaled to specify which weight is used.
[0267] Weighted prediction (WP) is a tool for efficiently encoding images including fade. According to weighted prediction, weighted parameters (weights and offsets) may be signaled for each reference picture included in each of the reference picture lists L0 and L1. Then, when motion compensation is performed, the weights and offsets may be applied to the corresponding reference pictures. Weighted prediction and BCW may be used for different types of images. To avoid interaction between weighted prediction and BCW, for a CU using weighted prediction, the BCW weight index may not be signaled. In this case, the weight may be inferred as 4. That is, equal weights may be applied.
[0268] In the case of a CU to which the merge mode is applied, the weight index may be inferred from neighboring blocks based on the merge candidate index. This may be applied to both the general merge mode and the inherited affine merge mode.
[0269] When constructing the affine merge mode, the affine motion information can be configured based on the motion information of up to three blocks. In this case, for the CU using the constructed affine merge mode, the following process can be performed to derive the BCW weight index.
[0270] (1) First, the range of BCW weight indices {0, 1, 2, 3, 4} can be divided into three groups: {0}, {1, 2, 3}, and {4}. When the BCW weight indices of all CPs are from the same group, the BCW weight index can be derived through step (2) below. Otherwise, the BCW weight index can be set to 2.
[0271] (2) When at least two CPs have the same BCW weight index, the same BCW weight index can be assigned as the weight index of the constructed affine merge candidate. Otherwise, the weight index of the constructed affine merge candidate can be set to 2.
[0272] Bidirectional Optical Flow (BDOF)
[0273] According to the present disclosure, BDOF can be used to refine the dual prediction signal. BDOF generates a prediction sample by calculating refined motion information when dual prediction is applied to the current block (e.g., CU). Therefore, the process of calculating refined motion information by applying BDOF can be included in the above motion information derivation step.
[0274] For example, BDOF can be applied at the 4×4 sub-block level. That is, BDOF can be performed for each 4×4 sub-block within the current block.
[0275] For example, BODF can be applied to CUs that meet the following conditions.
[0276] 1) The height of the CU is not 4 and the size of the CU is not 4×8
[0277] 2) The CU is not in the affine mode or the ATMVP merge mode
[0278] 3) The CU is encoded in the true dual prediction mode, that is, one of the two reference pictures is before the current picture in time order, and the other is after the current picture in time order
[0279] In addition, BDOF can be applied only to the luma component. However, the present disclosure is not limited thereto, and BDOF can be applied to the chroma component or both the luma component and the chroma component.
[0280] The BDOF mode is based on the concept of optical flow. That is, it is assumed that the motion of the object is smooth. When BDOF is applied, for each 4×4 sub-block, the motion refinement (v x , v y) Motion refinement can be calculated by minimizing the difference between the L0 prediction sample and the L1 prediction sample. Motion refinement can be used to adjust the dual prediction sample values within a 4×4 sub-block.
[0281] Hereinafter, the process of performing BDOF will be described in more detail.
[0282] First, the horizontal gradient and the vertical gradient of two prediction signals can be calculated. In this case, k can be 0 or 1. The gradient can be calculated by directly calculating the difference between two adjacent samples as shown in Equation 4 below.
[0283] [Equation 4]
[0284]
[0285] In Equation 4 above, I (k) (i, j) represents the sample value at the coordinates (i, j) of the prediction signal in list k (k = 0, 1). For example, I (0) (i, j) can represent the sample value at position (i, j) in the L0 prediction block, and I (1) (i, j) can represent the sample value at position (i, j) in the L1 prediction block.
[0286] In Equation 4 above, the difference between two samples is shifted right by 4. However, the present disclosure is not limited thereto, and shift1 can be determined based on the bit depth of the luminance component. For example, when the bit depth of the luminance component is bitDepth, shift1 can be determined as max(6, bitDepth - 6), or can be simply determined as a fixed value 6. In Equation 4 above, for gradient calculation, first the difference between two samples is calculated, and then a right shift operation is applied to the difference. However, the present disclosure is not limited thereto, and the gradient can be calculated by applying a right shift operation to two sample values and then calculating the difference between the right shifted values.
[0287] As described above, after calculating the gradient, the autocorrelation and cross-correlation S1, S2, S3, S5, and S6 between the gradients can be calculated as follows.
[0288] [Equation 5]
[0289] S1 = ∑ (i,j)∈Ω ψ x (i, j)·ψ x (i, j), S3 = ∑ (i,j)∈Ω θ(i, j)·ψ x (i, j)
[0290]
[0291] where
[0292]
[0293] θ(i,j) = (I (1) (i,j) >> n b ) - (I (0) (i,j) >> n b )
[0294] where Ω is a 6×6 window around the 4×4 sub-block.
[0295] The motion refinement (v x , v y ) can be derived as follows using the autocorrelation and cross-correlation between the above gradients.
[0296] [Equation 6]
[0297]
[0298]
[0299] where th′ BIO = 2 13-BD . is the floor function.
[0300] Based on the derived motion refinement and gradients, the following adjustment can be performed for each sample in the 4×4 sub-block.
[0301] [Equation 7]
[0302]
[0303] Finally, the predicted sample pred of the CU to which BDOF is applied can be calculated by adjusting the dual-predicted samples of the CU as follows BDOF .
[0304] [Equation 8]
[0305] pred BDOF (x,y) = (I (0) (x,y) + I (1) (x,y) + b(x,y) + o offset ) >> shift
[0306] In the above equation, n a , n b and n S2 can be 3, 6, and 12 respectively. These values can be selected such that the multiplier does not exceed 15 bits in BDOF processing and the bit-width of the intermediate parameters is maintained within 32 bits.
[0307] To derive the gradient value, prediction samples I that exist outside the current CU in the list k (k = 0, 1) can be generated (k) (i, j). Figure 21 is a view of the CU that exemplifies the extension to perform BDOF.
[0308] As Figure 21 shown, to perform BDOF, rows / columns extending around the boundary of the CU can be used. To control the computational complexity of generating prediction samples outside the boundary, prediction samples in the extended region ( Figure 21 the white region in) can be generated using a bilinear filter, and prediction samples in the CU ( Figure 21 the gray region in) can be generated using a normal 8-tap motion compensation interpolation filter. The sample values at the extended positions can be used only for gradient calculation. When sample values and / or gradient values outside the CU boundary are needed to perform the remaining steps of the BDOF process, the nearest neighbor sample values and / or gradient values can be filled (repeated) and used.
[0309] When the width and / or height of the CU is greater than 16 luma samples, the corresponding CU can be divided into sub-blocks with a width and / or height of 16 luma samples. The boundaries of the sub-blocks can be processed in the same manner as the CU boundary described above. The maximum unit size for performing the BDOF process can be limited to 16×16.
[0310] When BCW is available for the current block, for example, when the BCW weight index specifies unequal weights, BDOF may not be applied. Similarly, when WP is available for the current block, for example, when luma_weight_lx_flag of at least one of the two reference pictures is 1, BDOF may not be applied. In this case, luma_weight_lx_flag can be information specifying the weighting factor of the WP for the luma component indicating whether there is an lx prediction (x is 0 or 1) in the bitstream or information specifying whether WP is applied to the luma component of the lx prediction. When the CU is encoded in the SMVD mode, BDOF may not be applied.
[0311] Optical Flow Prediction Refinement (PROF)
[0312] Hereinafter, a method for refining a sub-block-based affine motion compensation prediction block by applying optical flow will be described. The prediction samples generated by performing sub-block-based affine motion compensation can be refined based on the difference derived from the optical flow equation. The refinement of these prediction samples in the present disclosure can be referred to as optical flow prediction refinement (PROF). Through PROF, inter-frame prediction at the pixel-level granularity can be achieved without increasing the bandwidth of memory access.
[0313] The parameters of the affine motion model can be used to derive the motion vectors of each pixel in the CU. However, since pixel-based affine motion compensation prediction results in high complexity and an increase in memory access bandwidth, sub-block-based affine motion compensation prediction can be performed. When performing sub-block-based affine motion compensation prediction, the CU can be divided into 4×4 sub-blocks, and motion vectors can be determined for each sub-block. In this case, the motion vectors of each sub-block can be derived from the CPMV of the CU. Sub-block-based affine motion compensation has a trade-off relationship between coding efficiency and complexity and memory access bandwidth. Since motion vectors are derived on a sub-block basis, the complexity and memory access bandwidth are reduced, but the prediction accuracy is reduced.
[0314] Therefore, optical flow can be applied to sub-block-based affine motion compensation prediction to achieve motion compensation with a refined granularity through refinement.
[0315] As described above, the luminance prediction samples can be refined by adding the difference derived from the optical flow equation after performing sub-block-based affine motion compensation. More specifically, PROF can be performed in the following four steps.
[0316] Step 1) Generate a predicted sub-block I(i,j) by performing sub-block-based affine motion compensation.
[0317] Step 2) Calculate the spatial gradients g x (i,j) and g y (i,j) of the predicted sub-block at each sample position. In this case, a 3-tap filter can be used, and the filter coefficients can be [-1, 0, 1]. For example, the spatial gradients can be calculated as follows.
[0318] [Equation 9]
[0319] g x (i,j) = I(i + 1,j) - I(i - 1,j)
[0320] g y (i,j) = I(i,j + 1) - I(i,j - 1)
[0321] To calculate the gradients, the predicted sub-block can be extended by one pixel on each side. In this case, to reduce the memory bandwidth and complexity, the pixels at the extended boundaries can be copied from the nearest integer pixels in the reference picture. Therefore, additional interpolation of the padding area can be skipped.
[0322] Step 3) Calculate the luminance prediction refinement (ΔI(i,j)) through the optical flow equation. For example, the following formula can be used.
[0323] [Equation 10]
[0324] ΔI(i,j) = g x (i,j) * Δv x (i,j) + g y (i,j) * Δv y (i,j)
[0325] In the above formula, Δv(i,j) represents the difference between the pixel motion vector (pixel MV, v(i,j)) calculated at the sample position (i,j) and the sub-block MV of the sub-block to which the sample (i,j) belongs.
[0326] Figure 22 is a view illustrating the relationship between Δv(i,j), v(i,j), and the sub-block motion vector.
[0327] In Figure 22 the example shown, for example, the difference between the motion vector v(i,j) at the upper left sample position of the current sub-block and the motion vector v SB of the current sub-block can be represented by a thick dashed arrow, and the vector represented by the thick dashed arrow can correspond to Δv(i,j).
[0328] The affine model parameters and pixel positions from the center of the sub-block do not change. Therefore, Δv(i,j) can be calculated only for the first sub-block and reused for other sub-blocks in the same CU. Assuming that the horizontal and vertical offsets from the pixel position to the center of the sub-block are x and y respectively, Δv(x,y) can be derived as follows.
[0329] [Equation 11]
[0330]
[0331] For the 4-parameter affine model,
[0332]
[0333] For the 6-parameter affine model,
[0334]
[0335] In the above, (v 0x , v 0y ), (v 1x , v 1y ) and (v 2x , v 2y ) correspond to the upper left CPMV, upper right CPMV, and lower left CPMV respectively, and w and h represent the width and height of the CU respectively.
[0336] Step 4) Finally, the final prediction block I'(i,j) can be generated based on the calculated luminance prediction refinement ΔI(i,j) and the prediction sub-block I(i,j). For example, the final prediction block I' can be generated as follows.
[0337] [Equation 12]
[0338] I′(i,j) = I(i,j) + ΔI(i,j)
[0339] As described above, by applying BDOF in the inter-frame prediction process to refine the reference samples in the motion compensation process, the compression performance of the image can be increased. BDOF can be executed in the normal mode. That is, BDOF is not executed in the case of the affine mode, GPM mode, or CIIP mode.
[0340] As a method similar to BDOF, PROF can be performed on the blocks encoded in the affine mode. As described above, by refining the reference samples in each 4×4 sub-block via PROF, the compression performance of the image can be increased.
[0341] Since both PROF and BDOF use the characteristics of optical flow, whether to apply PROF can be determined according to conditions similar to the application conditions of BDOF. According to the present disclosure, various embodiments of WP and BCW can be provided.
[0342] In the present disclosure, setting or deriving any information (e.g., a flag) as true can mean that the corresponding information is derived as a first value (e.g., "1"). Additionally, when any information is set as true, it can indicate that the process specified by the corresponding information (e.g., BDOF, PROF, WP, etc.) is executed. Conversely, in the present disclosure, setting or deriving any information (e.g., a flag) as false can mean that the corresponding information is derived as a second value (e.g., "0"). Additionally, when any information is set as false, it can indicate that the process specified by the corresponding information is not executed.
[0343] When the following various conditions are met, BDOF can be executed in the motion compensation process by setting bdofFlag as true.
[0344] [Table 1]
[0345]
[0346] The conditions for executing BDOF described in Table 1 above can be as described in Table 2 below.
[0347] [Table 2]
[0348]
[0349] However, the conditions for performing BDOF are not limited to the examples in Table 1 and Table 2 above, and some conditions can be omitted. Additionally, conditions other than the above-mentioned conditions can be considered.
[0350] When BDOF is not applied to the current block according to the above conditions, for example, when the prediction mode of the current block is the affine mode, PROF is similar to the application of BDOF. For example, when the prediction mode of the current block is the affine mode, it can be determined whether to apply PROF (cbProfFlagLX), and when cbProfFlagLX is true, PROF can be executed.
[0351] In BDOF, the characteristics of the optical flow are used to determine the offset of the samples. Therefore, when the luminance values of the reference pictures are different, that is, when BCW or weighted prediction (WP) is applied, BDOF is not executed. However, although the characteristics of the optical flow are used to derive the offset of the samples, PROF can be executed regardless of whether BCW or WP is applied.
[0352] According to an embodiment of the present disclosure, from a design perspective, in order to coordinate between BDOF and PROF, PROF may not be applied to the blocks to which BCW or WP is applied. For example, when BcwIdx is not 0, when luma_weight_l0_flag[refIdxL0] is 1, or when luma_weight_l1_flag[refIdxL1] is 1, the information cbProfFlagLX specifying whether to apply PROF can be set to false. BcwIdx not being 0 may mean that BCW is applied to the current block, and luma_weight_lX_flag[refIdxLX] (X = 0 or 1) being 1 may mean that WP is applied to the current block. In the present disclosure, BcwIdx being 0 may mean that equal weights are applied, that is, a dual-prediction block is generated by the average sum of the L0 prediction block and the L1 prediction block. Therefore, when setting cbProfFlagLX, if BCW or WP is applied to the current block, control can be executed by adding the above conditions not to apply PROF.
[0353] The following table shows an example of setting cbProfFlagLX according to the present disclosure, and the underlined part shows the added conditions.
[0354] [Table 3]
[0355]
[0356] The following table shows another example of setting cbProfFlagLX according to the present disclosure, and the underlined part shows the added conditions.
[0357] [Table 4]
[0358]
[0359] As described above, it can be determined whether to apply PROF to the current block. For example, cbProfFlagLX (X = 0 or 1) can indicate whether to apply PROF to the L0 prediction direction or the L1 prediction direction, and according to the method in Table 3, cbProfFlagLX can be determined based on at least one of bcwIdx, luma_wighted_l0_flag, and / or luma_wighted_l1_flag. As another example, according to the method in Table 4, cbProfFlagLX can be determined based on at least one of bcwIdx, slice_type, pps_weighted_pred_flag, and / or pps_weighted_bipred_flag. slice_type specifies the slice type of the current slice to which the current picture belongs, pps_weighted_pred_flag is a picture parameter set (PPS) parameter that specifies whether to apply WP to the P slice referring to the corresponding PPS, and pps_weighted_bipred_flag is a picture parameter set (PPS) parameter that specifies whether to apply WP to the B slice referring to the corresponding PPS.
[0360] According to another embodiment of the present disclosure, PROF can be applied when performing BCW or WP (explicit weighted prediction).
[0361] Generally, the BCW weight index (bcw_idx) is signaled only when WP is not available. Therefore, the bcw_idx and the weight factor of WP are not signaled simultaneously. The following table shows an example of the syntax structure for signaling the bcw_idx.
[0362] [Table 5]
[0363]
[0364] According to Table 5 above, as a condition for signaling bcw_idx, check whether all weighted prediction flags (e.g., luma_weight_l0_flag, luma_weight_l1_flag, chroma_weight_l0_flag, chroma_weight_l1_flag) of the reference pictures specified by the reference picture indices (ref_idx_l0, ref_idx_l1) of the corresponding CU are 0. Therefore, according to Table 5, even if the WP application flags (e.g., pps_weighted_pred_flag and / or pps_weighted_bipred_flag) sent at the PPS are true, when the weighted prediction flag for a specific reference picture index is 0, bcw_idx can be signaled. That is, according to the example in Table 5, even if WP is applied, bcw_idx can be signaled.
[0365] Figure 23 FIG. is a flowchart illustrating an example of performing PROF, BCW, WP, and / or averaging according to the present disclosure.
[0366] Refer to Figure 23 , first, a cbProfFlag (e.g., cbProfFlagLX) that specifies whether PROF is applied to the current block can be derived (S2310). The cbProfFlag can be derived based on various methods of the present disclosure.
[0367] Thereafter, in step S2320, it can be checked whether the cbProfFlag is true, and if true, PROF can be performed on the current block in step S2330. PROF can be performed according to the above method, and as a result of performing PROF, refined prediction samples of the current block can be obtained. In step S2320, when the cbProfFlag is false, step S2330 can be skipped.
[0368] Thereafter, in step S2340, a weightedPredFlag that specifies whether weighted prediction (WP) is applied to the current block can be derived, and its value can be checked whether it is true. The method for deriving the weightedPredFlag will be described later. When the weightedPredFlag is true, it can be determined that weighted prediction is applied to the current block, and weighted prediction can be performed on the current block (S2350). The weighted prediction of the current block can be performed based on the weighted parameters (weights and offsets) of the reference picture of the current block. As described above, the weighted parameters of the reference picture can be explicitly signaled through the bitstream.
[0369] When weightedPredFlag is false, it can be determined that weighted prediction is not applied to the current block and it can be checked whether bcwIdx is 0 (S2360). bcwIdx can be derived differently according to the prediction mode of the current block. For example, when the prediction mode of the current block is the skip mode or the merge mode, the bcwIdx of the current block can be derived as the bcwIdx of the merge candidate specified by the merge candidate index of the current block. When the prediction mode of the current block is not the merge mode (e.g., the MVP mode), the bcwIdx of the current block can be reconstructed by parsing the syntax element bcw_idx signaled through the bitstream. If bcw_idx is not signaled through the bitstream, the value of bcwIdx can be inferred as 0. When bcwIdx is 0, it can be specified that BCW may not be applied to the current block. As described above, bcwIdx being 0 can mean applying equal weights, that is, meaning generating a bi-prediction block from the average sum of the L0 prediction block and the L1 prediction block.
[0370] In step S2360, when bcwIdx is not 0, it can be determined that BCW is applied to the current block, and BCW can be performed on the current block based on the weights specified by bcwIdx (S2370). In step S2360, when bcwIdx is 0, it can be determined that BCW is not applied to the current block and an average sum can be performed on the current block (S2380).
[0371] Table 6 shows an example of deriving weightedPredFlag according to the present disclosure and performing WP or BCW accordingly.
[0372] [Table 6]
[0373]
[0374] According to the method of Table 6, weightedPredFlag can be derived based on the slice type of the current slice to which the current block belongs and the WP application flags signaled through the PPS (e.g., pps_weighted_pred_flag, pps_weighted_bipred_flag). Specifically, when the slice type of the current block is a P slice, weightedPredflag can be determined as the value of pps_weighted_pred_flag. Additionally, when the slice type of the current block is a B slice, weightedPredflag can be determined as the value of pps_weighted_bipred_flag.
[0375] According to the method of Table 6, when the weightedPredFlag determined as described above is false, default weighted prediction is performed, and when bcwIdx is not 0, BCW, or when bcwIdx is 0, average sum can be performed. Additionally, when weightedPredFlag is true, explicit weighted prediction can be performed based on the signaled weighted parameters.
[0376] According to the method of Table 6, weightedPredFlag is determined only by the slice type of the current slice and PPS information rather than the weighted prediction flags of individual reference pictures (e.g., luma_weight_l0_flag, luma_weight_l1_flag, chroma_weight_l0_flag, chroma_weight_l1_flag). However, bcw_idx can be signaled based on the weighted prediction flags of individual reference pictures (e.g., luma_weight_l0_flag, luma_weight_l1_flag, chroma_weight_l0_flag, chroma_weight_l1_flag), as shown in Table 5.
[0377] Therefore, even when weightedPredFlag is determined to be true according to the method of Table 6, bcw_idx can be signaled. In this case, even when bcw_idx is not 0 (default), WP (explicit weighted prediction) is always performed.
[0378] Refer to Figure 23 The embodiments described with reference to Table 5 and Table 6 have the problem of performing WP even when bcw_idx is not the default. Hereinafter, another embodiment of the present disclosure for solving the above problem will be described.
[0379] Figure 24 is a flowchart exemplifying another example of performing PROF, BCW, WP, and / or average sum according to the present disclosure.
[0380] Refer to Figure 24 , first, a cbProfFlag (S2410) that specifies whether to apply PROF to the current block can be derived. The cbProfFlag can be derived based on various methods of the present disclosure.
[0381] Thereafter, in step S2420, it is possible to check whether cbProfFlag is true, and when it is true, in step S2430, PROF can be performed on the current block. PROF can be performed according to the method described above, and as a result of performing PROF, a refined prediction sample of the current block can be obtained. In step S2420, when cbProfFlag is false, step S2430 can be skipped.
[0382] Thereafter, in step S2440, it is possible to check whether bcwIdx is 0. bcwIdx can be derived differently according to the prediction mode of the current block. For example, when the prediction mode of the current block is the skip mode or the merge mode, the bcwIdx of the current block can be derived as the bcwIdx of the merge candidate specified by the merge candidate index of the current block. When the prediction mode of the current block is not the merge mode (e.g., the MVP mode), the bcwIdx of the current block can be reconstructed by parsing the syntax element bcw_idx signaled through the bitstream. If bcw_idx is not signaled through the bitstream, the value of bcwIdx can be inferred as 0. In step S2440, when bcwIdx is not 0, it can be determined to apply BCW to the current block, and BCW can be performed on the current block based on the weight specified by bcwIdx (S2450).
[0383] In step S2440, when bcwIdx is 0, it can be determined not to apply BCW to the current block. Thereafter, in step S2460, it is possible to check whether weightedPredFlag, which specifies whether weighted prediction (WP) is applied to the current block, is true.
[0384] In step S2460, when weightedPredFlag is true, it can be determined to apply weighted prediction to the current block, and weighted prediction can be performed on the current block (S2470). The weighted prediction of the current block can be performed based on the weighted parameters (weights and offsets) of the reference picture of the current block. As described above, the weighted parameters of the reference picture can be explicitly signaled through the bitstream.
[0385] When weightedPredFlag is false in step S2460, it can be determined not to apply weighted prediction to the current block, and average sum can be performed on the current block (S2480).
[0386] According to the embodiment described with reference to Figure 24 When both explicit weighted prediction (WP) and BCW are applicable to the current block, BCW can be preferentially applied.
[0387] Table 7 shows another example of deriving weightedPredFlag according to the present disclosure and performing WP or BCW accordingly.
[0388] [Table 7]
[0389]
[0390] The method for deriving weightedPredFlag is the same as the methods of Table 6 and Table 7, and thus its detailed description will be omitted. According to the method of Table 7, when the weightedPredFlag determined as described above is false or bcwIdx is not 0, default weighted prediction can be performed, and BCW or average sum can be performed according to the value of bcwIdx. Additionally, when weightedPredFlag is true and bcwIdx is 0, explicit weighted prediction can be performed based on the signaled weighted parameters.
[0391] According to the method of Table 7, when weightedPredFlag is true and bcwIdx is not 0, i.e., when both explicit weighted prediction (WP) and BCW are applicable to the current block, BCW can be preferentially applied.
[0392] Figure 25 is a flowchart illustrating an example of performing BCW or WP according to the method of Table 7.
[0393] First, the slice type of the current slice to which the current block belongs can be determined (S2510). When the slice type is a P slice, weightedPredFlag can be derived as pps_weighted_pred_flag (S2520). When the slice type is a B slice, weightedPredFlag can be derived as pps_weighted_bipred_flag (S2530).
[0394] Thereafter, in step S2540, it can be determined whether weightedPredFlag is 0 or bcwIdx is not 0. bcwIdx can be derived differently according to the prediction mode of the current block. For example, when the prediction mode of the current block is the skip mode or the merge mode, the bcwIdx of the current block can be derived as the bcwIdx of the merge candidate specified by the merge candidate index of the current block. When the prediction mode of the current block is not the merge mode (e.g., the MVP mode), the bcwIdx of the current block can be reconstructed by parsing the syntax element bcw_idx signaled through the bitstream. If bcw_idx is not signaled through the bitstream, the value of bcwIdx can be inferred as 0.
[0395] When weightedPredFlag is 0 or when BcwIdx is not 0, default weighted prediction (S2550) can be performed. In this case, when BcwIdx is 0, the average sum described in step S2380 or S2480 can be performed. When BcwIdx is not 0, the BCW described in step S2370 or step S2450 can be performed.
[0396] When weightedPredFlag is not 0 and BcwIdx is 0, explicit weighted prediction (S2560) can be performed. In this case, the WP described in step S2350 or step S2470 can be performed.
[0397] Table 8 shows another example of deriving weightedPredFlag according to the present disclosure and accordingly performing WP or BCW.
[0398] [Table 8]
[0399]
[0400] According to the method of Table 8, weightedPredFlag can be derived by further considering bcwIdx. Specifically, when the bcwIdx of the current block is not 0, weightedPredFlag can be derived as false. When the bcwIdx of the current block is 0, weightedPredFlag can be derived based on the slice type of the current slice to which the current block belongs and the WP application flag signaled by PPS (e.g., pps_weighted_pred_flag, pps_weighted_bipred_flag) according to the method of Table 6. According to the method of Table 8, when weightedPredFlag determined as described above is false (the second value, e.g., 0), default weighted prediction is performed, and BCW or average sum can be performed according to the value of bcwIdx. Additionally, when weightedPredFlag is true (the first value, e.g., 1), explicit weighted prediction can be performed based on the signaled weighted parameters.
[0401] According to the method of Table 8, when bcwIdx is not 0, by deriving weightedPredFlag as false, when both explicit weighted prediction (WP) and BCW are applicable to the current block, BCW can be preferentially applied.
[0402] Figure 26 is a flowchart illustrating an example of performing BCW or WP according to the method of Table 8.
[0403] First, it can be determined whether BcwIdx is not 0 (S2610). BcwIdx can be derived differently according to the prediction mode of the current block. For example, when the prediction mode of the current block is the skip mode or the merge mode, the bcwIdx of the current block can be derived as the bcwIdx of the merge candidate specified by the merge candidate index of the current block. When the prediction mode of the current block is not the merge mode (e.g., the MVP mode), the bcwIdx of the current block can be reconstructed by parsing the syntax element bcw_idx signaled through the bitstream. If bcw_idx is not signaled through the bitstream, the value of bcwIdx can be inferred as 0.
[0404] When BcwIdx is not 0, weightedPredFlag can be derived as 0 (S2620). Thereafter, default weighted prediction can be performed through the determination in step S2660 (S2670). Alternatively, step S2620 and step S2660 can be skipped, and step S2670 can be immediately performed. Step S2670 is performed in the same manner as step S2550, so its detailed description will be omitted.
[0405] When BcwIdx is 0, the slice type of the current slice to which the current block belongs can be determined (S2630). When the slice type is a P slice, weightedPredFlag can be derived as pps_weighted_pred_flag (S2640). When the slice type is a B slice, weightedPredFlag can be derived as pps_weighted_bipred_flag (S2650).
[0406] Thereafter, in step S2660, it can be determined whether weightedPredFlag is 0.
[0407] When weightedPredFlag is 0, default weighted prediction can be performed (S2670). When weightedPredFlag is not 0, explicit weighted prediction can be performed (S2680). Step S2670 and step S2680 are performed in the same manner as step S2550 and step S2560 respectively, so their detailed descriptions will be omitted.
[0408] Table 9 shows another example of deriving weightedPredFlag according to the present disclosure and accordingly performing WP or BCW.
[0409] [Table 9]
[0410]
[0411] According to the method in Table 9, the weightedPredFlag can be derived by considering the slice type of the current slice to which the current block belongs and the weighted prediction flags (e.g., luma_weight_l0_flag, luma_weight_l1_flag, chroma_weight_l0_flag, chroma_weight_l1_flag) of the reference pictures specified by the reference picture indexes (ref_idx_l0, ref_idx_l1) of the current block. Specifically, when the current slice to which the current block belongs is a P slice and both weighted prediction flags in the L0 direction (e.g., luma_weight_l0_flag, chroma_weight_l0_flag) are 0, the weightedPredFlag can be derived as 0. Additionally, when the current slice to which the current block belongs is a P slice and at least one of the weighted prediction flags in the L0 direction (e.g., luma_weight_l0_flag, chroma_weight_l0_flag) is not 0, the weightedPredFlag can be derived as 1.
[0412] When the current slice to which the current block belongs is a B slice and all weighted prediction flags in the L0 direction and the L1 direction (e.g., luma_weight_l0_flag, luma_weight_l1_flag, chroma_weight_l0_flag, chroma_weight_l1_flag) are 0, the weightedPredFlag can be derived as 0. Additionally, when the current slice to which the current block belongs is a B slice and at least one of the weighted prediction flags in the L0 direction and the L1 direction (e.g., luma_weight_l0_flag, luma_weight_l1_flag, chroma_weight_l0_flag, chroma_weight_l1_flag) is not 0, the weightedPredFlag can be derived as 1.
[0413] According to the method in Table 9, when the weightedPredFlag determined as described above is false (the second value, e.g., 0), default weighted prediction can be performed, and BCW or average sum can be performed according to the value of bcwIdx. Additionally, when the weightedPredFlag is true (the first value, e.g., 1), explicit weighted prediction can be performed based on the signaled weighted parameters.
[0414] According to the method of Table 9, weightedPredFlag is derived based on the condition of signaling bcw_idx. That is, when bcw_idx is signaled, weightedPredFlag is derived as false, and when both explicit weighted prediction (WP) and BCW are applicable to the current block, BCW can be preferentially applied.
[0415] Figure 27 It is a flowchart illustrating an example of performing BCW or WP according to the method of Table 9.
[0416] First, the slice type of the current slice to which the current block belongs can be determined (S2710). When the slice type is a P slice, weightedPredFlag can be derived based on luma_weight_l0_flag and / or chroma_weight_l0_flag as described with reference to Table 9 (S2720). When the slice type is a B slice, weightedPredFlag can be derived based on luma_weight_l0_flag, chroma_weight_l0_flag, luma_weight_l1_flag, and / or chroma_weight_l1_flag (S2730).
[0417] Thereafter, in step S2740, it can be determined whether weightedPredFlag is 0.
[0418] When weightedPredFlag is 0, default weighted prediction can be performed (S2750). When weightedPredFlag is not 0, explicit weighted prediction can be performed (S2760). Step S2750 and step S2760 are respectively performed in the same manner as step S2550 and step S2560, so their detailed descriptions will be omitted.
[0419] Table 10 shows another example of deriving weightedPredFlag according to the present disclosure and accordingly performing WP or BCW.
[0420] [Table 10]
[0421]
[0422] The method for deriving weightedPredFlag is the same as the methods in Table 9 and Table 10, and thus its detailed description will be omitted. According to the method in Table 10, when the weightedPredFlag determined as above is false (a second value, e.g., 0) or bcwIdx is not 0, default weighted prediction can be performed, and BCW or average sum can be performed according to the value of bcwIdx. Additionally, when weightedPredFlag is true (a first value, e.g., 1) and bcwIdx is 0, explicit weighted prediction can be performed based on the signaled weighted parameter.
[0423] According to the method in Table 10, weightedPredFlag is derived based on the condition of signaling bcw_idx. That is, when bcw_idx is signaled, weightedPredFlag is derived as false, and when both explicit weighted prediction (WP) and BCW are applicable to the current block, BCW can be preferentially applied.
[0424] Additionally, according to the method in Table 10, when weightedPredFlag is true and bcwIdx is not 0, that is, when both explicit weighted prediction (WP) and BCW are applicable to the current block, BCW can be preferentially applied.
[0425] Figure 28 is a flowchart illustrating an example of performing BCW or WP according to the method in Table 10.
[0426] Figure 28 Steps S2810 to S2830 of Figure 27 are the same as steps S2710 to S2730 of
[0427] Refer to Figure 28 Hereafter, in step S2840, it can be determined whether weightedPredFlag is 0 or bcwIdx is not 0, and default weighted prediction in step S2850 or explicit weighted prediction in step S2860 can be performed based on the determination result. Figure 28 Steps S2840 to S2860 of Figure 25 are the same as steps S2540 to S2560 of
[0428] As described above, the embodiments described with reference to Figure 23 as well as Tables 5 and 6 have the problem of performing WP even when bcw_idx is not the default. Hereinafter, another embodiment of the present disclosure for solving the above problem will be described.
[0429] PROF can be applied even when BCW or WP (explicit weighted prediction) is executed. Generally, since bcw_idx is parsed from the bitstream only when WP is not available, BCW and WP should not coexist. However, as shown in the syntax structure in Table 5, in order to parse bcw_idx from the bitstream, only the weighted prediction flags (e.g., luma_weight_l0_flag, luma_weight_l1_flag, chroma_weight_l0_flag, chroma_weight_l1_flag) of the reference pictures specified by the reference picture index of the current block are checked. Additionally, as shown in Table 6, the weightedPredFlag is derived based on the slice type of the current slice to which the current block belongs and the WP application flag signaled by PPS (e.g., pps_weighted_pred_flag, pps_weighted_bipred_flag). Therefore, even when the WP application flag is true, if the weighted prediction flag of a specific reference picture is 0, bcw_idx is parsed. As a result, WP and BCW can be applied simultaneously.
[0430] Table 11 shows a modified syntax structure for parsing bcw_idx according to another example of the present disclosure.
[0431] [Table 11]
[0432]
[0433] In the syntax structure of Table 11, since the parsing condition of bcw_idx is modified considering the derivation condition of weightedPredFlag in Table 6, the situation where WP and BCW are applied simultaneously can be eliminated by deriving weightedPredFlag as false when parsing bcw_idx and deriving weightedPredFlag as true only when bcw_idx is not parsed. For example, the implementation described with reference to Figure 23 and Tables 5 and 6 can solve the above problem by replacing Table 5 with Table 11.
[0434] Figure 29 is a flowchart exemplifying the execution of BDOF, PROF, BCW, WP, and / or averaging based on the values of bdofFlag, weight index BcwIdx, and weightedPredFlag.
[0435] Referring to Figure 29 , first, bdofFlag specifying whether BDOF is applied to the current block can be derived (S2910).
[0436] In step S2920, when bdofFlag is true, it can be determined that BDOF is applied to the current block, and BDOF can be performed on the current block (S2922). Thereafter, average sum can be performed on the current block (S2924).
[0437] Otherwise, when bdofFlag is false in step S2920, it can be determined that BDOF is not applied to the current block, and cbProfFlag that specifies whether PROF is applied to the current block can be derived (S2930). cbProfFlag can be derived based on various methods of the present disclosure.
[0438] Thereafter, in step S2940, it can be checked whether cbProfFlag is true, and when it is true, in step S2950, PROF can be performed on the current block. PROF can be performed according to the above method, and as a result of performing PROF, a refined prediction sample of the current block can be obtained. In step S2940, when cbProfFlag is false, step S2950 can be skipped.
[0439] Thereafter, in step S2960, it can be checked whether bcwIdx is 0. bcwIdx can be derived differently according to the prediction mode of the current block. For example, when the prediction mode of the current block is skip mode or merge mode, bcwIdx of the current block can be derived as the bcwIdx of the merge candidate specified by the merge candidate index of the current block. When the prediction mode of the current block is not merge mode (e.g., MVP mode), bcwIdx of the current block can be reconstructed by parsing the syntax element bcw_idx signaled through the bitstream. If bcw_idx is not signaled through the bitstream, the value of bcwIdx can be inferred as 0. In step S2960, when bcwIdx is not 0, it can be determined that BCW is applied to the current block, and BCW can be performed on the current block based on the weight specified by bcwIdx (S2962).
[0440] In step S2960, when bcwIdx is 0, it can be determined that BCW is not applied to the current block. Thereafter, in step S2970, it can be checked whether weightedPredFlag that specifies whether weighted prediction (WP) is applied to the current block is true.
[0441] In step S2970, when weightedPredFlag is true, it can be determined that weighted prediction is applied to the current block, and weighted prediction can be performed on the current block (S2972). The weighted prediction of the current block can be performed based on the weighted parameters (weights and offsets) of the reference picture of the current block. As described above, the weighted parameters of the reference picture can be explicitly signaled through the bitstream.
[0442] When weightedPredFlag is false in step S2970, it can be determined that weighted prediction is not applied to the current block, and an average sum can be performed on the current block (S2974).
[0443] According to the Figure 29 described embodiment, when both explicit weighted prediction (WP) and BCW are applicable to the current block, BCW can be preferentially applied.
[0444] Although, for clarity of description, the exemplary methods of the present disclosure are shown as a series of operations, it is not intended to limit the order of performing the steps, and these steps can be performed simultaneously or in a different order when necessary. To implement the method according to the present invention, the described steps can further include other steps, can include the remaining steps except for some steps, or can include other additional steps except for some steps.
[0445] In the present disclosure, an image encoding device or an image decoding device that performs a predetermined operation (step) can perform an operation (step) of confirming the execution condition or situation of the corresponding operation (step). For example, if it is described that a predetermined operation is performed when a predetermined condition is satisfied, the image encoding device or the image decoding device can perform the predetermined operation after determining whether the predetermined condition is satisfied.
[0446] The various embodiments of the present disclosure are not a list of all possible combinations and are intended to describe representative aspects of the present disclosure, and the matters described in the various embodiments can be applied independently or in combinations of two or more.
[0447] The various embodiments of the present disclosure can be implemented in hardware, firmware, software, or a combination thereof. In the case where the present disclosure is implemented by hardware, the present disclosure can be implemented by an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a general purpose processor, a controller, a microcontroller, a microprocessor, etc.
[0448] In addition, an image decoding device and an image encoding device according to an embodiment of the present disclosure may be included in a multimedia broadcast transmission and reception device, a mobile communication terminal, a home theater video device, a digital cinema video device, a surveillance camera, a video chat device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camera, a video-on-demand (VoD) service providing device, an over-the-top (OTT) video device, an Internet streaming service providing device, a three-dimensional (3D) video device, a video phone video device, a medical video device, etc., and may be used to process video signals or data signals. For example, an OTT video device may include a game console, a Blu-ray player, an Internet-connected TV, a home theater system, a smart phone, a tablet PC, a digital video recorder (DVR), etc.
[0449] Figure 30 is a view showing a content streaming system to which an embodiment of the present disclosure can be applied.
[0450] As Figure 30 shown, a content streaming system to which an embodiment of the present disclosure is applied may mainly include an encoding server, a streaming server, a network server, a media storage device, a user device, and a multimedia input device.
[0451] The encoding server compresses content input from a multimedia input device such as a smart phone, a camera, a video camera, etc. into digital data to generate a bitstream and sends the bitstream to the streaming server. As another example, when a multimedia input device such as a smart phone, a camera, a video camera, etc. directly generates a bitstream, the encoding server may be omitted.
[0452] The bitstream may be generated by an image encoding method or an image encoding device according to an embodiment of the present disclosure, and the streaming server may temporarily store the bitstream during the process of sending or receiving the bitstream.
[0453] The streaming server sends multimedia data to the user device based on a request from the user via the network server, and the network server serves as a medium for informing the user of the service. When the user requests a required service from the network server, the network server may deliver it to the streaming server, and the streaming server may send multimedia data to the user. In this case, the content streaming system may include a separate control server. In this case, the control server is used to control commands / responses between devices in the content streaming system.
[0454] The streaming server may receive content from the media storage device and / or the encoding server. For example, when receiving content from the encoding server, the content may be received in real time. In this case, in order to provide a smooth streaming service, the streaming server may store the bitstream for a predetermined time.
[0455] Examples of user devices may include mobile phones, smartphones, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head-mounted displays), digital TVs, desktop computers, digital signage, etc.
[0456] Each server in the content streaming system may operate as a distributed server, in which case the data received from each server may be distributed.
[0457] The scope of the present disclosure includes software or machine-executable commands (e.g., operating systems, applications, firmware, programs, etc.) for enabling the operation of methods according to various embodiments to be executed on a device or computer, and non-transitory computer-readable media having such software or commands stored thereon and executable on the device or computer.
[0458] Industrial Applicability
[0459] Embodiments of the present disclosure may be used to encode or decode images.
Claims
1. An image decoding method performed by an image decoding device, the image decoding method comprising the following steps: Receiving a bitstream; Deriving a first flag for bidirectional optical flow BDOF for a current block regarding a picture based on the received bitstream; Specifying to apply the BDOF to the current block based on the first flag: performing BDOF processing on the current block; Specifying not to apply the BDOF to the current block based on the first flag: Checking, based on the received bitstream, (i) a second flag for weighted prediction regarding the current block and (ii) a weight index BcwIdx for performing bi-prediction BCW with CU-level weights on the current block, And determining whether to perform default weighted prediction or explicit weighted prediction on the current block based on the second flag and the weight index BcwIdx; Wherein, based on the second flag being equal to 0 or the weight index BcwIdx not being equal to 0, performing the default weighted prediction on the current block, Wherein, based on the second flag being equal to 1 and the weight index BcwIdx being equal to 0, performing the explicit weighted prediction on the current block, and Wherein, the second flag is determined independently of the weight prediction flag for the reference picture regarding the current block; Wherein, based on the weight index BcwIdx being equal to 0, the default weighted prediction performs an average sum on the current block, Wherein, based on the weight index BcwIdx not being equal to 0, the default weighted prediction performs the BCW on the current block, and Wherein, the explicit weighted prediction is performed based on a weighted parameter for the reference picture regarding the current block, the weighted parameter including a weight sum and an offset.
2. The image decoding method according to claim 1, wherein, The second flag is determined differently based on the slice type of the current slice to which the current block belongs.
3. The image decoding method according to claim 2, Among them, Based on the slice type of the current slice being a P slice, the second flag is derived as the value of pps_weighted_pred_flag signaled in the picture parameter set PPS, and Wherein, based on the slice type of the current slice being a B slice, the second flag is derived as the value of pps_weighted_bipred_flag signaled in the PPS.
4. The image decoding method according to claim 1, Among them, Deriving the weight index BcwIdx based on a syntax element bcw_idx signaled through the bitstream, and Wherein, based on the absence of the syntax element bcw_idx in the bitstream, the weight index BcwIdx is derived as 0.
5. The image decoding method according to claim 4, wherein, Parsing the syntax element bcw_idx from the bitstream based on the weighted prediction flag for the reference picture of the current block.
6. The image decoding method according to claim 4, wherein, Parsing the syntax element bcw_idx from the bitstream based on all weighted prediction flags for the reference picture of the current block being equal to 0.
7. The image decoding method according to claim 1, wherein, Explicitly signaling the weighted parameter through the bitstream.
8. An image encoding method performed by an image encoding device, the image encoding method comprising the following steps: Derive a first flag for the bidirectional optical flow BDOF of the current block with respect to the picture; Based on the first flag, specify to apply the BDOF to the current block: perform BDOF processing on the current block; Based on the first flag, specify not to apply the BDOF to the current block: Check (i) a second flag for weighted prediction with respect to the current block and (ii) a weight index BcwIdx for performing dual prediction BCW with CU-level weights on the current block; and Based on the second flag and the weight index BcwIdx, determine whether to perform default weighted prediction or explicit weighted prediction on the current block, wherein, based on the second flag being equal to 0 or the weight index BcwIdx not being equal to 0, perform the default weighted prediction on the current block, wherein, based on the second flag being equal to 1 and the weight index BcwIdx being equal to 0, perform the explicit weighted prediction on the current block, and wherein, the second flag is determined independently of the weight prediction flag of the reference picture for the current block; Encode information related to the first flag, the second flag, and the weight index BcwIdx, wherein, based on the weight index BcwIdx being equal to 0, the default weighted prediction performs an average sum on the current block, wherein, based on the weight index BcwIdx not being equal to 0, the default weighted prediction performs the BCW on the current block, and wherein, the explicit weighted prediction is performed based on the weighted parameters of the reference picture for the current block, and the weighted parameters include a weight sum and an offset.
9. A method for transmitting a bitstream, the method comprising the steps of: Transmit a bitstream generated by the following steps: Derive a first flag for the bidirectional optical flow BDOF of the current block with respect to the picture; Based on the first flag, specify to apply the BDOF to the current block: perform BDOF processing on the current block; Based on the first flag, specify not to apply the BDOF to the current block: Check (i) a second flag for weighted prediction with respect to the current block and (ii) a weight index BcwIdx for performing dual prediction BCW with CU-level weights on the current block; and Based on the second flag and the weight index BcwIdx, determine whether to perform default weighted prediction or explicit weighted prediction on the current block, wherein, based on the second flag being equal to 0 or the weight index BcwIdx not being equal to 0, perform the default weighted prediction on the current block, wherein, based on the second flag being equal to 1 and the weight index BcwIdx being equal to 0, perform the explicit weighted prediction on the current block, and wherein, the second flag is determined independently of the weight prediction flag of the reference picture for the current block; Encode information related to the first flag, the second flag, and the weight index BcwIdx, wherein, based on the weight index BcwIdx being equal to 0, the default weighted prediction performs an average sum on the current block, Wherein, based on the weight index BcwIdx being not equal to 0, the default weighted prediction performs the BCW on the current block, and Wherein, the explicit weighted prediction is performed based on weighted parameters of a reference picture for the current block, and the weighted parameters include a weight and an offset.