Image decoding and encoding method and data transmission method for image
By deducing the weight index information in bidirectional prediction and generating predictive samples, the problem of low image/video encoding efficiency in the prior art is solved, and efficient image/video compression is achieved.
Patent Information
- Application Number
- CN202510123145.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-14
- Filing Date
- 2020-06-10
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art is difficult to effectively improve the encoding efficiency of images/videos, especially when processing high resolution and high quality images/videos, the transmission and storage costs are high.
By deducing the weight index information about the bidirectional prediction in inter prediction, a merged candidate list of the current block is generated, and based on the selected candidates, the L0 and L1 prediction samples are generated, and the prediction samples of the current block are finally generated based on these samples and weight information.
The overall image/video compression efficiency is improved, motion vector candidates are constructed efficiently, and weight-based bidirectional prediction is realized.
Smart Images

Figure CN119946296A_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application with the original application number 202080051014.X (International application number: PCT / KR2020 / 007523, application date: June 10, 2020, invention name: Image decoding method and device for deriving weight index information for bidirectional prediction). Technical Field
[0002] The present disclosure relates to an image decoding method and apparatus for deriving weight index information about bidirectional prediction. Background Art
[0003] Recently, there is an increasing demand for high-resolution, high-quality images / videos such as 4K or 8K ultra-high definition (UHD) images / videos in various fields. As the image / video resolution or quality becomes higher, relatively more information or bits are transmitted compared to conventional image / video data. Therefore, if the image / video data is transmitted via a medium such as an existing wired / wireless broadband line or is stored in a conventional storage medium, the cost of transmission and storage is easily increased.
[0004] In addition, there is growing interest and demand for virtual reality (VR) and artificial reality (AR) content and immersive media such as holograms; and there is also growing broadcasting of images / videos that exhibit image / video characteristics that are different from actual images / videos (e.g., game images / videos).
[0005] Therefore, highly efficient image / video compression technology is required to effectively compress and transmit, store, or play high-resolution, high-quality images / videos showing various characteristics as described above. Summary of the invention
[0006] Technical issues
[0007] The present disclosure provides a method and an apparatus for improving image coding efficiency.
[0008] The present disclosure also provides a method and apparatus for deriving weight index information regarding bidirectional prediction in inter-frame prediction.
[0009] The present disclosure also provides a method and apparatus for deriving weight index information about candidates in an affine merge candidate list during bi-directional prediction.
[0010] Technical Solution
[0011] According to one embodiment of the present disclosure, there is provided an image decoding method performed by a decoding device. The method comprises the following steps: receiving image information including inter-frame prediction mode information and inter-frame prediction type information through a bitstream; generating a merge candidate list of a current block based on the inter-frame prediction mode information; selecting a candidate from the candidates included in the merge candidate list; deriving the inter-frame prediction type of the current block as a bidirectional prediction based on the inter-frame prediction type information; deriving motion information about the current block based on the selected candidate; generating L0 prediction samples and L1 prediction samples of the current block based on the motion information; and generating prediction samples of the current block based on the L0 prediction samples, the L1 prediction samples and weight information, wherein the weight information is based on the information about the selected candidate. The method is derived from the weight index information of the selected candidate, wherein the candidate includes an affine merge candidate, and the affine merge candidate includes a control point motion vector (CPMV), when the affine merge candidate includes the CPMV of control point 0 (CP0) located on the upper left side of the current block, the weight index information about the affine merge candidate is derived based on the weight index information of a specific block among the neighboring blocks of the CP0, and when the affine merge candidate does not include the CPMV of the CP0 located on the upper left side of the current block, the weight index information about the affine merge candidate is derived based on the weight index information of a specific block among the neighboring blocks of control point 1 (CP1) located on the upper right side of the current block.
[0012] According to another embodiment of the present disclosure, a method for image encoding performed by an encoding device is provided. The method includes the following steps: determining the inter-frame prediction mode of the current block and generating inter-frame prediction mode information indicating the inter-frame prediction mode; generating a merge candidate list of the current block based on the inter-frame prediction mode information; generating selection information indicating one of the candidates included in the merge candidate list; generating inter-frame prediction type information indicating that the inter-frame prediction type of the current block is bidirectional prediction; and encoding image information including the inter-frame prediction mode information, the selection information and the inter-frame prediction type information, wherein the candidate includes an affine merge candidate, and the affine merge candidate includes an affine merge candidate. The affine merge candidate includes a control point motion vector (CPMV), and when the affine merge candidate includes the CPMV of control point 0 (CP0) located on the upper left side of the current block, the weight index information about the affine merge candidate is indicated based on the weight index information about a specific block among the neighboring blocks of the CP0, and when the affine merge candidate does not include the CPMV of the CP0 located on the upper left side of the current block, the weight index information about the affine merge candidate is indicated based on the weight index information about a specific block among the neighboring blocks of the control point 1 (CP1) located on the upper right side of the current block.
[0013] According to another embodiment of the present disclosure, a computer-readable storage medium storing coding information that causes an image decoding device to perform an image decoding method is provided. The image decoding method includes the following steps: receiving image information including inter-frame prediction mode information and inter-frame prediction type information through a bitstream; generating a merge candidate list of a current block based on the inter-frame prediction mode information; selecting a candidate from the candidates included in the merge candidate list; deriving the inter-frame prediction type of the current block as a bidirectional prediction based on the inter-frame prediction type information; deriving motion information about the current block based on the selected candidate; generating L0 prediction samples and L1 prediction samples of the current block based on the motion information; and generating prediction samples of the current block based on the L0 prediction samples, the L1 prediction samples and the weight information, wherein the weight information is based on information about The method is derived based on weight index information of a selected candidate, wherein the candidate includes an affine merge candidate, and the affine merge candidate includes a control point motion vector (CPMV), and when the affine merge candidate includes the CPMV of control point 0 (CP0) located on the upper left side of the current block, the weight index information about the affine merge candidate is derived based on weight index information about a specific block among the neighboring blocks of the CP0, and when the affine merge candidate does not include the CPMV of the CP0 located on the upper left side of the current block, the weight index information about the affine merge candidate is derived based on weight index information about a specific block among the neighboring blocks of control point 1 (CP1) located on the upper right side of the current block.
[0014] Technical Effects
[0015] According to the present disclosure, the overall image / video compression efficiency can be improved.
[0016] According to the present disclosure, motion vector candidates can be efficiently constructed during inter prediction.
[0017] According to the present disclosure, weight-based bidirectional prediction can be performed efficiently. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is a diagram schematically illustrating an example of a video / image encoding system to which an embodiment of the present disclosure can be applied.
[0019] Figure 2 is a diagram schematically illustrating a configuration of a video / image encoding device to which an embodiment of the present disclosure can be applied.
[0020] Figure 3 is a diagram schematically illustrating a configuration of a video / image decoding device to which an embodiment of the present disclosure can be applied.
[0021] Figure 4is a diagram for describing the merge mode in inter prediction.
[0022] Figure 5a and Figure 5b is a diagram exemplarily illustrating CPMV for affine motion prediction.
[0023] Figure 6 is a diagram exemplarily illustrating a case in which an affine MVF is determined in units of sub-blocks.
[0024] Figure 7 is a diagram for describing the affine merge mode in inter prediction.
[0025] Figure 8 is a diagram for describing the positions of candidates in the affine merge mode.
[0026] Fig. 9 is a diagram for describing SbTMVP in inter-frame prediction.
[0027] Fig.10 and Fig.11 is a diagram schematically illustrating an example of a video / image decoding method and related components according to an embodiment of the present disclosure.
[0028] Fig.12 and Fig.13 is a diagram schematically illustrating an example of an image / video encoding method and related components according to an embodiment of the present disclosure.
[0029] Fig.14 is a diagram illustrating an example of a content streaming system to which an embodiment disclosed in the present disclosure can be applied. DETAILED DESCRIPTION
[0030] The present disclosure can be variously modified and has several exemplary embodiments. Therefore, the specific exemplary embodiments of the present disclosure will be illustrated in the drawings and described in detail. However, this is not intended to limit the present disclosure to specific embodiments. The terms used in this specification are only used to describe specific exemplary embodiments, rather than to limit the present disclosure. Unless the context clearly indicates otherwise, the singular form is intended to include the plural form. It will be understood that the terms "including", "having" etc. used in this specification specify the presence of features, numbers, steps, operations, components, parts or combinations thereof set forth in this specification, but do not exclude the presence or addition of one or more other features, numbers, steps, operations, components, parts or combinations thereof.
[0031] In addition, for the convenience of describing different feature functions, each component in the drawings described in the present disclosure is illustrated independently, which does not mean that each component is implemented as separate hardware or separate software. For example, two or more components among the components can be combined to form a component, or a component can be divided into multiple components. Embodiments in which each component is integrated and / or separated are also included in the scope of the present disclosure.
[0032] In the present disclosure, "A or B" may mean "only A", "only B", or "both A and B". In other words, "A or B" in the present disclosure may be interpreted as "A and / or B". For example, in the present disclosure, "A, B or C" means "only A", "only B", "only C", or "any one of A, B, and C and any combination thereof".
[0033] A slash ( / ) or a comma (,) used in the present disclosure may mean "and / or". For example, "A / B" may mean "A and / or B". Thus, "A / B" may mean "only A", "only B", or "both A and B". For example, "A, B, C" may mean "A, B, or C".
[0034] In this document, "at least one of A and B" may mean "only A", "only B", or "both A and B". In addition, in this document, the expression "at least one of A or B" or "at least one of A and / or B" may be interpreted as the same as "at least one of A and B".
[0035] In addition, in this document, "at least one of A, B, and C" may mean "only A", "only B", "only C", or "any combination of A, B, and C". In addition, "at least one of A, B, or C" or "at least one of A, B and / or C" may mean "at least one of A, B, and C".
[0036] In addition, brackets used in this document may mean "for example". Specifically, in the case of expressing "prediction (intra-frame prediction)", it may indicate that "intra-frame prediction" is proposed as an example of "prediction". In other words, the term "prediction" in this document is not limited to "intra-frame prediction", and it may indicate that "intra-frame prediction" is proposed as an example of "prediction". In addition, even in the case of expressing "prediction (i.e., intra-frame prediction)", it may indicate that "intra-frame prediction" is proposed as an example of "prediction".
[0037] In this document, technical features independently described in one drawing may be implemented independently or may be implemented simultaneously.
[0038] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. In addition, throughout the accompanying drawings, similar reference numerals are used to indicate similar elements, and the same description of similar elements may be omitted.
[0039] Figure 1 An example of a video / image encoding system to which an embodiment of the present disclosure can be applied is illustrated.
[0040] Reference Figure 1 The video / image coding system may include a first device (source device) and a second device (receiving device). The source device may send the coded video / image information or data to the receiving device in the form of a file or stream transmission via a digital storage medium or a network.
[0041] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.
[0042] The video source may acquire the video / image by capturing, synthesizing or generating a video / image process. The video source may include a video / image capturing device, and / or a video / image generating device. For example, the video / image capturing device may include one or more cameras, a video / image archive including previously captured videos / images, etc. For example, the video / image generating device may include a computer, a tablet computer, and a smart phone, and may generate the video / image (electronically). For example, a virtual video / image may be generated by a computer, etc. In this case, the video / image capturing process may be replaced by a process that generates relevant data.
[0043] The encoding device can encode the input video / image. For compression and encoding efficiency, the encoding device can perform a series of processes such as prediction, transformation and quantization. The encoded data (encoded video / image information) can be output in the form of a bit stream.
[0044] The transmitter may transmit the encoded image / image information or data output in the form of a bit stream to a receiver of a receiving device in the form of a file or stream via a digital storage medium or a network. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter may include an element for generating a media file in a predetermined file format, and may include an element for transmission via a broadcast / communication network. The receiver may receive / extract a bit stream and transmit the received bit stream to a decoding device.
[0045] The decoding device may decode a video / image by performing a series of processes corresponding to the operations of the encoding device, such as inverse quantization, inverse transformation, and prediction.
[0046] The renderer may render the decoded video / image. The rendered video / image may be displayed through a display.
[0047] The present disclosure relates to video / image coding. For example, the methods / implementations disclosed in the present disclosure may be applied to methods disclosed in the Versatile Video Coding (VVC) standard, the Essential Video Coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the second generation audio video coding standard (AVS2), or the next generation video / image coding standard (e.g., H.267 or H.268, etc.).
[0048] This document proposes various embodiments of video / image coding, and unless otherwise mentioned, the above embodiments may also be performed in combination with each other.
[0049] In this document, video may refer to a series of images over time. A picture generally refers to a unit representing an image in a specific time frame, and a slice / tile refers to a unit that constitutes a part of a picture from a coding perspective. A slice / tile may include one or more coding tree units (CTUs). A picture may consist of one or more slices / tiles.
[0050] A tile is a rectangular area of a CTU within a specific tile column and a specific tile row in a picture. A tile column is a rectangular area of a CTU whose height is equal to the height of the picture and whose width is specified by a syntax element in a picture parameter set. A tile row is a rectangular area of a CTU whose height is specified by a syntax element in a picture parameter set and whose width is equal to the width of the picture. Tile scanning is a specific sequential ordering of CTUs of a partitioned picture, where CTUs are continuously ordered in tiles by CTU raster scanning, and tiles in a picture are continuously ordered by raster scanning of tiles of the picture. A slice may include multiple complete tiles of a picture that may be contained in one NAL unit or multiple consecutive CTU rows in a tile. In this document, tile groups and slices may be used interchangeably. For example, in this document, a tile group / tile group header may be referred to as a slice / slice header.
[0051] In addition, a picture can be divided into two or more sub-pictures. A sub-picture can be a rectangular area of one or more slices within a picture.
[0052] A pixel or a pel may refer to the smallest unit constituting a picture (or image). In addition, a "sample" may be used as a term corresponding to a pixel. A sample may generally represent a pixel or a value of a pixel, and may represent only a pixel / pixel value of a luminance component, or only a pixel / pixel value of a chrominance component.
[0053] A unit may represent a basic unit of image processing. A unit may include at least one of a specific area of a picture and information related to the area. A unit may include a luminance block and two chrominance (e.g., cb, cr) blocks. In some cases, a unit may be used interchangeably with terms such as a block or an area. In general, an M×N block may include M columns and N rows of samples (or sample arrays) or a set (or array) of transform coefficients. Alternatively, a sample may mean a pixel value in a spatial domain, and when such a pixel value is transformed into a frequency domain, it may mean a transform coefficient in a frequency domain.
[0054] Figure 2 Schematic diagram of the configuration of a video / image encoding device to which the embodiments of the present document can be applied. Hereinafter, the so-called encoding device may include an image encoding device and / or a video encoding device. In addition, the so-called image encoding method / device may include a video encoding method / device. Alternatively, the so-called video encoding method / device may include an image encoding method / device.
[0055] Reference Figure 2 , the encoding device 200 includes an image segmenter 210, a predictor 220, a residual processor 230 and an entropy encoder 240, an adder 250, a filter 260 and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, an inverse quantizer 234 and an inverse transformer 235. The residual processor 230 may also include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstructed block generator. According to an embodiment, the image segmenter 210, the predictor 220, the residual processor 230, the entropy encoder 240, the adder 250 and the filter 260 may be configured by at least one hardware component (e.g., an encoder chipset or processor). In addition, the memory 270 may include a decoded picture buffer (DPB), or may be configured by a digital storage medium. The hardware component may also include the memory 270 as an internal / external component.
[0056] The image divider 210 may divide the input image (or picture or frame) input to the encoding device 200 into one or more processors. For example, the processor may be referred to as a coding unit (CU). In this case, the coding unit may be recursively divided from a coding tree unit (CTU) or a maximum coding unit (LCU) according to a quadtree binary tree ternary tree (QTBTTT) structure. For example, one coding unit may be divided into a plurality of coding units with a deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary structure. In this case, for example, a quadtree structure may be applied first, and a binary tree structure and / or a ternary structure may be applied later. Alternatively, a binary tree structure may be applied first. The encoding process according to the present disclosure may be performed based on a final coding unit that is no longer divided. In this case, the maximum coding unit may be used as the final coding unit based on coding efficiency, etc. according to image characteristics, or if necessary, the coding unit may be recursively divided into coding units with a deeper depth and a coding unit with an optimal size may be used as the final coding unit. Here, the encoding process may include a process of prediction, transformation, and reconstruction (to be described later). As another example, the processor may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may be split or divided from the above-mentioned final coding unit. The prediction unit may be a unit for sample prediction, and the transform unit may be a unit for deriving a transform coefficient and / or a unit for deriving a residual signal from the transform coefficient.
[0057] In some cases, a unit may be used interchangeably with terms such as a block or region. In general, an M×N block may represent a set of samples or transform coefficients consisting of M columns and N rows. A sample may generally represent a pixel or a pixel value, and may represent only a pixel / pixel value of a luma component or only a pixel / pixel value of a chroma component. A sample may be used as a term corresponding to one picture (or image) of a pixel or a pixel element.
[0058] The encoding device 200 can subtract the prediction signal (prediction block, prediction sample array) output from the inter predictor 221 or the intra predictor 222 from the input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is sent to the transformer 232. In this case, as illustrated, the unit for subtracting the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) in the encoding device (encoder) 200 can be referred to as a subtractor 231. The predictor can perform prediction on the processing target block (hereinafter, referred to as the current block) and generate a prediction block including the prediction sample of the current block. The predictor can determine whether to apply intra prediction or inter prediction in units of the current block or CU. The predictor can generate various information about prediction such as prediction mode information, and send the generated information to the entropy encoder 240, as described below when describing each prediction mode. The information about the prediction can be encoded by the entropy encoder 240 and output in the form of a bitstream.
[0059] The intra-frame predictor 222 may predict the current block with reference to samples in the current picture. Depending on the prediction mode, the referenced samples may be located near the current block or may be spaced apart. In intra-frame prediction, the prediction mode may include multiple non-directional modes and multiple directional modes. For example, the non-directional mode may include a DC mode and a plane mode. For example, depending on the level of detail of the prediction direction, the directional mode may include 33 directional prediction modes or 65 directional prediction modes. However, this is only an example, and more or fewer directional prediction modes may be used depending on the settings. The intra-frame predictor 222 may use the prediction mode applied to the neighboring blocks to determine the prediction mode applied to the current block.
[0060] The inter-frame predictor 221 may derive a prediction block of the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. Here, in order to reduce the amount of motion information sent in the inter-frame prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may also include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, the neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block may be referred to as a collocated reference block, a collocated CU (colCU), etc., and the reference picture including the temporal neighboring block may be referred to as a collocated picture (colPic). For example, the inter-frame predictor 221 may configure a motion information candidate list based on the neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter-frame prediction may be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the inter-frame predictor 221 may use the motion information of the neighboring block as the motion information of the current block. In the skip mode, unlike the merge mode, the residual signal may not be transmitted. In the case of the motion vector prediction (MVP) mode, the motion vector of the neighboring block may be used as a motion vector predictor, and the motion vector of the current block may be indicated by signaling the motion vector difference.
[0061] The predictor 220 may generate a prediction signal based on various prediction methods described below. For example, the predictor may apply intra prediction or inter prediction to the prediction of a block, and may apply intra prediction and inter prediction at the same time. This may be referred to as a combination of inter and intra prediction (CIIP). In addition, the predictor may predict a block based on an intra-block copy (IBC) prediction mode or based on a palette mode. The IBC prediction mode or the palette mode may be used for image / video encoding of content such as games, for example, screen content coding (SCC). IBC basically performs prediction within the current picture, but may be performed similarly to inter prediction in terms of deriving a reference block within the current picture. That is, IBC may use at least one of the inter prediction techniques described in this document. The palette mode may be considered an example of intra coding or intra prediction. When the palette mode is applied, the sample values in the picture may be signaled based on information about the palette table and the palette index.
[0062] The prediction signal generated by the predictor (including the inter-frame predictor 221 and / or the intra-frame predictor 222) can be used to generate a reconstruction signal or can be used to generate a residual signal. The transformer 232 can generate a transform coefficient by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loeve transform (KLT), a graph-based transform (GBT), or a conditional nonlinear transform (CNT). Here, GBT refers to a transform obtained from a curve graph when the relationship information between pixels is represented by a curve graph. CNT refers to a transform obtained based on a prediction signal generated using all previously reconstructed pixels. In addition, the transform process can be applied to a block of pixels of the same size as a square, or can be applied to a block of variable size rather than a square.
[0063] The quantizer 233 quantizes the transform coefficient and sends it to the entropy encoder 240, and the entropy encoder 240 encodes the quantized signal (information about the quantized transform coefficient) and outputs the encoded signal as a bitstream. Information about the quantized transform coefficient may be referred to as residual information. The quantizer 233 may rearrange the quantized transform coefficients in block form into a one-dimensional vector form based on the coefficient scanning order, and may generate information about the transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. The entropy encoder 240 may perform various encoding methods such as (for example) exponential Golomb, context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC). The entropy encoder 240 may encode information necessary for video / image reconstruction (for example, the value of a syntax element, etc.) other than the quantized transform coefficients together or separately. The encoded information (for example, the encoded video / image information) may be transmitted or stored in units of network abstraction layer (NAL) units in the form of a bitstream. The video / image information may also include information about various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may also include conventional constraint information. In this document, information and / or syntax elements sent / signaled to a decoding device from an encoding device may be included in the video / image information. The video / image information may be encoded and included in a bitstream through the above-mentioned encoding process. The bitstream may be sent through a network, or may be stored in a digital storage medium. Here, the network may include a broadcast network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. A sending unit (not shown) and / or a storage unit (not shown) for sending or storing a signal output from the entropy encoder 240 may be configured as an internal / external element of the encoding device 200, or the sending unit may be included in the entropy encoder 240.
[0064] The quantized transform coefficients output from the quantizer 233 can be used to generate a prediction signal. For example, the residual signal (residual block or residual sample) can be reconstructed by applying inverse quantization and inverse transform to the quantized transform coefficients using the inverse quantizer 234 and the inverse transformer (inverse transform unit) 235. The adder 250 can add the reconstructed residual signal to the prediction signal output from the inter-frame predictor 221 or the intra-frame predictor 222 to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). When there is no residual of the processing target block, such as when the skip mode is applied, the prediction block can be used as a reconstructed block. The adder 250 can be referred to as a recovery unit or a recovery block generator. The generated reconstructed signal can be used for intra-frame prediction of the next processing target block in the current picture, and can be used for inter-frame prediction of the next picture after filtering as described below.
[0065] Furthermore, luma mapping with chroma scaling (LMCS) may be applied during the picture encoding and / or reconstruction process.
[0066] The filter 260 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 260 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and store the modified reconstructed picture in the memory 270, specifically, in the DPB of the memory 270. Various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filter 260 can generate various types of information related to filtering, and transmit the generated information to the entropy encoder 240, as described later when each filtering method is described. The information related to filtering can be encoded by the entropy encoder 240 and output in the form of a bit stream.
[0067] The modified reconstructed picture transmitted to the memory 270 may be used as a reference picture in the inter predictor 221. When inter prediction is applied by the encoding device, prediction mismatch between the encoding device 200 and the decoding device may be avoided, and encoding efficiency may be improved.
[0068] The DPB of the memory 270 may store the modified reconstructed picture for use as a reference picture in the inter-frame predictor 221. The memory 270 may store the motion information of the block from which the motion information in the current picture is derived (or encoded) and / or the motion information of the block in the reconstructed picture. The stored motion information may be transmitted to the inter-frame predictor 221 to be used as the motion information of the spatial neighboring block or the motion information of the temporal neighboring block. The memory 270 may store the reconstructed samples of the reconstructed block in the current picture and may transmit the reconstructed samples to the intra-frame predictor 222.
[0069] In addition, in this document, at least one of quantization / dequantization and / or transform / inverse transform may be omitted. When quantization / dequantization is omitted, the quantized transform coefficient may be referred to as a transform coefficient. When transform / inverse transform is omitted, the transform coefficient may be referred to as a coefficient or a residual coefficient, or may still be referred to as a transform coefficient for consistency of expression.
[0070] In addition, in this document, the quantized transform coefficient and the transform coefficient may be referred to as a transform coefficient and a scaled transform coefficient, respectively. In this case, the residual information may include information about the transform coefficient, and the information about the transform coefficient may be signaled by a residual coding syntax. The transform coefficient may be derived based on the residual information (or information about the transform coefficient), and the scaled transform coefficient may be derived by inverse transforming (scaling) the transform coefficient. The residual sample may be derived based on the inverse transform (transform) of the scaled transform coefficient. This may also be applied / expressed in other parts of this document.
[0071] Figure 3 is a diagram schematically illustrating a configuration of a video / image decoding device to which the disclosure of this document can be applied.
[0072] Reference Figure 3 , the decoding device 300 may include and be configured with an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an intra-frame predictor 331 and an inter-frame predictor 332. The residual processor 320 may include an inverse quantizer 321 and an inverse transformer 322. According to an embodiment, the entropy decoder 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 described above may be configured by one or more hardware components (e.g., a decoder chipset or a processor). In addition, the memory 360 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may also include a memory 360 as an internal / external component.
[0073] When a bit stream including video / image information is input, the decoding device 300 may respond to the Figure 2 The image is reconstructed by processing video / image information in the encoding device illustrated in . For example, the decoding device 300 can derive the unit / block based on the block segmentation related information obtained from the bit stream. The decoding device 300 can perform decoding using a processing unit applied to the encoding device. Therefore, the processing unit for decoding can be, for example, a coding unit, and the coding unit can be divided from the coding tree unit or the maximum coding unit according to a quadtree structure, a binary tree structure and / or a ternary tree structure. One or more transform units can be derived from the coding unit. In addition, the reconstructed image signal decoded and output by the decoding device 300 can be reproduced by a reproduction device.
[0074] The decoding device 300 may receive the Figure 2 The signal output by the encoding device of the video image processing apparatus 310 can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can parse the bitstream to derive the information (e.g., video / image information) necessary for image reconstruction (or picture reconstruction). The video / image information may also include information about various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may also include conventional constraint information. The decoding device may also decode the picture based on the information about the parameter set and / or the conventional constraint information. The signaled / received information and / or syntax elements described in this document may then be decoded by a decoding process and obtained from the bitstream. For example, the entropy decoder 310 decodes the information in the bitstream based on a coding method such as exponential Golomb coding, context adaptive variable length coding (CAVLC), or context adaptive binary arithmetic coding (CABAC), and outputs syntax elements required for image reconstruction and quantized values of transform coefficients for the residual. More specifically, the CABAC entropy decoding method can receive a bin corresponding to each syntax element in a bitstream, determine a context model by using information of a decoded target syntax element, information of a decoded target block, or information of a symbol / bin decoded in a previous stage, and perform arithmetic decoding on the bin by predicting the probability of the occurrence of the bin according to the determined context model, and generate a symbol corresponding to the value of each syntax element. In this case, the CABAC entropy decoding method can update the context model by using information of a decoded symbol / bin of a context model for the next symbol / bin after determining the context model. Information related to prediction among the information decoded by the entropy decoder 310 can be provided to a predictor (inter-frame predictor 332 and intra-frame predictor 331), and residual values (i.e., quantized transform coefficients and related parameter information) that have been entropy decoded in the entropy decoder 310 can be input to a residual processor 320.
[0075] The dequantizer 321 may dequantize the quantized transform coefficient and output the transform coefficient. The dequantizer 321 may rearrange the quantized transform coefficient in a two-dimensional block form. In this case, the rearrangement may be performed based on the coefficient scanning order performed in the encoding device. The dequantizer 321 may perform dequantization on the quantized transform coefficient using a quantization parameter (e.g., quantization step size information) and obtain the transform coefficient.
[0076] The inverse transformer 322 inversely transforms the transform coefficients to obtain a residual signal (residual block, residual sample array).
[0077] The predictor may perform prediction on the current block and generate a prediction block including prediction samples of the current block. The predictor may determine whether to apply intra prediction or inter prediction to the current block based on information on prediction output from the entropy decoder 310 and may determine a specific intra / inter prediction mode.
[0078] The predictor 330 may generate a prediction signal based on various prediction methods to be described later. For example, the predictor may apply intra prediction or inter prediction to the prediction of a block, and may apply intra prediction and inter prediction at the same time. This may be referred to as a combination of inter and intra prediction (CIIP). In addition, the predictor may predict a block based on an intrablock copy (IBC) prediction mode or based on a palette mode. The IBC prediction mode or the palette mode may be used for image / video encoding of content such as games, for example, screen content coding (SCC). IBC may basically perform prediction within the current picture, but may be performed similarly to inter prediction in terms of deriving a reference block within the current picture. That is, IBC may use at least one of the inter prediction techniques described in this document. The palette mode may be considered an example of intra coding or intra prediction. When the palette mode is applied, information about the palette table and the palette index may be included in the video / image information and signaled.
[0079] The intra-frame predictor 331 can predict the current block by referring to samples in the current picture. Depending on the prediction mode, the referenced sample can be located near the current block, or its location can be separated from the current block. In intra-frame prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. The intra-frame predictor 331 can determine the prediction mode to be applied to the current block by using the prediction mode applied to the neighboring block.
[0080] The inter-frame predictor 332 may derive a prediction block of the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, in order to reduce the amount of motion information sent in the inter-frame prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may also include information about the inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, the neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. For example, the inter-frame predictor 332 may construct a motion information candidate list based on the neighboring blocks, and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter-frame prediction may be performed based on various prediction modes, and information about the prediction may include information indicating a mode of inter-frame prediction for the current block.
[0081] The adder 340 can generate a reconstructed signal (reconstructed picture, reconstructed block or reconstructed sample array) by adding the obtained residual signal to the prediction signal (prediction block or prediction sample array) output from the predictor (including the inter-frame predictor 332 and / or the intra-frame predictor 331). If the residual of the target block is not processed (such as when the skip mode is applied), the prediction block can be used as the reconstructed block.
[0082] The adder 340 may be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal may be used for intra prediction of the next block to be processed in the current picture, and as described later, may also be output through filtering or may also be used for inter prediction of the next picture.
[0083] In addition, luminance mapping with chroma scaling (LMCS) can also be applied to the picture decoding process.
[0084] The filter 350 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 350 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and store the modified reconstructed picture in the memory 360, specifically, in the DPB of the memory 360. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0085] The (modified) reconstructed picture stored in the DPB of the memory 360 may be used as a reference picture in the inter-frame predictor 332. The memory 360 may store the motion information of the block from which the motion information in the current picture is derived (or decoded) and / or the motion information of the block in the reconstructed picture. The stored motion information may be transmitted to the inter-frame predictor 332 so as to be used as the motion information of the spatial neighboring block or the motion information of the temporal neighboring block. The memory 360 may store the reconstructed samples of the reconstructed block in the current picture and transmit the reconstructed samples to the intra-frame predictor 331.
[0086] In the present disclosure, the embodiments described in the filter 260 , the inter predictor 221 , and the intra predictor 222 of the encoding apparatus 200 may be applied equally or correspondingly to the filter 350 , the inter predictor 332 , and the intra predictor 331 .
[0087] In addition, as described above, when performing video encoding, prediction is performed to improve compression efficiency. Accordingly, a prediction block including prediction samples for a current block as a block to be encoded (i.e., an encoding target block) can be generated. Here, the prediction block includes prediction samples in a spatial domain (or a pixel domain). The prediction block is derived in the same manner in the encoding device and the decoding device, and the encoding device can signal information about the residual between the original block and the prediction block (residual information) rather than the original sample value of the original block to the decoding device, thereby improving the image coding efficiency. The decoding device can derive a residual block including residual samples based on the residual information, add the residual block and the prediction block to generate a reconstructed block including reconstructed samples, and generate a reconstructed picture including the reconstructed block.
[0088] Residual information can be generated through a transformation and quantization process. For example, the encoding device can derive a residual block between the original block and the prediction block, perform a transformation process on the residual samples (residual sample array) included in the residual block to derive a transformation coefficient, perform a quantization process on the transformation coefficient to derive a quantized transformation coefficient, and signal the relevant residual information to the decoding device (through a bitstream). Here, the residual information may include value information, position information, transformation technology, transformation kernel, quantization parameter, etc. of the quantized transformation coefficient. The decoding device can perform an inverse quantization / inverse transformation process based on the residual information and derive residual samples (or residual blocks). The decoding device can generate a reconstructed picture based on the prediction block and the residual block. In addition, in order to provide a reference for inter-frame prediction of subsequent pictures, the encoding device can also perform inverse quantization / inverse transformation on the quantized transformation coefficient to derive a residual block and generate a reconstructed picture based on this.
[0089] Figure 4 is a diagram for describing the merge mode in inter prediction.
[0090] When the merge mode is applied, the motion information about the current prediction block is not sent directly, but the motion information about the neighboring prediction blocks is used to derive the motion information about the current prediction block. Therefore, the motion information about the current prediction block can be indicated by sending flag information indicating the use of the merge mode and a merge index indicating which nearby prediction block is used. The merge mode can be referred to as the regular merge mode. For example, when the value of the regular_merge_flag syntax element is 1, the merge mode can be applied.
[0091] In order to perform the merge mode, the encoding device needs to search for a merge candidate block for deriving motion information about the current prediction block. For example, up to five merge candidate blocks can be used, but the embodiments of the present disclosure are not limited to this. In addition, the maximum number of merge candidate blocks can be sent in a slice header or a tile group header, but the embodiments of the present disclosure are not limited to this. After finding the merge candidate block, the encoding device can generate a merge candidate list, and can select the merge candidate block with the smallest cost among the merge candidate blocks as the final merge candidate block.
[0092] The present disclosure may provide various implementations of merging candidate blocks constituting a merging candidate list.
[0093] For example, the merge candidate list may use five merge candidate blocks. For example, four spatial merge candidates and one temporal merge candidate may be used. As a specific example, in the case of spatial merge candidates, Figure 4 The blocks illustrated in may be used as spatial merge candidates. Hereinafter, a spatial merge candidate or a spatial MVP candidate to be described later may be referred to as an SMVP, and a temporal merge candidate or a temporal MVP candidate to be described later may be referred to as a TMVP.
[0094] For example, the merge candidate list of the current block may be constructed based on the following process.
[0095] The encoding device / decoding device may search for spatial neighboring blocks of the current block and insert the derived spatial merge candidates into the merge candidate list. For example, the spatial neighboring blocks may include the lower left neighboring block, the left neighboring block, the upper right neighboring block, the upper neighboring block, and the upper left neighboring block of the current block. However, this is an example, and in addition to the above-mentioned spatial neighboring blocks, additional neighboring blocks such as the right neighboring block, the lower neighboring block, and the lower right neighboring block may be further used as spatial neighboring blocks. The encoding device may detect available blocks by searching spatial neighboring blocks based on priority, and may derive motion information about the detected blocks as spatial merge candidates. For example, the encoding device or decoding device may be configured to perform a search such as A 1 →B 1 →B 0 →A 0 →B 2 Search in this order Figure 4 The five blocks illustrated in FIG. 3 can be used to index the available candidates in sequence to form a merge candidate list.
[0096] The encoding device may search for temporally neighboring blocks of a current block and insert the derived temporal merge candidates into a merge candidate list. The temporally neighboring blocks may be at a reference picture that is a different picture from the current picture in which the current block is located. The reference picture in which the temporally neighboring blocks are located may be referred to as a co-located picture or a collocated picture. The temporally neighboring blocks may be searched for on the collocated picture in the order of the bottom-right neighboring block and the bottom-right center block of the co-located block of the current block. In addition, when motion data compression is applied, specific motion information may be stored as representative motion information in each predetermined storage unit in the collocated picture. In this case, it is not necessary to store the motion information about all blocks in the predetermined storage unit, and accordingly, a motion data compression effect can be obtained. In this case, the predetermined storage unit may be predetermined as, for example, a unit of 16×16 samples or a unit of 8×8 samples, or the size information about the predetermined storage unit may be signaled from the encoding device to the decoding device. When motion data compression is applied, the motion information about the temporally neighboring blocks may be replaced with the representative motion information on the predetermined storage unit where the temporally neighboring blocks are located. That is, in this case, from the perspective of implementation, instead of the prediction block at the coordinates of the temporally neighboring blocks, the motion information of the prediction block covering the arithmetic left-shift position after arithmetic right-shift by a certain value based on the coordinates (top-left sample position) of the temporally neighboring blocks may be used to derive the temporal merge candidates. For example, when the predetermined storage unit is a unit of 2n×2n samples, if the coordinates of the temporally neighboring block are (xTnb, yTnb), the motion information of the prediction block at the corrected position ((xTnb>>n)<<n), (yTnb>>n)<<n)) may be used for the temporal merge candidates. Specifically, when the predetermined storage unit is a unit of 16×16 samples, if the coordinates of the temporally neighboring block are (xTnb, yTnb), the motion information of the prediction block at the corrected position ((xTnb>>4)<<4), (yTnb>>4)<<4)) may be used for the temporal merge candidates. Alternatively, when the predetermined storage unit is a unit of 8×8 samples, if the coordinates of the temporally neighboring block are (xTnb, yTnb), the motion information of the prediction block at the corrected position (xTnb>>3)<<3), (yTnb>>3)<<3)) may be used for the temporal merge candidates.
[0097] The encoding device may check whether the number of current merge candidates is less than the maximum number of merge candidates. The maximum number of merge candidates may be predefined or signaled from the encoding device to the decoding device. For example, the encoding device may generate and encode information about the maximum number of merge candidates and send the information to the decoder in the form of a bitstream. When the maximum number of merge candidates is full, the subsequent candidate addition process may not be performed.
[0098] As a result of the check, when the current number of merge candidates is less than the maximum number of merge candidates, the encoding device may insert an additional merge candidate into the merge candidate list. For example, the additional merge candidate may include at least one of a history-based merge candidate, a pairwise average merge candidate, an ATMVP, a combined bidirectional prediction merge candidate (when the slice / tile group type of the current slice / tile group is type B), and / or a zero vector merge candidate to be described later.
[0099] As a result of the check, when the current number of merge candidates is not less than the maximum number of merge candidates, the encoding device may terminate the construction of the merge candidate list. In this case, the encoding device may select the optimal merge candidate from the merge candidates constituting the merge candidate list based on the rate-distortion (RD) cost, and signal the selection information (e.g., merge index) indicating the selected merge candidate to the decoding device. The decoding device may select the optimal merge candidate based on the merge candidate list and the selection information.
[0100] As described above, the motion information about the selected merge candidate may be used as the motion information about the current block, and the prediction sample of the current block may be derived based on the motion information about the current block. The encoding device may derive the residual sample of the current block based on the prediction sample, and may signal the residual information about the residual sample to the decoding device. As described above, the decoding device may generate a reconstructed sample based on the residual sample derived based on the residual information and the prediction sample, and may generate a reconstructed picture based on this.
[0101] When skip mode is applied, motion information about the current block may be derived in the same manner as when merge mode is applied. However, when skip mode is applied, the residual signal of the corresponding block is omitted, and thus the prediction sample may be directly used as the reconstructed sample. For example, when the value of the cu_skip_flag syntax element is 1, skip mode may be applied.
[0102] In addition, pairwise average merge candidates may be referred to as pairwise average candidates or pairwise candidates. Pairwise average candidates may be generated by averaging predefined candidate pairs in an existing merge candidate list. In addition, predefined pairs may be defined as {(0,1), (0,2), (1,2), (0,3), (1,3), (2,3)}. Here, the number may indicate the merge index of the merge candidate list. The averaged motion vector may be calculated separately for each reference list. For example, when two motion vectors are available in a list, they may be averaged even if the two motion vectors point to different reference pictures. For example, when only one motion vector is available, one motion vector may be used directly. For example, when no available motion vector exists, the list may remain invalid.
[0103] For example, when the merge candidate list is not full even after adding the pairwise average merge candidate, that is, when the current number of merge candidates in the merge candidate list is less than the maximum number of merge candidates, a zero vector (zero MVP) may be inserted last until the maximum number of merge candidates appears. That is, a zero vector may be inserted until the number of current merge candidates in the merge candidate list becomes the maximum number of merge candidates.
[0104] In addition, conventionally, only one motion vector can be used to represent the motion of the coding block. That is, a translation motion model can be used. However, although this method can represent the optimal motion in units of blocks, it is not actually the optimal motion for each sample, and if the optimal motion vector can be determined in units of samples, the coding efficiency can be improved. For this purpose, an affine motion model can be used. The affine motion prediction method for encoding using an affine motion model can be as follows.
[0105] The affine motion prediction method can use two, three or four motion vectors to represent the motion vector per sample of the block. For example, the affine motion model can represent four types of motion. The affine motion model that represents three types of motion (translation, scaling and rotation) among the motions that the affine motion model can represent can be called a similar (or simplified) affine motion model. However, the affine motion model is not limited to the above motion models.
[0106] Figure 5a and Figure 5b is a diagram exemplarily illustrating CPMV for affine motion prediction.
[0107] Affine motion prediction may use two or more control point motion vectors (CPMVs) to determine a motion vector of a sample position included in a block. In this case, a set of motion vectors may be referred to as an affine motion vector field (MVF).
[0108] For example, Figure 5a A case of using two CPMVs, which may be referred to as a 4-parameter affine model, may be shown. In this case, for example, a motion vector at a sample position (x, y) may be determined as in Equation 1.
[0109] [Formula 1]
[0110]
[0111] For example, Figure 5b A case of using three CPMVs, which may be referred to as a 6-parameter affine model, may be exemplified. In this case, for example, a motion vector at a sample position (x, y) may be determined as in Equation 2.
[0112] [Formula 2]
[0113]
[0114] In equations 1 and 2, {v x ,v y} can represent the motion vector at position (x, y). In addition, {v 0x ,v 0y} may indicate the CPMV of the control point (CP) at the upper left corner of the coding block, and {v 1x ,v 1y} can indicate the CPMV of the CP at the upper right corner, {v 2x ,v 2y} may indicate the CPMV of the CP at the lower left corner position. In addition, W may indicate the width of the current block, and H may indicate the height of the current block.
[0115] Figure 6 is a diagram exemplarily illustrating a case in which an affine MVF is determined in units of sub-blocks.
[0116] In the encoding / decoding process, the affine MVF can be determined in units of samples or in units of previously defined sub-blocks. For example, when determined in units of samples, a motion vector can be obtained based on each sample value. Alternatively, for example, when determined in units of sub-blocks, the motion vector of the corresponding block can be obtained based on the sample value of the center of the sub-block (i.e., the lower right side of the center, i.e., the lower right sample among the four samples in the center). That is, in affine motion prediction, the motion vector of the current block can be derived in units of samples or sub-blocks.
[0117] exist Figure 6 In the case of , the affine MVF is determined in units of 4×4 sub-blocks, but the size of the sub-block may be modified differently.
[0118] That is, when affine prediction is available, the three motion models applicable to the current block may include a translation motion model, a 4-parameter affine motion model, and a 6-parameter affine motion model. Here, the translation motion model may represent a model in which an existing block unit motion vector is used, the 4-parameter affine motion model may represent a model in which two CPMVs are used, and the 6-parameter affine motion model may represent a model in which three CPMVs are used.
[0119] Furthermore, affine motion prediction may include an affine MVP (or affine inter) mode or an affine merge mode.
[0120] Figure 7 is a diagram for describing the affine merge mode in inter prediction.
[0121] For example, in the affine merge mode, the CPMV may be determined based on the affine motion model of the neighboring blocks encoded by affine motion prediction. For example, the neighboring blocks encoded as affine motion prediction in search order may be used for the affine merge mode. That is, when at least one of the neighboring blocks is encoded in the affine motion prediction, the current block may be encoded in the affine merge mode. Here, the refined merge mode may be referred to as AF_MERGE.
[0122] When the affine merge mode is applied, the CPMV of the neighboring block can be used to derive the CPMV of the current block. In this case, the CPMV of the neighboring block can be used as the CPMV of the current block as it is, and the CPMV of the neighboring block can be modified based on the size of the neighboring block and the size of the current block and used as the CPMV of the current block.
[0123] On the other hand, in the case of an affine merge mode in which a motion vector (MV) is derived in units of sub-blocks, this may be referred to as a sub-block merge mode, and the sub-block merge mode may be indicated based on a sub-block merge flag (or merge_subblock_flag syntax element). Alternatively, when the value of the merge_subblock_flag syntax element is 1, it may be indicated that the sub-block merge mode is applied. In this case, the affine merge candidate list to be described later may be referred to as a sub-block merge candidate list. In this case, the sub-block merge candidate list may also include a candidate derived by the SbTMVP to be described later. In this case, the candidate derived by the SbTMVP may be used as a candidate of index 0 of the sub-block merge candidate list. In other words, the candidate derived from the SbTMVP may be before the inherited affine candidate or the constructed affine candidate to be described later in the sub-block merge candidate list.
[0124] When the affine merge mode is applied, an affine merge candidate list can be constructed to derive the CPMV of the current block. For example, the affine merge candidate list may include at least one of the following candidates. 1) Inherited affine merge candidate. 2) Constructed affine merge candidate. 3) Zero motion vector candidate (or zero vector). Here, the inherited affine merge candidate is a candidate derived based on the CPMV of a neighboring block when the neighboring block is encoded in the affine mode, the constructed affine merge candidate is a candidate derived by constructing the CPMV based on the MV of the neighboring block of the corresponding CP in units of each CPMV, and the zero motion vector candidate may indicate a candidate consisting of a CPMV whose value is 0.
[0125] For example, the affine merge candidate list can be constructed as follows.
[0126] There may be up to two inherited affine candidates, and the inherited affine candidates may be derived from the affine motion model of the neighboring blocks. The neighboring blocks may include a left neighboring block and an upper neighboring block. The candidate blocks may be as follows: Figure 4 The scanning order of the left predictor can be A 1 →A 0 , and the scanning order of the upper predictor can be B 1 →B 0 →B 2 Only one inherited candidate may be selected from each of the left side and the upper side. Pruning check may not be performed between two inherited candidates.
[0127] When checking the neighboring affine blocks, the control point motion vector of the checked block can be used to derive the CPMVP candidate in the affine merge list of the current block. Here, the neighboring affine block can indicate a block encoded in the affine prediction mode among the neighboring blocks of the current block. For example, referring to Figure 7 , when the lower left neighboring block A is encoded in affine prediction mode, the motion vectors v2, v3 and v4 of the upper left corner, upper right corner and lower left corner of the neighboring block A can be obtained. When the neighboring block A is encoded using a 4-parameter affine motion model, the two CPMVs of the current block can be calculated based on v2 and v3. When the neighboring block A is encoded using a 6-parameter affine motion model, the two CPMVs of the current block can be calculated based on v2, v3 and v4.
[0128] Figure 8 is a diagram for describing the positions of candidates in the affine merge mode.
[0129] The constructed affine candidate may mean a candidate constructed by combining translation motion information around each control point. Motion information about a control point may be derived from a specified spatial perimeter and temporal perimeter. CPMVk (k=0, 1, 2, 3) may represent the kth control point.
[0130] Reference Figure 8 For CPMV0, press B 2 →B 3 →A 2 The blocks are checked in the order of B and the motion vector of the first available block can be used. 1 →B 0 Check the blocks in order, and for CPMV2, you can press A 1 →A 0 If available, the temporal motion vector predictor (TMVP) can be used with CPMV3.
[0131] After obtaining the motion vectors of the four control points, an affine merge candidate can be generated based on the acquired motion information. The combination of control point motion vectors can correspond to any one of {CPMV0, CPMV1, CPMV2}, {CPMV0, CPMV1, CPMV3}, {CPMV0, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV3}, {CPMV0, CPMV1} and {CPMV0, CPMV2}.
[0132] A combination of three CPMVs may constitute a 6-parameter affine merge candidate, and a combination of two CPMVs may constitute a 4-parameter affine merge candidate. To avoid motion scaling processing, relevant combinations of control point motion vectors may be discarded when the reference indices of the control points are different.
[0133] Fig. 9 is a diagram for describing SbTMVP in inter-frame prediction.
[0134] In addition, a sub-block-based temporal motion vector prediction (SbTMVP) method may also be used. For example, SbTMVP may be referred to as advanced temporal motion vector prediction (ATMVP). SbTMVP may use the motion field in a collocated picture to improve the motion vector prediction and merge mode of the CU in the current picture. Here, a collocated picture may be referred to as a collocated picture.
[0135] For example, SbTMVP can predict motion at the sub-block (or sub-CU) level. In addition, SbTMVP can apply motion shifting before obtaining temporal motion information from the collocated picture. Here, the motion shifting can be obtained from the motion vector of one of the spatial neighboring blocks of the current block.
[0136] SbTMVP can predict the motion vector of a sub-block (or sub-CU) in the current block (or CU) according to two steps.
[0137] In the first step, we can Figure 4 The order in A 1 , B 1 , B 0 and A 0 To test the spatial neighboring blocks. The first spatial neighboring block with a motion vector that uses the collocated picture as its reference picture may be checked and this motion vector may be selected as the motion shift to be applied. When no such motion is detected from the spatial neighboring blocks, the motion shift may be set to (0, 0).
[0138] In the second step, the motion shift checked in the first step can be applied to obtain sub-block level motion information (motion vector and reference index) from the collocated picture. For example, the motion shift can be added to the coordinates of the current block. For example, the motion shift can be set to Figure 4 A 1 In this case, for each sub-block, motion information about the sub-block may be derived using motion information about the corresponding block in the collocated picture. Temporal motion scaling may be applied to align the reference picture of the temporal motion vector with the reference picture of the current block.
[0139] A sub-block based merge list including a combination of both SbTVMP candidates and affine merge candidates can be used for signaling of affine merge mode. Here, affine merge mode may be referred to as sub-block based merge mode. Depending on the flag included in the sequence parameter set (SPS), SbTVMP mode may be available or unavailable. When SbTMVP mode is available, the SbTMVP predictor may be added as the first entry in the sub-block based merge candidate list, and the affine merge candidate may follow. The maximum allowable size of affine merge candidates may be 5.
[0140] The size of the sub-CU (or sub-block) used in SbTMVP can be fixed to 8×8, and as in affine merge mode, SbTMVP mode can be applied only to blocks with both width and height of 8 or more. The encoding logic of the additional SbTMVP merge candidate can be the same as the encoding logic of other merge candidates. That is, for each CU in a P or B slice, an RD check using an additional rate-distortion (RD) cost can be performed to determine whether to use the SbTMVP candidate.
[0141] In addition, the prediction block of the current block may be derived based on the motion information derived according to the prediction mode. The prediction block may include prediction samples (prediction sample arrays) of the current block. When the motion vector of the current block indicates a fractional sample unit, an interpolation process may be performed. Accordingly, the prediction sample of the current block may be derived based on the fractional sample unit reference sample in the reference picture. When affine inter prediction (affine prediction mode) is applied to the current block, the prediction sample may be generated based on the sample / sub-block unit MV. When bidirectional prediction is applied, the prediction sample may be used as the prediction sample derived by the weighted sum (according to the phase) or weighted average of the prediction sample derived based on the L0 prediction (i.e., prediction using the reference picture in the reference picture list L0 and MVL0) and the prediction sample derived based on the L1 prediction (i.e., prediction using the reference picture in the reference picture list L1 and MVL1). Here, the motion vector in the L0 direction may be referred to as the L0 motion vector or MVL0, and the motion vector in the L1 direction may be referred to as the L1 motion vector or MVL1. In the case of applying bidirectional prediction, when the reference picture for L0 prediction and the reference picture for L1 prediction are in different temporal directions relative to the current picture (ie, corresponding to the case of bidirectional direction or bidirectional prediction), this may be referred to as true bidirectional prediction.
[0142] In addition, as described above, reconstructed samples and reconstructed pictures may be generated based on the derived prediction samples, and then, processes such as in-loop filtering may be performed.
[0143] In addition, when bidirectional prediction is applied to the current block, the prediction sample can be derived based on weighted average. For example, bidirectional prediction using weighted average can be referred to as bidirectional prediction with CU-level weight (BCW), bidirectional prediction with weighted average (BWA), or bidirectional prediction with weighted average.
[0144] Conventionally, a bidirectional prediction signal (i.e., a bidirectional prediction sample) can be derived by a simple average of an L0 prediction signal (L0 prediction sample) and an L1 prediction signal. That is, a bidirectional prediction sample can be derived as an average of an L0 prediction sample based on an L0 reference picture and MVL0 and an L1 prediction sample based on an L1 reference picture and MVL1. However, when bidirectional prediction is applied, a bidirectional prediction signal (bidirectional prediction sample) can be derived by a weighted average of an L0 prediction signal and an L1 prediction signal as follows. For example, a bidirectional prediction signal (bidirectional prediction sample) can be derived as in Formula 3.
[0145] [Formula 3]
[0146] P bi-pred =((8-w)*P 0 +w*P 1 +4)>>3
[0147] In Formula 3, Pbi-pred may indicate the value of a bidirectional prediction signal, that is, a predicted sample value derived by applying bidirectional prediction, and w may indicate a weight. In addition, P0 may indicate the value of an L0 prediction signal, that is, a predicted sample value derived by applying L0 prediction, and P1 may indicate the value of an L1 prediction signal, that is, a predicted sample value derived by applying L1 prediction.
[0148] For example, 5 weights may be allowed in weighted average bidirectional prediction. For example, the five weights w may include -2, 3, 4, 5, or 10. That is, the weight w may be determined as one of the weight candidates including -2, 3, 4, 5, or 10. For each CU to which bidirectional prediction is applied, the weight w may be determined by one of two methods. In the first method, the weight index may be signaled after the motion vector difference of the non-merged CU. In the second method, the weight index of the merged CU may be inferred from the neighboring blocks based on the merge candidate index.
[0149] For example, weighted average bidirectional prediction can be applied to a CU with 256 or more luma samples. That is, when the product of the width and height of the CU is greater than or equal to 256, weighted average bidirectional prediction can be applied. In the case of a low-latency picture, five weights can be used, and in the case of a non-low-latency picture, three weights can be used. For example, the three weights may include 3, 4, or 5.
[0150] For example, in an encoding device, a fast search algorithm can be applied to find a weight index without significantly increasing the complexity of the encoding device. The algorithm can be summarized as follows. For example, when the current picture is a low-delay picture when combined with adaptive motion vector resolution (AMVR) (when AMVR is used as an inter-frame prediction mode), non-equal weights can be conditionally checked for 1 pixel and 4 pixel motion vector precision. For example, when combined with affine (when the affine prediction mode is used as the inter-frame prediction mode), when the affine prediction mode is currently selected as the best mode, affine motion estimation (ME) can be performed with non-equal weights. For example, when two reference pictures of bidirectional prediction are the same, non-equal weights can be conditionally checked. For example, when specific conditions are met according to the POC distance between the current picture and the reference picture, the encoding quantization parameter (QP), and the time level, non-equal weights may not be searched.
[0151] For example, the BCW weight index may be encoded using one context coding bin followed by a bypass coding bin. The first context coding bin may indicate whether the same weights are used. When non-equal weights are used based on the first context coding bin, bypass coding may be used to signal additional bins to indicate the non-equal weights to be used.
[0152] Furthermore, when bidirectional prediction is applied, weight information for generating a prediction sample may be derived based on weight index information about a candidate selected from among candidates included in the merge candidate list.
[0153] According to an embodiment of the present disclosure, when constructing a motion vector candidate for a merge mode, weight index information about a temporal motion vector candidate may be derived as follows. For example, when a temporal motion vector candidate uses bidirectional prediction, weight index information about weighted averaging may be derived. That is, when the inter-frame prediction type is bidirectional prediction, weight index information about a temporal merge candidate in a merge candidate list (or a temporal motion vector candidate) may be derived.
[0154] For example, the weight index information about the weighted average related to the temporal motion vector candidate may always be derived as 0. Here, the weight index information being 0 may mean that the weight of each reference direction (ie, the L0 prediction direction and the L1 prediction direction in the bidirectional prediction) is the same. For example, the process for deriving the motion vector of the luma component for the merge mode may be shown in Table 1 below.
[0155] [Table 1]
[0156]
[0157]
[0158]
[0159] Referring to Table 1, gbiIdx may indicate a bidirectional prediction weight index, and gbiIdxCol may indicate a bidirectional prediction weight index for a temporal merge candidate (e.g., a temporal motion vector candidate in a merge candidate list). In the process of deriving a motion vector for a luma component in a merge mode (Table 3 of Directory 8.4.2.2), gbiIdxCol may be derived as 0. That is, the weight index of the temporal motion vector candidate may be derived as 0.
[0160] Alternatively, the weight index of the weighted average of the temporal motion vector candidates may be derived based on the weight index information about the collocated block. Here, the collocated block may be referred to as a collocated block, a co-located block, or a co-located reference block, and the collocated block may indicate a block at the same position as the current block on the reference picture. For example, the process for deriving the motion vector for the luminance component of the merge mode may be shown in Table 2 below.
[0161] [Table 2]
[0162]
[0163]
[0164]
[0165]
[0166] Referring to Table 2, gbiIdx may indicate a bidirectional prediction weight index, and gbiIdxCol may indicate a bidirectional prediction weight index for a temporal merge candidate (e.g., a temporal motion vector candidate in a merge candidate list). In the process of deriving a motion vector for a luma component of a merge mode, when the slice type or tile group type is B (Table 4 of Directory 8.4.2.2), gbiIdxCol may be derived as gbiIdxCol. That is, the weight index of the temporal motion vector candidate may be derived as the weight index of the collocated block.
[0167] In addition, according to another embodiment of the present disclosure, when the motion vector candidates for the merge mode are constructed in units of sub-blocks, the weight index of the weighted average of the temporal motion vector candidates can be derived. Here, the merge mode in units of sub-blocks can be referred to as an affine merge mode (in units of sub-blocks). The temporal motion vector candidate can indicate a sub-block-based temporal motion vector candidate and can be referred to as an SbTMVP (or ATMVP) candidate. That is, when the inter-frame prediction type is bidirectional prediction, weight index information about the SbTMVP candidate (or sub-block-based temporal motion vector candidate) in the affine merge candidate list or the sub-block merge candidate list can be derived.
[0168] For example, the weight index information about the weighted average of the sub-block-based temporal motion vector candidates may always be derived as 0. Here, the weight index information of 0 may mean that the weights of the reference directions (i.e., the L0 prediction direction and the L1 prediction direction in the bidirectional prediction) are the same. For example, the process for deriving the motion vector and the reference index in the sub-block merge mode and the process for deriving the sub-block-based temporal merge candidate may be as shown in Tables 3 and 4.
[0169] [Table 3]
[0170]
[0171]
[0172]
[0173]
[0174]
[0175]
[0176] [Table 4]
[0177]
[0178]
[0179] Referring to Tables 3 and 4 above, gbiIdx may indicate a bidirectional prediction weight index, gbiIdxSbCol may indicate a bidirectional prediction weight index of a subblock-based temporal merge candidate (e.g., a temporal motion vector candidate in a subblock-based merge candidate list), and in a process for deriving a subblock-based temporal merge candidate (8.4.4.3), gbiIdxSbCol may be derived as 0. That is, the weight index of the subblock-based temporal motion vector candidate may be derived as 0.
[0180] Alternatively, weight index information of a weighted average of sub-block-based temporal motion vector candidates may be derived based on weight index information about a temporal center block. For example, the temporal center block may indicate a sub-block or sample at a collocated block or at the center of a collocated block, and specifically, may indicate a sub-block or sample at the lower right side of four center sub-blocks or samples of a collocated block. For example, in this case, a process for deriving a motion vector and a reference index in a sub-block merge mode, a process for deriving a sub-block-based temporal merge candidate, and a process for deriving basic motion information for sub-block-based temporal merge may be shown in Tables 5, 6, and 7.
[0181] [Table 5]
[0182]
[0183]
[0184]
[0185]
[0186]
[0187] [Table 6]
[0188]
[0189]
[0190]
[0191] [Table 7]
[0192]
[0193]
[0194]
[0195] Referring to Tables 5, 6, and 7, gbiIdx may indicate a bidirectional prediction weight index, and gbiIdxSbCol may indicate a bidirectional prediction weight index of a sub-block-based temporal merge candidate (e.g., a temporal motion vector candidate in a sub-block-based merge candidate list). In the process for deriving basic motion information about sub-block-based temporal merging (8.4.4.4), gbiIdxSbCol may be derived as gbiIdxcolCb. That is, the weight index of the sub-block-based temporal motion vector candidate may be derived as the weight index of the temporal center block. For example, the temporal center block may indicate a sub-block or sample at the center of a collocated block or a collocated block, and specifically, may indicate a sub-block or sample at the lower right side of the four center sub-blocks or samples of the collocated block.
[0196] Alternatively, weight index information of a weighted average of sub-block-based temporal motion vector candidates may be derived based on weight index information in units of each sub-block, and when a sub-block is not available, the weight index information may be derived based on weight index information of a temporal center block. For example, a temporal center block may indicate a sub-block or sample at a collocated block or at the center of a collocated block, and specifically, may indicate a sub-block or sample at the lower right side of four center sub-blocks or samples of a collocated block. For example, in this case, a process for deriving a motion vector and a reference index in a sub-block merge mode, a process for deriving a sub-block-based temporal merge candidate, and a process for deriving basic motion information for sub-block-based temporal merge may be shown in Tables 8, 9, and 10.
[0197] [Table 8]
[0198]
[0199]
[0200]
[0201]
[0202]
[0203]
[0204] [Table 9]
[0205]
[0206]
[0207]
[0208] [Table 10]
[0209]
[0210]
[0211]
[0212] Referring to Tables 8, 9, and 10, gbiIdx may indicate a bidirectional prediction weight index, and gbiIdxSbCol may indicate a bidirectional prediction weight index for a subblock-based temporal merge candidate (e.g., a temporal motion vector candidate in a subblock-based merge candidate list). In a process (8.4.4.3) for deriving basic motion information about subblock-based temporal merge, gbiIdxSbCol may be derived as gbiIdxcolCb. Alternatively, in a process (8.4.4.3) for deriving basic motion information about subblock-based temporal merge according to a condition (e.g., when both availableFlagL0SbCol and availableFlagL1SbCol are 0), gbiIdxSbCol may be derived as ctrgbiIdx, and in a process (8.4.4.4) for deriving basic motion information about subblock-based temporal merge, ctrgbiIdx may be derived as gbiIdxSbCol. That is, the weight index of the sub-block-based temporal motion vector candidate may be derived as a weight index in units of each sub-block, or when the sub-block is not available, may be derived as a weight index of the temporal center block. For example, the temporal center block may indicate a sub-block or sample at the center of a collocated block or a collocated block, and specifically, may indicate a sub-block or sample at the lower right side of four center sub-blocks or samples of the collocated block.
[0213] In addition, according to another embodiment of the present disclosure, when constructing motion vector candidates for merge mode, weight index information about paired candidates can be derived. For example, paired candidates can be included in a merge candidate list, and weight index information about the weighted average of paired candidates can be derived. Paired candidates can be derived based on other merge candidates in the merge candidate list, and when the paired candidates use bidirectional prediction, the weight index of the weighted average can be derived. That is, when the inter-frame prediction type is bidirectional prediction, weight index information about paired candidates in the merge candidate list can be derived.
[0214] The pair candidate may be derived based on the other two merge candidates (eg, cand0 and cand1) among the candidates included in the merge candidate list.
[0215] For example, the weight index information about the paired candidate can be derived based on the weight index information about either of the two merge candidates (e.g., merge candidate cand0 or merge candidate cand1). For example, the weight index information about the candidate using bidirectional prediction among the two merge candidates can be used to derive the weight index information about the paired candidate.
[0216] Alternatively, when the weight index information about each of the other two merge candidates is the same as the first weight index information, the weight index information about the paired candidate can be derived based on the first weight index information. In addition, when the weight index information about each of the other two merge candidates is not the same, the weight index information about the paired candidate can be derived based on the default weight index information. The default weight index information may correspond to weight index information for assigning the same weight to each of the L0 prediction sample and the L1 prediction sample.
[0217] Alternatively, when the weight index information about each of the other two merge candidates is the same as the first weight index information, the weight index information about the paired candidate can be derived based on the first weight index information. In addition, when the weight index information about each of the other two merge candidates is not the same, the weight index information about the paired candidate includes default weight index information among the weight index information about each of the other two candidates. The default weight index information may correspond to weight index information for assigning the same weight to each of the L0 prediction sample and the L1 prediction sample.
[0218] In addition, according to another embodiment of the present disclosure, when the motion vector candidates for the merge mode are constructed in units of sub-blocks, weight index information of the weighted average of the temporal motion vector candidates can be derived. Here, the merge mode in units of sub-blocks can be referred to as an affine merge mode (in units of sub-blocks). The temporal motion vector candidate may indicate a temporal motion vector candidate based on a sub-block, and may be referred to as an SbTMVP (or ATMVP) candidate. The weight index information about the SbTMVP candidate can be derived based on the weight index information about the left neighboring block of the current block. That is, when the candidate derived by SbTMVP uses bidirectional prediction, the weight index of the left neighboring block of the current block can be derived as a weight index of the merge mode based on the sub-block.
[0219] For example, since the SbTMVP candidate can derive the collocated block based on the spatially adjacent left block (or left neighboring block) of the current block, the weight index of the left neighboring block can be considered reliable. Therefore, the weight index of the SbTMVP candidate can be derived as the weight index of the left neighboring block.
[0220] In addition, according to another embodiment of the present disclosure, when constructing motion vector candidates for affine merge mode, when the affine merge candidate uses bidirectional prediction, weight index information about weighted averaging can be derived. That is, when the inter prediction type is bidirectional prediction, weight index information about the candidates in the affine merge candidate list or the subblock merge candidate list can be derived.
[0221] For example, among the affine merge candidates, the constructed affine merge candidates can derive CP0, CP1, CP2, or CP3 candidates based on the spatial neighboring blocks or temporal neighboring blocks of the current block to indicate the candidate for deriving the MVF as an affine model. For example, CP0 can indicate a control point located at the upper left sample position of the current block, CP1 can indicate a control point located at the upper right sample position of the current block, and CP2 can indicate the lower left sample position of the current block. In addition, CP3 can indicate a control point located at the lower right sample position of the current block.
[0222] For example, among the affine merge candidates, a constructed affine merge candidate can be generated based on a combination of each control point of the current block like {CP0, CP1, CP2}, {CP0, CP1, CP3}, {CP0, CP2, CP3}, {CP1, CP2, CP3}, {CP0, CP1} and {CP0, CP2}. For example, the affine merge candidate may include at least one of {CPMV0, CPMV1, CPMV2}, {CPMV0, CPMV1, CPMV3}, {CPMV0, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV3}, {CPMV0, CPMV1} and {CPMV0, CPMV2}. CPMV0, CPMV1, CPMV2 and CPMV3 may correspond to motion vectors of CP0, CP1, CP2 and CP3, respectively.
[0223] In an embodiment, when an affine merge candidate includes a CPMV of a control point 0 (CP0) located at the upper left side of the current block, weight index information about the affine merge candidate may be derived based on weight index information about a specific block among neighboring blocks of CP0. That is, when an affine merge candidate includes a CPMV of a control point 0 (CP0) located at the upper left side of the current block, weight index information about the affine merge candidate may be derived based on the 0th weight index information about CP0. In this case, the specific block among the neighboring blocks of CP0 corresponds to a block used to derive the CPMV of CP0, and the neighboring blocks of CP0 may include an upper left neighboring block of the current block, a left neighboring block adjacent to the lower side of the upper left neighboring block, and an upper neighboring block adjacent to the right side of the upper left neighboring block.
[0224] On the other hand, when the affine merge candidate does not include the CPMV of CP0 located at the upper left side of the current block, the weight index information about the affine merge candidate can be derived based on the weight index information of a specific block among the neighboring blocks of the control point 1 (CP1) located at the upper right position of the current block. That is, when the affine merge candidate does not include the CPMV of CP0 located at the upper left side of the current block, the weight index information about the affine merge candidate can be derived based on the first weight index information of the control point 1 (CP1) located at the upper right side of the current block. Among the neighboring blocks of CP1, the specific block corresponds to the block used to derive the CPMV of CP1, and the CP1 neighboring blocks may include the upper right neighboring block of the current block and the upper neighboring block adjacent to the left side of the upper right neighboring block.
[0225] According to the above method, weight index information about affine merge candidates can be derived based on the weight index information of the blocks used to derive {CPMV0, CPMV1, CPMV2}, {CPMV0, CPMV1, CPMV3}, {CPMV0, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV3}, {CPMV0, CPMV1} and {CPMV0, CPMV2}, respectively.
[0226] According to another embodiment of deriving weight index information about an affine merge candidate, when the weight index information about CP0 located on the upper left side of the current block is the same as the weight index information about CP1 located on the upper right side of the current block, the weight index information about the affine merge candidate can be derived based on the weight index information about a specific block among the neighboring blocks of CP0. In addition, when the weight index information about CP0 located on the upper left side of the current block is different from the weight index information about CP1 located on the upper right side of the current block, the weight index information about the affine merge candidate can be derived based on default weight index information. The default weight index information may correspond to weight index information for assigning the same weight to each of the L0 prediction sample and the L1 prediction sample.
[0227] According to another embodiment of deriving weight index information about affine merge candidates, weight index information about affine merge candidates can be derived as the weight index of the candidate with high occurrence frequency among the weight index of each candidate. For example, the weight index of the candidate block determined as the motion vector in CP0 among the CP0 candidate blocks, the weight index of the candidate block determined as the motion vector in CP1 among the CP1 candidate blocks, the weight index of the candidate block determined as the motion vector in CP2 among the CP2 candidate blocks, and / or the weight index of the candidate block determined as the motion vector in CP3 among the CP3 candidate blocks, which has the most overlap, can be derived as the weight index of the affine merge candidate.
[0228] For example, CP0 and CP1 may be used as control points, CP0, CP1, and CP2 may be used, and CP3 may not be used. However, for example, when a CP3 candidate of an affine block (a block encoded in an affine prediction mode) is to be used, the method of deriving a weight index in a temporal candidate block described in the above embodiment may be used.
[0229] Fig.10 and Fig.11 is a diagram schematically illustrating an example of a video / image decoding method and related components according to an embodiment of the present disclosure.
[0230] Fig.10 The method disclosed in can be Figure 2 or Fig.11 Specifically, for example, Fig.10 S1000 to S1030 in Fig.11 The predictor 220 of the encoding device 200 in the embodiment of the present invention is executed, and Fig.10 The S1040 in Fig.11 The entropy encoder 240 of the encoding device 200 performs. In addition, although Fig.10 Not shown in the example, but in Fig.11 In the present invention, the prediction sample or prediction related information can be derived by the predictor 220 of the encoding device 200, the residual information can be derived from the original sample or the prediction sample by the residual processor 230 of the encoding device 200, and the bitstream can be generated from the residual information or prediction related information by the entropy encoder 240 of the encoding device 200. Fig.10 The method disclosed in may include the embodiments described above in the present disclosure.
[0231] Reference Fig.10 , the encoding device may determine the inter prediction mode of the current block and generate inter prediction mode information indicating the inter prediction mode (S1000). For example, the encoding device may determine the merge mode, the affine (merge) mode, or the sub-block merge mode as the inter prediction mode to be applied to the current block, and may generate inter prediction mode information indicating the determined merge mode, the affine (merge) mode, or the sub-block merge mode.
[0232] The encoding device may generate a merge candidate list of the current block based on the inter prediction mode (S1010). For example, the encoding device may generate a merge candidate list according to the determined inter prediction mode. Here, when the determined inter prediction mode is an affine merge mode or a subblock merge mode, the merge candidate list may be referred to as an affine merge candidate list or a subblock merge candidate list, but may also be simply referred to as a merge candidate list.
[0233] For example, a candidate may be inserted into the merge candidate list until the number of candidates in the merge candidate list becomes the maximum number of candidates. Here, the candidate may indicate a candidate or candidate block for deriving motion information (or motion vector) of the current block. For example, the candidate block may be derived by searching the neighboring blocks of the current block. For example, the neighboring blocks may include spatial neighboring blocks and / or temporal neighboring blocks of the current block, and the spatial neighboring blocks may be preferentially searched to derive (spatial merging) candidates, and then the temporal neighboring blocks may be searched to derive (temporal merging) candidates, and the derived candidates may be inserted into the merge candidate list. For example, when the number of candidates in the merge candidate list is less than the maximum number of candidates in the merge candidate list even after the candidate is inserted, an additional candidate may be inserted. For example, the additional candidate includes at least one of a history-based merge candidate, a pairwise average merge candidate, an ATMVP, and a combined bidirectional prediction merge candidate (when the slice / tile group type of the current slice / tile group is type B) and / or a zero vector merge candidate.
[0234] Alternatively, for example, the candidate may be inserted into the affine merge candidate list until the number of candidates in the affine merge candidate list becomes the maximum number of candidates. Here, the candidate may include a control point motion vector (CPMV) of the current block. Alternatively, the candidate may indicate a candidate or candidate block for deriving the CPMV. The CPMV may indicate a motion vector at a control point (CP) of the current block. For example, the number of CPs may be 2, 3, or 4, and the CP may be at least a portion of the upper left side (or upper left corner), upper right side (or upper right corner), lower left side (or lower left corner), or lower right side (or lower right corner) of the current block, and there may be only one CP at each position.
[0235] For example, the candidate can be derived by searching the neighboring blocks of the current block (or the neighboring blocks of the CP of the current block). For example, the affine merge candidate list may include at least one of an inherited affine merge candidate, a constructed affine merge candidate, and a zero motion vector candidate. For example, in the affine merge candidate list, the inherited affine merge candidate may be inserted first, and then the constructed affine merge candidate may be inserted. In addition, when the number of candidates in the affine merge candidate list is less than the maximum number of candidates even if the affine merge candidate constructed in the affine merge candidate list is inserted, the remainder may be filled with a zero motion vector candidate. Here, the zero motion vector candidate may be referred to as a zero vector. For example, the affine merge candidate list may be a list of affine merge modes for deriving motion vectors in units of samples, or may be a list of affine merge modes for deriving motion vectors in units of sub-blocks. In this case, the affine merge candidate list may be referred to as a sub-block merge candidate list, and the sub-block merge candidate list may also include a candidate derived from the SbTMVP (or an SbTMVP candidate). For example, when the SbTMVP candidate is included in the sub-block merge candidate list, it may be before the inherited affine merge candidate and the constructed affine merge candidate in the sub-block merge candidate list.
[0236] The encoding device may generate selection information indicating one of the candidates included in the merge candidate list (S1020). For example, the merge candidate list may include at least some of spatial merge candidates, temporal merge candidates, paired candidates, or zero vector candidates, and one of these candidates may be selected for inter-frame prediction of the current block. Alternatively, for example, the sub-block merge candidate list may include at least some of inherited affine merge candidates, constructed affine merge candidates, SbTMVP candidates, or zero vector candidates, and one of these candidates may be selected for inter-frame prediction of the current block.
[0237] For example, the selection information may include index information indicating a selected candidate in the merge candidate list. For example, the selection information may be referred to as merge index information or sub-block merge index information.
[0238] The encoding device may generate inter-prediction type information indicating the inter-prediction type of the current block as bidirectional prediction (S1030). For example, the inter-prediction type of the current block may be determined as bidirectional prediction among L0 prediction, L1 prediction, or bidirectional prediction, and inter-prediction type information indicating this may be generated. Here, L0 prediction may indicate prediction based on reference picture list 0, L1 prediction may indicate prediction based on reference picture list 1, and bidirectional prediction may indicate prediction based on reference picture list 0 and reference picture list 1. For example, the encoding device may generate inter-prediction type information based on the inter-prediction type. For example, the inter-prediction type information may include an inter_pred_idc syntax element.
[0239] The encoding device may encode image information including inter-frame prediction mode information, selection information, and inter-frame prediction type information (S1040). For example, the image information may be referred to as video information. The image information may include various information according to the above-mentioned embodiments of the present disclosure. For example, the image information may include at least a portion of prediction-related information or residual-related information. For example, the prediction-related information may include at least a portion of inter-frame prediction mode information, selection information, and inter-frame prediction type information. For example, the encoding device may generate a bitstream or encoding information by encoding all or part of the image information including the above-mentioned information (or syntax elements). Alternatively, it may be output in the form of a bitstream. In addition, the bitstream or encoding information may be sent to a decoding device via a network or a storage medium.
[0240] Despite Fig.10 Although not illustrated in the figure, the encoding device may, for example, generate a prediction sample of the current block. Alternatively, for example, the encoding device may generate a prediction sample of the current block based on a selected candidate. Alternatively, for example, the encoding device may derive motion information based on the selected candidate, and may generate a prediction sample of the current block based on the motion information. For example, the encoding device may generate an L0 prediction sample and an L1 prediction sample according to a bidirectional prediction, and may generate a prediction sample of the current block based on the L0 prediction sample and the L1 prediction sample. In this case, the prediction sample of the current block may be generated from the L0 prediction sample and the L1 prediction sample using weight index information (or weight information) for bidirectional prediction. Here, the weight information may be displayed based on the weight index information.
[0241] In other words, for example, the encoding device may generate an L0 prediction sample and an L1 prediction sample of the current block based on the selected candidate. For example, when it is determined that the inter prediction type of the current block is bidirectional prediction, the current block may be predicted using reference picture list 0 and reference picture list 1. For example, the L0 prediction sample may represent a prediction sample of the current block derived based on reference picture list 0, and the L1 prediction sample may represent a prediction sample of the current block derived based on reference picture list 1.
[0242] For example, the candidate may include a spatial merge candidate. For example, when the selected candidate is a spatial merge candidate, L0 motion information and L1 motion information may be derived based on the spatial merge candidate, and L0 prediction samples and L1 prediction samples may be generated based thereon.
[0243] For example, the candidate may include a temporal merge candidate. For example, when the selected candidate is a temporal merge candidate, L0 motion information and L1 motion information may be derived based on the temporal merge candidate, and L0 prediction samples and L1 prediction samples may be generated based thereon.
[0244] For example, the candidate may include a paired candidate. For example, when the selected candidate is a paired candidate, L0 motion information and L1 motion information may be derived based on the paired candidate, and L0 prediction samples and L1 prediction samples may be generated based on this. For example, the paired candidate may be derived based on two other merge candidates among the candidates included in the merge candidate list.
[0245] Alternatively, for example, the merge candidate list may be a subblock merge candidate list, and an affine merge candidate, a subblock merge candidate, or an SbTMVP candidate may be selected. Here, an affine merge candidate in units of subblocks may be referred to as a subblock merge candidate.
[0246] For example, the candidate may include a subblock merge candidate. For example, when the selected candidate is a subblock merge candidate, L0 motion information and L1 motion information may be derived based on the subblock merge candidate, and L0 prediction samples and L1 prediction samples may be generated based thereon. For example, the subblock merge candidate may include a control point motion vector (CPMV), and the L0 prediction sample and the L1 prediction sample may be generated by performing prediction in units of subblocks based on the CPMV.
[0247] Here, the CPMV may be indicated based on one of the neighboring blocks of the control point (CP) of the current block. For example, the number of CPs may be 2, 3, or 4, the CP may be located at least a portion of the upper left side (or upper left corner), upper right side (or upper right corner), lower left side (or lower left corner), or lower right side (or lower right corner) of the current block, and only one CP may exist at each position.
[0248] For example, the CP may be CP0 located at the upper left side of the current block. In this case, the neighboring blocks may include the upper left neighboring block of the current block, the lower left neighboring block adjacent to the lower side of the upper left neighboring block, and the upper neighboring block adjacent to the right side of the upper left neighboring block. Alternatively, the neighboring blocks may include Figure 8 A in 2 Block, B 2 Block or B 3 piece.
[0249] Alternatively, for example, the CP may be CP1 located at the upper right side of the current block. In this case, the neighboring blocks may include the upper right corner neighboring block of the current block and the upper neighboring block adjacent to the left side of the upper right corner neighboring block. Alternatively, the neighboring blocks may include Figure 8 B 0 Block or B 1 piece.
[0250] Alternatively, for example, the CP may be CP2 located at the lower left side of the current block. In this case, the neighboring blocks may include the lower left corner neighboring block of the current block and the left neighboring block adjacent to the upper side of the lower left corner neighboring block. Alternatively, the neighboring blocks may include Figure 8 A in 0 Block or A 1 piece.
[0251] Alternatively, for example, the CP may be CP3 located at the lower right side of the current block. Here, CP3 may also be referred to as RB. In this case, the neighboring block may include a collocated block of the current block or a lower right corner neighboring block of the collocated block. Here, the collocated block may include a block located at the same position as the current block in a reference picture different from the current picture in which the current block is located. Alternatively, the neighboring block may include Figure 8 The block T in .
[0252] Alternatively, for example, the candidate may include an SbTMVP candidate. For example, when the selected candidate is an SbTMVP candidate, L0 motion information and L1 motion information may be derived based on the left neighboring block of the current block, and based on this, L0 prediction samples and L1 prediction samples may be generated. For example, L0 prediction samples and L1 prediction samples may be generated by performing prediction in units of sub-blocks.
[0253] For example, the L0 motion information may include an L0 reference picture index, an L0 motion vector, etc., and the L1 motion information may include an L1 reference picture index, an L1 motion vector, etc. The L0 reference picture index may include information indicating a reference picture in reference picture list 0, and the L1 reference picture index may include information indicating a reference picture in reference picture list 1.
[0254] For example, the encoding device may generate a prediction sample of the current block based on the L0 prediction sample, the L1 prediction sample and the weight information. For example, the weight information may be displayed based on the weight index information. The weight index information may indicate weight index information about bidirectional prediction. For example, the weight information may include information about the weighted average of the L0 prediction sample or the L1 prediction sample. That is, the weight index information may indicate index information about the weight used for weighted averaging, and the weight index information may be generated in the process of generating the prediction sample based on the weighted average. For example, the weight index information may include information indicating any one of three or five weights. For example, the weighted average may represent a weighted average in bidirectional prediction (BCW) with CU-level weights or bidirectional prediction (BWA) with weighted average.
[0255] For example, the candidate may include a temporal merge candidate, and the weight index information about the temporal merge candidate may be represented by 0. That is, the weight index information about the temporal merge candidate may be represented by 0. Here, the weight index information of 0 may mean that the weight of each reference direction (i.e., the L0 prediction direction and the L1 prediction direction in the bidirectional prediction) is the same. Alternatively, for example, the candidate may include a temporal merge candidate, and the weight index information may be indicated based on the weight index information about the collocated block. That is, the weight index information about the temporal merge candidate may be indicated based on the weight index information about the collocated block. Here, the collocated block may include a block in the same position as the current block in a reference picture different from the current picture in which the current block is located.
[0256] For example, the candidate may include a paired candidate, and the weight index information may be indicated based on the weight index information about one of the other two candidates in the merged candidate list used to derive the paired candidate. That is, the weight index information about the paired candidate may be indicated based on the weight index information about one of the other two candidates in the merged candidate list used to derive the paired candidate.
[0257] For example, the candidate may include a paired candidate, and the paired candidate may be indicated based on the other two candidates among the candidates. When the weight index information about each of the other two candidates is the same as the first weight index information, the weight index information about the paired candidate may be indicated based on the first weight index information. When the weight index information about each of the other two candidates is different, the weight index information about the paired candidate may be indicated based on the default weight index information, and in this case, the default weight index information may correspond to the weight index information for assigning the same weight to each of the L0 prediction sample and the L1 prediction sample.
[0258] For example, the candidate may include a paired candidate, and the paired candidate may be indicated based on the other two candidates among the candidates. When the weight index information about each of the other two candidates is the same as the first weight index information, the weight index information about the paired candidate may be indicated based on the first weight index information. When the weight index information about each of the other two candidates is not the same, the weight index information may be indicated based on the weight index information of each of the other two candidates that is not the default weight index information. The default weight index information may correspond to weight index information for assigning the same weight to each of the L0 prediction sample and the L1 prediction sample.
[0259] For example, the merge candidate list may be a subblock merge candidate list, and an affine merge candidate, a subblock merge candidate, or an SbTMVP candidate may be selected. Here, an affine merge candidate in subblock units may be referred to as a subblock merge candidate.
[0260] For example, the candidate includes an affine merge candidate, and the affine merge candidate may include a control point motion vector (CPMV).
[0261] For example, when the affine merge candidate includes the CPMV of the control point 0 (CP0) located at the upper left side of the current block, the weight index information about the affine merge candidate may be indicated based on the weight index information about a specific block among the neighboring blocks of CP0. When the affine merge candidate does not include the CPMV of CP0 located at the upper left side of the current block, the weight index information about the affine merge candidate may be indicated based on the weight index information about a specific block among the neighboring blocks of the control point 1 (CP1) located at the upper right position of the current block.
[0262] A specific block among the neighboring blocks of CP0 corresponds to a block used to derive the CPMV of CP0, and the neighboring blocks of CP0 may include an upper left neighboring block of the current block, a left neighboring block adjacent to the lower side of the upper left neighboring block, and an upper neighboring block adjacent to the right side of the upper left neighboring block.
[0263] Among the neighboring blocks of CP1, a specific block corresponds to a block for deriving the CPMV of CP1, and the CP1 neighboring blocks may include an upper right neighboring block of the current block and an upper neighboring block adjacent to the left side of the upper right neighboring block.
[0264] Alternatively, for example, the candidate may include an SbTMVP candidate, and the weight index information about the SbTMVP candidate may be indicated based on the weight index information about the left neighboring block of the current block. That is, the weight index information about the SbTMVP candidate may be indicated based on the weight index information about the left neighboring block.
[0265] Alternatively, for example, the candidate may include an SbTMVP candidate, and the weight index information about the SbTMVP candidate may be represented by 0. That is, the weight index information about the SbTMVP candidate may be represented by 0. Here, the weight index information of 0 may mean that the weight of each reference direction (ie, the L0 prediction direction and the L1 prediction direction in the bidirectional prediction) is the same.
[0266] Alternatively, for example, the candidate may include an SbTMVP candidate, and the weight index information may be indicated based on the weight index information about the center block in the collocated block. That is, the weight index information about the SbTMVP candidate may be indicated based on the weight index information about the center block in the collocated block. Here, the collocated block may include a block in the same position as the current block in a reference picture different from the current picture in which the current block is located, and the center block may include a lower right sub-block among four sub-blocks located at the center of the collocated block.
[0267] Alternatively, for example, the candidate may include an SbTMVP candidate, and the weight index information may be indicated based on the weight index information about each subblock in the collocated block. That is, the weight index information about the SbTMVP candidate may be indicated based on the weight index information about each subblock in the collocated block.
[0268] Alternatively, although Fig.10 Although not illustrated in the example, the encoding device may, for example, derive residual samples based on the predicted samples and the original samples. In this case, residual related information may be derived based on the residual samples. Residual samples may be derived based on residual related information. Reconstructed samples may be generated based on the residual samples and the predicted samples. Reconstructed blocks and reconstructed pictures may be derived based on the reconstructed samples. Alternatively, for example, the encoding device may encode image information including residual related information or prediction related information.
[0269] For example, the encoding device can generate a bit stream or coding information by encoding all or part of the image information including the above information (or syntax elements). Alternatively, it can be output in the form of a bit stream. In addition, the bit stream or coding information can be sent to a decoding device via a network or a storage medium. Alternatively, the bit stream or coding information can be stored in a computer-readable storage medium, and the bit stream or coding information can be generated by the above image coding method.
[0270] Fig.12 and Fig.13 is a diagram schematically illustrating an example of a video / image decoding method and related components according to an embodiment of the present disclosure.
[0271] Fig.12 The method disclosed in can be Figure 3 or Fig.13 Specifically, for example, Fig.12 The S1200 in Fig.13 The entropy decoder 310 of the decoding device 300 in the embodiment of the present invention is executed, and Fig.12 S1210 to S1260 in Fig.13 The predictor 330 of the decoding device 300 in FIG. Fig.12 Not shown in the example, but in Fig.13 In the present invention, the entropy decoder 310 of the decoding device 300 can derive prediction-related information or residual information from the bitstream, the residual processor 320 of the decoding device 300 can derive residual samples from the residual information, the predictor 330 of the decoding device 300 can derive prediction samples from the prediction-related information, and the adder 340 of the decoding device 300 can derive a reconstructed block or a reconstructed picture from the residual samples or the prediction samples. Fig.12The method disclosed in may include the embodiments described above in the present disclosure.
[0272] Reference Fig.12 , the decoding device may receive image information including inter-frame prediction mode information and inter-frame prediction type information through a bit stream (S1200). For example, the image information may be referred to as video information. The image information may include various information according to the above-mentioned embodiments of the present disclosure. For example, the image information may include at least a portion of prediction-related information or residual-related information.
[0273] For example, the prediction-related information may include inter-frame prediction mode information or inter-frame prediction type information. For example, the inter-frame prediction mode information may include information indicating at least some of the various inter-frame prediction modes. For example, various modes such as merge mode, skip mode, motion vector prediction (MVP) mode, affine mode, sub-block merge mode, or merge with MVD (MMVD) mode may be used. In addition, the motion vector refinement (DMVR) mode, adaptive motion vector resolution (AMVR) mode, bidirectional prediction with CU-level weights or bidirectional optical flow (BDOF) on the decoder side, etc. may be used additionally or alternatively as auxiliary modes. For example, the inter-frame prediction type information may include an inter_pred_idc syntax element. Alternatively, the inter-frame prediction type information may include information indicating any one of L0 prediction, L1 prediction, and bidirectional prediction.
[0274] The decoding device may generate a merge candidate list of the current block based on the inter prediction mode information (S1210). For example, the decoding device may determine the inter prediction mode of the current block as a merge mode, an affine (merge) mode, or a sub-block merge mode based on the inter prediction mode information, and generate a merge candidate list according to the determined inter prediction mode. Here, when the inter prediction mode is an affine merge mode or a sub-block merge mode, the merge candidate list may be referred to as an affine merge candidate list or a sub-block merge candidate list, but may also be simply referred to as a merge candidate list.
[0275] For example, a candidate may be inserted into the merge candidate list until the number of candidates in the merge candidate list becomes the maximum number of candidates. Here, the candidate may indicate a candidate or candidate block for deriving motion information (or motion vector) of the current block. For example, the candidate block may be derived by searching the neighboring blocks of the current block. For example, the neighboring blocks may include spatial neighboring blocks and / or temporal neighboring blocks of the current block, and the spatial neighboring blocks may be preferentially searched to derive (spatial merging) candidates, and then the temporal neighboring blocks may be searched to derive (temporal merging) candidates, and the derived candidates may be inserted into the merge candidate list. For example, when the number of candidates in the merge candidate list is less than the maximum number of candidates in the merge candidate list even after the candidate is inserted, an additional candidate may be inserted. For example, the additional candidate includes at least one of a history-based merge candidate, a pairwise average merge candidate, an ATMVP, and a combined bidirectional prediction merge candidate (when the slice / tile group type of the current slice / tile group is type B) and / or a zero vector merge candidate.
[0276] Alternatively, for example, the candidate may be inserted into the affine merge candidate list until the number of candidates in the affine merge candidate list becomes the maximum number of candidates. Here, the candidate may include a control point motion vector (CPMV) of the current block. Alternatively, the candidate may indicate a candidate or candidate block for deriving the CPMV. The CPMV may indicate a motion vector at a control point (CP) of the current block. For example, the number of CPs may be 2, 3, or 4, and the CP may be at least a portion of the upper left side (or upper left corner), upper right side (or upper right corner), lower left side (or lower left corner), or lower right side (or lower right corner) of the current block, and there may be only one CP at each position.
[0277] For example, the candidate block can be derived by searching the neighboring blocks of the current block (or the neighboring blocks of the CP of the current block). For example, the affine merge candidate list may include at least one of an inherited affine merge candidate, a constructed affine merge candidate, and a zero motion vector candidate. For example, in the affine merge candidate list, the inherited affine merge candidate may be inserted first, and then the constructed affine merge candidate may be inserted. In addition, when the number of candidates in the affine merge candidate list is less than the maximum number of candidates even if the affine merge candidate constructed in the affine merge candidate list is inserted, the remainder may be filled with a zero motion vector candidate. Here, the zero motion vector candidate may be referred to as a zero vector. For example, the affine merge candidate list may be a list of affine merge modes for deriving motion vectors in units of samples, or may be a list of affine merge modes for deriving motion vectors in units of sub-blocks. In this case, the affine merge candidate list may be referred to as a sub-block merge candidate list, and the sub-block merge candidate list may also include a candidate derived from the SbTMVP (or an SbTMVP candidate). For example, when the SbTMVP candidate is included in the sub-block merge candidate list, it may be before the inherited affine merge candidate and the constructed affine merge candidate in the sub-block merge candidate list.
[0278] The decoding device may generate selection information indicating one of the candidates included in the merge candidate list (S1220). For example, the merge candidate list may include at least some of spatial merge candidates, temporal merge candidates, paired candidates, or zero vector candidates, and one of these candidates may be selected for inter-frame prediction of the current block. Alternatively, for example, the sub-block merge candidate list may include at least some of inherited affine merge candidates, constructed affine merge candidates, SbTMVP candidates, or zero vector candidates, and one of these candidates may be selected for inter-frame prediction of the current block. For example, the selected candidate may be selected from the merge candidate list based on the selection information. For example, the selection information may include index information indicating the selected candidate in the merge candidate list. For example, the selection information may be referred to as merge index information or sub-block merge index information. For example, the selection information may be included in the image information. Alternatively, the selection information may be included in the inter-frame prediction mode information.
[0279] The decoding device may derive the inter prediction type of the current block as bidirectional prediction based on the inter prediction type information (S1230). For example, the inter prediction type of the current block may be derived as bidirectional prediction among L0 prediction, L1 prediction, or bidirectional prediction based on the inter prediction type information. Here, L0 prediction may indicate prediction based on reference picture list 0, L1 prediction may indicate prediction based on reference picture list 1, and bidirectional prediction may indicate prediction based on reference picture list 0 and reference picture list 1. For example, the inter prediction type information may include an inter_pred_idc syntax element.
[0280] The decoding device may derive motion information about the current block based on the selected candidate (S1240). For example, the decoding device may derive L0 motion information and L1 motion information based on the candidate of the inter-frame prediction type derived as bidirectional prediction. For example, the L0 motion information may include an L0 reference picture index, an L0 motion vector, etc., and the L1 motion information may include an L1 reference picture index, an L1 motion vector, etc. The L0 reference picture index may include information indicating a reference picture in reference picture list 0, and the L1 reference picture index may include information indicating a reference picture in reference picture list 1.
[0281] The decoding device may generate an L0 prediction sample and an L1 prediction sample of the current block based on the motion information (S1250). For example, when the inter prediction type of the current block is derived as bidirectional prediction, the current block may be predicted using reference picture list 0 and reference picture list 1. For example, the L0 prediction sample may represent a prediction sample of the current block derived based on reference picture list 0, and the L1 prediction sample may represent a prediction sample of the current block derived based on reference picture list 1.
[0282] For example, the candidate may include a spatial merge candidate. For example, when the selected candidate is a spatial merge candidate, L0 motion information and L1 motion information may be derived based on the spatial merge candidate, and L0 prediction samples and L1 prediction samples may be generated based thereon.
[0283] For example, the candidate may include a temporal merge candidate. For example, when the selected candidate is a temporal merge candidate, L0 motion information and L1 motion information may be derived based on the temporal merge candidate, and L0 prediction samples and L1 prediction samples may be generated based thereon.
[0284] For example, the candidate may include a paired candidate. For example, when the selected candidate is a paired candidate, L0 motion information and L1 motion information may be derived based on the paired candidate, and L0 prediction samples and L1 prediction samples may be generated based on this. For example, the paired candidate may be derived based on two other candidates among the candidates included in the merge candidate list.
[0285] Alternatively, for example, the merge candidate list may be a subblock merge candidate list, and an affine merge candidate, a subblock merge candidate, or an SbTMVP candidate may be selected. Here, an affine merge candidate in units of subblocks may be referred to as a subblock merge candidate.
[0286] For example, the candidate may include an affine merge candidate. For example, when the selected candidate is an affine merge candidate, L0 motion information and L1 motion information may be derived based on the affine merge candidate, and L0 prediction samples and L1 prediction samples may be generated based on this. For example, the affine merge candidate may include a control point motion vector (CPMV), and the L0 prediction sample and the L1 prediction sample may be generated by performing prediction in units of sub-blocks based on the CPMV.
[0287] Here, the CPMV may be derived based on one of the neighboring blocks of the control point (CP) of the current block. For example, the number of CPs may be 2, 3, or 4, the CP may be located at least a portion of the upper left side (or upper left corner), upper right side (or upper right corner), lower left side (or lower left corner), or lower right side (or lower right corner) of the current block, and there may be only one CP at each position.
[0288] For example, the CP may be CP0 located at the upper left side of the current block. In this case, the neighboring blocks may include the upper left neighboring block of the current block, the left neighboring block adjacent to the lower side of the upper left neighboring block, and the upper neighboring block adjacent to the right side of the upper left neighboring block. Alternatively, the neighboring blocks may include Figure 8 A in 2 Block, B 2 Block or B 3 piece.
[0289] Alternatively, for example, the CP may be CP1 located at the upper right side of the current block. In this case, the neighboring blocks may include the upper right corner neighboring block of the current block and the upper neighboring block adjacent to the left side of the upper right corner neighboring block. Alternatively, the neighboring blocks may include Figure 8 B 0 Block or B 1 piece.
[0290] Alternatively, for example, the CP may be CP2 located at the lower left side of the current block. In this case, the neighboring blocks may include the lower left corner neighboring block of the current block and the left neighboring block adjacent to the upper side of the lower left corner neighboring block. Alternatively, the neighboring blocks may include Figure 8 A in 0 Block or A 1 piece.
[0291] Alternatively, for example, the CP may be CP3 located at the lower right side of the current block. Here, CP3 may also be referred to as RB. In this case, the neighboring block may include a collocated block of the current block or a lower right corner neighboring block of the collocated block. Here, the collocated block may include a block located at the same position as the current block in a reference picture different from the current picture in which the current block is located. Alternatively, the neighboring block may include Figure 8 The block T in .
[0292] Alternatively, for example, the candidate may include an SbTMVP candidate. For example, when the selected candidate is an SbTMVP candidate, L0 motion information and L1 motion information may be derived based on the left neighboring block of the current block, and based on this, L0 prediction samples and L1 prediction samples may be generated. For example, L0 prediction samples and L1 prediction samples may be generated by performing prediction in units of sub-blocks.
[0293] For example, the decoding device may generate a prediction sample of the current block based on the L0 prediction sample, the L1 prediction sample and the weight information (S1260). For example, the weight information may be derived based on the weight index information of the candidate selected from the candidates included in the merge candidate list. For example, the weight information may include information about the weighted average of the L0 prediction sample or the L1 prediction sample. That is, the weight index information may indicate index information about the weight used for weighted averaging, and the weighted average may be performed based on the weight index information. For example, the weight index information may include information indicating any one of three or five weights. For example, the weighted average may represent a weighted average in a bidirectional prediction (BCW) with a CU-level weight or a bidirectional prediction (BWA) with a weighted average.
[0294] For example, the candidate may include a temporal merge candidate, and the weight index information about the temporal merge candidate may be derived as 0. That is, the weight index information about the temporal merge candidate may be derived as 0. Here, the weight index information of 0 may mean that the weight of each reference direction (ie, the L0 prediction direction and the L1 prediction direction in the bidirectional prediction) is the same.
[0295] For example, the candidate may include a temporal merge candidate, and the weight index information about the temporal merge candidate may be derived based on the weight index information about the collocated block. That is, the weight index information about the temporal merge candidate may be derived based on the weight index information about the collocated block. Here, the collocated block may include a block in a reference picture different from a current picture in which the current block is located that is in the same position as the current block.
[0296] For example, the candidate may include a paired candidate, and the weight index information may be derived as weight index information about one of the other two candidates in the merged candidate list used to derive the paired candidate. That is, the weight index information about the paired candidate may be derived as weight index information about one of the other two candidates in the merged candidate list used to derive the paired candidate.
[0297] For example, the candidate may include a paired candidate, and the paired candidate may be derived based on the other two candidates among the candidates. When the weight index information about each of the other two candidates is the same as the first weight index information, the weight index information for the paired candidate may be derived based on the first weight index information. When the weight index information about each of the other two candidates is different, the weight index information about the paired candidate may be derived based on the default weight index information, and in this case, the default weight index information may correspond to the weight index information for assigning the same weight to each of the L0 prediction sample and the L1 prediction sample.
[0298] For example, the candidate may include a paired candidate, and the paired candidate may be derived based on the other two candidates among the candidates. When the weight index information about each of the other two candidates is the same as the first weight index information, the weight index information for the paired candidate may be derived based on the first weight index information. When the weight index information about each of the other two candidates is different, the weight index information may be derived based on the weight index information of each of the other two candidates that is not the default weight index information. The default weight index information may correspond to weight index information for assigning the same weight to each of the L0 prediction sample and the L1 prediction sample.
[0299] For example, the merge candidate list may be a subblock merge candidate list, and an affine merge candidate, a subblock merge candidate, or an SbTMVP candidate may be selected. Here, an affine merge candidate in subblock units may be referred to as a subblock merge candidate.
[0300] For example, the candidate includes an affine merge candidate, and the affine merge candidate may include a control point motion vector (CPMV).
[0301] For example, when the affine merge candidate includes the CPMV of the control point 0 (CP0) located at the upper left side of the current block, the weight index information about the affine merge candidate can be derived based on the weight index information of the specific block among the neighboring blocks of CP0. When the affine merge candidate does not include the CPMV of CP0 located at the upper left side of the current block, the weight index information about the affine merge candidate can be derived based on the weight index information of the specific block among the neighboring blocks of the control point 1 (CP1) located at the upper right position of the current block.
[0302] A specific block among the neighboring blocks of CP0 corresponds to a block used to derive the CPMV of CP0, and the neighboring blocks of CP0 may include an upper left neighboring block of the current block, a left neighboring block adjacent to the lower side of the upper left neighboring block, and an upper neighboring block adjacent to the right side of the upper left neighboring block.
[0303] Among the neighboring blocks of CP1, a specific block corresponds to a block for deriving the CPMV of CP1, and the CP1 neighboring blocks may include an upper right neighboring block of the current block and an upper neighboring block adjacent to the left side of the upper right neighboring block.
[0304] Alternatively, for example, the candidate may include an SbTMVP candidate, and the weight index information about the SbTMVP candidate may be derived based on the weight index information about the left neighboring block of the current block. That is, the weight index information about the SbTMVP candidate may be derived based on the weight index information about the left neighboring block.
[0305] Alternatively, for example, the candidate may include an SbTMVP candidate, and the weight index information about the SbTMVP candidate may be derived as 0. That is, the weight index information about the SbTMVP candidate may be derived as 0. Here, the weight index information of 0 may mean that the weight of each reference direction (ie, the L0 prediction direction and the L1 prediction direction in the bidirectional prediction) is the same.
[0306] Alternatively, for example, the candidate may include an SbTMVP candidate, and the weight index information may be derived based on the weight index information about the center block in the collocated block. That is, the weight index information about the SbTMVP candidate may be derived based on the weight index information about the center block in the collocated block. Here, the collocated block may include a block in the same position as the current block in a reference picture different from the current picture in which the current block is located, and the center block may include a lower right sub-block among the four sub-blocks located at the center of the collocated block.
[0307] Alternatively, for example, the candidate may include an SbTMVP candidate, and the weight index information may be derived based on the weight index information about each subblock in the collocated block. That is, the weight index information about the SbTMVP candidate may be derived based on the weight index information about each subblock of the collocated block.
[0308] Despite Fig.12 Although not illustrated in the figure, the decoding device may, for example, derive residual samples based on residual related information included in the image information. In addition, the decoding device may generate reconstructed samples based on the predicted samples and the residual samples. A reconstructed block and a reconstructed picture may be derived based on the reconstructed samples.
[0309] For example, the decoding device can obtain video / image information including all or part of the above-mentioned multiple pieces of information (or syntax elements) by decoding the bit stream or coding information. In addition, the bit stream or coding information can be stored in a computer-readable storage medium, and can cause the above-mentioned decoding method to be executed.
[0310] Although the method has been described based on a flowchart that lists steps or blocks in sequence in the above-mentioned embodiments, the steps of this document are not limited to a specific order, and specific steps can be performed in different steps or in a different order or simultaneously relative to the above-mentioned steps. In addition, it will be understood by a person of ordinary skill in the art that the steps in the flowchart are not exclusive, and another step may be included therein, or one or more steps in the flowchart may be deleted without affecting the scope of the present disclosure.
[0311] The above-mentioned method according to the present disclosure may be in the form of software, and the encoding device and / or decoding device according to the present disclosure may be included in a device for performing image processing (e.g., TV, computer, smart phone, set-top box, display device, etc.).
[0312] When the embodiments of the present disclosure are implemented by software, the above-mentioned methods can be implemented by modules (processing or functions) that perform the above-mentioned functions. The module can be stored in a memory and executed by a processor. The memory can be installed inside or outside the processor and can be connected to the processor via various well-known devices. The processor may include an application-specific integrated circuit (ASIC), other chipsets, logic circuits and / or data processing devices. The memory may include a read-only memory (ROM), a random access memory (RAM), a flash memory, a memory card, a storage medium and / or other storage devices. In other words, according to the embodiments of the present disclosure, it can be implemented and executed on a processor, a microprocessor, a controller or a chip. For example, the functional units illustrated in the corresponding figures can be implemented and executed on a computer, a processor, a microprocessor, a controller or a chip. In this case, information about the implementation (e.g., information about instructions) or an algorithm can be stored in a digital storage medium.
[0313] In addition, the decoding device and the encoding device of the embodiment of the present document can be included in a multimedia broadcast transceiver, a mobile communication terminal, a home theater video device, a digital theater video device, a surveillance camera, a video chat device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a portable camera, a video on demand (VoD) service provider, an over-the-top (OTT) video device, an Internet streaming service provider, a 3D video device, a virtual reality (VR) device, an augmented reality (AR) device, an image phone video device, a vehicle terminal (e.g., a vehicle (including an autonomous vehicle) terminal, an aircraft terminal or a ship terminal) and a medical video device; and can be used to process image signals or data. For example, an OTT video device can include a game console, a Blueray player, a networked TV, a home theater system, a smart phone, a tablet PC, and a digital video recorder (DVR).
[0314] In addition, the processing method of the embodiment of the application of this document can be generated in the form of a program executed by a computer and can be stored in a computer-readable recording medium. The multimedia data with a data structure according to the embodiment of this document can also be stored in a computer-readable recording medium. The computer-readable recording medium includes all kinds of storage devices and distributed storage devices storing computer-readable data. The computer-readable recording medium may include, for example, a Blu-ray disc (BD), a universal serial bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. The computer-readable recording medium also includes a medium implemented in the form of a carrier wave (e.g., transmission on the Internet). In addition, the bit stream generated by the encoding method can be stored in a computer-readable recording medium, or can be transmitted through a wired or wireless communication network.
[0315] In addition, the embodiments of the present document can be implemented as a computer program product based on a program code, and the program code can be executed on a computer according to the embodiments of the present document. The program code can be stored on a computer readable carrier.
[0316] Fig.14 An example of a content streaming system to which the embodiments of this document can be applied is shown.
[0317] Reference Fig.14 A content streaming system to which the embodiments of this document are applied may generally include an encoding server, a streaming server, a network server, a media storage, a user device, and a multimedia input device.
[0318] The encoding server is used to compress the content input from the multimedia input device such as a smart phone, a camera, a camcorder, etc. into digital data, generate a bit stream, and transmit it to the streaming server. As another example, in the case where the multimedia input device such as a smart phone, a camera, a camcorder, etc. directly generates a bit stream, the encoding server can be omitted.
[0319] The bitstream may be generated by the encoding method or the bitstream generating method to which the embodiments of this document are applied, and the streaming server may temporarily store the bitstream during the process of transmitting or receiving the bitstream.
[0320] The streaming server transmits multimedia data to the user device based on the user's request through the network server, which acts as a tool to inform the user of what services exist. When the user requests the service the user wants, the network server transfers the request to the streaming server, and the streaming server transmits the multimedia data to the user. In this regard, the content streaming system may include a separate control server, and in this case, the control server is used to control the commands / responses between the various devices in the content streaming system.
[0321] The streaming server may receive content from a media storage device and / or an encoding server. For example, in the case where the content is received from the encoding server, the content may be received in real time. In this case, the streaming server may store the bitstream for a predetermined period of time to smoothly provide a streaming service.
[0322] For example, user devices may include mobile phones, smart phones, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation, tablet PCs, tablet PCs, ultrabooks, wearable devices (e.g., watch-type terminals (smart watches), glasses-type terminals (smart glasses), head-mounted displays (HMDs)), digital TVs, desktop computers, digital signage, etc.
[0323] Each server in the content streaming system may be operated as a distributed server, and in this case, data received by each server may be processed in a distributed manner.
[0324] The claims in this specification can be combined in various ways. For example, the technical features in the method claims of this specification can be combined to be implemented or performed in a device, and the technical features in the device claims can be combined to be implemented or performed in a method. In addition, the technical features in the method claims and the device claims can be combined to be implemented or performed in a device. In addition, the technical features in the method claims and the device claims can be combined to be implemented or performed in a method.
Claims
1. An image decoding method performed by a decoding device, the image decoding method comprising the following steps: Obtaining image information including inter-frame prediction mode information through a bit stream; deriving inherited affine candidates and constructed affine candidates based on neighboring blocks of the current block and the inter prediction mode information; generating a merge candidate list of the current block based on the inherited affine candidates and the constructed affine candidates; Selecting a candidate from among the candidates included in the merge candidate list; deriving motion information about the current block based on the selected candidate; Generate an L0 prediction sample and an L1 prediction sample of the current block based on the motion information; as well as generating a prediction sample of the current block based on the L0 prediction sample, the L1 prediction sample and weight information, wherein the weight information is weight information for bidirectional prediction and is derived based on a weight index of the candidate for selection, Wherein, the constructed affine candidates include control point motion vectors CPMV, and Wherein, based on the situation that the constructed affine candidate is generated based on {CP0, CP1, CP3}, the weight index for the constructed affine candidate is derived based on the weight index for the CP0, wherein the CP0 is related to the upper left corner of the current block, the CP1 is related to the upper right corner of the current block, and the CP3 is related to the lower right corner of the current block.
2. An image encoding method performed by an encoding device, the image encoding method comprising the following steps: determining an inter prediction mode of a current block, and generating inter prediction mode information indicating the inter prediction mode; deriving inherited affine candidates and constructed affine candidates based on neighboring blocks of the current block and the inter prediction mode information; generating a merge candidate list of the current block based on the inherited affine candidates and the constructed affine candidates; generating selection information indicating a selected one of the candidates included in the merge candidate list; deriving motion information about the current block based on the selected candidate; Generate an L0 prediction sample and an L1 prediction sample of the current block based on the motion information; generating a prediction sample of the current block based on the L0 prediction sample, the L1 prediction sample and weight information, wherein the weight information is weight information for bidirectional prediction and is derived based on a weight index of the candidate for selection; and encoding the image information including the inter-frame prediction mode information and the selection information, Wherein, the constructed affine candidates include control point motion vectors CPMV, and Wherein, based on the situation that the constructed affine candidate is generated based on {CP0, CP1, CP3}, the weight index for the constructed affine candidate is derived based on the weight index for the CP0, wherein the CP0 is related to the upper left corner of the current block, the CP1 is related to the upper right corner of the current block, and the CP3 is related to the lower right corner of the current block.
3. A method for transmitting image data, the method comprising the following steps: Obtaining a bitstream for the image, wherein the bitstream is generated based on the following steps: determining an inter prediction mode of a current block and generating inter prediction mode information indicating the inter prediction mode, deriving inherited affine candidates and constructed affine candidates based on neighboring blocks of the current block and the inter prediction mode information, generating a merge candidate list of the current block based on the inherited affine candidates and the constructed affine candidates, generating selection information indicating a selected one of the candidates included in the merge candidate list, deriving motion information about the current block based on the selected candidate, generating L0 prediction samples and L1 prediction samples of the current block based on the motion information, generating prediction samples of the current block based on the L0 prediction samples, the L1 prediction samples and weight information, wherein the weight information is weight information for bidirectional prediction and is derived based on a weight index for the selected candidate, and encoding image information including the inter prediction mode information and the selection information; and sending said data comprising said bit stream, Wherein, the constructed affine candidates include control point motion vectors CPMV, and Wherein, based on the situation that the constructed affine candidate is generated based on {CP0, CP1, CP3}, the weight index for the constructed affine candidate is derived based on the weight index for the CP0, wherein the CP0 is related to the upper left corner of the current block, the CP1 is related to the upper right corner of the current block, and the CP3 is related to the lower right corner of the current block.