Image encoding / decoding method, bitstream transmission method, and recording medium storing bitstream
The image encoding/decoding method addresses the high cost of transmitting and storing high-resolution images by using bidirectional prediction units to enhance encoding/decoding efficiency and improve compression through adaptive motion vector differential signaling.
Patent Information
- Application Number
- JP2024547653
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-02-15
- Filing Date
- 2023-02-15
- Publication Date
- 2025-09-17
- Estimated Expiration
- 2043-02-15
AI Technical Summary
The increasing demand for high-resolution, high-quality images leads to a significant increase in transmission and storage costs due to the higher amount of information required, necessitating highly efficient image compression techniques.
An image encoding/decoding method that employs bidirectional prediction based on different types of prediction units, including sub-block-unit and non-sub-block-unit prediction types, to improve encoding/decoding efficiency and adaptively derive motion vector differentials.
This method enhances encoding/decoding efficiency by adaptively deriving motion information for each prediction direction, improving compression efficiency and bit efficiency for signaling motion vector differentials.
Smart Images

Figure 0007741336000024 
Figure 0007741336000025 
Figure 0007741336000026
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an image encoding / decoding method, a bitstream transmission method, and a recording medium storing a bitstream, and more particularly, to bidirectional prediction based on different types of prediction units during inter prediction. [Background technology]
[0002] Recently, demand for high-resolution, high-quality images, such as HD (High Definition) images and UHD (Ultra High Definition) images, has been increasing in various fields. As image data becomes higher in resolution and quality, the amount of information or bits to be transmitted increases relatively compared to conventional image data. The increase in the amount of information or bits to be transmitted results in an increase in transmission costs and storage costs.
[0003] This requires highly efficient image compression techniques for effectively transmitting, storing, and reproducing high-resolution, high-quality image information. Summary of the Invention [Problem to be solved by the invention]
[0004] An object of the present disclosure is to provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.
[0005] Another object of the present disclosure is to provide a method for performing bidirectional prediction based on prediction units of different types.
[0006] Another object of the present disclosure is to provide various methods for deriving prediction candidates for bidirectional prediction.
[0007] Another object of the present disclosure is to provide a method for encoding and obtaining motion vector differentials.
[0008] Another object of the present disclosure is to provide a non-transitory computer-readable recording medium for storing a bitstream generated by the image encoding method according to the present disclosure.
[0009] Another object of the present disclosure is to provide a non-transitory computer-readable recording medium for storing a bitstream that is received by an image decoding device according to the present disclosure, decoded, and used to restore an image.
[0010] Another object of the present disclosure is to provide a method for transmitting a bitstream generated by the image encoding method according to the present disclosure.
[0011] The technical problems to be solved by the present disclosure are not limited to the above-mentioned technical problems, and other technical problems not mentioned above will be clearly understood by a person having ordinary skill in the technical field to which the present disclosure pertains from the following description. [Means for solving the problem]
[0012] An image decoding method according to one aspect of the present disclosure may be an image decoding method performed by an image decoding device, and may include: a step of obtaining first information based on the prediction direction of a current block being bidirectional; a step of deriving at least one bidirectional motion information candidate corresponding to the different prediction types from a reference block based on the first information indicating that different prediction types are applied to each prediction direction of the current block; and a step of bidirectionally predicting the current block with the different prediction types based on one of the at least one bidirectional motion information candidate, wherein the different prediction types are a sub-block-unit (sub PU) prediction type and a non-sub-block-unit (non-sub PU) prediction type.
[0013]
[0016] According to another aspect of the present invention, an image encoding method is an image encoding method performed by an image encoding device, the image encoding method including: determining, based on the prediction direction of the current block being bidirectional, whether different prediction types are applied to each prediction direction of the current block - first information indicating whether different prediction types are applied to each prediction direction of the current block is encoded into a bitstream; deriving at least one bidirectional motion information candidate corresponding to the different prediction types from a reference block based on the different prediction types being applied to each prediction direction of the current block; and bidirectionally predicting the current block with the different prediction types based on one of the at least one bidirectional motion information candidate, wherein the different prediction types are a sub-block unit (sub PU) prediction type and a non-sub-block unit (non-sub PU) prediction type.
[0014] A computer-readable recording medium according to another aspect of the present disclosure can store a bitstream generated by the image encoding method or apparatus of the present disclosure.
[0015] A transmission method according to another aspect of the present disclosure can transmit a bitstream generated by the image coding method or apparatus of the present disclosure.
[0016] The features described above in this brief summary of the present disclosure are merely exemplary aspects of the detailed description of the present disclosure that follows and are not intended to limit the scope of the present disclosure. [Effects of the Invention]
[0017] According to the present disclosure, an image encoding / decoding method and apparatus with improved encoding / decoding efficiency can be provided.
[0018] Furthermore, according to the present disclosure, motion information is adaptively derived for each prediction direction, thereby improving compression efficiency.
[0019] Furthermore, according to the present disclosure, candidates for inducing different motion information for each prediction direction can be efficiently configured.
[0020] Additionally, the present disclosure may improve bit efficiency for signaling motion vector differentials.
[0021] The effects obtained by the present disclosure are not limited to the effects described above, and other effects not described above will be clearly understood by those having ordinary skill in the art to which the present disclosure pertains from the following description. [Brief explanation of the drawings]
[0022] [Figure 1] 1 is a diagram illustrating a video coding system to which embodiments of the present disclosure can be applied;
[0023] [Figure 2] 1 is a diagram schematically illustrating an image encoding device to which an embodiment of the present disclosure can be applied.
[0024] [Figure 3] FIG. 1 is a diagram schematically illustrating an image decoding device to which an embodiment of the present disclosure can be applied.
[0025] [Figure 4] FIG. 1 is a diagram illustrating an inter-prediction unit of an image encoding device.
[0026] [Figure 5] 1 is a flowchart illustrating a method for encoding an image based on inter prediction.
[0027] [Figure 6] FIG. 1 is a diagram illustrating an inter prediction unit of an image decoding device.
[0028] [Figure 7] 1 is a flowchart illustrating a method for decoding an image based on inter prediction.
[0029] [Figure 8] 1 is a flowchart illustrating an inter prediction method.
[0030] [Figure 9] 1 is a flowchart illustrating an image encoding method according to an embodiment of the present disclosure.
[0031] [Figure 10] 1 is a flowchart illustrating an image decoding method according to an embodiment of the present disclosure.
[0032] [Figure 11] FIG. 1 is a diagram illustrating an example of a bidirectional prediction method according to an embodiment of the present disclosure.
[0033] [Figure 12-15] FIG. 10 is a diagram illustrating an example of deriving bidirectional motion information candidates.
[0034] [Figure 16] 10 is a flowchart illustrating an image encoding method and an image decoding method according to another embodiment of the present disclosure.
[0035] [Figure 17] 10 is a flowchart illustrating an image encoding method and an image decoding method according to another embodiment of the present disclosure.
[0036] [Figure 18] 10A and 10B are diagrams illustrating examples of surrounding blocks used to derive bidirectional motion information candidates.
[0037] [Figure 19] 10 is a flowchart illustrating an image encoding method according to another embodiment of the present disclosure.
[0038] [Figure 20] 10 is a flowchart illustrating an image decoding method according to another embodiment of the present disclosure.
[0039] [Figure 21] 10 is a flowchart illustrating an image encoding method according to another embodiment of the present disclosure.
[0040] [Figure 22] 10 is a flowchart illustrating an image decoding method according to another embodiment of the present disclosure.
[0041] [Figure 23] FIG. 10 is a diagram illustrating an example of signaling motion vector differentials according to the present disclosure.
[0042] [Figure 24] 10A and 10B are diagrams illustrating various positions of a reference block;
[0043] [Figure 25] FIG. 1 is a diagram illustrating an exemplary content streaming system to which an embodiment of the present disclosure can be applied. DETAILED DESCRIPTION OF THE INVENTION
[0044] The present disclosure will be described in detail below with reference to the accompanying drawings, so that those skilled in the art can easily implement the present disclosure. However, the present disclosure may be embodied in various different forms and is not limited to the embodiments described herein.
[0045] In describing the embodiments of the present disclosure, if it is determined that a detailed description of a known configuration or function may obscure the gist of the present disclosure, the detailed description thereof will be omitted. In addition, in the drawings, parts that are not related to the description of the present disclosure will be omitted, and similar parts will be designated by similar reference numerals.
[0046] In this disclosure, when a component is referred to as being "coupled," "coupled," or "connected" to another component, this includes not only a direct connection, but also an indirect connection where another component exists between them. Furthermore, when a component is referred to as "including" or "having" another component, this does not mean that the other component is excluded, but that the component can further include the other component, unless otherwise specified.
[0047] In this disclosure, terms such as "first" and "second" are used only to distinguish one component from another component, and do not limit the order or importance of the components unless otherwise specified. Therefore, within the scope of this disclosure, a first component in one embodiment may be referred to as a second component in another embodiment, and similarly, a second component in one embodiment may be referred to as a first component in another embodiment.
[0048] In this disclosure, components that are distinguished from one another are used to clearly describe the characteristics of each component and do not necessarily mean that the components are separate. In other words, multiple components may be integrated into a single hardware or software unit, or a single component may be distributed into multiple hardware or software units. Therefore, even if not otherwise specified, such integrated or distributed embodiments are also included within the scope of this disclosure.
[0049] In this disclosure, the components described in various embodiments are not necessarily essential components, and some may be optional components. Therefore, an embodiment consisting of a subset of the components described in one embodiment is also within the scope of this disclosure. Furthermore, an embodiment including other components in addition to the components described in various embodiments is also within the scope of this disclosure.
[0050] The present disclosure relates to image encoding and decoding, and terms used in this disclosure may have their ordinary meaning in the technical field to which the present disclosure belongs unless they are newly defined in this disclosure.
[0051] In this disclosure, a "picture" generally refers to a unit representing any one image in a specific time period, and a slice / tile is a coding unit constituting a part of a picture, and one picture may be composed of one or more slices / tiles. Furthermore, a slice / tile may include one or more coding tree units (CTUs).
[0052] In this disclosure, "pixel" or "pel" may refer to the smallest unit constituting one picture (or image). Also, "sample" may be used as a term corresponding to pixel. A sample may generally indicate a pixel or a pixel value, and may indicate only a pixel / pixel value of a luma component, or may indicate only a pixel / pixel value of a chroma component.
[0053] In this disclosure, the term "unit" may refer to a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to that region. The term "unit" may be used interchangeably with terms such as "sample array," "block," or "area," depending on the situation. In general, an M×N block may include a set (or array) of samples or transform coefficients consisting of M columns and N rows.
[0054] In the present disclosure, a "current block" may refer to any one of a "current coding block," a "current coding unit," a "block to be coded," a "block to be decoded," or a "block to be processed." When prediction is performed, a "current block" may refer to a "current predicted block" or a "block to be predicted." When transformation (inverse transformation) / quantization (inverse quantization) is performed, a "current block" may refer to a "current transformed block" or a "block to be transformed." When filtering is performed, a "current block" may refer to a "block to be filtered."
[0055] Furthermore, in this disclosure, unless explicitly stated as a chroma block, the term "current block" may refer to a block including both a luma component block and a chroma component block, or the "luma block of the current block." The luma component block of the current block may be expressed explicitly as a "luma block" or a "current luma block," including the explicit description of the luma component block. Furthermore, the chroma component block of the current block may be expressed explicitly as a "chroma block" or a "current chroma block," including the explicit description of the chroma component block.
[0056] In the present disclosure, " / " and "," can be interpreted as "and / or." For example, "A / B" and "A, B" can be interpreted as "A and / or B." Also, "A / B / C" and "A, B, C" can mean "at least one of A, B, and / or C."
[0057] In this disclosure, "or" can be interpreted as "and / or." For example, "A or B" can mean 1) only "A," 2) only "B," or 3) "A and B." Alternatively, in this disclosure, "or" can mean "additionally or alternatively."
[0058] Video Coding System Overview
[0059] FIG. 1 is a diagram illustrating a video coding system to which embodiments of the present disclosure can be applied.
[0060] A video coding system according to one embodiment may include an encoding device 10 and a decoding device 20. The encoding device 10 may transmit encoded video and / or image information or data to the decoding device 20 in a file or streaming format via a digital storage medium or a network.
[0061] An encoding device 10 according to an embodiment may include a video source generation unit 11, an encoding unit 12, and a transmission unit 13. A decoding device 20 according to an embodiment may include a reception unit 21, a decoding unit 22, and a rendering unit 23. The encoding unit 12 may be referred to as a video / image encoding unit, and the decoding unit 22 may be referred to as a video / image decoding unit. The transmission unit 13 may be included in the encoding unit 12. The reception unit 21 may be included in the decoding unit 22. The rendering unit 23 may include a display unit, which may be configured as a separate device or an external component.
[0062] The video source generation unit 11 can acquire video / images through a video / image capture, synthesis, or generation process. The video source generation unit 11 can include a video / image capture device and / or a video / image generation device. The video / image capture device can include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. The video / image generation device can include, for example, a computer, a tablet, a smartphone, etc., and can (electronically) generate video / images. For example, virtual video / images can be generated via a computer, etc., in which case the video / image capture process can be replaced with a process in which related data is generated.
[0063] The encoder 12 may encode the input video / image. The encoder 12 may perform a series of steps such as prediction, transformation, and quantization for compression and coding efficiency. The encoder 12 may output the encoded data (encoded video / image information) in a bitstream format.
[0064] The transmitting unit 13 can acquire coded video / image information or data output in a bitstream format and transmit it to the receiving unit 21 of the decoding device 20 or another external object in a file or streaming format via a digital storage medium or a network. Digital storage media can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The transmitting unit 13 can include elements for generating a media file in a predetermined file format and elements for transmitting it via a broadcasting / communication network. The transmitting unit 13 can be provided as a transmitting device separate from the encoding device 12. In this case, the transmitting device can include at least one processor for acquiring coded video / image information or data output in a bitstream format and a transmitting unit for transmitting it in a file or stream format. The receiving unit 21 can extract / receive the bitstream from the storage medium or network and transmit it to the decoding unit 22.
[0065] The decoding unit 22 can decode the video / image by performing a series of steps such as inverse quantization, inverse transformation, and prediction corresponding to the operations of the encoding unit 12.
[0066] The rendering unit 23 can render the decoded video / images, and the rendered video / images can be displayed via the display unit.
[0067] Overview of the image encoding device
[0068] FIG. 2 is a diagram schematically illustrating an image encoding device to which an embodiment of the present disclosure can be applied.
[0069] 2, the image encoding device 100 may include an image division unit 110, a subtraction unit 115, a transform unit 120, a quantization unit 130, an inverse quantization unit 140, an inverse transform unit 150, an addition unit 155, a filtering unit 160, a memory 170, an inter prediction unit 180, an intra prediction unit 185, and an entropy encoding unit 190. The inter prediction unit 180 and the intra prediction unit 185 may be collectively referred to as a "prediction unit." The transform unit 120, the quantization unit 130, the inverse quantization unit 140, and the inverse transform unit 150 may be included in a residual processing unit. The residual processing unit may further include a subtraction unit 115.
[0070] Depending on the embodiment, all or at least some of the components constituting the image encoding device 100 may be realized by a single hardware component (e.g., an encoder or a processor). Also, the memory 170 may include a decoded picture buffer (DPB) and may be realized by a digital storage medium.
[0071] The image division unit 110 may divide an input image (or picture, frame) input to the image encoding device 100 into one or more processing units. As an example, the processing units may be called coding units (CUs). The coding units may be obtained by recursively dividing a coding tree unit (CTU) or a largest coding unit (LCU) using a QT / BT / TT (quad-tree / binary-tree / ternary-tree) structure. For example, one coding unit may be divided into multiple coding units at deeper depths based on a quad-tree structure, a binary-tree structure, and / or a ternary-tree structure. To divide the coding units, the quad-tree structure may be applied first, and then the binary-tree structure and / or the ternary-tree structure may be applied later. The coding procedure according to the present disclosure may be performed based on the final coding unit that is not further divided. The maximum coding unit may be used as the final coding unit, or a lower-depth coding unit obtained by dividing the maximum coding unit may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and / or reconstruction, which will be described later. As another example, a processing unit of the coding procedure may be a prediction unit (PU) or a transform unit (TU). The prediction unit and the transform unit may be divided or partitioned from the final coding unit, respectively. The prediction unit may be a unit of sample prediction, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.
[0072] The prediction unit (inter prediction unit 180 or intra prediction unit 185) may perform prediction on a current block (current block) to generate a predicted block including prediction samples for the current block. The prediction unit may determine whether intra prediction or inter prediction is applied to the current block or CU. The prediction unit may generate various information related to prediction of the current block and transmit it to the entropy coding unit 190. The prediction information may be coded by the entropy coding unit 190 and output in a bitstream format.
[0073] The intra prediction unit 185 may predict the current block by referring to samples in the current picture. The referenced samples may be located in the neighborhood of the current block or may be located far away from the current block according to the intra prediction mode and / or intra prediction technique. The intra prediction modes may include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes may include, for example, DC mode and Planar mode. The directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the granularity of the prediction direction. However, this is merely an example, and more or less directional prediction modes may be used depending on the settings. The intra prediction unit 185 may also determine the prediction mode to be applied to the current block using the prediction modes applied to neighboring blocks.
[0074] The inter prediction unit 180 may derive a predicted block for a current block based on a reference block (reference sample array) identified by a motion vector on a reference picture. To reduce the amount of motion information transmitted in inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation between the motion information of neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on the inter prediction direction (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block may be called a collocated reference block, a collocated CU (colCU), etc. The reference picture including the temporal neighboring block may be called a collocated picture (colPic). For example, the inter predictor 180 may construct a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to derive a motion vector and / or a reference picture index for the current block. Inter prediction may be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the inter predictor 180 may use motion information of neighboring blocks as motion information for the current block. In the case of skip mode, unlike in merge mode, a residual signal may not be transmitted.In the case of a motion vector prediction (MVP) mode, the motion vector of a neighboring block is used as a motion vector predictor, and the motion vector of the current block can be signaled by encoding a motion vector difference and an indicator for the motion vector predictor. The motion vector difference may mean the difference between the motion vector of the current block and the motion vector predictor.
[0075] The predictor may generate a prediction signal based on various prediction methods and / or prediction techniques, which will be described later. For example, the predictor may apply intra prediction or inter prediction to predict the current block, or may simultaneously apply intra prediction and inter prediction. A prediction method that simultaneously applies intra prediction and inter prediction to predict the current block may be referred to as combined inter and intra prediction (CIIP). The predictor may also perform intra block copy (IBC) to predict the current block. Intra block copy can be used for content image / video coding, such as screen content coding (SCC), for games. IBC is a method of predicting a current block using an already reconstructed reference block in a current picture that is located a predetermined distance away from the current block. When IBC is applied, the position of the reference block in the current picture may be coded as a vector (block vector) corresponding to the predetermined distance. IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that a reference block is derived within the current picture. That is, the IBC may use at least one of the inter prediction techniques described in this disclosure.
[0076] The prediction signal generated by the prediction unit may be used to generate a restored signal or a residual signal. The subtraction unit 115 may subtract the prediction signal (predicted block, predicted sample array) output from the prediction unit from the input image signal (original block, original sample array) to generate a residual signal (residual signal, residual block, residual sample array). The generated residual signal may be transmitted to the conversion unit 120.
[0077] The transform unit 120 may generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loeve transform (KLT), a graph-based transform (GBT), or a conditionally non-linear transform (CNT). Here, the GBT refers to a transform obtained from a graph representing inter-pixel relationship information. The CNT refers to a transform obtained based on a predicted signal generated using all previously reconstructed pixels. The transform process may be applied to pixel blocks having the same square size or to non-square blocks of variable size.
[0078] The quantization unit 130 may quantize the transform coefficients and transmit the quantized transform coefficients to the entropy coding unit 190. The entropy coding unit 190 may encode the quantized signal (information about the quantized transform coefficients) and output the encoded signal in a bitstream format. The information about the quantized transform coefficients may be referred to as residual information. The quantization unit 130 may rearrange the quantized transform coefficients in a block format into a one-dimensional vector format based on a coefficient scan order, and may generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector format.
[0079] The entropy coding unit 190 may perform various coding methods, such as exponential Golomb, context-adaptive variable length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy coding unit 190 may also code information necessary for video / image restoration (e.g., values of syntax elements) together with or separately from the quantized transform coefficients. The coded information (e.g., coded video / image information) may be transmitted or stored in a bitstream format in network abstraction layer (NAL) units. The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / image information may also include general constraint information. The signaling information, transmitted information and / or syntax elements mentioned in this disclosure may be encoded through the above-described encoding procedure and included in the bitstream.
[0080] The bitstream may be transmitted via a network or stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as a USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitting unit (not shown) that transmits and / or a storing unit (not shown) that stores the signal output from the entropy encoding unit 190 may be provided as an internal / external element of the image encoding device 100, or the transmitting unit may be provided as a component of the entropy encoding unit 190.
[0081] The quantized transform coefficients output from the quantization unit 130 can be used to generate a residual signal. For example, the residual signal (residual block or residual sample) can be reconstructed by applying inverse quantization and inverse transform to the quantized transform coefficients via the inverse quantization unit 140 and the inverse transform unit 150.
[0082] The adder 155 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter prediction unit 180 or the intra prediction unit 185. When there is no residual for the current block to be processed, such as when a skip mode is applied, the predicted block may be used as the reconstructed block. The adder 155 may be referred to as a reconstruction unit or a reconstructed block generation unit. The generated reconstructed signal may be used for intra prediction of the next current block to be processed in the current picture, and may also be used for inter prediction of the next picture after filtering, as will be described later.
[0083] The filtering unit 160 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 160 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture and store the modified reconstructed picture in the memory 170, specifically, in the DPB of the memory 170. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, etc. The filtering unit 160 may generate various information related to filtering and transmit it to the entropy coding unit 190, as will be described later in connection with each filtering method. The filtering information may be coded by the entropy coding unit 190 and output in a bitstream format.
[0084] The modified reconstructed picture transmitted to the memory 170 can be used as a reference picture in the inter prediction unit 180. When inter prediction is applied through this, the image encoding device 100 can avoid a prediction mismatch between the image encoding device 100 and the image decoding device, and can also improve encoding efficiency.
[0085] The DPB in the memory 170 may store modified reconstructed pictures for use as reference pictures in the inter predictor 180. The memory 170 may store motion information of blocks from which motion information in the current picture is derived (or coded) and / or motion information of already reconstructed intra-picture blocks. The stored motion information may be transmitted to the inter predictor 180 to be used as motion information of spatially surrounding blocks or temporally surrounding blocks. The memory 170 may store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra predictor 185.
[0086] Overview of the image decoding device
[0087] FIG. 3 is a diagram schematically illustrating an image decoding device to which an embodiment of the present disclosure can be applied.
[0088] 3, the image decoding apparatus 200 may include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an adder 235, a filtering unit 240, a memory 250, an inter prediction unit 260, and an intra prediction unit 265. The inter prediction unit 260 and the intra prediction unit 265 may be collectively referred to as a "prediction unit." The inverse quantization unit 220 and the inverse transform unit 230 may be included in a residual processing unit.
[0089] Depending on the embodiment, all or at least some of the components constituting the image decoding device 200 may be realized by a single hardware component (e.g., a decoder or a processor). Also, the memory 170 may include a DPB and may be realized by a digital storage medium.
[0090] The image decoding device 200, which receives a bitstream including video / image information, can reconstruct an image by performing a process corresponding to the process performed by the image encoding device 100 of FIG. 2. For example, the image decoding device 200 can perform decoding using a processing unit applied in the image encoding device. Therefore, the decoding processing unit can be, for example, a coding unit. The coding unit can be obtained by dividing a coding tree unit or a maximum coding unit. The reconstructed image signal decoded and output by the image decoding device 200 can be reproduced by a reproduction device (not shown).
[0091] The image decoding apparatus 200 may receive a signal output from the image encoding apparatus of FIG. 2 in a bitstream format. The received signal may be decoded via an entropy decoding unit 210. For example, the entropy decoding unit 210 may parse the bitstream to derive information (e.g., video / image information) necessary for image reconstruction (or picture reconstruction). The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / image information may also include general constraint information. The image decoding apparatus may further use the information on the parameter sets and / or the general constraint information to decode an image. The signaling information, received information, and / or syntax elements referred to in the present disclosure may be obtained from the bitstream by being decoded via the decoding procedure. For example, the entropy decoding unit 210 may decode information in a bitstream based on a coding method such as Exponential-Golomb coding, CAVLC, or CABAC, and output values of syntax elements required for image restoration and quantized values of transform coefficients related to residuals. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element from the bitstream, determines a context model using information on the syntax element to be decoded and decoded information on neighboring blocks and the block to be decoded, or information on symbols / bins decoded in a previous step, predicts the occurrence probability of the bins based on the determined context model, and performs arithmetic decoding of the bins to generate symbols corresponding to the values of each syntax element. After determining the context model, the CABAC entropy decoding method may update the context model using information on the decoded symbol / bin for the context model of the next symbol / bin.Among the information decoded by the entropy decoding unit 210, information related to prediction is provided to the prediction units (inter prediction unit 260 and intra prediction unit 265), and residual values entropy decoded by the entropy decoding unit 210, i.e., quantized transform coefficients and related parameter information, may be input to the inverse quantization unit 220. Also, among the information decoded by the entropy decoding unit 210, information related to filtering may be provided to the filtering unit 240. Meanwhile, a receiving unit (not shown) for receiving a signal output from the image encoding device may be further provided as an internal / external element of the image decoding device 200, or the receiving unit may be provided as a component of the entropy decoding unit 210.
[0092] Meanwhile, the image decoding apparatus according to the present disclosure may be referred to as a video / image / picture decoding apparatus. The image decoding apparatus may include an information decoder (video / image / picture information decoder) and / or a sample decoder (video / image / picture sample decoder). The information decoder may include an entropy decoding unit 210, and the sample decoder may include at least one of an inverse quantization unit 220, an inverse transform unit 230, an adder 235, a filtering unit 240, a memory 250, an inter prediction unit 260, and an intra prediction unit 265.
[0093] The inverse quantization unit 220 may inverse quantize the quantized transform coefficients and output the transform coefficients. The inverse quantization unit 220 may rearrange the quantized transform coefficients in a two-dimensional block format. In this case, the rearrangement may be performed based on the coefficient scanning order performed in the image encoding device. The inverse quantization unit 220 may perform inverse quantization on the quantized transform coefficients using a quantization parameter (e.g., quantization step size information) to obtain transform coefficients.
[0094] The inverse transform unit 230 can inversely transform the transform coefficients to obtain a residual signal (residual block, residual sample array).
[0095] The prediction unit may perform prediction on a current block and generate a predicted block including prediction samples for the current block. The prediction unit may determine whether intra prediction or inter prediction is applied to the current block based on information about the prediction output from the entropy decoding unit 210, and may determine a specific intra / inter prediction mode (prediction technique).
[0096] The prediction unit can generate a prediction signal based on various prediction methods (techniques) described below, as described in the description of the prediction unit of the image encoding device 100.
[0097] The intra predictor 265 may predict the current block by referring to samples in the current picture. The description of the intra predictor 185 may also be applied to the intra predictor 265.
[0098] The inter prediction unit 260 may derive a predicted block for a current block based on a reference block (reference sample array) identified by a motion vector on a reference picture. To reduce the amount of motion information transmitted in inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on correlations between motion information of neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on an inter prediction direction (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. For example, the inter prediction unit 260 may construct a motion information candidate list based on the neighboring blocks and derive a motion vector and / or a reference picture index for the current block based on received candidate selection information. Inter prediction may be performed based on various prediction modes (techniques), and the prediction information may include information indicating the inter prediction mode (technique) for the current block.
[0099] The adder 235 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to a prediction signal (predicted block, predicted sample array) output from a prediction unit (including the inter prediction unit 260 and / or intra prediction unit 265). When there is no residual for the current block, such as when a skip mode is applied, the predicted block can be used as the reconstructed block. The description of the adder 155 also applies to the adder 235. The adder 235 may also be referred to as a reconstruction unit or a reconstructed block generator. The generated reconstructed signal may be used for intra prediction of the next current block in the current picture, and may also be used for inter prediction of the next picture via filtering, as described below.
[0100] The filtering unit 240 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 240 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture, and may store the modified reconstructed picture in the memory 250, specifically, in a DPB of the memory 250. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, etc.
[0101] The (modified) reconstructed picture stored in the DPB of the memory 250 can be used as a reference picture in the inter predictor 260. The memory 250 can store motion information of a block from which motion information in the current picture is derived (or decoded) and / or motion information of a block in an already reconstructed picture. The stored motion information can be transmitted to the inter predictor 260 to be used as motion information of a spatially surrounding block or a temporally surrounding block. The memory 250 can store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra predictor 265.
[0102] In this specification, the embodiments described for the filtering unit 160, inter prediction unit 180 and intra prediction unit 185 of the image encoding device 100 can also be applied in a similar or corresponding manner to the filtering unit 240, inter prediction unit 260 and intra prediction unit 265 of the image decoding device 200, respectively.
[0103] Inter Prediction
[0104] The prediction units of the image encoding device 100 and the image decoding device 200 may perform inter prediction on a block-by-block basis to derive prediction samples. Inter prediction may be derived in a manner dependent on data elements (e.g., sample values or motion information) of pictures other than the current picture. When inter prediction is applied to the current block, a predicted block (prediction sample array) for the current block may be derived based on a reference block (reference sample array) identified by a motion vector in a reference picture indicated by a reference picture index. In this case, to reduce the amount of motion information transmitted in the inter prediction mode, motion information of the current block may be predicted in block, sub-block, or sample units based on correlation between motion information of neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on an inter prediction type (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). When inter prediction is applied, the neighboring blocks may include spatial neighboring blocks in the current picture and temporal neighboring blocks in the reference picture. The reference picture including the reference block and the reference picture including the temporal peripheral block may be the same or different. The temporal peripheral block may be called a collocated reference block, a collocated CU (colCU), etc., and the reference picture including the temporal peripheral block may be called a collocated picture (colPic). For example, a motion information candidate list may be constructed based on the peripheral blocks of the current block, and flag or index information indicating which candidate is selected (used) to derive the motion vector and / or reference picture index of the current block may be signaled.Inter prediction may be performed based on various prediction modes. For example, in skip mode and merge mode, the motion information of the current block may be the same as the motion information of the selected neighboring block. In skip mode, unlike merge mode, a residual signal may not be transmitted. In motion vector prediction (MVP) mode, the motion vector of a selected neighboring block may be used as a motion vector predictor, and a motion vector difference may be signaled. In this case, the motion vector of the current block may be derived using the sum of the motion vector predictor and the motion vector difference.
[0105] The motion information may include L0 motion information and / or L1 motion information depending on the inter-prediction type (such as L0 prediction, L1 prediction, or Bi prediction). A motion vector in the L0 direction may be referred to as an L0 motion vector or MVL0, and a motion vector in the L1 direction may be referred to as an L1 motion vector or MVL1. Prediction based on an L0 motion vector may be referred to as L0 prediction, prediction based on an L1 motion vector may be referred to as L1 prediction, and prediction based on both the L0 motion vector and the L1 motion vector may be referred to as bi-prediction (Bi-prediction). Here, the L0 motion vector may represent a motion vector associated with a reference picture list L0 (L0), and the L1 motion vector may represent a motion vector associated with a reference picture list L1 (L1). The reference picture list L0 may include, as reference pictures, pictures that precede the current picture in output order, and the reference picture list L1 may include pictures that follow the current picture in output order. The previous pictures may be referred to as forward (reference) pictures, and the following pictures may be referred to as backward (reference) pictures. The reference picture list L0 may further include, as reference pictures, pictures that follow the current picture in output order. In this case, the previous picture may be indexed first in the reference picture list L0, and the subsequent picture may be indexed next. The reference picture list L1 may further include, as reference pictures, pictures that follow the current picture in output order. In this case, the subsequent picture may be indexed first in the reference picture list L1, and the previous picture may be indexed next. Here, the output order may correspond to a picture order count (POC) order.
[0106] FIG. 4 is a diagram illustrating the inter prediction unit (180) of the image decoding device 100, and FIG. 5 is a flowchart illustrating a method for encoding an image based on inter prediction.
[0107] The image encoding device 100 may perform inter prediction on a current block (S510). The image encoding device 100 may derive an inter prediction mode and motion information of the current block and generate a prediction sample for the current block. Here, the steps of determining the inter prediction mode, deriving the motion information, and generating the prediction sample may be performed simultaneously, or one step may be performed before the other steps. For example, the inter prediction unit 180 of the image encoding device 100 may include a prediction mode determination unit 181, a motion information derivation unit 182, and a prediction sample derivation unit 183. The prediction mode determination unit 181 may determine a prediction mode for the current block, the motion information derivation unit 182 may derive motion information of the current block, and the prediction sample derivation unit 183 may derive a prediction sample for the current block. For example, the inter prediction unit 180 of the image encoding device 100 may search for blocks similar to the current block within a certain region (search region) of a reference picture through motion estimation and derive a reference block whose difference from the current block is minimum or equal to or less than a certain criterion. Based on this, a reference picture index indicating a reference picture in which the reference block is located may be derived, and a motion vector may be derived based on the position difference between the reference block and the current block. The image encoding device 100 may determine a mode to be applied to the current block from various prediction modes. The image encoding device 100 may compare RD costs for the various prediction modes and determine an optimal prediction mode for the current block.
[0108] For example, when a skip mode or a merge mode is applied to the current block, the image encoding device 100 may construct a merge candidate list (described below) and derive a reference block whose difference from the current block is minimum or equal to or less than a certain criterion among reference blocks indicated by merge candidates included in the merge candidate list. In this case, a merge candidate associated with the derived reference block may be selected, and merge index information indicating the selected merge candidate may be generated and signaled to a decoding device. Motion information of the current block may be derived using motion information of the selected merge candidate.
[0109] As another example, when the (A)MVP mode is applied to the current block, the image encoding apparatus 100 may construct an (A)MVP candidate list (described later) and use a motion vector of a selected MVP (motion vector predictor) candidate from among the MVP candidates included in the (A)MVP candidate list as the MVP of the current block. In this case, for example, a motion vector pointing to a reference block derived by the above-described motion estimation may be used as the motion vector of the current block, and an MVP candidate having a motion vector with the smallest difference from the motion vector of the current block may be the selected MVP candidate. A motion vector difference (MVD), which is a difference obtained by subtracting the MVP from the motion vector of the current block, may be derived. In this case, information regarding the MVD may be signaled to the image decoding apparatus 200. Furthermore, when the (A)MVP mode is applied, the value of the reference picture index may be configured as reference picture index information and separately signaled to the image decoding apparatus 200.
[0110] The image encoding apparatus 100 may derive residual samples based on the predicted samples (S520). The image encoding apparatus 100 may derive the residual samples by comparing the original samples of the current block with the predicted samples.
[0111] The image encoding device 100 may encode image information including prediction information and residual information (S530). The image encoding device 100 may output the encoded image information in a bitstream format. The prediction information may include prediction mode information (e.g., a skip flag, a merge flag, or a mode index) and information on motion information as information related to the prediction procedure. The information on the motion information may include candidate selection information (e.g., a merge index, an MVP flag, or an MVP index) that is information for deriving a motion vector. The information on the motion information may also include the above-mentioned information on MVD and / or reference picture index information. The information on the motion information may also include information indicating whether L0 prediction, L1 prediction, or bi-prediction is applied. The residual information is information on the residual sample. The residual information may include information on quantized transform coefficients for the residual sample.
[0112] The output bitstream may be stored in a (digital) storage medium and then transmitted to the image decoding device 200, or may be transmitted to the image decoding device 200 via a network.
[0113] Meanwhile, as described above, the image coding apparatus 100 can generate a reconstructed picture (including reconstructed samples and reconstructed blocks) based on the reference samples and the residual samples. This is because the image coding apparatus 100 derives the same prediction result as that performed in the image decoding apparatus 200, thereby improving coding efficiency. Therefore, the image coding apparatus 100 can store the reconstructed picture (or reconstructed samples, reconstructed blocks) in a memory and use it as a reference picture for inter prediction. As described above, an in-loop filtering procedure can be further applied to the reconstructed picture.
[0114] FIG. 6 is a diagram schematically illustrating the inter prediction unit 260 of the image decoding device 200, and FIG. 7 is a flowchart showing a method for decoding an image based on inter prediction.
[0115] The image decoding apparatus 200 may perform operations corresponding to those performed by the image encoding apparatus 100. The image decoding apparatus 200 may perform prediction on a current block based on received prediction information and derive predicted samples.
[0116] Specifically, the image decoding apparatus 200 may determine a prediction mode for the current block based on the received prediction information (S710). The image decoding apparatus 200 may determine which inter prediction mode is applied to the current block based on prediction mode information in the prediction information.
[0117] For example, it may determine whether the merge mode or (A)MVP mode is applied to the current block based on the merge flag. Alternatively, it may select one of various inter prediction mode candidates based on the mode index. The inter prediction mode candidates may include skip mode, merge mode, and / or (A)MVP mode, or may include various inter prediction modes described below.
[0118] The image decoding apparatus 200 may derive motion information of the current block based on the determined inter prediction mode (S720). For example, when a skip mode or a merge mode is applied to the current block, the image decoding apparatus 200 may construct a merge candidate list (described below) and select one merge candidate from among the merge candidates included in the merge candidate list. The selection may be performed based on the above-described selection information (merge index). Motion information of the selected merge candidate may be used to derive motion information of the current block. The motion information of the selected merge candidate may be used as motion information of the current block.
[0119] As another example, when the (A)MVP mode is applied to the current block, the image decoding apparatus 200 may construct an (A)MVP candidate list (described later) and use a motion vector of a selected MVP (motion vector predictor) candidate from among the MVP candidates included in the (A)MVP candidate list as the MVP of the current block. The selection may be performed based on the selection information (MVP flag or MVP index) described above. In this case, the MVD of the current block may be derived based on information related to the MVD, and a motion vector of the current block may be derived based on the MVP of the current block and the MVD. Furthermore, a reference picture index of the current block may be derived based on the reference picture index information. A picture pointed to by the reference picture index in the reference picture list for the current block may be derived as a reference picture referenced for inter-prediction of the current block.
[0120] Meanwhile, as will be described later, the motion information of the current block may be derived without constructing a candidate list, and in this case, the motion information of the current block may be derived according to a procedure disclosed in a prediction mode, which will be described later. In this case, the candidate list construction as described above may be omitted.
[0121] The image decoding apparatus 200 may generate prediction samples for the current block based on the motion information of the current block (S730). In this case, the reference picture may be derived based on a reference picture index of the current block, and the prediction samples of the current block may be derived using samples of a reference block to which the motion vector of the current block points on the reference picture. In this case, as described below, a prediction sample filtering procedure may further be performed on all or some of the prediction samples of the current block, as necessary.
[0122] For example, the inter prediction unit 260 of the image decoding device 200 may include a prediction mode determination unit 261, a motion information derivation unit 262, and a prediction sample derivation unit 263, and may determine a prediction mode for the current block based on prediction mode information received by the prediction mode determination unit 181, derive motion information (such as a motion vector and / or a reference picture index) of the current block based on information regarding the motion information received by the motion information derivation unit 182, and derive a prediction sample of the current block by the prediction sample derivation unit 183.
[0123] The image decoding apparatus 200 may generate residual samples for the current block based on the received residual information (S740). The image decoding apparatus 200 may generate reconstructed samples for the current block based on the predicted samples and the residual samples, and may generate a reconstructed picture based on the reconstructed samples (S750). Thereafter, an in-loop filtering procedure may be further applied to the reconstructed picture, as described above.
[0124] 8, as described above, the inter prediction procedure may include an inter prediction mode determination step (S810), a motion information derivation step (S820) based on the determined prediction mode, and a prediction (prediction sample generation) step (S830) based on the derived motion information. The inter prediction procedure may be performed in the image encoding device 100 and the image decoding device 200, as described above.
[0125] Inter prediction mode determination
[0126] Various inter-prediction modes can be used to predict a current block in a picture. For example, various modes can be used, such as merge mode, skip mode, motion vector prediction (MVP) mode, affine mode, sub-block merge mode, and merge with MVD (MVVD) mode. Decoder side motion vector refinement (DMVR) mode, adaptive motion vector resolution (AMVR) mode, bi-prediction with CU-level weight (BCW), and bi-directional optical flow (BDOF) can be used in addition to or instead of the accompanying modes. Affine mode may also be referred to as affine motion prediction mode. MVP mode may also be referred to as advanced motion vector prediction mode. Motion information candidates derived by some modes and / or some modes in this document may also be included as one of the motion information-related candidates of other modes. For example, an HMVP candidate may be added as a merge candidate in the merge / skip mode, or as an MVP candidate in the MVP mode.
[0127] Prediction mode information indicating the inter prediction mode of the current block may be signaled from the image encoding apparatus 100 to the image decoding apparatus 200. The prediction mode information may be included in a bitstream and received by the image decoding apparatus 200. The prediction mode information may include index information indicating one of multiple candidate modes. Alternatively, the inter prediction mode may be indicated through hierarchical signaling of flag information. In this case, the prediction mode information may include one or more flags. For example, a skip flag may be signaled to indicate whether or not to apply the skip mode. If the skip mode is not applied, a merge flag may be signaled to indicate whether or not to apply the merge mode. If the merge mode is not applied, an MVP mode may be indicated to be applied. Alternatively, a flag for additional classification may be further signaled. The affine mode may be signaled in an independent mode or in a mode dependent on the merge mode or MVP mode. For example, the affine mode may include an affine merge mode and an affine MVP mode.
[0128] Derivation of motion information
[0129] Inter prediction may be performed using motion information of the current block. The image encoding device 100 may derive optimal motion information for the current block through a motion estimation procedure. For example, the image encoding device 100 may search for a similar reference block with high correlation using an original block in an original picture for the current block in fractional pixel units within a predetermined search range in the reference picture, thereby deriving motion information. Block similarity may be derived based on a phase-based sample value difference. For example, block similarity may be calculated based on the SAD between the current block (or a template of the current block) and the reference block (or a template of the reference block). In this case, motion information may be derived based on the reference block with the smallest SAD within the search range. The derived motion information may be signaled to the image decoding device 200 according to various methods based on the inter prediction mode.
[0130] Generating prediction samples
[0131] A predicted block for the current block may be derived based on motion information derived according to a prediction mode. The predicted block may include prediction samples (prediction sample array) of the current block. If the motion vector of the current block points to a fractional sample unit, an interpolation procedure may be performed, thereby deriving prediction samples of the current block based on reference samples in fractional sample units within a reference picture. If affine inter-prediction is applied to the current block, prediction samples may be generated based on sample / sub-block unit MVs. If bi-prediction is applied, prediction samples derived through a weighted sum or weighted average (by phase) of prediction samples derived based on L0 prediction (i.e., prediction using a reference picture in a reference picture list L0 and MVL0) and prediction samples derived based on L1 prediction (i.e., prediction using a reference picture in a reference picture list L1 and MVL1) may be used as prediction samples of the current block. When bi-prediction is applied, if the reference picture used for L0 prediction and the reference picture used for L1 prediction are located in different temporal directions relative to the current picture (i.e., if it is bi-predictive but also corresponds to bidirectional prediction), this can be called true bi-prediction.
[0132] As described above, reconstructed samples and reconstructed pictures can be generated based on the derived predicted samples, and then procedures such as in-loop filtering can be performed.
[0133] Affine prediction
[0134] Existing video coding systems use a single motion vector (MV) to represent the motion of a coding block (using a translation motion model). However, while the above method can represent optimal motion on a block-by-block basis, if an optimal MV can be determined on a pixel-by-pixel basis, rather than the optimal motion of each actual pixel, coding efficiency can be improved. To this end, this embodiment describes an affine motion prediction method that uses an affine motion model for coding. The fine motion prediction method can represent MVs on a pixel-by-pixel basis in a block using two, three, or four MVs.
[0135] An affine motion model can express four types of motion. Of the motions that can be expressed by an affine motion model, an affine motion model that expresses three types of motion (translation, scale, and rotate) is called a similarity (or simplified) affine motion model. Here, the proposed method will be described based on the similarity (or simplified) affine motion model. However, the method is not limited to this motion model.
[0136] Affine motion prediction can determine the MV of a pixel position included in a block using two or more control point motion vectors (CPMVs). In this case, a collection of MVs is called an affine motion vector field (MVF), and in the encoding / decoding process, the affine MVF can be determined in pixel units or in predefined sub-block units. When determined in pixel units, the MV is obtained based on the value of each pixel. When determined in sub-block units, the MV of the block can be obtained based on the pixel value at the center of the sub-block (bottom right of the center, i.e., the bottom right sample of the four central samples).
[0137] When affine prediction is available, the motion models applicable to the current block can include the following three: a translational motion model, a 4-parameter affine motion model, and a 6-parameter affine motion model. Here, the translational motion model can indicate a model in which an existing block-based motion vector is used, the 4-parameter affine motion model can indicate a model in which two CPMVs are used, and the 6-parameter affine motion model can indicate a model in which three CPMVs are used.
[0138] The affine motion prediction may include an affine MVP (or affine inter) mode and an affine merge. In the affine motion prediction, the motion vector of the current block may be derived in units of samples or in units of sub-blocks.
[0139] When the affine merge mode is applied, the CPMV of the current block can be derived using the CPMV of a neighboring block. In this case, the CPMV of the neighboring block can be used as the CPMV of the current block as is, or the CPMV of the neighboring block can be modified based on the size of the neighboring block and the size of the current block, and then used as the CPMV of the current block.
[0140] Meanwhile, in the case of affine merging in which MVs are derived on a subblock basis, this may be referred to as subblock merging mode, which may be indicated based on merge_subblock_flag (value 1). In this case, an affine merging candidate list (described later) may also be referred to as a subblock merging candidate list. In this case, the subblock merging candidate list may further include candidates derived using sbTMVP (described later). In this case, the candidate derived using sbTMVP may be used as the candidate with the 0th index in the subblock merging candidate list. In other words, the candidate derived using sbTMVP may be positioned before inherited affine candidates and constructed affine candidates (described later) in the subblock merging candidate list.
[0141] In affine MVP mode, after determining two or more control point motion vector predictions (CPMVPs) and CPMVs for the current block, a control point motion vector difference (CPMVD) corresponding to the difference value can be transmitted from the image encoding device 100 to the image decoding device 200.
[0142] For example, when the value of affine merge flag or merge_subblock_flag is 0, the affine MVP mode can be applied. Or, for example, when the value of inter_affine_flag is 1, the affine MVP mode can be applied. The affine MVP mode may be called an affine CP MVP mode. Or, the affine MVP mode may be called an affine inter mode or an inter affine mode. The affine MVP candidate list may be called a control point motion vectors predictor candidate list.
[0143] Example
[0144] The present disclosure relates to bidirectional motion compensation in an inter prediction process, and more particularly, to a method for performing bidirectional prediction of a current block by applying different prediction types for each prediction direction (L0 or L1).
[0145] The different prediction types may be a sub-block unit (sub PU) prediction type and a non-sub-block unit (non-sub PU) prediction type. The sub PU prediction type may be referred to as a sub PU prediction mode, and may include a sub-block merge mode and an affine mode in which MVs are derived in sub-block units. The non-sub PU prediction type is a prediction type that does not correspond to the sub PU prediction type and may be referred to as a non-sub PU prediction mode.
[0146] In the conventional bidirectional motion compensation process, when the current block is a sub-PU prediction type, motion compensation is performed based on the sub-PU prediction type in both directions (L0, L1), and when the current block is a non-sub-PU prediction type, motion compensation is performed based on the non-sub-PU prediction type in both directions to generate a predicted block.
[0147] For example, determining a prediction mode, deriving motion information, and generating prediction samples through bidirectional prediction in a conventional merge mode can derive bidirectional motion information from prediction candidates. Specific examples of performing bidirectional prediction in a conventional merge mode can be classified as follows: That is, the conventional merge mode does not support a method of predicting one prediction direction using a sub PU prediction type and predicting the other prediction direction using a non-sub PU prediction type.
[0148] -If merge_subblock_flag is 1, prediction is performed based on the sub PU prediction type in both directions.
[0149] -If merge_subblock_flag is 0, prediction is performed based on the non-sub PU prediction type in both directions.
[0150] Furthermore, in the conventional merge mode, information used for prediction varies depending on the prediction type, i.e., the conventional merge mode has a restriction that the same information must be used in both directions during bidirectional prediction.
[0151] When applying the sub PU prediction type, if the neighboring block (reference block) is a block predicted in affine mode, the stored CPMV of the neighboring block is used.
[0152] -When applying non-sub PU prediction type, use the saved MV of surrounding blocks
[0153] As another example, the AMVP mode that signals MVD, MVP index, reference picture index, etc. to guide motion information can be classified as follows: That is, the conventional AMVP mode also does not support a method in which one prediction direction is predicted using the sub PU prediction type and the other prediction direction is predicted using the non-sub PU prediction type.
[0154] -If inter_affine_flag is 1, prediction is performed based on the sub PU prediction type in both directions.
[0155] -If inter_affine_flag is 0, prediction is performed based on non-sub PU in both directions.
[0156] In addition, in the conventional AMVP mode, the information used for prediction varies depending on the prediction type, i.e., the conventional AMVP mode has a restriction that the same information must be used in both directions during bidirectional prediction.
[0157] When applying the sub PU prediction type, if the neighboring block (reference block) is a block predicted in affine mode, the stored CPMV of the neighboring block is used.
[0158] -When applying non-sub PU prediction type, use the saved MV of surrounding blocks
[0159] Furthermore, in conventional AMVP modes, the amount of information signaled via the bitstream differs for each prediction type.
[0160] -When applying the sub PU prediction type, signal two or more MVDs (xoffset, yoffset) required to derive the MV of each CP
[0161] -When applying a non-sub PU prediction type, signal the MVD required to generate the prediction block
[0162] The method proposed by the present disclosure may correspond to a method of simultaneously supporting different prediction types (i.e., sub PU prediction type and non-sub PU prediction type) in the process of generating a prediction block in each prediction direction when performing bidirectional prediction for one prediction block. Therefore, according to the method proposed by the present disclosure, different prediction types are adaptively induced for each prediction direction, thereby improving compression efficiency.
[0163] In addition, the method proposed by the present disclosure performs motion compensation using different prediction types for each prediction direction, and in this process, can adaptively refer to necessary information for each prediction direction. Therefore, according to the method proposed by the present disclosure, bit efficiency required for signaling necessary information can be improved.
[0164] Example 1
[0165] FIG. 9 is a flowchart illustrating an image coding method according to the first embodiment.
[0166] 9, the image coding device 100 may determine whether the prediction direction of the current block is bidirectional (S900). For example, whether the prediction direction of the current block is bidirectional may be determined based on whether the value of inter_pred_idc, which is a syntax element indicating the prediction direction of the current block, indicates bidirectional prediction (PRED_BI).
[0167] When the prediction direction of the current block is bidirectional, the image encoding apparatus 100 may determine whether different prediction types are applied to each prediction direction (S910). Here, the different prediction types may be a sub PU prediction type and a non-sub PU prediction type. That is, step S910 may be a step of determining whether a specific direction is predicted based on the sub PU prediction type and whether another direction is predicted based on the non-sub PU prediction type. First information (merge_mixed_pu_flag or inter_mixed_pu_flag), which is information indicating whether different prediction types are applied to each prediction direction of the current block, may be coded into a bitstream.
[0168] When different prediction types are applied to the respective prediction directions, the image encoding apparatus 100 may derive at least one bidirectional motion information candidate corresponding to the different prediction types from the reference block of the current block (S920).
[0169] For example, when the L0 direction of the current block is a sub-PU prediction type and the L1 direction is a non-sub-PU prediction type, a bidirectional motion information candidate having sub-PU-based motion information in the L0 direction and non-sub-PU-based motion information in the L1 direction may be derived. As another example, when the L0 direction of the current block is a non-sub-PU prediction type and the L1 direction is a sub-PU prediction type, a bidirectional motion information candidate having non-sub-PU-based motion information in the L0 direction and sub-PU-based motion information in the L1 direction may be derived. Here, the sub-PU-based motion information may be motion information used for the sub-PU prediction type, and the non-sub-PU-based motion information may be motion information used for the non-sub-PU prediction type. At least one bidirectional motion information candidate may constitute a motion information candidate list.
[0170] The image encoding apparatus 100 may bidirectionally predict the current block with different prediction types based on one of the derived bidirectional motion information candidates, that is, based on the motion information candidate list (S930).
[0171] FIG. 10 is a flowchart of an image decoding method according to the first embodiment.
[0172] 10, the image decoding apparatus 200 may determine whether the prediction direction of the current block is bidirectional (S1000). For example, whether the prediction direction of the current block is bidirectional may be determined based on whether the value of inter_pred_idc, which is a syntax element indicating the prediction direction of the current block, indicates bidirectional prediction (PRED_BI).
[0173] When the prediction direction of the current block is bidirectional, the image decoding device 200 can obtain first information (merge_mixed_pu_flag or inter_mixed_pu_flag), which is information indicating whether different prediction types are applied to each prediction direction of the current block, from the bitstream (S1010).
[0174] The image decoding apparatus 200 may determine whether different prediction types are applied to the prediction directions of the current block based on the first information (S1020). Furthermore, when the first information indicates that different prediction types are applied to the prediction directions, the image decoding apparatus 200 may derive at least one bidirectional motion information candidate corresponding to the different prediction types from a reference block (S1030).
[0175] For example, when the L0 direction of the current block is a sub-PU prediction type and the L1 direction is a non-sub-PU prediction type, a bidirectional motion information candidate having sub-PU-based motion information in the L0 direction and non-sub-PU-based motion information in the L1 direction may be derived. As another example, when the L0 direction of the current block is a non-sub-PU prediction type and the L1 direction is a sub-PU prediction type, a bidirectional motion information candidate having non-sub-PU-based motion information in the L0 direction and sub-PU-based motion information in the L1 direction may be derived. Here, the sub-PU-based motion information may be motion information used for the sub-PU prediction type, and the non-sub-PU-based motion information may be motion information used for the non-sub-PU prediction type. At least one bidirectional motion information candidate may constitute a motion information candidate list.
[0176] The image decoding apparatus 200 may perform bidirectional prediction of the current block with different prediction types based on one of the guided bidirectional motion information candidates, that is, based on the motion information candidate list (S1040).
[0177] Fig. 11 is a diagram illustrating an example of a bidirectional prediction method according to Example 1. In Fig. 11, it is assumed that the prediction type of the L0 direction (Reference picture list (L0)) of a current block (Current picture) is the sub-PU prediction type, and the prediction type of the L1 direction (Reference picture list (L1)) is the non-sub-PU prediction type.
[0178] Under this assumption, a bidirectional motion information candidate whose prediction type in the L0 direction is the sub PU prediction type and whose prediction type in the L1 direction is the non-sub PU prediction type can be used for bidirectional prediction of the current block.
[0179] Therefore, sub-PU based motion compensation (MC) is performed in one direction (e.g., L0), and non-sub-PU based motion compensation is performed in the opposite direction (e.g., L1), thereby generating a prediction block for the current block.
[0180] Example 2
[0181] As mentioned above, in merge mode, bidirectional motion information of the current block must be derived using only the motion information of the prediction candidate (motion information candidate) without any signaling for separate syntax (except for indexes), so the prediction candidate in this disclosure must satisfy the following two conditions.
[0182] - Can induce two-way movement information
[0183] - It is possible to induce the necessary motion information for each prediction direction (e.g., L0:CPMV, L1:MV)
[0184] For example, when a reference block is encoded / decoded by the method proposed in the present disclosure or when the reference block is bidirectionally predicted by a sub PU prediction type, bidirectional motion information for applying the method proposed in the present disclosure is stored in the reference block. Therefore, in such cases, bidirectional predictions of different prediction types can be performed.
[0185] As another example, if the reference block is not encoded / decoded by the method proposed in the present disclosure, or if the reference block is bidirectionally predicted by a non-sub PU prediction type, bidirectional motion information for applying the method proposed in the present disclosure may not be stored in the reference block. Therefore, in such cases, the following solution may be performed.
[0186] - Reconstruct the motion information required for each prediction direction to derive bidirectional motion information candidates
[0187] -Do not use the motion information of the reference block to guide bidirectional motion information candidates
[0188] Thus, when the method proposed through the present disclosure is applied to merge mode, in the process of constructing a prediction candidate list (motion information candidate list) for the current block, if the reference candidate satisfies the following conditions, the reference candidate can be configured as a prediction candidate, or prediction information regenerated based on the reference candidate can be configured as a prediction candidate for the current block.
[0189] Bidirectional prediction information (bidirectional motion information) can be derived from motion information of a decoded neighboring block / a block in a decoded previous picture (i.e., a reference block), and sub-PU-based motion information can be derived from at least one direction of the bidirectional motion information.
[0190] The motion information of the reference block can be decomposed into prediction direction units and recombined to derive bidirectional motion information, and sub-PU-based motion information can be derived from at least one of the recombined bidirectional motion information.
[0191] For any prediction direction, the motion information of the current block can be combined with the motion information of the reference block so that it can be derived in sub-PU units.
[0192] On the other hand, when the method proposed through this disclosure is applied to AMVP mode, prediction blocks must be generated using different prediction types for each prediction direction, and therefore motion information candidates to be referenced for prediction of each prediction direction must be adaptively derived.
[0193] For example, as shown in the left side of Fig. 12(a), if the current block is coded using the proposed method (e.g., L0: non-sub PU, L1: sub PU),
[0194] Motion information candidates can be compensated for with the signaled MVD to derive motion information for the current block.
[0195] In this process, information on a reference block that has already been decoded in the vicinity can be used as a motion information candidate, or motion information reconstructed using information on the reference block can be used as a motion information candidate.
[0196] That is, the proposed method performs prediction based on different prediction types for each prediction direction, and therefore, reference information differs for each prediction direction. Therefore, the proposed method can be applied only when necessary information is given.
[0197] For example, if a sub PU prediction type is applied to the L0 direction of the current block and a non-sub PU prediction type is applied to the L1 direction, two or more MVDs for CPMV must be signaled to induce motion information in the L0 direction, and fewer than two MVDs must be signaled to induce motion information in the L1 direction.
[0198] However, if the L0-direction CPMV is not derived from the prediction candidate (reference block), the L0-direction CPMV of the current block will not be derived. For example, as shown on the right side of Figure 12(a), if the reference block is bidirectionally predicted, but both directions are predicted using the non-sub PU prediction type, or if the reference block is not bidirectionally predicted, appropriate motion information for the current block cannot be derived.
[0199] Next, we summarize the basic operation of the method proposed by this disclosure when applied to AMVP mode.
[0200] As illustrated in FIG. 12(b), when the L0 direction is predicted based on the sub PU prediction type and the L1 direction is predicted based on the non-sub PU prediction type, an MVP for each CP can be derived with reference to the L0 direction CPMV of a prediction candidate (reference block), and the L0 direction CPMV (motion information) of the current block can be derived by considering each derived MVP and the MVD of each signaled CP. That is, an L0 direction motion information candidate of the current block can be derived by applying an affine model to the L0 direction motion information of the prediction candidate. In addition, the L0 direction motion information candidate of the current block can be derived by considering the L1 direction motion information of the prediction candidate and the signaled MVD.
[0201] In the present disclosure, the induced bidirectional motion information candidates can form a single motion information candidate list. Therefore, compared to the conventional AMVP mode in which a separate motion information candidate list is formed for each prediction direction and separate index information is signaled, the present disclosure can easily form the motion information candidate list and reduce the amount of bits required for signaling index information.
[0202] Based on the reasons and details explained above, the second embodiment proposes a method of deriving appropriate bidirectional motion information candidates from reference blocks and configuring them in a prediction candidate list (motion information candidate list).
[0203] Example 2-1
[0204] Example 2-1 is an example in which, when bidirectional motion information can be derived from motion information of a reference block and sub-PU-based motion information can be derived from at least one direction of the bidirectional motion information, bidirectional motion information candidates for a current block are derived from the bidirectional motion information. Here, "when sub-PU-based motion information can be derived from at least one direction" may mean that at least one or more sub-PU-based motion information is included in the bidirectional motion information of the reference block.
[0205] 13 and 14 are diagrams for explaining an example of deriving bidirectional motion information candidates when Example 2-1 is applied to the merge mode.
[0206] As illustrated in Figure 13(a), if the reference block of the current block is coded using the proposed method (L0: non-sub PU, L1: sub PU), or if the bidirectional motion information of the reference block includes at least one sub-PU-based motion information, it can be induced as a bidirectional motion information candidate.
[0207] Specifically, as illustrated in Figure 13(b), the L0-direction motion information of the current block (Cur) can be derived by applying an affine model based on the L0-direction CPMV of the reference block, and the L1-direction motion information of the current block can be derived from the L1-direction motion information of the reference block. Therefore, the bidirectional motion information of the reference block can be derived as bidirectional motion information candidates.
[0208] As another example, as illustrated in Figure 14(a), when a reference block is bidirectionally predicted and both directions are predicted using the sub PU prediction type, the bidirectional motion information of the reference block stores the CPMV of the surrounding blocks (left side of Figure 14(b)), and at the same time, the motion information (MV) of the reference block itself is also stored (right side of Figure 14(b)).
[0209] 14(c), in this case, the L0 direction motion information of the current block can be derived by applying an affine model based on the CPMV of the reference block, and the MV of the reference block can be used as the L1 direction motion information of the current block. Therefore, the bidirectional motion information of such a reference block can be derived as a bidirectional motion information candidate.
[0210] FIG. 15 is a diagram for explaining an example of deriving bidirectional motion information candidates when Example 2-1 is applied to the AMVP mode.
[0211] If the reference block stores CPMV, as when decoded using the proposed method or affine mode, bidirectional motion information candidates can be derived by combining the motion information (CPMV) with motion information (MV) stored at the position of the reference block that is not CPMV.
[0212] For example, as illustrated in Figure 15(a), different prediction types may be applied to each prediction direction of a current block (L0: non-sub PU, L1: sub PU), and positions of reference blocks for constructing bidirectional motion information candidates may be defined as A to G. In addition, motion information for each position of the reference block (candidate position) is as shown in Table 1.
[0213] [Table 1]
[0214] In this example, among the bidirectional motion information of the reference block, bidirectional motion information suitable for applying the prediction method (prediction type) of the current block may be induced as a bidirectional motion information candidate. That is, as shown in Table 2, bidirectional motion information in which the prediction type in the L0 direction is the non-sub PU prediction type and the prediction type in the L1 direction is the sub PU prediction type may be induced as a bidirectional motion information candidate.
[0215] [Table 2]
[0216] 15(b), the CPMV (CP) of the neighboring block is stored at position D, and at the same time, the motion information (D) of the reference block itself is also stored, so that the CP MV can be used for the sub PU prediction type, and the motion information of the reference block itself can be referenced for the non-sub PU prediction type. Therefore, the motion information coded by the sub PU at position D (i.e., sub PU-based motion information) can be processed like non-sub PU-based motion information, and through this process, the list in Table 2 can be modified as shown in Table 3 to form a bidirectional motion information candidate list.
[0217] [Table 3]
[0218] As another example, in the example of FIG. 15(a), the motion information for each position (candidate position) of the reference block is as shown in Table 4.
[0219] [Table 4]
[0220] In this case, among the motion information of the reference blocks, motion information suitable for applying different prediction types of the current block can be derived as bidirectional motion information candidates as shown in Table 5.
[0221] [Table 5]
[0222] Example 2-2
[0223] Example 2-2 is an example in which bidirectional motion information can be derived by decomposing the motion information of the reference block in units of prediction directions and recombining the decomposed information, and when sub-PU-based motion information can be derived from at least one direction of the recombined bidirectional motion information, bidirectional motion information candidates for the current block are derived from the decomposed information. Here, "when sub-PU-based motion information can be derived from at least one direction" may mean that at least one or more sub-PU-based motion information is included in the recombined bidirectional motion information.
[0224] FIG. 16 is a flowchart illustrating an image coding method and an image decoding method according to Example 2-2.
[0225] 16, the image encoding device 100 and the image decoding device 200 may determine the prediction direction of a reference block (S1600). If the prediction direction of the reference block is unidirectional, the image encoding device 100 and the image decoding device 200 may combine motion information of the reference block in units of prediction directions to derive combined bidirectional motion information (S1610). In addition, the image encoding device 100 and the image decoding device 200 may derive bidirectional motion information candidates based on motion information including one or more sub-PU-based motion information among the combined bidirectional motion information (S1620).
[0226] According to an embodiment, steps S1610 and S1620 may be performed independently of step S1610. That is, combined bidirectional motion information and bidirectional motion information candidates may be derived without determining the prediction direction of the reference block.
[0227] Next, an example of deriving bidirectional motion information candidates when embodiment 2-2 is applied to the merge mode will be described.
[0228] As illustrated in Figure 15(a), different prediction types are applied to each prediction direction of the current block (L0: non-sub PU, L1: sub PU), and the positions of the reference block for constructing bidirectional motion information candidates are defined as A to G. In addition, the motion information for each position (candidate position) of the reference block may be unidirectional motion information as shown in Table 6.
[0229] [Table 6]
[0230] In this case, the unidirectional motion information of the reference block is recombined to derive recombined bidirectional motion information, and bidirectional motion information candidates as shown in Table 7 can be derived from the recombined bidirectional motion information.
[0231] [Table 7]
[0232] Next, an example of deriving bidirectional motion information candidates when embodiment 2-2 is applied to the AMVP mode will be described.
[0233] When the position of a reference block is as shown in Figure 15(a) and the motion information of the reference block is as shown in Table 1, the motion information of the reference block is separated for each prediction direction and recombined to derive recombined bidirectional motion information as shown in Table 8. The recombined bidirectional motion information can be used as bidirectional motion information candidates for bidirectional prediction of the current block.
[0234] [Table 8]
[0235] Example 2-3
[0236] If the motion information of the reference block is all non-sub PU-based motion information, the motion information for deriving the sub PU-based motion information of the current block cannot be referenced from the reference block. For example, as shown in Figure 15(a), if different prediction types are applied to each prediction direction of the current block and the motion information of the reference blocks (A to G) is all non-sub PU motion information, the CPMV for supporting the L1-direction sub PU prediction type of the current block cannot be derived.
[0237] To solve this problem, Example 2-3 proposes a method of deriving sub-PU-based motion information of a current block from neighboring blocks located around the current block, rather than from a reference block.
[0238] For this purpose, reference blocks are divided into adjacent blocks and peripheral blocks. An adjacent block may correspond to a block adjacent to a current block. For example, in FIG. 15(a), blocks A to G adjacent to the current block may correspond to the adjacent blocks. In addition, the adjacent blocks may be predefined blocks used to derive motion information of the current block. A peripheral block is a block located around the current block and may or may not be adjacent to the current block. A peripheral block may be a block that does not correspond to an adjacent block. The reference block may include an adjacent block and a peripheral block.
[0239] FIG. 17 is a flowchart showing an image coding method and an image decoding method according to Example 2-3.
[0240] 17, the image encoding device 100 and the image decoding device 200 may determine the prediction type of a reference block (S1700). If the prediction types of all reference blocks are non-sub PU, the image encoding device 100 and the image decoding device 200 may combine the motion information of the reference blocks in units of prediction directions to derive combined bidirectional motion information (S1710). Furthermore, the image encoding device 100 and the image decoding device 200 may derive bidirectional motion information candidates based on the combined bidirectional motion information (S1720).
[0241] The neighboring blocks may include bidirectional motion information. Even if the motion information of the neighboring blocks is not bidirectional, the combined bidirectional motion information can be derived by adaptively referring to the motion information in each prediction direction unit.
[0242] An example of the peripheral blocks is shown in Figure 18. As illustrated in Figure 18, the peripheral blocks may be reconstruction blocks located around the CP position of the current block. That is, A0, A1, B0, and B1 may be peripheral blocks located around the CP position at the bottom left of the current block, A2, A3, A4, B2, B3, and B4 may be peripheral blocks located around the CP position at the top left of the current block, and A5, A6, B5, and B6 may be peripheral blocks located around the CP position at the top right of the current block.
[0243] In the example of Figure 18, one motion information for the non-sub PU prediction type can be induced in the L0 direction, and two or more CPMVs for the sub PU prediction type can be induced in the L1 direction. As a result, bidirectional motion information candidates (prediction candidate list) can be configured as shown in Table 9.
[0244] [Table 9]
[0245] In Table 9, the information in parentheses of MV() and CPMV() may indicate the positions of the neighboring blocks shown in Figure 18. That is, as shown in Table 9, non-sub PU-based motion information for the L0 direction may be derived using motion information of the neighboring blocks, and sub PU-based motion information for the L1 direction may be derived using CPMV combined with motion information of the neighboring blocks.
[0246] Example 3
[0247] In order to apply different prediction types to each prediction direction according to the method of the present disclosure, a syntax element for supporting this needs to be signaled. Example 3 is an example of a method for signaling a syntax element.
[0248] FIG. 19 is a flowchart illustrating an image coding method according to the third embodiment.
[0249] 19, the image encoding device 100 may determine whether a first condition (apply_proposed_method_condition) is satisfied (S1900). The first condition is a condition for whether to apply the method of the present disclosure, and may include whether the prediction direction of the current block is bidirectional. Specific details of the first condition will be described later.
[0250] If a first condition is satisfied, the image encoding apparatus 100 may encode the first information into a bitstream (S1910). Here, the first information may be encoded as merge_mixed_pu_flag or inter_mixed_pu_flag. The image encoding apparatus 100 may determine whether different prediction types are applied to the prediction directions of the current block (S1920).
[0251] Also, if it is determined that different prediction types are applied to the prediction directions of the current block, the image encoding device 100 may encode the second information (inter_mixed_pu_l0_flag) and the third information (cu_subpu_type_flag) (S1930). According to an embodiment, the step S1930 may be performed only when the prediction mode of the current block is the AMVP mode. That is, the second information and the third information may be encoded when the prediction mode of the current block is the AMVP mode.
[0252] FIG. 20 is a flowchart illustrating an image decoding method according to the third embodiment.
[0253] 20, the image decoding apparatus 200 may determine whether a first condition (apply_proposed_method_condition) is satisfied (S2000). The first condition is a condition for whether to apply the method of the present disclosure, and may include whether the prediction direction of the current block is bidirectional. Specific details of the first condition will be described later.
[0254] If a first condition is satisfied, the image decoding apparatus 200 may obtain first information (merge_mixed_pu_flag or inter_mixed_pu_flag) from a bitstream (S2010). The image decoding apparatus 200 may determine whether different prediction types are applied to the prediction directions of the current block based on the first information (S2020).
[0255] When the first information indicates that different prediction types are applied to the prediction directions of the current block, the image decoding apparatus 200 may obtain the second information (inter_mixed_pu_l0_flag) and the third information (cu_subpu_type_flag) from the bitstream (S2030). According to an embodiment, the step S2030 may be performed only when the prediction mode of the current block is the AMVP mode. That is, the second information and the third information may be obtained when the prediction mode of the current block is the AMVP mode.
[0256] Example 3 can be divided into 1) an example (Example 3-1) in which no other information is signaled (except index information) and is performed like the conventional merge mode, and 2) an example (Example 3-2) in which other information is signaled and is performed like the conventional AMVP mode.
[0257] Example 3-1
[0258] An example of the syntax structure according to Example 3-1 is shown in Table 10.
[0259] [Table 10-1] [Table 10-2]
[0260] merge_idx may indicate a merge candidate in the merge candidate list. merge_mixed_pu_flag is first information and may be encoded and obtained if a first condition (apply_proposed_method_condition) is satisfied. The first information may indicate whether different prediction types are applied to the prediction directions of the current block. merge_mixed_pu_flag=1 may indicate that different prediction types are applied to the prediction directions of the current block, and merge_mixed_pu_flag=0 may indicate that different prediction types are not applied to the prediction directions of the current block.
[0261] When merge_mixed_pu_flag=1, a prediction block is generated by applying the method proposed through the present disclosure, and therefore, information such as merge_subblock_flag or regular_merge_flag, which is additionally signaled to support the conventional method, does not need to be signaled. That is, merge_subblock_flag may be signaled (merge_mixed_pu_flag=0) when the method proposed through the present disclosure is not applied.
[0262] The regular_merge_flag may indicate whether a merge mode (regular merge mode) is applied to the current block, where regular_merge_flag=1 may indicate that the regular merge mode is used to generate inter prediction parameters for the current block.
[0263] merge_subblock_flag may indicate whether a sub-block merge mode (or an affine merge mode) is applied to the current block. That is, merge_subblock_flag may indicate whether a sub-block merge mode is used to generate sub-block-based inter prediction parameters for the current block. merge_subblock_idx may indicate a merge candidate index in a sub-block-based merge candidate list.
[0264] The index information (merge_mixed_pu_idx) may indicate a bidirectional motion information candidate used for predicting a current block among at least one bidirectional motion information candidate. The index information may be signaled when the maximum number (MaxNumCombinedSubblockMergeCand) of prediction candidates (bidirectional motion information candidates) for the method proposed through the present disclosure exceeds 1. That is, when the maximum number of prediction candidates is 1 or less, prediction may be performed using one bidirectional motion information candidate without signaling the index information.
[0265] The mmvd_merge_flag may indicate whether the MMVD (merge mode with motion vector difference) mode is applied to the current block. mmvd_merge_flag=1 may indicate that the MMVD mode is used to generate inter-prediction parameters for the current block. The mmvd_cand_flag may indicate whether the first candidate (0) or the second candidate (1) in a merging candidate list is used with the MVD derived from the mmvd_distance_idx and mmvd_direction_idx. The mmvd_distance_idx may indicate an index used to derive the variable MmvdDistance[x0][y0], and the mmvd_direction_idx may indicate an index used to derive the variable MmvdSign[x0][y0].
[0266] ciip_flag may indicate whether a combined inter-picture merge and intra-picture prediction mode is applied to the current block. merge_gpm_partition_idx may indicate a partition shape of a GPM mode (geometric partitioning merge mode), and merge_gpm_idx0 and merge_gpm_idx1 may indicate the first and second candidates in a GPM-based motion compensation candidate list, respectively.
[0267] Another example of the syntax structure according to Example 3-1 is shown in Table 11.
[0268] [Table 11-1] [Table 11-2]
[0269] The syntax structure in Table 11 is an example in which the method proposed by the present disclosure is syntactically integrated with a conventional coding tool (e.g., subblock merge mode). That is, when the subblock merge mode is applied (merge_subblock_flag=1), it is determined whether the first condition (apply_proposed_method_condition) is satisfied, and if the first condition is satisfied, it is determined whether the method proposed by the present disclosure is applied (merge_mixed_pu_flag). Also, in the example of Table 11, the function of the index information (merge_mixed_pu_idx) can be performed by merge_subblock_idx.
[0270] Another example of the syntax structure according to Example 3-1 is shown in Table 12.
[0271] [Table 12]
[0272] The syntax structure of Table 12 is an example in which index information for selecting bidirectional motion information candidates is not signaled. Specifically, according to Table 12, if(merge_mixed_pu_flag==1) and if(MaxNumCombinedSubblockMergeCand>1) in Table 11 are not performed, and merge_mixed_pu_idx in Table 11 does not need to be signaled. In this case, the bidirectional motion information candidate used for predicting the current block is determined based on the cost of template matching, and may be determined as the bidirectional motion information candidate with the smallest cost, for example.
[0273] Example 3-2
[0274] An example of the syntax structure according to Example 3-2 is shown in Table 13.
[0275] [Table 13-1] [Table 13-2]
[0276] inter_mixed_pu_flag is first information that can be coded and acquired when a first condition (apply_proposed_method_condition) is satisfied. The first information may indicate whether different prediction types are applied to the prediction directions of the current block. inter_mixed_pu_flag=1 may indicate that different prediction types are applied to the prediction directions of the current block, and inter_mixed_pu_flag=0 may indicate that different prediction types are not applied to the prediction directions of the current block.
[0277] The second information (inter_mixed_pu_l0_flag) may be signaled (inter_mixed_pu_flag=1) when different prediction types are applied to each prediction direction of the current block. inter_mixed_pu_l0_flag may indicate different prediction types applied to each prediction direction. For example, inter_mixed_pu_l0_flag=1 may indicate that the prediction type applied to the L0 direction is the sub-PU prediction type, and inter_mixed_pu_l0_flag=0 may indicate that the prediction type applied to the L0 direction is the non-sub-PU prediction type. As another example, inter_mixed_pu_l0_flag=1 may indicate that the prediction type applied to the L0 direction is the non-sub-PU prediction type, and inter_mixed_pu_l0_flag=0 may indicate that the prediction type applied to the L0 direction is the sub-PU prediction type.
[0278] The third information (cu_subpu_type_flag) can be signaled when different prediction types are applied to each prediction direction of the current block (inter_mixed_pu_flag=1). The cu_subpu_type_flag can be used to determine a parameter model to be applied to each prediction direction, and the determination of the parameter model will be described later.
[0279] The inter_affine_flag may indicate whether the affine mode is applied to the current block. The cu_affine_type_flag may indicate the parameter model applied to the current block. For example, cu_affine_type_flag=1 may indicate that a 6-parameter model is applied, and cu_affine_type_flag=0 may indicate that a 4-parameter model is applied.
[0280] sym_mvd_flag may indicate whether symmetric MVD is used in MVD coding. For example, sym_mvd_flag=1 may indicate that ref_idx_l0, ref_idx_l1, and mvd_coding are not present. ref_idx_l0 may indicate a reference picture in reference picture list 0 (L0 direction), and ref_idx_l1 may indicate a reference picture in reference picture list 1 (L1 direction).
[0281] mvp_l0_flag may indicate the candidate in MVP candidate list 0 that is selected to derive the MVP of the current block when MVP mode is applied, and mvp_l1_flag may indicate the candidate in MVP candidate list 1 that is selected to derive the MVP of the current block when MVP mode is applied.
[0282] Another example of the syntax structure according to Example 3-2 is shown in Table 14.
[0283] [Table 14-1] [Table 14-2]
[0284] The syntax structure in Table 14 is an example in which the method proposed in the present disclosure is syntactically integrated with a conventional coding tool (e.g., affine mode). If inter_affine_flag=1, a sub PU prediction mode may be applied to the current block. Therefore, if the prediction direction of the current block is bidirectional (inter_pred_idc==PRED_BI), it may be signaled whether different prediction types are applied to each prediction direction (inter_mixed_pu_flag) and what prediction type is applied to each prediction direction. Also, in the example of Table 13, the function of the third information (cu_subpu_type_flag) may be performed by cu_affine_type_flag.
[0285] An example of a method for determining different prediction types for each prediction direction based on the second information (inter_mixed_pu_l0_flag) is shown in Table 15.
[0286] [Table 15]
[0287] Referring to Table 15, if inter_mixed_pu_flag=0, inter_mixed_pu_l0_flag is not signaled, and the bidirectional prediction of the current block can be determined as the sub PU prediction type. If inter_mixed_pu_flag=1 and inter_mixed_pu_l0_flag=0, the L0 direction of the current block can be determined as the sub PU prediction type, and the L1 direction of the current block can be determined as the non-sub PU prediction type. If inter_mixed_pu_flag=1 and inter_mixed_pu_l0_flag=1, the L0 direction of the current block can be determined as the non-sub PU prediction type, and the L1 direction of the current block can be determined as the sub PU prediction type.
[0288] According to an embodiment, the first information may indicate whether different prediction types are applied to each prediction direction, as well as the different prediction types to be applied to each prediction direction. That is, the first information may further indicate the different prediction types to be applied to each prediction direction. To this end, the first information may be expressed in the form of an index (inter_mixed_pu_idc) as shown in Table 16, instead of the above-described flag form.
[0289] [Table 16]
[0290] Table 17 shows an example of a method for determining whether different prediction types are applied to each prediction direction and different prediction types for each prediction direction based on the first information in the form of an index.
[0291] [Table 17]
[0292] Referring to Table 17, when inter_mixed_PU_idc=0, the bidirectional prediction type of the current block may be determined as the sub PU prediction type. When inter_mixed_PU_idc=10, the L0 direction of the current block may be determined as the sub PU prediction type, and the L1 direction of the current block may be determined as the non-sub PU prediction type. When inter_mixed_pu_idc=11, the L0 direction of the current block may be determined as the non-sub PU prediction type, and the L1 direction of the current block may be determined as the sub PU prediction type.
[0293] Example 4
[0294] Example 4 is an example for various examples of the first condition (apply_proposed_method_condition).
[0295] The first condition may include the following conditions: That is, if the following conditions are met, the first information (merge_mixed_pu_flag or inter_mixed_pu_flag) can be encoded and obtained.
[0296] Condition 1: Whether the current block generates a predicted block using Bi-prediction
[0297] Condition 1 is whether the prediction direction of the current block is bidirectional. For example, condition 1 is whether inter_pred_idc[x0][y0]==PRED_BI is satisfied.
[0298] Condition 2: Whether the size of the current block satisfies the predefined size
[0299] Condition 2 is whether the size of the current block (cbWidth, cbHeight) satisfies a predetermined size. For example, Condition 2 may include: cbWidth >= minimum width (MIN_COMB_PU_WIDTH), cbHeight >= minimum height (MIN_COMB_PU_HEIGHT), cbWidth <= maximum width (MAX_COMB_PU_WIDTH), and cbHeight <= maximum height (MAX_COMB_PU_HEIGHT).
[0300] MIN_COMB_PU_WIDTH, MIN_COMB_PU_HEIGHT, MAX_COMB_PU_WIDTH, and MAX_COMB_PU_HEIGHT may be predefined values or may be values signaled from a higher level of the bitstream. In addition, the values of MIN_COMB_PU_WIDTH, MIN_COMB_PU_HEIGHT, MAX_COMB_PU_WIDTH, and MAX_COMB_PU_HEIGHT may have values within a range of the minimum / maximum block size supported by the compression technique.
[0301] Condition 3: Whether or not the method proposed by this disclosure is supported at a higher level (SPS, PPS, PH, SH, etc.)
[0302] Whether the method proposed by the present disclosure is supported can be defined by a syntax element (comb_subpu_enable_flag) signaled from a higher level, and condition 3 can be satisfied when comb_subpu_enable_flag=1.
[0303] Condition 4: Whether or not MVD signaling of a specific prediction direction is supported at a higher level (SPS, PPS, PH, SH, etc.)
[0304] Whether MVD signaling for a particular prediction direction is enabled can be defined by a syntax element (ph_mvd_l1_zero_flag) signaled from a higher level, and condition 4 is satisfied when ph_mvd_l1_zero_flag=0 (MVD is signaled).
[0305] Example 5
[0306] Example 5 is an example of a method for adaptively signaling MVD for each prediction direction.
[0307] In the conventional AMVP mode, a prediction candidate index (mvp_lx_flag), a reference picture index (ref_idx_lx), and motion information offsets (MvdLx, mvdCpLX) can be coded into a bitstream and signaled. When the prediction type of a specific direction of the current block is a sub-PU prediction type (e.g., affine), at least two or more MVDs of each control point (CP), i.e., mvdCpLX, can be signaled. Therefore, the amount of bits of signaled information may be significantly reduced compared to a non-sub-PU prediction type (e.g., non-affine). Furthermore, in the conventional method, the same prediction type is applied to each prediction direction, which may further increase the amount of signaled information in the case of bidirectional prediction.
[0308] The method proposed by the present disclosure applies different prediction types (sub-PU prediction type and non-sub-PU prediction type) for each prediction direction, thereby reducing the amount of bits signaled in the case of bidirectional prediction, thereby improving compression efficiency.
[0309] FIG. 21 is a flowchart illustrating an image coding method according to the fifth embodiment.
[0310] 21, the image coding apparatus 100 may determine different prediction types to be applied to each prediction direction of a current block (S2100). For example, the image coding apparatus 100 may determine the prediction type of the L0 direction as the sub-PU prediction type and the prediction type of the L1 direction as the non-sub-PU prediction type. As another example, the image coding apparatus 100 may determine the prediction type of the L0 direction as the non-sub-PU prediction type and the prediction type of the L1 direction as the sub-PU prediction type. Information regarding the determination of the prediction type may be coded in second information (inter_mixed_pu_l0_flag) and signaled.
[0311] The image encoding apparatus 100 may determine a parameter model to be applied to each prediction direction. The parameter model may include a translational motion model, a 4-parameter model (e.g., a 4-parameter affine motion model), and a 6-parameter model (e.g., a 6-parameter affine motion model). Information about the parameter model may be coded and signaled in third information (cu_subpu_type_flag).
[0312] The image coding apparatus 100 may determine a prediction type applied to each prediction direction (S2120) and signal an MVD for each prediction direction based on the determination (S2130, S2140). For example, if a non-sub PU prediction type is applied to the L0 direction and a sub PU prediction type is applied to the L1 direction, one MVD may be coded for the L0 direction and multiple MVDs may be coded for the L1 direction. Here, the number of MVDs to be coded for the prediction direction to which the sub PU prediction type is applied may be determined based on the parameter model determined through step S2110.
[0313] FIG. 22 is a flowchart of an image decoding method according to the fifth embodiment.
[0314] 22, the image decoding apparatus 200 may derive a parameter model (MotionModelIdc[lX][x][y]) of each prediction direction based on the signaled second information (inter_mixed_pu_l0_flag) and third information (cu_subpu_type_flag) (S2200). Here, MotionModelIdc[lx][x][y] is a variable indicating a motion model (parameter model) of a coding unit, and lx may indicate a reference picture list direction (prediction direction).
[0315] MotionModelIdc[lx][x][y] can be derived as follows:
[0316] -If general_merge_flag=1, MotionModelIdc[lx][x][y]=merge_subblock_flag[x0][y0] for lx=0, 1
[0317] -Not true (general_merge_flag=0),
[0318] If -inter_mixed_pu_flag=1,
[0319] MotionModelIdc[0][x][y]=inter_mixed_pu_l0_flag[x0][y0]? cu_subpu_type_flag[x0][y0]+1:0
[0320] MotionModelIdc[1][x][y]=inter_mixed_pu_l0_flag[x0][y0]?0: cu_subpu_type_flag[x0][y0]+1
[0321] If -inter_mixed_pu_flag=1,
[0322] MotionModelIdc[lx][x][y]=inter_affine_flag[x0][y0]+cu_affine_type_flag[x0][y0] for lx=0,1
[0323] The parameter model indicated by MotionModelIdc[lx][x][y] is as shown in Table 18.
[0324] [Table 18]
[0325] Since different prediction types are applied to each prediction direction, MotionModelIdc may indicate different values for each prediction direction. Also, since the proposed method is applied when inter_mixed_pu_flag=1, the first array of MotionModelIdc may be defined to indicate each prediction direction in order to define each prediction direction.
[0326] The image decoding apparatus 200 can obtain the MVD for each prediction direction by performing mvd_coding for each prediction direction (S2210 to S2250).
[0327] Specifically, the image decoding apparatus 200 may perform mvd_coding to obtain a first MVD (S2210) and may check the value of MotionModelIdc (S2220). If the value of MotionModelIdc indicates a 4-parameter model or a 6-parameter model (MotionModelIdc>0), the image decoding apparatus 200 may additionally perform mvd_coding to obtain a second MVD (S2230). On the other hand, if MotionModelIdc=0, additional mvd_coding may not be performed.
[0328] The image decoding apparatus 200 may additionally check the value of MotionModelIdc (S2240). As a result, if the value of MotionModelIdc indicates a 6-parameter model (MotionModelIdc>1), the image decoding apparatus 200 may additionally perform mvd_coding to obtain a third MVD (S2250). On the other hand, if MotionModelIdc=1, additional mvd_coding may not be performed.
[0329] The above-described method for adaptively encoding and obtaining MVD can be expressed in a syntax structure as shown in Table 19.
[0330] [Table 19]
[0331] In mvd_coding(x0, y0, lx, n) in Table 19, lx indicates the prediction direction, and n indicates the number of times mvd_coding is performed.
[0332] The image decoding apparatus 200 may derive bidirectional motion information of the current block based on at least one bidirectional motion information candidate and the MVD obtained for each prediction direction, and may perform bidirectional prediction based on the induced bidirectional motion information.
[0333] An example of bidirectionally predicting a current block with different prediction types using adaptively signaled MVDs is shown in Figure 23. In Figure 23, the L0 direction is a sub-PU prediction type, so multiple MVDs (e.g., as many as the number of CPs) are signaled, and the L1 direction is a non-sub prediction type, so a single MVD can be signaled.
[0334] Example 6
[0335] Example 6 is an example for various positions of the reference block.
[0336] 24 is a diagram illustrating various positions of a reference block that can derive motion information in more detail. As illustrated in FIG. 24, the reference block may include a block adjacent to a current block and a block that is not adjacent to the current block.
[0337] Therefore, the image encoding device 100 and the image decoding device 200 can refer not only to the motion information of blocks adjacent to the current block, but also to the motion information of blocks that are not adjacent to the current block, thereby enabling more detailed deriving of the motion information of the current block.
[0338] FIG. 25 is a diagram illustrating an exemplary content streaming system to which an embodiment of the present disclosure can be applied.
[0339] As shown in FIG. 25, a content streaming system to which an embodiment of the present disclosure is applied can broadly include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.
[0340] The encoding server compresses content input from a multimedia input device such as a smartphone, camera, or camcorder into digital data to generate a bitstream and transmits the bitstream to the streaming server. As another example, if a multimedia input device such as a smartphone, camera, or video camera directly generates a bitstream, the encoding server can be omitted.
[0341] The bitstream can be generated by an image encoding method and / or image encoding device to which an embodiment of the present disclosure is applied, and the streaming server can temporarily store the bitstream during the process of transmitting or receiving the bitstream.
[0342] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server serves as an intermediary for informing the user of available services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, which then transmits the multimedia data to the user. In this case, the content streaming system may include a separate control server, which may control commands and responses between devices in the content streaming system.
[0343] The streaming server may receive content from a media storage and / or an encoding server. For example, when receiving content from the encoding server, the content may be received in real time. In this case, the streaming server may store the bitstream for a certain period of time to provide a smooth streaming service.
[0344] Examples of the user device include a mobile phone, a smartphone, a laptop computer, a digital broadcasting terminal, a personal digital assistant (PDA), a portable multimedia player (PMP), a navigation system, a slate PC, a tablet PC, an ultrabook, a wearable device such as a smartwatch, smart glass, a head mounted display (HMD), a digital TV, a desktop computer, and digital signage.
[0345] Each server in the content streaming system can be operated as a distributed server, in which case data received from each server can be processed in a distributed manner.
[0346] The scope of the present disclosure includes software or machine-executable commands (e.g., operating systems, applications, firmware, programs, etc.) that cause operations according to the methods of various embodiments to be performed on a device or computer, and non-transitory computer-readable medium on which such software or commands can be stored and executed on a device or computer. [Industrial Applicability]
[0347] The embodiments of the present disclosure can be used to encode / decode images.
Claims
1. An image decoding method performed by an image decoding device, the image decoding method comprising: obtaining first information based on the prediction direction of the current block being bidirectional; deriving at least one bidirectional motion information candidate corresponding to each prediction type from a reference block based on the first information indicating that a different prediction type is applied to each prediction direction of the current block; bidirectionally predicting the current block with the different prediction types based on any one of the at least one bidirectional motion information candidate; The image decoding method, wherein the different prediction types are a sub-PU (sub-block unit) prediction type and a non-sub-PU (non-sub-block unit) prediction type.
2. the bidirectional motion information candidate is derived based on motion information including at least one sub-PU-based motion information among bidirectional motion information of the reference block; The image decoding method of claim 1 , wherein the sub-PU-based motion information is motion information used for the sub-PU prediction type.
3. the bidirectional motion information candidate is derived based on motion information including at least one sub-PU-based motion information among motion information obtained by combining motion information of the reference blocks based on a prediction direction; The image decoding method of claim 1 , wherein the sub-PU-based motion information is motion information used for the sub-PU prediction type.
4. The image decoding method according to claim 3 , wherein the combined motion information is derived based on the fact that a prediction direction of the reference block is unidirectional.
5. the reference blocks include neighboring blocks located around the current block and predetermined neighboring blocks adjacent to the current block; 2. The image decoding method according to claim 1, wherein, based on the prediction mode of the predetermined neighboring block being the non-sub PU prediction type, the bidirectional motion information candidate is derived based on motion information obtained by combining motion information of the surrounding blocks based on a prediction direction.
6. The obtaining step includes: determining a first condition regarding whether the current block is bi-directionally predicted with the different prediction type; acquiring the first information based on the first condition being satisfied; The image decoding method according to claim 1 , wherein the first condition includes a condition regarding the prediction direction of the current block and a predetermined size condition regarding a size of the current block.
7. The image decoding method according to claim 1 , wherein the first condition further includes a condition regarding whether a motion vector difference (mvd) in a particular direction among both directions is signaled.
8. obtaining second information based on the first information indicating that different prediction types are applied to each prediction direction of the current block; the different prediction types to be applied to each prediction direction are determined based on the second information; The image decoding method according to claim 1 , wherein the bidirectional motion information candidates correspond to the determined different prediction types.
9. The image decoding method according to claim 1 , wherein the first information further indicates different prediction types to be applied to each prediction direction.
10. obtaining third information based on the first information indicating that different prediction types are applied to each prediction direction of the current block; The image decoding method according to claim 8 , wherein a parameter model to be applied to each prediction direction is determined based on the third information.
11. Bidirectional motion information of the current block is derived based on any one of the at least one bidirectional motion information candidate and an mvd obtained for each prediction direction of the current block; The image decoding method of claim 10 , wherein the number of mvds obtained for each prediction direction of the current block is determined based on the second information and the third information.
12. The image decoding method of claim 1, wherein one MVD is obtained for a direction to which a non-sub PU-based prediction mode is applied among the prediction directions of the current block.
13. The image decoding method of claim 1 , wherein the first information is obtained based on whether a sub-block merging mode is applied to the current block or an affine prediction mode is applied to the current block.
14. An image coding method performed by an image coding device, the image coding method comprising: determining whether different prediction types are applied to each prediction direction of the current block based on whether the prediction direction of the current block is bidirectional, wherein first information indicating that different prediction types are applied to each prediction direction of the current block is coded into a bitstream; deriving at least one bidirectional motion information candidate corresponding to the different prediction types from a reference block based on the different prediction types being applied to each prediction direction of the current block; bidirectionally predicting the current block with the different prediction types based on any one of the at least one bidirectional motion information candidate; The image coding method, wherein the different prediction types are a sub-PU (sub-block unit) prediction type and a non-sub-PU (non-sub-block unit) prediction type.
15. A method for transmitting a bitstream generated by an image coding method, comprising: The image encoding method includes: determining whether different prediction types are applied to each prediction direction of the current block based on whether the prediction direction of the current block is bidirectional, wherein first information indicating that different prediction types are applied to each prediction direction of the current block is coded into a bitstream; deriving at least one bidirectional motion information candidate corresponding to the different prediction types from a reference block based on the different prediction types being applied to each prediction direction of the current block; bidirectionally predicting the current block with the different prediction types based on any one of the at least one bidirectional motion information candidate; The method, wherein the different prediction types are a sub-block unit (sub PU) prediction type and a non-sub-block unit (non-sub PU) prediction type.
Citation Information
Patent Citations
Methods of accessing affine history-based motion vector predictor buffer
US20200296383A1
Simplified coding of generalized bi-directional index
US20210092435A1
Moving image decoding device
WO2017195608A1
Inter-frame prediction method and device
WO2020232845A1
Affine prediction improvements for video coding
WO2021249375A1