Image encoding / decoding method and apparatus, and recording medium for storing bit stream
By exporting multiple prediction blocks of the current block and combining different prediction modes and condition selections, the problem of low compression efficiency for high-resolution and high-quality images is solved, achieving more efficient image encoding and decoding effects.
Patent Information
- Application Number
- CN202480026322.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-04-22
- Filing Date
- 2024-04-22
- Publication Date
- 2025-11-18
AI Technical Summary
Existing image compression techniques struggle to effectively address the compression efficiency issues of high-resolution and high-quality images, especially lacking efficient prediction modes and methods when exporting multiple prediction blocks from the current block.
By exporting multiple prediction blocks of the current block, using the first and second prediction blocks to generate the final prediction block, and selecting modification methods based on different prediction modes and conditions, such as bilateral matching or template matching, and combining inter-frame and intra-frame prediction modes, prediction performance and compression efficiency are improved.
It improves the efficiency and performance of image compression by using a multi-prediction block approach and an improved prediction mode, thereby enhancing the image encoding and decoding performance.
Smart Images

Figure CN120982086A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The disclosure relates to an image encoding / decoding method and apparatus and a recording medium storing a bitstream. BACKGROUND
[0002] Recently, the demand for high-resolution and high-quality images such as HD (High Definition) images and UHD (Ultra High Definition) images has been increasing in various application fields, and thus, efficient image compression techniques are being discussed.
[0003] There are various techniques such as an inter prediction technique of predicting pixel values included in a current picture from pictures before or after the current picture using a video compression technique, an intra prediction technique of predicting pixel values included in a current picture by using pixel information in the current picture, an entropy coding technique of assigning a short symbol to a value having a high frequency of occurrence and assigning a long symbol to a value having a low frequency of occurrence, etc., and these image compression techniques can be used to efficiently compress and transmit or store image data.
[0004] DISCLOSURE
[0005] [TECHNICAL PROBLEM]
[0006] The disclosure is to provide a method and apparatus for deriving a multi-prediction block of a current block.
[0007] The disclosure is to provide a method and apparatus of deriving a multi-prediction block according to a prediction mode.
[0008] [TECHNICAL SOLUTION]
[0009] The image decoding method and apparatus according to the disclosure can derive a plurality of prediction blocks including a first prediction block and a second prediction block for a current block, derive a final prediction block of the current block based on the first prediction block and the second prediction block, and reconstruct the current block based on the final prediction block of the current block.
[0010] In the image decoding method and apparatus according to the disclosure, the first prediction block can be derived based on at least one of a first L0 motion vector predictor or a second L1 motion vector predictor, and the second prediction block can be derived based on at least one of the first L1 motion vector predictor or the second L0 motion vector predictor. Here, at least one of the first L0 motion vector predictor or the first L1 motion vector predictor can be derived based on a first prediction mode, and at least one of the second L0 motion vector predictor or the second L1 motion vector predictor can be derived based on a second prediction mode.
[0011] In the image decoding method and apparatus according to the disclosure, the first prediction mode can be an AMVP mode, and the second prediction mode can be a merge mode.
[0012] In the image decoding method and apparatus according to the disclosure, the first L0 motion vector predictor can be derived based on the first prediction mode, and the first L1 motion vector predictor can be derived based on the pre-derived first L0 motion vector predictor.
[0013] In the image decoding method and apparatus according to the disclosure, at least one of the first L0 motion vector predictor or the second L1 motion vector predictor can be modified based on any one of a bilateral matching-based modification method or a template matching-based modification method.
[0014] In the image decoding method and apparatus according to the disclosure, any one of the bilateral matching-based modification method or the template matching-based modification method can be selected based on a pre-defined first condition.
[0015] In the image decoding method and apparatus according to the disclosure, at least one of the first L1 motion vector predictor or the second L0 motion vector predictor can be modified based on any one of a bilateral matching-based modification method or a template matching-based modification method.
[0016] In the image decoding method and apparatus according to the disclosure, any one of the bilateral matching-based modification method or the template matching-based modification method can be selected based on a pre-defined second condition.
[0017] In the image decoding method and apparatus according to the disclosure, whether the second condition is satisfied can be determined based on whether the first condition is satisfied.
[0018] In the image decoding method and apparatus according to the disclosure, the first prediction block can be derived based on an inter mode, and the second prediction block can be derived based on one or more intra prediction modes of the current block. Here, the one or more intra prediction modes of the current block can include at least one of a planar mode, an MPM, an MIP mode, a DIMD-based intra prediction mode, or a TIMD-based intra prediction mode.
[0019] In the image decoding method and apparatus according to the disclosure, the first prediction block can be derived based on an inter prediction block and an intra prediction block of the current block, and the second prediction block can be derived based on an additional intra prediction mode of the current block.
[0020] In the image decoding method and apparatus according to the disclosure, the first prediction block can be derived based on one or more block vectors derived from an IBC candidate list of the current block, and the second prediction block can be derived based on one or more intra prediction modes of the current block.
[0021] The image encoding method and apparatus according to the disclosure can derive a plurality of prediction blocks including a first prediction block and a second prediction block for a current block, derive a final prediction block of the current block based on the first prediction block and the second prediction block, derive a residual block of the current block based on the final prediction block of the current block, and encode the residual block of the current block.
[0022] A computer-readable digital storage medium storing encoded video / image information is provided, thereby causing an image decoding method to be performed by a decoding apparatus according to the disclosure.
[0023] A computer-readable digital storage medium storing video / image information generated according to an image encoding method according to the disclosure is provided.
[0024] A method and apparatus for transmitting video / image information generated according to an image encoding method according to the disclosure are provided.
[0025] [Advantageous Effects]
[0026] According to the disclosure, compression efficiency can be improved by deriving a multi-prediction block of a current block and deriving a final prediction block through a weighted sum thereof.
[0027] According to the disclosure, by proposing a method of deriving various multi-prediction blocks according to a prediction mode of a current block, prediction performance and compression efficiency can be improved. BRIEF DESCRIPTION OF DRAWINGS
[0028] Figure 1 A video / image encoding system according to the disclosure is shown.
[0029] Figure 2 A schematic block diagram of an encoding apparatus to which embodiments of the disclosure are applicable and which performs encoding of a video / image signal is shown.
[0030] Figure 3 A schematic block diagram of a decoding apparatus to which embodiments of the disclosure are applicable and which performs decoding of a video / image signal is shown.
[0031] Figure 4 An inter prediction method performed by a decoding apparatus according to an embodiment of the disclosure is shown.
[0032] Figure 5 A schematic configuration of an inter predictor 332 performing an inter prediction method according to the disclosure is shown.
[0033] Figure 6 An inter prediction method performed by an encoding apparatus 200 according to an embodiment of the disclosure is shown.
[0034] Figure 7An exemplary configuration of an inter-predictor 221 performing an inter-prediction method according to the disclosure is shown.
[0035] Figure 8 An example of a content streaming system to which embodiments of the disclosure can be applied is shown. DETAILED DESCRIPTION
[0036] Because the disclosure can be changed variously and has several embodiments, specific embodiments will be illustrated in the drawings and described in detail in the detailed description. However, it is not intended to limit the disclosure to specific embodiments, and it should be understood to include all changes, equivalents, and alternatives included in the spirit and technical scope of the disclosure. While each drawing is described, like reference numerals are used for like components.
[0037] Terms such as first, second, and the like can be used to describe various components, but the components should not be limited by the terms. The terms are used only to distinguish one component from other components. For example, without departing from the scope of the disclosure, a first component can be referred to as a second component, and similarly, a second component can also be referred to as a first component. The terms and / or combinations of any one or more related statement items included in a plurality of related statement items.
[0038] When a component is referred to as being "connected" or "linked" to another component, it should be understood that it can be directly connected or linked to the other component, but another component can also be present in the middle. On the other hand, when a component is referred to as being "directly connected" or "directly linked" to another component, it should be understood that there is no other component in the middle.
[0039] The terms used in this application are only used to describe specific embodiments and are not intended to limit the disclosure. Unless the context clearly indicates otherwise, a singular expression includes a plural expression. In this application, it should be understood that terms such as "include" or "have" are intended to designate the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, but do not exclude the possibility of existence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof in advance.
[0040] The present disclosure relates to video / image coding. For example, the methods / embodiments disclosed herein can be applied to methods disclosed in the Versatile Video Coding (VVC) standard. In addition, the methods / embodiments disclosed herein can be applied to methods disclosed in the Essential Video Coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the second generation of Audio Video Coding standard (AVS2), or the next generation video / image coding standard (e.g., H.267 or H.268, etc.).
[0041] This specification proposes various embodiments of video / image coding, and the embodiments can be combined with each other to be performed unless otherwise specified.
[0042] Here, a video can refer to a set of a series of images over time. A picture generally refers to a unit representing one image within a specific time period, and a slice / tile is a unit forming a part of a picture in coding. A slice / tile can include at least one coding tree unit (CTU). One picture can be composed of at least one slice / tile. A tile is a rectangular region composed of a plurality of CTUs within a specific tile column and a specific tile row of one picture. A tile column is a rectangular region of CTUs having the same height as a picture and a width assigned by syntax requirements of a picture parameter set. A tile row is a rectangular region of CTUs having a height assigned by a picture parameter set and a width the same as a width of a picture. CTUs within one tile can be arranged consecutively according to a CTU raster scan, and tiles within one picture can be arranged consecutively according to a raster scan of tiles. One slice can include an integer number of complete tiles or an integer number of consecutive complete CTU rows within tiles of one picture that can be exclusively included in a single NAL unit. Meanwhile, one picture can be divided into at least two sub-pictures. A sub-picture can be a rectangular region of at least one slice within a picture.
[0043] A pixel, a pel, or a picture element can refer to the smallest unit constituting one picture (or image). In addition, a "sample" can be used as a term corresponding to a pixel. A sample can generally represent a pixel or a pixel value, and can represent only a pixel / pixel value of a luma component or only a pixel / pixel value of a chroma component.
[0044] A unit can represent a basic unit of image processing. A unit can include at least one of a specific region of a picture and information related to the corresponding region. One unit can include one luma block and two chroma (e.g., cb, cr) blocks. In some cases, a unit can be used interchangeably with terms such as a block or a region. In general, an MxN block can include a set (or an array) of transform coefficients or samples (or a sample array) composed of M columns and N rows.
[0045] Here, "A or B" can refer to "only A", "only B", or "both A and B". In other words, here, "A or B" can be interpreted as "A and / or B". For example, here, "A, B, or C" can refer to "only A", "only B", "only C", or "any combination of A, B, and C".
[0046] A slash ( / ) or a comma used herein can refer to "and / or". For example, "A / B" can refer to "A and / or B". Thus, "A / B" can refer to "only A", "only B", or "both A and B". For example, "A, B, C" can refer to "A, B, or C".
[0047] Here, "at least one of A and B" can refer to "only A", "only B", or "both A and B". Also, herein, expressions such as "at least one of A or B" or "at least one of A and / or B" can be interpreted in the same manner as "at least one of A and B".
[0048] Also, here, "at least one of A, B, and C" can refer to "only A", "only B", "only C", or "any combination of A, B, and C". Also, "at least one of A, B, or C" or "at least one of A, B, and / or C" can refer to "at least one of A, B, and C".
[0049] Also, brackets used herein can refer to "for example". Specifically, when indicated as "prediction (intra prediction)", "intra prediction" can be proposed as an example of "prediction". In other words, "prediction" here is not limited to "intra prediction", and "intra prediction" can be proposed as an example of "prediction". Also, even when indicated as "prediction (i.e., intra prediction)", "intra prediction" can be proposed as an example of "prediction".
[0050] Here, technical features described separately in the drawings can be implemented separately or simultaneously.
[0051] Figure 1 A video / image encoding system according to the disclosure is illustrated.
[0052] Reference Figure 1A video / image encoding system can include a first device (a source device) and a second device (a sink device).
[0053] The source device can transmit the encoded video / image information or data in the form of a file or a stream to the sink device through a digital storage medium or a network. The source device can include a video source, an encoding apparatus, and a transmission unit. The sink device can include a reception unit, a decoding apparatus, and a renderer. The encoding apparatus can be referred to as a video / image encoding apparatus, and the decoding apparatus can be referred to as a video / image decoding apparatus. A transmitter can be included in the encoding apparatus. A receiver can be included in the decoding apparatus. The renderer can include a display unit, and the display unit can be composed of a separate device or an external component.
[0054] The video source can acquire a video / image through a process of capturing, synthesizing, or generating a video / image. The video source can include a device that captures a video / image and a device that generates a video / image. The device that captures a video / image can include at least one camera, a video / image archive including a previously captured video / image, or the like. The device that generates a video / image can include a computer, a tablet, a smartphone, or the like, and can (electronically) generate a video / image. For example, a virtual video / image can be generated through a computer or the like, and in this case, the process of capturing a video / image can be replaced by a process of generating related data.
[0055] The encoding apparatus can encode an input video / image. The encoding apparatus can perform a series of processes such as prediction, transformation, quantization, or the like for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0056] The transmission unit can transmit the encoded video / image information or data output in the form of a bitstream to the reception unit of the sink device in the form of a file or a stream through a digital storage medium or a network. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, or the like. The transmission unit can include an element for generating a media file through a predetermined file format and can include an element for transmission through a broadcast / communication network. The reception unit can receive / extract a bitstream and transmit it to the decoding apparatus.
[0057] The decoding apparatus can decode a video / image by performing a series of processes such as inverse quantization, inverse transformation, prediction, or the like corresponding to the operations of the encoding apparatus.
[0058] The renderer can render the decoded video / image. The rendered video / image can be displayed through a display unit.
[0059] Figure 2 A rough block diagram of an encoding apparatus to which embodiments of the present disclosure can be applied and which performs encoding of a video / image signal is illustrated.
[0060] Reference Figure 2 The encoding apparatus 200 can be composed of an image partitioner 210, a predictor 220, a residue processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 can include an inter-predictor 221 and an intra-predictor 222. The residue processor 230 can include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residue processor 230 can further include a subtractor 231. The adder 250 can be referred to as a reconstructor or a reconstructed block generator. According to an embodiment, the above-described image partitioner 210, predictor 220, residue processor 230, entropy encoder 240, adder 250, and filter 260 can be configured by at least one hardware component (e.g., an encoder chipset or a processor). In addition, the memory 270 can include a decoded picture buffer (DPB) and can be configured by a digital storage medium. The hardware component can further include the memory 270 as an internal / external component.
[0061] The image partitioner 210 can partition an input image (or picture, frame) input to the encoding apparatus 200 into at least one processing unit. As an example, the processing unit can be referred to as a coding unit (CU). In this case, the coding unit can be recursively partitioned from a coding tree unit (CTU) or a largest coding unit (LCU) according to a quad-tree binary-tree ternary (QTBTTT) structure.
[0062] For example, one coding unit can be partitioned into a plurality of coding units having a deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure can be applied first, and the binary-tree structure and / or the ternary structure can be applied later. Alternatively, the binary-tree structure can be applied before the quad-tree structure. The encoding process according to this specification can be performed based on a final coding unit that is no longer partitioned. In this case, based on the coding efficiency according to the characteristics of the image, etc., the largest coding unit can be directly used as the final coding unit, or if necessary, the coding unit can be recursively partitioned into coding units of a deeper depth, and the coding unit having the optimal size can be used as the final coding unit. Here, the encoding process can include processes such as prediction, transformation, and reconstruction, etc. which will be described later.
[0063] As another example, the processing unit can further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit can be divided or split from the above-described final encoding unit, respectively. The prediction unit can be a unit of sample prediction, and the transform unit can be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.
[0064] In some cases, the unit can be used interchangeably with terms such as a block or an area. In general, an MxN block can represent a set of transform coefficients or samples consisting of M columns and N rows. The sample can generally represent a pixel or a pixel value, and can represent only a pixel / pixel value of a luminance component or only a pixel / pixel value of a chrominance component. The sample can be used as a term to correspond to a pixel or a pel for one picture (or image).
[0065] The encoding apparatus 200 can subtract a prediction signal (a prediction block, a prediction sample array) output from the inter-predictor 221 or the intra-predictor 222 from an input image signal (an original block, an original sample array) to generate a residual signal (a residual signal, a residual sample array), and the generated residual signal is transmitted to the transformer 232. In this case, the unit that subtracts the prediction signal (the prediction block, the prediction sample array) from the input image signal (the original block, the original sample array) within the encoding apparatus 200 can be referred to as a subtractor 231.
[0066] The predictor 220 can perform prediction on a block to be processed (hereinafter, referred to as a current block), and generate a predicted block including predicted samples for the current block. The predictor 220 can determine whether to apply intra-prediction or inter-prediction in units of the current block or CU. The predictor 220 can generate various information about prediction, such as prediction mode information, etc., and transmit the same to the entropy encoder 240, as described later in the description of each prediction mode. The information about prediction can be encoded in the entropy encoder 240 and output in the form of a bitstream.
[0067] The intra-predictor 222 can predict the current block by referring to samples within the current picture. Depending on the prediction mode, the referred samples can be positioned in the vicinity of the current block or can be positioned at a distance away from the current block. In intra-prediction, the prediction mode can include at least one non-directional mode and a plurality of directional modes. The non-directional mode can include at least one of a DC mode or a planar mode. Depending on the level of detail of the prediction direction, the directional mode can include 33 directional modes or 65 directional modes. However, this is only an example, and more or less directional modes can be used according to the configuration. The intra-predictor 222 can determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring block.
[0068] The inter predictor 221 can derive a prediction block for a current block based on a reference block (a reference sample array) specified by a motion vector on a reference picture. In this case, to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in a block, sub-block, or sample unit based on the correlation of the motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). For inter prediction, the neighboring blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block can be the same or different. The temporal neighboring block can be referred to as a collocated reference block, a collocated CU (colCU), etc., and the reference picture including the temporal neighboring block can be referred to as a collocated picture (colPic). For example, the inter predictor 221 can configure a motion information candidate list based on the neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or the reference picture index of the current block. The inter prediction can be performed based on various prediction modes, and for example, for a skip mode and a merge mode, the inter predictor 221 can use the motion information of the neighboring blocks as the motion information of the current block. For the skip mode, unlike the merge mode, a residual signal can not be transmitted. For a motion vector prediction (MVP) mode, the motion vector of a surrounding block is used as a motion vector predictor, and a motion vector difference is signaled to indicate the motion vector of the current block.
[0069] The predictor 220 can generate a prediction signal based on various prediction methods described later. For example, the predictor can not only apply intra prediction or inter prediction to predict one block, but also simultaneously apply intra prediction and inter prediction. It can be referred to as a combined inter and intra prediction (CIIP) mode. In addition, the predictor can be based on an intra block copy (IBC) prediction mode or can be based on a palette mode for prediction for a block. The IBC prediction mode or the palette mode can be used for content image / video coding of games, etc., such as screen content coding (SCC), etc. The IBC basically performs prediction within a current picture, but it can be performed similarly to inter prediction in that it derives a reference block within the current picture. In other words, the IBC can use at least one of the inter prediction techniques described herein. The palette mode can be considered as an example of intra coding or intra prediction. When the palette mode is applied, sample values within a picture can be signaled based on information about a palette table and a palette index. The prediction signal generated by the predictor 220 can be used to generate a reconstructed signal or a residual signal.
[0070] The transformer 232 can generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique can include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loève transform (KLT), a graph-based transform (GBT), or a conditional non-linear transform (CNT). Here, the GBT refers to a transform obtained from a graph when relationship information between pixels is expressed as the graph. The CNT refers to a transform obtained based on a prediction signal generated by using all previously reconstructed pixels. In addition, the transform process can be applied to square pixel blocks of the same size or can be applied to non-square blocks of variable sizes.
[0071] The quantizer 233 can quantize the transform coefficients and transmit them to the entropy encoder 240, and the entropy encoder 240 can encode and output information about the quantized transform coefficients (about the quantized transform coefficients) as a bitstream. The information about the quantized transform coefficients can be referred to as residual information. The quantizer 233 can rearrange the quantized transform coefficients in a block form into a one-dimensional vector form based on a coefficient scan order, and can generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form.
[0072] The entropy encoder 240 can perform various encoding methods such as exponential Golomb, context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), or the like. The entropy encoder 240 can encode information necessary for video / video image reconstruction (for example, values of syntax elements, or the like) other than the transform coefficients quantized together or individually.
[0073] Encoded information (e.g., encoded video / image information) can be transmitted or stored in a form of a bitstream in units of network abstraction layer (NAL) units. The video / image information can further include information on various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information can further include general constraint information. Here, information and / or syntax elements transmitted / signaled from an encoding apparatus to a decoding apparatus can be included in the video / image information. The video / image information can be encoded through the above-described encoding process and included in a bitstream. The bitstream can be transmitted through a network or can be stored in a digital storage medium. Here, the network can include a broadcasting network and / or a communication network, etc., and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmission unit (not shown) for transmission and / or a storage unit (not shown) for storing a signal output from the entropy encoder 240 can be configured as an internal / external element of the encoding apparatus 200, or the transmission unit can also be included in the entropy encoder 240.
[0074] The quantized transform coefficients output from the quantizer 233 can be used to generate a prediction signal. For example, a dequantized and inverse-transformed can be applied to the quantized transform coefficients by the inverse quantizer 234 and the inverse transformer 235 to reconstruct a residual signal (a residual block or a residual sample). The adder 250 can add the reconstructed residual signal to the prediction signal output from the inter-predictor 221 or the intra-predictor 222 to generate a reconstructed signal (a reconstructed picture, a reconstructed block, a reconstructed sample array). When there is no residual of a block to be processed, such as when a skip mode is applied, a prediction block can be used as a reconstructed block. The adder 250 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-prediction of a next block to be processed within a current picture, and can also be used for inter-prediction of a next picture through filtering described later. Meanwhile, luma mapping with chroma scaling (LMCS) can be applied in a picture encoding and / or reconstructing process.
[0075] The filter 260 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 260 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and can store the modified reconstructed picture in the memory 270, particularly in the DPB of the memory 270. The various filtering methods can include a deblocking filter, a sample adaptive offset, an adaptive loop filter, a bilateral filter, etc. The filter 260 can generate and transmit various information on filtering to the entropy encoder 240. The information on filtering can be encoded in the entropy encoder 240 and output in the form of a bitstream.
[0076] The modified reconstructed picture transmitted to the memory 270 can be used as a reference picture in the inter-predictor 221. When inter-prediction is applied thereto, the encoding apparatus can avoid prediction mismatch in the encoding apparatus 200 and the decoding apparatus, and can also improve coding efficiency.
[0077] The DPB of the memory 270 can store the modified reconstructed picture to be used as a reference picture in the inter-predictor 221. The memory 270 can store motion information of a block from which motion information in the current picture is derived (or encoded) and / or motion information of a block in the pre-reconstructed picture. The stored motion information can be transmitted to the inter-predictor 221 to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory 270 can store reconstructed samples of a reconstructed block in the current picture and transmit them to the intra-predictor 222.
[0078] Figure 3 A rough block diagram of a decoding apparatus to which embodiments of the present disclosure can be applied and which performs decoding of a video / image signal is illustrated.
[0079] Reference Figure 3 The decoding apparatus 300 can be configured by including an entropy decoder 310, a residue processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 can include an inter-predictor 332 and an intra-predictor 331. The residue processor 320 can include a dequantizer 321 and an inverse transformer 321.
[0080] According to an embodiment, the above-described entropy decoder 310, residue processor 320, predictor 330, adder 340, and filter 350 can be configured by one hardware component (e.g., a decoder chipset or a processor). In addition, the memory 360 can include a decoded picture buffer (DPB) and can be configured by a digital storage medium. The hardware component can further include the memory 360 as an internal / external component.
[0081] When a bitstream including video / image information is input, the decoding apparatus 300 can reconstruct an image in response to a process of processing video / image information in the encoding apparatus. Figure 2 The decoding apparatus 300 can derive a unit / block based on related information of block partitioning obtained from the bitstream. The decoding apparatus 300 can perform decoding by using a processing unit applied in the encoding apparatus. Accordingly, the decoded processing unit can be an encoding unit, and the encoding unit can be partitioned from a coding tree unit or a largest coding unit according to a quad tree structure, a binary tree structure, and / or a ternary tree structure. At least one transform unit can be derived from the encoding unit. Also, the reconstructed image signal decoded and output by the decoding apparatus 300 can be played back by a playback device.
[0082] Decoding device 300 can receive data in bitstream form from... Figure 2 The signal output by the encoding device and the received signal can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can parse the bitstream to derive information (e.g., video / image information) necessary for image reconstruction (or picture reconstruction). The video / image information may further include information about various parameter sets such as adaptive parameter sets (APS), picture parameter sets (PPS), sequence parameter sets (SPS), or video parameter sets (VPS). In addition, the video / image information may further include general constraint information. The decoding device can further decode the picture based on the information about the parameter sets and / or the general constraint information. The information sent / received by the signal and / or the syntax elements described later herein can be decoded and obtained from the bitstream through the decoding process. For example, the entropy decoder 310 can decode the information in the bitstream based on encoding methods such as exponential Golomb coding, CAVLC, CABAC, etc., and output the values of the syntax elements necessary for image reconstruction and the quantized values of the transform coefficients of the residuals. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element from the bitstream, determine a context model using information about the syntax element to be decoded, decoding information of surrounding blocks and the block to be decoded, or information about symbols / bins decoded in the previous step, perform arithmetic decoding on the bins by predicting the occurrence probability of the bins based on the determined context model, and generate symbols corresponding to the value of each syntax element. In this case, after determining the context model, the CABAC entropy decoding method can update the context model by using information about the decoded symbols / bins for the context model used for the next symbol / bin. Among the information decoded in the entropy decoder 310, information about prediction is provided to the predictors (inter-frame predictor 332 and intra-frame predictor 331), and the residual values of entropy decoding performed on them in the entropy decoder 310, i.e., the quantized transform coefficients and related parameter information, can be input to the residual processor 320. The residual processor 320 can derive residual signals (residual blocks, residual samples, residual sample arrays). In addition, information about filtering among the information decoded in the entropy decoder 310 can be provided to the filter 350. Meanwhile, the receiving unit (not shown) that receives the signal output from the encoding device can be further configured as an internal / external element of the decoding device 300 or the receiving unit can be a component of the entropy decoder 310.
[0083] Meanwhile, a decoding apparatus according to this specification can be referred to as a video / image / picture decoding apparatus, and the decoding apparatus can be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder can include the entropy decoder 310, and the sample decoder can include at least one of the inverse quantizer 321, the inverse transformer 322, the adder 340, the filter 350, the memory 360, the inter predictor 332, and the intra predictor 331.
[0084] The inverse quantizer 321 can inverse-quantize the quantized transform coefficients and output the transform coefficients. The inverse quantizer 321 can rearrange the quantized transform coefficients into a two-dimensional block form. In this case, the rearrangement can be performed based on a coefficient scan order performed in the encoding apparatus. The inverse quantizer 321 can perform inverse quantization on the quantized transform coefficients by using a quantization parameter (e.g., quantization step length information) and obtain the transform coefficients.
[0085] The inverse transformer 322 inverse-transforms the transform coefficients to obtain a residual signal (a residual block, a residual sample array).
[0086] The predictor 320 can perform prediction on the current block and generate a prediction block including prediction samples for the current block. The predictor 320 can determine whether to apply intra prediction or inter prediction to the current block based on information about prediction output from the entropy decoder 310, and determine a specific intra / inter prediction mode.
[0087] The predictor 320 can generate a prediction signal based on various prediction methods described later. For example, the predictor 320 can not only apply intra prediction or inter prediction to predict one block, but also simultaneously apply intra prediction and inter prediction. It can be referred to as a combined inter and intra prediction (CIIP) mode. In addition, the predictor can be based on an intra block copy (IBC) prediction mode or can be based on a palette mode for prediction of a block. The IBC prediction mode or the palette mode can be used for content image / video encoding of games and the like, such as screen content coding (SCC) and the like. The IBC basically performs prediction within a current picture, but it can be performed similarly to inter prediction in that it derives a reference block within the current picture. In other words, the IBC can use at least one of the inter prediction techniques described herein. The palette mode can be considered as an example of intra coding or intra prediction. When the palette mode is applied, information about a palette table and a palette index can be included in the video / image information and signaled.
[0088] The intra predictor 331 can predict the current block by referring to samples within the current picture. Depending on the prediction mode, the referred samples can be located in the vicinity of the current block or can be located at a distance away from the current block. In intra prediction, the prediction mode can include at least one non-directional mode and a plurality of directional modes. The intra predictor 331 can determine the prediction mode applied to the current block by using the prediction mode applied to a neighboring block.
[0089] The inter predictor 332 can derive a prediction block for the current block based on a reference block (a reference sample array) specified by a motion vector on a reference picture. In this case, to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of a block, a sub-block, or a sample based on the correlation of the motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). For inter prediction, the neighboring blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter predictor 332 can configure a motion information candidate list based on the neighboring blocks, and derive a motion vector and / or a reference picture index of the current block based on received candidate selection information. The inter prediction can be performed based on various prediction modes, and information about the prediction can include information indicating an inter prediction mode for the current block.
[0090] The adder 340 can add the obtained residual signal to a prediction signal (a prediction block, a prediction sample array) output from the predictor (including the inter predictor 332 and / or the intra predictor 331) to generate a reconstructed signal (a reconstructed picture, a reconstructed block, a reconstructed sample array). When there is no residual of a block to be processed, as when a skip mode is applied, the prediction block can be used as the reconstructed block.
[0091] The adder 340 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra prediction of a next block to be processed in the current picture, can be output through filtering described later, or can be used for inter prediction of a next picture. Meanwhile, luma mapping with chroma scaling (LMCS) can be applied in the picture decoding process.
[0092] The filter 350 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 350 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and transmit the modified reconstructed picture to the memory 360, specifically, the DPB of the memory 360. The various filtering methods can include deblocking filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, etc.
[0093] The (modified) reconstructed pictures stored in the DPB of the memory 360 can be used as reference pictures in the inter prediction unit 332. The memory 360 can store motion information of a block from which the motion information in the current picture is derived (or decoded) and / or the motion information of a block in the pre-reconstructed picture. The stored motion information can be sent to the inter predictor 332 to be used as the motion information of a spatial neighboring block or the motion information of a temporal neighboring block. The memory 360 can store the reconstructed samples of a reconstructed block in the current picture and send them to the intra predictor 331.
[0094] Here, the embodiments described in the filter 260, the inter predictor 221 and the intra predictor 222 of the encoding apparatus 200 can also be applied to the filter 350, the inter predictor 332 and the intra predictor 331 of the decoding apparatus 300, respectively, equally or correspondingly.
[0095] Figure 4 A method of inter prediction performed by the decoding apparatus 300 according to an embodiment of the disclosure is illustrated.
[0096] Reference Figure 4 A plurality of prediction blocks of the current block can be derived S400.
[0097] The plurality of prediction blocks of the current block can include a first prediction block and a second prediction block. The first prediction block can be derived based on a first prediction mode, and the second prediction block can be derived based on a second prediction mode. Here, the first prediction mode and the second prediction mode can be the same prediction mode, or can be different prediction modes. Hereinafter, a method for deriving the first prediction block and the second prediction block will be described.
[0098] Example 1
[0099] The disclosure relates to a case where the first prediction mode and the second prediction mode are AMVP-MERGE modes. Specifically, it relates to a case where motion information for uni-prediction is obtained for each of the AMVP mode and the merge mode.
[0100] In the case of the AMVP-MERGE mode, bi-directional motion information of the current block can be obtained. Here, the bi-directional motion information can contain motion information in a first prediction direction obtained based on the AMVP mode and motion information in a second prediction direction obtained based on the merge mode. The first prediction direction can be expressed as an LX direction, and the second prediction direction can be expressed as an L(1-X) direction. X can have a value of 0 or 1. The same meaning can be interpreted in the embodiments described later.
[0101] In particular, at least one of an MVP index, a reference picture index, or MVD information can be signaled for the first prediction direction. The MVP index can specify any one of a plurality of MVP candidates belonging to an MVP candidate list. The reference picture index can specify any one of reference pictures in a reference picture list belonging to the first prediction direction. The MVD information can refer to information on a motion vector difference. Based on the MVP candidate specified by the MVP index, a motion vector predictor (MVP[0][X]) of the first prediction direction of the current block can be derived. In MVP[0][X], [0] denotes an AMVP mode, and [X] denotes an LX direction. In other words, MVP[0][X] can denote a motion vector predictor in the LX direction derived based on the AMVP mode. MVP[0][X] can refer to a motion vector of the MVP candidate specified by the MVP index, or can refer to a sum of a difference between a motion vector of the MVP candidate specified by the MVP index and a motion vector derived based on the MVD information.
[0102] A merge index can be signaled for the second prediction direction. The merge index can specify any one of a plurality of merge candidates belonging to a merge candidate list. Based on a motion vector of the second prediction direction of the merge candidate specified by the merge index, a motion vector predictor (MVP[1][1-X]) of the second prediction direction of the current block can be derived. In MVP[1][1-X], [1] denotes a merge mode, and [1-X] denotes an L(1-X) direction. In other words, MVP[1][1-X] can denote a motion vector predictor in the L(1-X) direction derived based on the merge mode.
[0103] A first prediction block of the current block can be derived based on the motion vector predictor (MVP[0][X]) for the first prediction direction. A second prediction block of the current block can be derived based on the motion vector predictor (MVP[1][1-X]) for the second prediction direction.
[0104] At least one of the motion vector predictors (MVP[0][X], MVP[1][1-X]) for bi-prediction of the current block can be modified based on any one of a bilateral matching based modification method or a template matching based modification method, which will be described later. The first prediction block and the second prediction block of the current block can be derived based on the modified motion vector predictors, respectively.
[0105] 1. Bilateral matching based modification method
[0106] The cost array can be calculated by performing a search based on a position of a reference block of the current block within a predetermined search range.
[0107] The reference block of the current block can include a reference block in a first prediction direction (hereinafter referred to as a first reference block) and a reference block in a second prediction direction (hereinafter referred to as a second reference block). The first reference block can be specified based on motion information of the first prediction direction. The second reference block can be specified by motion information of the second prediction direction. The motion information of the first prediction direction can be obtained based on any one of an AMVP mode or a merge mode, and the motion information of the second prediction direction can be obtained based on the other of the AMVP mode or the merge mode.
[0108] The cost array can be composed of a plurality of costs calculated for each search position within the search range. Each cost can be calculated as a sample difference between at least two blocks searched in two directions. As an example, the cost can be calculated as a sum of absolute differences (SAD) between at least two blocks searched in two directions. Here, the block searched in the first prediction direction is referred to as an LX block, and the block searched in the second prediction direction is referred to as an L(1-X) block. The cost can be calculated based on all samples belonging to the LX and L(1-X) blocks, or can be calculated based on some samples within the LX and L(1-X) blocks.
[0109] Here, some samples refer to sub-blocks of the LX and L(1-X) blocks, and at least one of a width or a height of the sub-blocks can be half of a width or a height of the LX and L(1-X) blocks. In other words, the LX and L(1-X) blocks have a size of W x H, and the some samples above can be sub-blocks having a size of W x H / 2, W / 2 x H, or W / 2 x H / 2. In this case, when the some samples are the W x H / 2 sub-blocks, the some samples can be top sub-blocks (or bottom sub-blocks) within the LX and L(1-X) blocks. When the some samples are the W / 2 x H sub-blocks, the some samples can be left sub-blocks (or right sub-blocks) within the LX and L(1-X) blocks. When the some samples are the W / 2 x H / 2 sub-blocks, the some samples can be top-left sub-blocks within the LX and L(1-X) blocks, but are not limited thereto.
[0110] Alternatively, the some samples can be defined as at least one of even-numbered sample lines or at least one of odd-numbered sample lines of the LX and L(1-X) blocks. In this case, the sample lines can be vertical sample lines or horizontal sample lines.
[0111] Alternatively, the some samples can refer to samples at positions equally predefined for the encoding / decoding apparatus. For example, the samples at the predefined positions can refer to at least one of a top-left sample, a top-right sample, a bottom-left sample, a bottom-right sample, a center sample, a center sample of a sample column / row adjacent to a boundary of the current block, or a sample located on a diagonal line within the current block located within the LX and L(1-X) blocks.
[0112] The search position within the search range can be a position shifted by p in an x-axis direction and by q in a y-axis direction from a position of a reference block of the current block. For example, when p and q are integers belonging to a range of -1 to 1, the number of search positions in integer pixel units within the search range can be up to 9. Alternatively, when p and q are integers belonging to a range of -2 to 2, the number of search positions in integer pixel units within the search range can be up to 25. However, without being limited thereto, p and q can belong to integers having a size (or absolute value) greater than 2, and a search in fractional pixel units can be performed.
[0113] The search position within the search range can be determined based on an offset equally predefined for the encoding apparatus and the decoding apparatus. In other words, the offset can be defined as a disparity vector between a position of a reference block of the current block and the search position. The offset can include at least one of a non-directional offset or a directional offset. The directional offset can include an offset for at least one of a left, right, top, bottom, top-left, top-right, bottom-left, or bottom-right direction. The non-directional offset can refer to an offset having a size of 0, and the directional offset can refer to an offset having a size (or absolute value) of at least one of an x component or a y component of the offset greater than or equal to 1.
[0114] As an example, the offset can be defined as shown in Table 1 below.
[0115] [Table 1]
[0116] Index (i) 0 1 2 3 4 5 6 7 8 dX[i] -1 0 1 -1 0 1 -1 0 1 dY[i] -1 -1 -1 0 0 0 1 1 1
[0117] Table 1 defines an offset specifying a search position for each index, and dX[i] can refer to an x component of the i-th offset and dY[i] can refer to a y component of the i-th offset. The offset according to Table 1 can include a non-directional offset of (0, 0) and 8 directional offsets. However, the index in Table 1 is only to distinguish the offsets, and does not limit a position of the offset corresponding to the index nor a priority between the offsets. In addition, Table 1 represents a case where the sizes of the x component and the y component of the offset are 1, but this is only an example, and an offset in which at least one of the x component or the y component has a size greater than or equal to 2 can be defined.
[0118] The offset can be defined for the L0 direction and the L1 direction, respectively. The offset for the search in the L1 direction can be determined depending on the offset for the search in the L0 direction. As an example, when the offset for the search in the L0 direction is (p, q), the offset for the search in the L1 direction can be set to (-p, -q) by mirroring. In addition, the offset for the search in the L1 direction can be determined independently of the offset for the search in the L0 direction.
[0119] Information on the size and / or direction of the offset described above can be equally predefined for the encoding and decoding apparatuses, or can be encoded in the encoding apparatus and signaled to the decoding apparatus. The information can also be variably determined by taking into account the block properties described above.
[0120] By the above-described method, a cost corresponding to each search position (or predefined offset) within a search range can be calculated. A cost having a minimum value among a plurality of costs belonging to a cost array can be identified, and a delta motion vector can be determined based on an offset corresponding to the identified cost. A motion vector predictor in the L0 direction can be modified based on the delta motion vector (deltaMV), and a motion vector predictor in the L1 direction can be modified based on a mirror delta motion vector (-deltaMV).
[0121] 2. Modification method based on template matching
[0122] The cost array can be calculated by performing a search based on a position of a reference block of a current block within a predetermined search range. Here, the reference block of the current block is the same as described above in the "modification method based on bilateral matching". The cost array can be composed of a plurality of costs calculated for each search position within the search range. Each cost can be calculated as a sample difference between a template region of the current block and a template region at the search position. As an example, the cost can be calculated as a sum of absolute differences (SAD) between the template region of the current block and the template region at the search position. The template region at the search position can refer to a template region of a block having a sample as a top-left sample at the search position.
[0123] The cost can be calculated based on all samples belonging to the template region of the current block and the template region at the search position, or can be calculated based on some samples within the template region.
[0124] Here, the some samples can refer to at least one of even-numbered sample lines or at least one of odd-numbered sample lines within the template region. In this case, the sample line can be a vertical sample line or a horizontal sample line.
[0125] Alternatively, the template region can include at least one of a top region, a left region, a top-left region, a bottom-left region, or a top-right region adjacent to a block (i.e., the current block or the block at the search position). In this case, the some samples can be limited to samples belonging to a region at a specific position within the template region. As an example, the region at the specific position can include at least one of a top region or a left region.
[0126] Alternatively, the template region can include neighboring sample lines and / or at least one non-neighboring sample line adjacent to the block. In this case, some samples can be limited to samples belonging to a sample line at a specific position within the template region. As an example, the sample line at the specific position is equally predefined for the encoding device and the decoding device, and can include at least one of a neighboring sample line, a non-neighboring sample line spaced apart from a boundary of the block by 1 sample, or a non-neighboring sample line spaced apart from a boundary of the block by 2 samples. Information indicating the sample line at the specific position can be signaled through the bitstream.
[0127] The search position within the search range can be a position shifted by p from a position of the reference block of the current block in the x-axis direction and shifted by q from the position of the reference block of the current block in the y-axis direction, and can be determined based on an offset equally predefined for the encoding device and the decoding device. The same as described above in the “Modification method based on bilateral matching”, and overlapping description is omitted here.
[0128] By the above-described method, a cost array can be calculated for each of the first prediction direction and the second prediction direction. A cost having a minimum value among a plurality of costs belonging to the cost array in the first prediction direction can be identified, and an incremental motion vector in the first prediction direction can be determined based on an offset corresponding to the identified cost. Based on the determined incremental motion vector in the first prediction direction, the motion vector predictor in the first prediction direction can be modified. Similarly, a cost having a minimum value among a plurality of costs belonging to the cost array in the second prediction direction can be identified, and an incremental motion vector in the second prediction direction can be determined based on an offset corresponding to the identified cost. Based on the determined incremental motion vector in the second prediction direction, the motion vector predictor in the second prediction direction can be modified.
[0129] According to whether the current block satisfies a predetermined first condition, either of the above-described modification method based on bilateral matching or the modification method based on template matching can be selectively used. The first condition according to the present disclosure can include at least one of a condition that the current block performs bi-prediction, a condition that bi-directional reference pictures of the current block exist in a time order before and / or after the current picture, or a condition that a POC difference between the current picture and each reference picture is the same. The time order can refer to an output order (picture order count, POC) or an encoding order.
[0130] When the first condition is true, the modification method based on bilateral matching can be used, and when the first condition is false, the modification method based on template matching can be used. Alternatively, when all of the first conditions are true, the modification method based on bilateral matching can be used, and when even one of the first conditions is false, the modification method based on template matching can be used. Alternatively, when one of the first conditions is true, the modification method based on bilateral matching can be used, and when all of the first conditions are false, the modification method based on template matching can be used.
[0131] As an example, when the current block performs bi-prediction, the modification method based on bilateral matching can be used, and otherwise, the modification method based on template matching can be used.
[0132] Alternatively, when the POC difference (diffPOC0) between the current picture including the current block and the reference picture in the first prediction direction is the same as the POC difference (diffPOC1) between the current picture and the reference picture in the second prediction direction, the modification method based on bilateral matching can be used, and otherwise, the modification method based on template matching can be used.
[0133] Alternatively, when the current block performs bi-prediction and diffPOC is the same as diffPOC1, the modification method based on bilateral matching can be used, and otherwise (i.e., when the current block does not perform bi-prediction or diffPOC is not the same as diffPOC1), the modification method based on template matching can be used.
[0134] Alternatively, when the current block performs bi-prediction, the modification method based on bilateral matching can be used even when diffPOC is not the same as diffPOC1. On the other hand, when the current block does not perform bi-prediction, the modification method based on template matching can be used.
[0135] The AMVP-MERGE mode according to the present disclosure can be applied even if the current block is a block coded in IBC (Intra Block Copy) mode. When the inter prediction mode of the current block is the IBC mode, the current picture to which the current block belongs can be used as the reference picture, and otherwise, the method of Embodiment 1 described above can be applied in the same manner.
[0136] Example 2
[0137] The present disclosure relates to a case where the first prediction mode and the second prediction mode are the AMVP-MERGE mode. Specifically, it relates to a case where motion information for bi-prediction is obtained for each of the AMVP mode and the merge mode.
[0138] In the case of the AMVP-MERGE mode, two or more bi-directional motion information of the current block can be obtained. As an example, the bi-directional motion information can be obtained based on the AMVP mode, and the bi-directional motion information can be obtained based on the merge mode. However, it is not limited thereto, uni-directional motion information can be obtained based on the AMVP mode, and bi-directional motion information can be obtained based on the merge mode. Alternatively, bi-directional motion information can be obtained based on the AMVP mode, and uni-directional motion information can be obtained based on the merge mode. Hereinafter, it is assumed that bi-directional motion information is obtained for each of the AMVP mode and the merge mode for convenience of description.
[0139] Motion information can be obtained based on the AMVP mode for each prediction direction of the current block. To this end, at least one of information on whether to perform bi-prediction, an MVP index, a reference picture index, or MVD information can be signaled. When the current block performs bi-prediction, the MVP index, the reference picture index, and the MVD information can be signaled for each prediction direction.
[0140] In particular, an MVP candidate list can be configured for each prediction direction of a current block. The MVP candidate list can contain multiple MVP candidates. A motion vector predictor (MVP[0][X], MVP[0][1-X]) in a corresponding prediction direction can be derived from the MVP candidate list in the prediction direction. In MVP[0][X], [0] represents an AMVP mode, and [X] represents an LX direction. In other words, MVP[0][X] can represent a motion vector predictor in the LX direction derived based on the AMVP mode. MVP[0][X] can be derived based on an MVP candidate specified by an MVP index in the LX direction. As an example, MVP[0][X] can refer to a motion vector of the MVP candidate specified by the MVP index in the LX direction. Alternatively, MVP[0][X] can refer to a motion vector derived based on a motion vector of the MVP candidate specified by the MVP index in the LX direction and a motion vector difference. Here, the motion vector difference can be derived based on MVD information in the LX direction. Similarly, in MVP[0][1-X], [0] represents the AMVP mode, and [1-X] represents an L(l-X) direction. In other words, MVP[0][1-X] can represent a motion vector predictor in the L(l-X) direction derived based on the AMVP mode. MVP[0][1-X] can be derived based on an MVP candidate specified by an MVP index in the L(l-X) direction. As an example, MVP[0][1-X] can refer to a motion vector of the MVP candidate specified by the MVP index in the L(l-X) direction. Alternatively, MVP[0][1-X] can refer to a motion vector derived based on a motion vector of the MVP candidate specified by the MVP index in the L(l-X) direction and a motion vector difference. Here, the motion vector difference can be derived based on MVD information in the L(l-X) direction.
[0141] Meanwhile, motion information of a current block can be obtained based on a merge mode. To this end, at least one of a merge index or MVD information specifying any one of a plurality of merge candidates belonging to a merge candidate list can be signaled. Motion information of the current block can be obtained based on motion information of the merge candidate specified by the merge index. Here, the motion information can contain at least one of a motion vector (or a motion vector predictor) or a reference picture index. When the specified merge candidate performs bi-prediction, bi-directional motion information of the current block can be obtained. However, when the specified merge candidate performs uni-prediction, uni-directional motion information of the current block can be obtained.
[0142] In particular, a merge candidate list including a plurality of merge candidates can be configured for the current block. A bi-directional motion vector predictor (MVP[1][X], MVP[1][1-X]) can be derived from the merge candidate list. In MVP[1][X], [1] indicates the merge mode and [X] indicates the LX direction. In other words, MVP[1][X] can indicate a motion vector predictor in the LX direction derived based on the merge mode. MVP[1][X] can be derived based on a merge candidate specified by a merge index. As an example, MVP[1][X] can refer to a motion vector in the LX direction of a merge candidate specified by the merge index. Alternatively, MVP[1][X] can refer to a motion vector derived based on a motion vector in the LX direction of a merge candidate specified by the merge index and a motion vector difference. Here, the motion vector difference can be derived based on MVD information. Similarly, in MVP[1][1-X], [1] indicates the merge mode and [1-X] indicates the L(1-X) direction. In other words, MVP[1][1-X] can indicate a motion vector predictor in the L(1-X) direction derived based on the merge mode. MVP[1][1-X] can be derived based on a merge candidate specified by a merge index. As an example, MVP[1][1-X] can refer to a motion vector in the L(1-X) direction of a merge candidate specified by the merge index. Alternatively, MVP[1][1-X] can refer to a motion vector derived based on a motion vector in the L(1-X) direction of a merge candidate specified by the merge index and a motion vector difference. Here, the motion vector difference can be derived based on MVD information.
[0143] The bi-directional motion vector predictors (MVP[0][X], MVP[0][1-X]) derived based on the AMVP mode and the bi-directional motion vector predictors (MVP[1][X], MVP[1][1-X]) derived based on the merge mode can be paired by using a predefined method. One or more bi-directional motion vector predictors can be derived for the current block by the pairing.
[0144] As an example, one bi-directional motion vector predictor can be derived for the current block by the pairing. In this case, the bi-directional motion vector predictor of the current block can be {MVP[0][X], MVP[1][1-X]}. In other words, the bi-directional motion vector predictor for the current block can be derived as a combination of a motion vector predictor in the LX direction according to the AMVP mode and a motion vector predictor in the L(1-X) direction according to the merge mode. Alternatively, the bi-directional motion vector predictor of the current block can be {MVP[0][1-X], MVP[1][X]}. In other words, the bi-directional motion vector predictor for the current block can be derived as a combination of a motion vector predictor in the L(1-X) direction according to the AMVP mode and a motion vector predictor in the LX direction according to the merge mode.
[0145] Alternatively, two bi-directional motion vector predictors for the current block can be derived by pairing. In this case, the bi-directional motion vector predictors for the current block can be {MVP[0][X], MVP[1][1-X]} and {MVP[0][1-X], MVP[1][X]}. In other words, either one of the two bi-directional motion vector predictors can be derived as a combination of a motion vector predictor in the LX direction according to the AMVP mode and a motion vector predictor in the L(1-X) direction according to the merge mode. The other one of the two bi-directional motion vector predictors can be derived as a combination of a motion vector predictor in the L(1-X) direction according to the AMVP mode and a motion vector predictor in the LX direction according to the merge mode.
[0146] When one bi-directional motion vector predictor is derived for the current block, a first prediction block for the current block can be derived based on MVP[0][X] and a second prediction block for the current block can be derived based on MVP[1][1-X]. Alternatively, when one bi-directional motion vector predictor is derived for the current block, a first prediction block for the current block can be derived based on MVP[0][1-X] and a second prediction block for the current block can be derived based on MVP[1][X].
[0147] When two bi-directional motion vector predictors are derived for the current block, a first prediction block for the current block can be derived based on MVP[0][X] and MVP[1][1-X]. As an example, the first prediction block for the current block can be derived based on a weighted sum between a reference block within an LX reference picture specified by MVP[0][X] and a reference block within an L(1-X) reference picture specified by MVP[1][1-X]. Further, a second prediction block for the current block can be derived based on MVP[0][1-X] and MVP[1][X]. As an example, the first prediction block for the current block can be derived based on a weighted sum between a reference block within an L(1-X) reference picture specified by MVP[0][1-X] and a reference block within an LX reference picture specified by MVP[1][X].
[0148] Alternatively, when two bi-directional motion vector predictors are derived for the current block, either one of the two bi-directional motion vector predictors can be selectively used. To this end, an index specifying either one of the two bi-directional motion predictors can be used. The index can be signaled through a bitstream or can be derived based on a condition equally predefined for the encoding apparatus and the decoding apparatus. A first prediction block and a second prediction block for the current block can be derived based on the bi-directional motion vector predictor specified by the index, respectively.
[0149] Hereinafter, for the convenience of description, any one of {MVP[0][X], MVP[1][1-X]} and {MVP[0][1-X], MVP[1][X]} derived by the above pairing is referred to as a first motion pair, and the other is referred to as a second motion pair.
[0150] At least one of the first motion pair and the second motion pair can be modified based on any one of the above modification method based on bilateral matching or the modification method based on template matching. In this case, the first prediction block and / or the second prediction block of the current block can be derived based on the modified motion pair.
[0151] The first motion pair can be modified based on any one of the modification method based on bilateral matching or the modification method based on template matching according to whether the current block satisfies a predetermined first condition. The first condition here is the same as in Embodiment 1, and the repeated description will be omitted here.
[0152] The second motion pair can be modified based on any one of the modification method based on bilateral matching or the modification method based on template matching according to whether the current block satisfies a predetermined second condition.
[0153] The second condition of the second motion pair can be defined independently of the first condition of the first motion pair. As an example, the second condition can include at least one of a condition that the current block performs bi-prediction, a condition that bi-directional reference pictures of the current block exist in time order before and / or after the current picture, or a condition that a POC difference between the current picture and each reference picture is the same. The second condition of the second motion pair can be the same as the first condition of the first motion pair. Alternatively, the second condition of the second motion pair can be defined differently from the first condition of the first motion pair.
[0154] Alternatively, the second condition of the second motion pair can include at least one of a condition that the first condition of the first motion pair is false, a condition that the current block performs bi-prediction, a condition that a POC difference between the current picture and each reference picture is the same, a condition that a cost of the first motion pair calculated by the modification method based on bilateral matching or the modification method based on template matching is less than a first threshold value, a condition that a POC difference between the current picture and a reference picture of the second motion pair is less than or equal to a POC difference between the current picture and a reference picture of the first motion pair, or a condition that a difference between the first motion pair (or the modified first motion pair) and the second motion pair is greater than a second threshold value.
[0155] When the second condition is true, the modification method based on bilateral matching can be used, and when the second condition is false, the modification method based on template matching can be used. Alternatively, when all the second conditions are true, the modification method based on bilateral matching can be used, and when even one of the second conditions is false, the modification method based on template matching can be used. Alternatively, when one of the second conditions is true, the modification method based on bilateral matching can be used, and when all the second conditions are false, the modification method based on template matching can be used.
[0156] As an example, when the first condition among the second conditions for the first motion pair is true or when all the remaining conditions among the second conditions are not true, the modification method based on template matching can be used, otherwise, the modification method based on bilateral matching can be used.
[0157] The inter prediction method according to the present disclosure can be changed and applied as follows.
[0158] In the AMVP-MERGE mode, the bi-directional motion information can be restricted to be derived for the current block. As an example, only the merge candidates with bi-directional motion information can be included in the merge candidate list. Alternatively, when the merge candidate specified by the merge index has uni-directional motion information, the bi-directional motion information can be derived based on the corresponding uni-directional motion information.
[0159] Alternatively, in the AMVP-MERGE mode, either of the MVP[0][X] or the MVP[0][1-X] can be derived based on the other one of the pre-derived MVP[0][X] or the MVP[0][1-X]. As an example, the pre-derived MVP[0][X] or the MVP[0][1-X] can be mirrored to the reference picture in the opposite direction based on the current picture to derive the MVP with opposite sign.
[0160] Alternatively, in the AMVP-MERGE mode, either of the MVP[1][X] or the MVP[1][1-X] can be derived based on the other one of the pre-derived MVP[1][X] or the MVP[1][1-X]. As an example, the pre-derived MVP[1][X] or the MVP[1][1-X] can be mirrored to the reference picture in the opposite direction based on the current picture to derive the MVP with opposite sign.
[0161] The AMVP-MERGE mode according to the present disclosure can be applied even if the current block is a block coded in IBC (Intra Block Copy) mode. When the inter prediction mode of the current block is IBC mode, the current picture to which the current block belongs can be taken as the reference picture, otherwise the method of the above embodiment 2 can be applied in the same way.
[0162] Example 3
[0163] The present disclosure relates to a case where a first prediction mode and a second prediction mode are a combined inter and intra prediction (CIIP) mode. When the CIIP mode is applied to a current block, a first prediction block and a second prediction block can be derived for the current block, respectively.
[0164] The first prediction block according to the present disclosure can be derived based on inter prediction.
[0165] As an example, the first prediction block can be derived based on uni- or bi-directional motion information of the current block. Here, the uni- or bi-directional motion information can be derived based on an AMVP mode or a merge mode, which is the same as described in Embodiments 1 and 2.
[0166] The second prediction block according to the present disclosure can be derived based on intra prediction.
[0167] As an example, one intra prediction mode can be determined for the current block, and the second prediction block can be derived based on the corresponding intra prediction mode. The one intra prediction mode can be any one of a planar mode, a most probable mode (MPM), a matrix-based intra prediction (MIP) mode, a DIMM-based intra prediction mode to be described below, or a TIM-based intra prediction mode to be described below. The MPM can be one or more MPM candidates among a plurality of MPM candidates belonging to an MPM candidate list, which can be specified by an MPM index signaled through a bitstream.
[0168] One intra prediction mode can be predefined in the encoding apparatus and the decoding apparatus. Alternatively, any one of predefined intra prediction modes for the CIIP mode can be selected, and the selected intra prediction mode can be set as the one intra prediction mode for the current block. Here, the predefined intra prediction modes for the CIIP mode can include at least one of a planar mode, an MPM, a MIP mode, a DIMM-based intra prediction mode, or a TIM-based intra prediction mode. An index specifying any one of the predefined intra prediction modes can be used for the selection. The index can be signaled through a bitstream, or can be derived based on a condition equally predefined for the encoding apparatus and the decoding apparatus.
[0169] 1. Decoder-side intra mode derivation (DIMD) method
[0170] The gradient can be calculated based on at least two samples belonging to a neighboring region of the current block. Here, the gradient can include at least one of a horizontal gradient or a vertical gradient. An intra prediction mode of the current block can be derived based on at least one of the calculated gradient or a gradient magnitude. Here, the gradient magnitude can be determined based on a sum of the horizontal gradient and the vertical gradient. Through this derivation method, one intra prediction mode or at least two intra prediction modes of the current block can be derived.
[0171] As an example, the gradient can be calculated in units of a window having a predetermined size. An angle representing a directionality of samples within a corresponding window can be calculated based on the calculated gradient. The calculated angle can correspond to any one of the plurality of predefined intra prediction modes described above. A magnitude of the gradient can be stored / updated for an intra prediction mode corresponding to the calculated angle. Through this process, an intra prediction mode corresponding to the calculated gradient can be determined for each window, and the magnitude of the gradient can be stored / updated for the determined intra prediction mode. A top T intra prediction mode having a largest magnitude among the stored gradient magnitudes can be selected, and the selected intra prediction mode can be set as the intra prediction mode of the current block. Here, T can be an integer of 1, 2, 3, or more.
[0172] The neighboring region for calculating the gradient is a region that is pre-reconstructed before the current block, which can include at least one of a left region, an above region, a top-left region, a bottom-left region, and a top-right region neighboring the current block. The neighboring region can include at least one of a neighboring sample line adjacent to the current block, a first non-neighboring sample line spaced apart from the current block by 1 sample, or a second non-neighboring sample line spaced apart from the current block by 2 samples. However, it is not limited thereto, and can also include a non-neighboring sample line spaced apart from the current block by N samples, N can be an integer greater than or equal to 3.
[0173] The neighboring region can be a region that is equally predefined for calculating the gradient by the encoding apparatus and the decoding apparatus. Alternatively, the neighboring region can be variably determined based on information specifying a position of the neighboring region. In this case, the information specifying the position of the neighboring region can be signaled through a bitstream. Alternatively, the position of the neighboring region can be determined based on at least one of whether the current block is positioned at a boundary of a coding tree unit, a size (e.g., width, height, ratio of width to height, product of width and height) of the current block, a partition type of the current block, a prediction mode of the neighboring region, or availability of the neighboring region.
[0174] As an example, when the current block is located at a top boundary of a coding tree unit, at least one of a top region, a top-left region, or a top-right region of the current block can not be referred to for calculating the gradient. When a width of the current block is greater than a height, either one of the top region or the left region (e.g., the top region) can be referred to for calculating the gradient, and the other (e.g., the left region) can not be referred to for calculating the gradient. Conversely, when the width of the current block is less than the height, either one of the top region or the left region (e.g., the left region) can be referred to for calculating the gradient, and the other (e.g., the top region) can not be referred to for calculating the gradient. When the current block is generated by horizontal block splitting, the top region can not be referred to for calculating the gradient. Conversely, when the current block is generated by vertical block splitting, the left region can not be referred to for calculating the gradient. When neighboring regions of the current block are coded in an inter mode, the corresponding neighboring regions can not be referred to for calculating the gradient. However, this is not limited thereto, and the corresponding neighboring regions can be referred to for calculating the gradient regardless of the prediction mode of the neighboring regions.
[0175] 2. Template-based intra mode derivation (TIMD) method
[0176] From a decoder side, an intra prediction mode can be derived based on a template region neighboring the current block, which will be described in detail below.
[0177] A cost of each of the predetermined candidate modes can be calculated.
[0178] The predetermined candidate modes can refer to a plurality of intra prediction modes that are equally predefined for the encoding apparatus and the decoding apparatus. Alternatively, for the derivation based on the template region, a candidate list consisting of the candidate modes can be generated, and a cost of the candidate modes belonging to the candidate list can be calculated. Alternatively, the cost can be calculated only for the first N candidate modes within the generated candidate list. Here, N can be a value that is equally predefined for the encoding apparatus and the decoding apparatus. As an example, N can be an integer of 2, 3, 4, 5, or more.
[0179] The candidate list for the derivation based on the template region can be configured in the same manner as the above-described MPM list. Alternatively, the candidate list can correspond to the above-described first MPM list or second MPM list. Alternatively, the candidate list can be configured as a combination of the above-described first and second groups (i.e., the MPM list), or can be configured as a combination of subgroups of the first and second groups (i.e., the first MPM list or the second MPM list).
[0180] The cost can be calculated as a sum of absolute differences (SAD) between the reconstructed samples and the predicted samples within the template region. Alternatively, the cost can also be calculated as a sum of absolute transformed differences (SATD) between the reconstructed samples and the predicted samples within the template region. Here, the SATD can refer to the SAD transformed into the frequency domain. As an example of the transformation, the Hadamard transformation can be used, but is not limited thereto. The predicted samples of the template region can be generated based on the candidate modes described above.
[0181] The template region for calculating the cost can be a pre-reconstructed region adjacent to the current block. As an example, the template region can include at least one of a top-adjacent region, a left-adjacent region, an upper-left adjacent region, a lower-left adjacent region, or an upper-right adjacent region of the current block.
[0182] The template region can be a region equally predefined for calculating the cost by the encoding apparatus and the decoding apparatus. Alternatively, the template region can be variably determined based on information specifying a position of the template region. In this case, the information specifying the position of the template region can be signaled through a bitstream. Alternatively, the position of the template region can be determined based on at least one of whether the current block is positioned at a boundary of a coding tree unit, a size (e.g., width, height, ratio of width to height, product of width and height) of the current block, a partition type of the current block, a prediction mode of an adjacent region, or availability of the adjacent region.
[0183] As an example, when the current block is positioned at a top boundary of a coding tree unit, at least one of a top-adjacent region, an upper-left adjacent region, or an upper-right adjacent region of the current block can not be referred to for calculating the cost. When a width of the current block is greater than a height, any one of the top-adjacent region or the left-adjacent region (e.g., the top-adjacent region) can be referred to for calculating the cost, and the other (e.g., the left-adjacent region) can not be referred to for calculating the gradient. Conversely, when the width of the current block is less than the height, any one of the top-adjacent region or the left-adjacent region (e.g., the left-adjacent region) can be referred to for calculating the cost, and the other (e.g., the top-adjacent region) can not be referred to for calculating the cost. When the current block is generated by a horizontal block partitioning, the top-adjacent region can not be referred to for calculating the cost. Conversely, when the current block is generated by a vertical block partitioning, the left-adjacent region can not be referred to for calculating the cost. When the adjacent regions of the current block are coded in an inter mode, the corresponding adjacent region can not be referred to for calculating the cost. However, this is not limited thereto, and the corresponding adjacent region can be referred to for calculating the cost regardless of the prediction mode of the adjacent region.
[0184] The template region can be constituted by N reference sample lines. Here, N can be an integer of 1, 2, 3, 4, or more. The number of reference sample lines configuring the template region can be the same regardless of the position of the neighboring region, or can vary depending on the position of the neighboring region. The cost can be calculated based on all samples belonging to the template region. Alternatively, the cost can be calculated by using only the reference sample lines at predetermined positions within the template region. Alternatively, the cost can be calculated based on all samples belonging to the reference sample lines at the predetermined positions, or can be calculated by using only the samples at predetermined positions on the reference sample lines at the predetermined positions. The positions of the samples and / or the reference sample lines used for the cost calculation can be determined based on at least one of whether the current block is positioned at a boundary of a coding tree unit, the size (e.g., width, height, ratio of width to height, product of width and height) of the current block, the partition type of the current block, the prediction mode of the neighboring region, or the availability of the neighboring region. Alternatively, information indicating the positions of the reference sample lines used for the cost calculation can be signaled through a bitstream.
[0185] One candidate mode having a minimum cost among the costs calculated for the candidate modes can be selected. As an example, the costs of five candidate modes within a candidate list can be calculated, respectively. The five candidate modes in the candidate list can be reordered in ascending order of the calculated costs. A first candidate mode can be selected among the 5 reordered candidate modes.
[0186] Alternatively, at least two candidate modes having minimum costs among the costs calculated for the candidate modes can also be selected. As an example, the costs of five candidate modes within a candidate list can be calculated, respectively. The five candidate modes in the candidate list can be reordered in ascending order of the calculated costs. A first two candidate modes can be selected among the 5 reordered candidate modes.
[0187] The one or more candidate modes selected through the above-described process can be set as the intra prediction mode of the current block.
[0188] Alternatively, when at least two candidate modes are selected through the above-described process, the intra prediction mode of the current block can be derived based on a comparison between the selected candidate modes and / or a comparison between at least one of the selected candidate modes and a threshold. As an example, the intra prediction mode of the current block can be derived based on whether the selected candidate modes satisfy the following condition.
[0189] [Condition] costMode2 < (K x costMode1)
[0190] In this condition, costMode1 can refer to a cost calculated based on any one of the selected candidate modes, and costMode2 can refer to a cost calculated based on another one of the selected candidate modes. As an example, costMode1 can refer to a cost calculated based on a candidate mode having a smaller cost among the selected candidate modes, and costMode2 can refer to a cost calculated based on a candidate mode having a larger cost among the selected candidate modes. In this condition, K denotes a predetermined comparison factor, which can be a value equally predefined for the encoding apparatus and the decoding apparatus. As an example, K can be an integer of 1, 2, or more, or can refer to a real number such as 1 / 2 or 1 / 4.
[0191] When the condition is satisfied, the selected candidate mode can be set as the intra prediction mode of the current block. On the other hand, when the condition is not satisfied, the candidate mode having the cost of costMode1 can be set as the intra prediction mode of the current block, and the candidate mode having the cost of costMode2 can not be used as the intra prediction mode of the current block.
[0192] Alternatively, two or more intra prediction modes can be determined for the current block, and a second prediction block can be derived based on the two or more determined intra prediction modes. In other words, prediction blocks can be derived based on the two or more intra prediction modes, respectively, and a second prediction block can be derived based on a weighted sum between the two or more derived prediction blocks. As an example, when two intra prediction modes of the current block are determined, a second prediction block can be derived as shown in Equation 1 below.
[0193] [Equation 1]
[0194] P intra = ((4 – w1) * P intra0 + w1 * P intra1 + offset) >> shift
[0195] In Equation 1, P intra denotes the second prediction block (or a sample of the second prediction block). Among them, P intra0 denotes a prediction block (or a sample of the prediction block) derived based on any one of the two intra prediction modes, and P intra1represents a prediction block (or samples of the prediction block) derived based on another one of the two intra prediction modes. w1 represents a weight of a weighted sum between the two prediction blocks. As an example, when the shift is 2 and the offset is 2, w1 can be determined as an integer in a range of 1 to 3. Alternatively, when the shift is 3, 4, or 5, the offset can be 4, 8, or 16, respectively, and w1 can be determined as an integer in a range of 1 to 7, 1 to 15, or 1 to 32, respectively.
[0196] The two or more intra prediction modes can include at least two of a planar mode, an MPM, an MIP mode, a DIMM-based intra prediction mode, or a TIMD-based intra prediction mode. A number of the intra prediction modes used to derive the second prediction block can be determined according to a predefined number.
[0197] Example 4
[0198] The disclosure relates to a case where the first prediction mode is a CIIP mode and the second prediction mode is an intra mode. In this case, the first prediction block and the second prediction block can be derived for the current block, respectively.
[0199] The first prediction block according to the disclosure can be derived based on an inter prediction and an intra prediction.
[0200] As an example, an inter prediction block can be derived based on the inter prediction, and an intra prediction block can be derived based on the intra prediction. The first prediction block can be derived based on the derived inter prediction block and the intra prediction block.
[0201] The inter prediction block can be derived based on uni-directional or bi-directional motion information of the current block. Here, the uni-directional or bi-directional motion information can be derived based on an AMVP mode or a merge mode, which is the same as described in Embodiments 1 and 2.
[0202] In addition, one intra prediction mode can be determined for the current block, and the intra block can be derived based on the corresponding intra prediction mode. The one intra prediction mode can be any one of a planar mode, an MPM, an MIP mode, a DIMM-based intra prediction mode, or a TIMD-based intra prediction mode. The method of determining one intra prediction mode for the current block is the same as in Embodiment 3, and a repeated description will be omitted here.
[0203] The first prediction block can be derived by a weighted sum between the derived inter prediction block and the intra prediction block. The first prediction block can be derived as shown in Equation 2 below.
[0204] [Equation 2]
[0205] P temp = ((4 – w2) * Pinter + w2 * P intra + offset) >> shift
[0206] In Equation 2, P temp denotes the first prediction block (or samples of the first prediction block). P inter denotes the inter prediction block (or, samples of the inter prediction block), and P intra denotes the intra prediction block (or, samples of the intra prediction block). w2 denotes the weight of the weighted sum between the inter and intra prediction blocks. The values of shift, offset and weight in Equation 2 are the same as described in Equation 1.
[0207] The second prediction block according to the present disclosure can be derived based on intra prediction.
[0208] As an example, an additional intra prediction mode can be determined for the current block, and the second prediction block can be derived based on the corresponding intra prediction mode. The additional intra prediction mode is an intra prediction mode different from the intra prediction mode of the intra prediction block of the first prediction block, which can be any one of the planar mode, the MPM, the MIP mode, the DIMM-based intra prediction mode or the TIM-based intra prediction mode. The additional intra prediction mode can be determined based on the method of determining one intra prediction mode for the current block described in Embodiment 3, and the repeated description will be omitted here.
[0209] However, according to the condition predefined equally for the encoding device and the decoding device, the derivation of the second prediction block (or the determination of the additional intra prediction mode) can also be omitted.
[0210] Example 5
[0211] The present disclosure relates to the case where the first prediction mode is the CIIP mode and the second prediction mode is the intra mode. In this case, the first prediction block and the second prediction block can be derived for the current block, respectively.
[0212] The first prediction block according to the present disclosure can be derived based on inter prediction and intra prediction.
[0213] As an example, the inter prediction block can be derived based on inter prediction, and the intra prediction block can be derived based on intra prediction. The first prediction block can be derived based on a weighted sum between the derived inter prediction block and the intra prediction block.
[0214] The inter prediction block can be derived based on the uni-directional or bi-directional motion information of the current block. Here, the uni-directional or bi-directional motion information can be derived based on the AMVP mode or the merge mode, which is the same as described in Embodiments 1 and 2.
[0215] In addition, two or more intra prediction modes can be determined for the current block, and the intra block can be derived based on any one of the two or more intra prediction modes. The two or more intra prediction modes can include at least two of a planar mode, an MPM, an MIP mode, a DIMM-based intra prediction mode, or a TIM-based intra prediction mode. The method of determining the two or more intra prediction modes of the current block is the same as described in Embodiment 3, and the repeated description will be omitted here.
[0216] The first prediction block can be derived by a weighted sum between the derived inter prediction block and the intra prediction block. The first prediction block can be derived as shown in Equation 3 below.
[0217] [Equation 3]
[0218] P temp = ((4 – w3) * P inter + w3 * P intra + offset) >> shift
[0219] In Equation 3, P temp represents the first prediction block (or a sample of the first prediction block). P inter represents the inter prediction block (or, a sample of the inter prediction block), and P intra represents the intra prediction block (or, a sample of the intra prediction block). w3 represents a weight of the weighted sum between the inter and intra prediction blocks. The values of the shift, the offset, and the weight in Equation 3 are the same as described in Equation 1.
[0220] The second prediction block according to the present disclosure can be derived based on an intra prediction.
[0221] As an example, the second prediction block can be derived based on another one of the two or more intra prediction modes determined for the current block.
[0222] Example 6
[0223] The present disclosure relates to a case where the first prediction mode is an IBC mode and the second prediction mode is an intra mode. In this case, the first prediction block and the second prediction block can be derived for the current block, respectively.
[0224] The first prediction block according to the present disclosure can be derived based on an IBC mode.
[0225] An IBC candidate list can be configured for the current block. The IBC candidate list can include a plurality of IBC candidates, and the plurality of IBC candidates can be derived based on neighboring blocks (or block vectors of the neighboring blocks) adjacent to the current block. One or more block vectors can be derived from the IBC candidate list. The first prediction block can be derived based on the derived block vector.
[0226] As an example, any one of the multiple IBC candidates belonging to the IBC candidate list of the current block can be selected. An index specifying any one of the multiple IBC candidates can be used for the selection. Here, the index can be signaled through the bitstream. The block vector of the current block can be derived based on the block vector of the selected IBC candidate. The reference block specified based on the derived block vector can be set as the first prediction block. Here, the reference block can belong to the current picture to which the current block belongs.
[0227] Alternatively, at least two of the multiple IBC candidates belonging to the IBC candidate list of the current block can be selected. To this end, a first index specifying any one of the multiple IBC candidates and a second index specifying another one of the multiple IBC candidates can be used. The first index and the second index can be signaled through the bitstream. Alternatively, the first index can be signaled through the bitstream, and the second index can be derived based on the signaled first index. A first block vector of the current block can be derived based on the block vector of the IBC candidate selected based on the first index. A second block vector of the current block can be derived based on the block vector of the IBC candidate selected based on the second index. The first prediction block can be derived by a weighted sum of a first reference block specified based on the first block vector and a second reference block specified based on the second block vector. Here, the first reference block and the second reference block can belong to the current picture to which the current block belongs.
[0228] Meanwhile, two different IBC candidates can be selected among the multiple IBC candidates based on the first index and the second index. For example, an index can be assigned to each of the multiple IBC candidates belonging to the IBC candidate list. The IBC candidate having the same index as the first index can be selected from the IBC candidate list. When the second index is smaller than the first index, the IBC candidate having the same index as the second index can be selected from the IBC candidate list. On the other hand, when the second index is greater than or equal to the first index, the IBC candidate having the same index as a value obtained by adding 1 to the second index can be selected from the IBC candidate list.
[0229] The second prediction block according to the present disclosure can be derived based on an intra prediction.
[0230] As an example, one intra prediction mode can be determined for the current block, and the second prediction block can be derived based on the corresponding intra prediction mode. The one intra prediction mode can be any one of a planar mode, an MPM, an MIP mode, a DIMM-based intra prediction mode, or a TIMD-based intra prediction mode. It is the same as described in Embodiment 3.
[0231] Alternatively, two or more intra prediction modes can be determined for the current block, and a second prediction block can be derived based on the two or more determined intra prediction modes. In other words, prediction blocks can be derived based on the two or more intra prediction modes, respectively, and the second prediction block can be derived based on a weighted sum between the two or more derived prediction blocks. This is the same as described in Embodiment 3, and the repeated description will be omitted here.
[0232] Referring to Figure 4 A final prediction block of the current block can be derived based on the multiple prediction blocks of the current block S410.
[0233] In other words, a prediction block of the current block can be derived based on a weighted sum of the first and second prediction blocks of the current block. As an example, the prediction block of the current block can be derived as shown in Equation 4 below.
[0234] [Equation 4]
[0235] P final = ((4 - w4) * P0 + w4 * P1 + offset) » shift
[0236] In Equation 4, P final denotes a prediction block (or a sample of the prediction block) of the current block. P0 denotes a first prediction block (or, a sample of the first prediction block), and P1 denotes a second prediction block (or, a sample of the second prediction block). w4 denotes a weight of a weighted sum between the first and second prediction blocks. As an example, when the shift is 2 and the offset is 2, w4 can be determined as an integer in a range of 1 to 3. Alternatively, when the shift is 3, 4, or 5, the offset can be 4, 8, or 16, respectively, and w4 can be determined as an integer in a range of 1 to 7, 1 to 15, or 1 to 32, respectively.
[0237] Figure 5 A schematic configuration of an inter predictor 332 performing an inter prediction method according to the present disclosure is shown.
[0238] Referring to Figure 5 The inter predictor 332 can include a multiple prediction block deriver 500 and a multiple prediction block weighter 510.
[0239] The multiple prediction block deriver 500 can derive multiple prediction blocks of the current block. The multiple prediction blocks can include a first prediction block and a second prediction block. The first prediction block can be derived based on a first prediction mode, and the second prediction block can be derived based on a second prediction mode. Here, the first prediction mode and the second prediction mode can be the same prediction mode, or can be different prediction modes. The first prediction block and the second prediction block according to the first prediction mode and the second prediction mode can be derived based on any one of the above-described embodiments 1 to 6, and detailed descriptions will be omitted here.
[0240] The multiple prediction block weighter 510 can derive a final prediction block of the current block based on the multiple prediction blocks of the current block. In other words, the multiple prediction block weighter 510 can derive a prediction block of the current block based on a weighted sum of the first prediction block and the second prediction block of the current block. This is the same as described with reference to Figure 4
[0241] Figure 6 An inter prediction method performed by the encoding apparatus 200 according to an embodiment of the disclosure is illustrated.
[0242] Referring to Figure 6 , multiple prediction blocks of the current block can be derived S600.
[0243] The multiple prediction blocks can include a first prediction block and a second prediction block. The first prediction block can be derived based on a first prediction mode, and the second prediction block can be derived based on a second prediction mode. Here, the first prediction mode and the second prediction mode can be the same prediction mode, or can be different prediction modes. The first prediction block and the second prediction block according to the first prediction mode and the second prediction mode can be derived based on any one of the above-described embodiments 1 to 6, and detailed descriptions will be omitted here.
[0244] Referring to Figure 6 , a final prediction block of the current block can be derived based on the multiple prediction blocks of the current block S610. In other words, a prediction block of the current block can be derived based on a weighted sum of the first prediction block and the second prediction block of the current block, which is the same as described with reference to Figure 4
[0245] Figure 7 A schematic configuration of the inter predictor 221 performing an inter prediction method according to the disclosure is illustrated.
[0246] Referring to Figure 7 , the inter predictor 221 can include a multiple prediction block deriver 700 and a multiple prediction block weighter 710.
[0247] The multiple prediction block deriver 700 can derive multiple prediction blocks of the current block. The multiple prediction blocks can include a first prediction block and a second prediction block. The first prediction block can be derived based on a first prediction mode, and the second prediction block can be derived based on a second prediction mode. Here, the first prediction mode and the second prediction mode can be the same prediction mode or different prediction modes. The first prediction block and the second prediction block according to the first prediction mode and the second prediction mode can be derived based on any one of the above-described embodiments 1 to 6, and detailed descriptions will be omitted here.
[0248] The multiple prediction block weighter 710 can derive a final prediction block of the current block based on the multiple prediction blocks of the current block. In other words, the multiple prediction block weighter 510 can derive the prediction block of the current block based on a weighted sum of the first prediction block and the second prediction block of the current block. This is the same as described with reference to Figure 4
[0249] In the above-described embodiments, the method is described as a series of steps or blocks based on the flowcharts, but the corresponding embodiments are not limited to the order of the steps, and some steps can occur simultaneously or in a different order from other steps as described above. In addition, it can be understood by those skilled in the art that the steps shown in the flowcharts are not exclusive, and other steps can be included or one or more steps in the flowcharts can be deleted without affecting the scope of the embodiments of the present disclosure.
[0250] The above-described method according to the embodiments of the present disclosure can be implemented in the form of software, and an encoding apparatus and / or a decoding apparatus according to the present disclosure can be included in a device that performs image processing, such as a TV, a computer, a smartphone, a set-top box, a display device, etc.
[0251] In the present disclosure, when the embodiments are implemented as software, the above-described method can be implemented as a module (process, function, etc.) that performs the above-described functions. The module can be stored in a memory and can be executed by a processor. The memory can be located inside or outside the processor, and can be connected to the processor by various well-known means. The processor can include an application-specific integrated circuit (ASIC), another chip set, a logic circuit, and / or a data processing device. The memory can include a read-only memory (ROM), a random access memory (RAM), a flash memory, a memory card, a storage medium, and / or other storage devices. In other words, the embodiments described herein can be executed by being implemented on a processor, a microprocessor, a controller, or a chip. For example, the functional units shown in each of the drawings can be executed by being implemented on a computer, a processor, a microprocessor, a controller, or a chip. In this case, information (e.g., information on instructions) or algorithms for implementation can be stored in a digital storage medium.
[0252] In addition, the decoding apparatus and the encoding apparatus to which the embodiments of the disclosure are applied can be included in a multimedia broadcast transmitting and receiving device, a mobile communication terminal, a home theater video device, a digital theater video device, a surveillance camera, a video session device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camcorder, a device for providing a video on demand (VoD) service, an over-the-top video (OTT) device, a device for providing an Internet streaming service, a three-dimensional (3D) video device, a virtual reality (VR) device, an augmented reality (AR) device, a video phone device, a transportation terminal (e.g., a vehicle (including an autonomous vehicle) terminal, an airplane terminal, a ship terminal, etc.), and a medical video device, etc., and can be used to process a video signal or a data signal. For example, the over-the-top video (OTT) device can include a game console, a Blu-ray player, a networked TV, a home theater system, a smartphone, a tablet, a digital video recorder (DVR), etc.
[0253] In addition, the processing method to which the embodiments of the disclosure are applied can be generated in the form of a program executed by a computer, and can be stored in a computer-readable recording medium. Multimedia data having a data structure according to the embodiments of the disclosure can also be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices that store computer-readable data. The computer-readable recording medium can include, for example, a Blu-ray disc (BD), a universal serial bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disc, and an optical media storage device. In addition, the computer-readable recording medium includes a medium implemented in a carrier wave form (e.g., transmission via the Internet). In addition, a bitstream generated by an encoding method can be stored in a computer-readable recording medium or can be transmitted through a wired or wireless communication network.
[0254] In addition, the embodiments of the disclosure can be implemented by a computer program product through program codes, and the program codes can be executed on a computer by the embodiments of the disclosure. The program codes can be stored on a computer-readable carrier.
[0255] Figure 8 An example of a content streaming system to which the embodiments of the disclosure can be applied is shown.
[0256] Reference Figure 8 A content streaming system to which the embodiments of the disclosure are applied can mainly include an encoding server, a streaming server, a web server, a media store, a user device, and a multimedia input device.
[0257] The encoding server generates a bitstream by compressing content input from a multimedia input device such as a smartphone, a camera, a camcorder, or the like into digital data, and transmits it to the streaming server. As another example, when a multimedia input device such as a smartphone, a camera, a camcorder, or the like directly generates a bitstream, the encoding server can be omitted.
[0258] A bitstream can be generated by applying the encoding method or the bitstream generation method of the embodiment of the disclosure, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.
[0259] The streaming server transmits multimedia data to a user device through a web server based on a request of a user, and the web server serves as a medium that informs the user of what service is available. When the user requests a desired service to the web server, the web server delivers it to the streaming server, and the streaming server transmits multimedia data to the user. In this case, the content streaming system can include a separate control server, and in this case, the control server controls commands / responses between each device in the content streaming system.
[0260] The streaming server can receive content from a media storage and / or an encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store a bitstream for a certain period of time.
[0261] Examples of the user device can include a mobile phone, a smartphone, a laptop computer, a digital broadcasting terminal, a personal digital assistant (PDA), a portable multimedia player (PMP), a navigation, a tablet PC, a tablet, an ultrabook, a wearable device (for example, a smartwatch, smartglasses, a head-mounted display (HMD), a digital TV, a desktop, a digital signage, etc.).
[0262] Each server in the content streaming system can be operated as a distributed server, and in this case, data received from each server can be distributed and processed.
[0263] The claims set forth herein can be combined in various ways. For example, the technical features of the method claims of the disclosure can be combined and implemented as an apparatus, and the technical features of the apparatus claims of the disclosure can be combined and implemented as a method. In addition, the technical features of the method claims of the disclosure and the technical features of the apparatus claims of the disclosure can be combined and implemented as an apparatus, and the technical features of the method claims of the disclosure and the technical features of the apparatus claims of the disclosure can be combined and implemented as a method.
Claims
1. An image decoding method, comprising: Export multiple prediction blocks for the current block, wherein the multiple prediction blocks include a first prediction block and a second prediction block; Based on the first prediction block and the second prediction block, derive the final prediction block for the current block; and The current block is reconstructed based on the final predicted block of the current block.
2. The method according to claim 1, wherein, The first prediction block is derived based on at least one of a first L0 motion vector predictor or a second L1 motion vector predictor. The second prediction block is derived based on at least one of a first L1 motion vector predictor or a second L0 motion vector predictor, and Specifically, at least one of the first L0 motion vector predictor or the first L1 motion vector predictor is derived based on a first prediction mode, and at least one of the second L0 motion vector predictor or the second L1 motion vector predictor is derived based on a second prediction mode.
3. The method according to claim 2, wherein, The first prediction mode is the AMVP mode, and The second prediction mode is the merging mode.
4. The method according to claim 1, wherein, The first L0 motion vector predictor is derived based on the first prediction mode, and Specifically, the first L1 motion vector predictor is derived based on the pre-derived first L0 motion vector predictor.
5. The method according to claim 2, wherein, At least one of the first L0 motion vector predictor or the second L1 motion vector predictor is modified based on either a bilateral matching-based modification method or a template matching-based modification method, and Specifically, based on a predefined first condition, either the bilateral matching-based modification method or the template matching-based modification method is selected.
6. The method according to claim 5, wherein, At least one of the first L1 motion vector predictor or the second L0 motion vector predictor is modified based on either a bilateral matching-based modification method or a template matching-based modification method, and Specifically, based on a predefined second condition, either the bilateral matching-based modification method or the template matching-based modification method is selected.
7. The method according to claim 6, wherein, Whether the second condition is met is determined based on whether the first condition is met.
8. The method according to claim 1, wherein, The first prediction block is derived based on the inter-frame pattern, and The second prediction block is derived based on one or more intra-frame prediction modes used for the current block.
9. The method according to claim 8, wherein, The one or more intra-prediction modes used for the current block include at least one of planar mode, MPM, MIP mode, DIMM-based intra-prediction mode, or TIM-based intra-prediction mode.
10. The method according to claim 1, wherein, The first prediction block is derived based on the inter-frame prediction blocks and intra-frame prediction blocks of the current block, and The second prediction block is derived based on an additional intra-frame prediction mode used for the current block.
11. The method according to claim 1, wherein, The first prediction block is derived based on one or more block vectors derived from the IBC candidate list of the current block, and The second prediction block is derived based on one or more intra-frame prediction modes used for the current block.
12. An image encoding method, comprising: Export multiple prediction blocks for the current block, wherein the multiple prediction blocks include a first prediction block and a second prediction block; The final prediction block of the current block is derived based on the first prediction block and the second prediction block; The residual block of the current block is derived based on the final predicted block of the current block; and The residual block is encoded for the current block.
13. A computer-readable storage medium for storing a bitstream generated by an image encoding method, the image encoding method comprising: Export multiple prediction blocks for the current block, wherein the multiple prediction blocks include a first prediction block and a second prediction block; The final prediction block of the current block is derived based on the first prediction block and the second prediction block; The residual block of the current block is derived based on the final predicted block of the current block; and The residual block is encoded for the current block.
14. A method for sending data, comprising: Obtain a bitstream for image information, wherein the bitstream is generated based on: deriving a plurality of prediction blocks for the current block, including a first prediction block and a second prediction block; deriving a final prediction block for the current block based on the first prediction block and the second prediction block; deriving a residual block for the current block based on the final prediction block; and encoding the residual block for the current block; and Send data including the bit stream.