Decoding method, encoding method, and method for transmitting image information

By using a multi-reference mode-based inter-frame prediction and coding method, the final prediction block of the image is generated, which solves the problem of increased data volume in high-resolution images and achieves efficient image compression and cost reduction.

CN121844561APending Publication Date: 2026-04-10LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
LG ELECTRONICS INC
Filing Date
2024-09-13
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

As the resolution and quality of high-definition and ultra-high-definition images improve, the amount of information transmitted in image data increases, leading to higher transmission and storage costs and necessitating efficient image compression technologies.

Method used

An inter-frame prediction method based on a multi-reference mode is adopted. The final prediction block is generated by a weighted average of the regular prediction block and the additional prediction block, and the image information is sent through an encoding method.

Benefits of technology

It achieves efficient image compression, reduces transmission and storage costs, and improves image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121844561A_ABST
    Figure CN121844561A_ABST
Patent Text Reader

Abstract

The decoding method comprises the steps of: obtaining information on a basic prediction; generating a basic prediction block by performing the basic prediction based on the information on the basic prediction; obtaining information about the additional prediction; generating an additional prediction block by performing additional prediction based on the information on the additional prediction; and generating a final prediction block for the current block based on a combination of the basic prediction block and the additional prediction block, where the basic prediction may include one of intra prediction, inter prediction, and Intra Block Copy (IBC), and the additional prediction may include at least one of intra prediction, inter prediction, and Intra Block Copy.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The disclosure relates to a decoding method, an encoding method, and a method for transmitting image information, and to a method for processing a multi-reference block. BACKGROUND

[0002] Recently, in various fields, the demand for high-resolution and high-quality images such as high definition (HD) images and ultra-high definition (UHD) images is increasing. As the resolution and quality of image data increase, the amount of transmitted information or bits relatively increases compared to existing image data. The increase in the amount of transmitted information or bits results in an increase in transmission costs and storage costs.

[0003] Therefore, an efficient image compression technique is needed to effectively transmit, store, and reproduce information about high-resolution and high-quality images. SUMMARY

[0004] TECHNICAL PROBLEM

[0005] The disclosure provides a method and an apparatus for performing inter prediction based on a multi-reference mode.

[0006] The disclosure also provides a signaling method and an apparatus for determining a multi-reference mode.

[0007] The disclosure also provides a derivation method and an apparatus for determining a multi-reference mode.

[0008] TECHNICAL SOLUTION

[0009] According to an embodiment of the disclosure, a decoding method includes the steps of obtaining information about a regular prediction, generating a regular prediction block by performing a regular prediction based on the information about the regular prediction, obtaining information about an additional prediction, and generating an additional prediction block by performing an additional prediction based on the information about the additional prediction, generating a final prediction block for a current block based on a weighted average of the regular prediction block and the additional prediction block, wherein the regular prediction includes any one of intra prediction, inter prediction, or intra block copy (IBC), and wherein the additional prediction includes at least one of intra prediction, inter prediction, or IBC.

[0010] The step of obtaining the information about the additional prediction can include obtaining additional prediction availability information indicating whether the additional prediction is performed, obtaining mode index information indicating an additional prediction mode based on the additional prediction availability information, and obtaining additional block generation information for generating the additional prediction block based on the mode index information, wherein the additional block generation information includes at least one of intra block generation information, inter block generation information, or IBC block generation information.

[0011] The step of obtaining the additional block generation information for generating the additional prediction block can comprise obtaining the additional block generation information for generating the additional prediction block based on satisfying a preset condition, wherein the preset condition comprises at least one of whether the additional prediction mode is allowed for a regular prediction mode, whether the additional prediction mode is allowed for a tool applied to the regular prediction mode, whether the additional prediction mode is allowed for a width or a height of the current block, whether the additional prediction mode is allowed for a quantization parameter (QP) of the current block, or whether the additional prediction mode is allowed for a temporal layer of the current block.

[0012] The step of obtaining the information on the additional prediction can comprise obtaining additional prediction availability information indicating whether the additional prediction is performed, and inheriting the additional block generation information for generating the additional prediction block based on the additional prediction availability information, and wherein the additional block generation information comprises at least one of intra block generation information, inter block generation information, or IBC block generation information.

[0013] The step of inheriting the additional block generation information for generating the additional prediction block can comprise inheriting the additional block generation information of a neighboring block having a regular prediction mode identical to a regular prediction mode for the current block.

[0014] The step of inheriting the additional block generation information for generating the additional prediction block can comprise deriving an additional prediction mode of a neighboring block having a regular prediction mode identical to a regular prediction mode for the current block, and inheriting the additional block generation information of the additional prediction mode of the neighboring block selected based on an inheritance priority of the additional prediction mode.

[0015] The step of inheriting the additional block generation information for generating the additional prediction block can comprise deriving an additional prediction mode of a neighboring block having a regular prediction mode identical to a regular prediction mode for the current block, and inheriting the additional block generation information of the additional prediction mode of the neighboring block allowed based on the regular prediction mode for the current block.

[0016] The step of obtaining the information on the additional prediction can comprise obtaining additional prediction availability information indicating whether the additional prediction is performed, and deriving the additional block generation information for generating the additional prediction block based on the additional prediction availability information, and wherein the additional block generation information comprises at least one of intra block generation information, inter block generation information, or IBC block generation information.

[0017] The step of inheriting the additional block generation information for generating the additional prediction block can comprise deriving an additional prediction mode of a neighboring block having a regular prediction mode identical to a regular prediction mode for the current block, and inheriting the additional block generation information of the additional prediction mode of the neighboring block selected based on an inheritance priority of the additional prediction mode.

[0018] The step of inheriting additional block generation information for generating an additional prediction block can include deriving an additional prediction mode of a neighboring block having a same regular prediction mode as a regular prediction mode for the current block, and inheriting additional block generation information of the additional prediction mode of the neighboring block allowed based on the regular prediction mode for the current block.

[0019] The step of obtaining information on additional prediction can include obtaining additional prediction availability information indicating whether to perform additional prediction, deriving additional block generation information for generating an additional prediction block based on the additional prediction availability information, and wherein the additional block generation information includes at least one of intra block generation information, inter block generation information, or IBC block generation information.

[0020] The step of deriving additional prediction execution information for performing additional prediction can include deriving the additional prediction execution information using a template matching method.

[0021] The step of deriving additional prediction execution information for performing additional prediction can include deriving the additional prediction execution information using at least one of a template-based intra mode derivation (TIMD) method or a directional intra mode derivation (DIMD) method.

[0022] The information on additional prediction can include information on a first additional prediction and information on a second additional prediction, the additional prediction block includes a first additional prediction block and a second additional prediction block, and the generation of the final prediction block includes generating the final prediction block based on a weighted average of the regular prediction block, the first additional prediction block, and the second additional prediction block.

[0023] According to an embodiment of the disclosure, an encoding method includes the steps of encoding information on regular prediction for generating a regular prediction block by performing regular prediction, encoding information on additional prediction for generating an additional prediction block by performing additional prediction, and encoding information on weighted average for generating a final prediction block for a current block based on a weighted average of the regular prediction block and the additional prediction block, wherein the regular prediction includes any one of intra prediction, inter prediction, or intra block copy (IBC), and wherein the additional prediction includes at least one of intra prediction, inter prediction, or IBC.

[0024] According to embodiments of the present disclosure, a method of transmitting image information includes encoding information on a regular prediction for generating a regular prediction block by performing a regular prediction, encoding information on an additional prediction for generating an additional prediction block by performing an additional prediction, encoding information on a weighted average for generating a final prediction block for a current block based on a weighted average of the regular prediction block and the additional prediction block, and transmitting the image information including the encoded information on the regular prediction, the encoded information on the additional prediction, and the encoded information on the weighted average, wherein the regular prediction includes any one of intra prediction, inter prediction, or intra block copy (IBC), and wherein the additional prediction includes at least one of intra prediction, inter prediction, or IBC.

[0025] Advantageous Effects

[0026] According to the present disclosure, a method and apparatus for performing inter prediction based on a multi-reference mode can be provided.

[0027] According to the present disclosure, a signaling method and apparatus for determining a multi-reference mode can be provided.

[0028] According to the present disclosure, a derivation method and apparatus for determining a multi-reference mode can be provided

[0029] Effects the present disclosure can achieve are not limited to what has been described above. Other effects not described above can become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0030] Figure 1 FIG. 1 is a diagram schematically illustrating a video encoding system to which embodiments of the present disclosure are applicable.

[0031] Figure 2 FIG. 2 is a diagram schematically illustrating an image encoding apparatus to which embodiments of the present disclosure are applicable.

[0032] Figure 3 FIG. 3 is a diagram schematically illustrating an image decoding apparatus to which embodiments of the present disclosure are applicable.

[0033] Figure 4 An example of a video / image encoding method based on intra template matching prediction to which embodiments of the present disclosure can be applied is exemplified.

[0034] Figure 5 An example of a video / image encoding method based on inter prediction to which embodiments of the present disclosure can be applied is exemplified.

[0035] Figure 6 An example of a video / image encoding method based on inter prediction to which embodiments of the present disclosure can be applied is exemplified.

[0036] Figure 7 An inter prediction process to which embodiments of the present disclosure can be applied is exemplified.

[0037] Figure 8 is a diagram exemplifying a spatial candidate used in an inter prediction process to which embodiments of the present disclosure can be applied.

[0038] Figure 9 is a diagram for describing an affine motion prediction used in an inter prediction process to which embodiments of the present disclosure can be applied.

[0039] Figure 10 is a diagram for describing a template matching to which embodiments of the present disclosure can be applied.

[0040] Figure 11 An inter prediction method performed by a decoding device according to embodiments of the present disclosure is exemplified.

[0041] Figure 12 is a diagram exemplifying a reference block used in a multi-reference mode according to embodiments of the present disclosure.

[0042] Figure 13 is a diagram exemplifying a method of signaling information about an additional reference block used in a multi-reference mode according to embodiments of the present disclosure.

[0043] Figure 14 is a diagram exemplifying a combination between a base block mode and an additional reference block mode according to embodiments of the present disclosure.

[0044] Figure 15 is a diagram exemplifying a method of signaling information about an additional reference block used in a multi-reference mode according to embodiments of the present disclosure.

[0045] Figure 16 is a diagram exemplifying a method of deriving an additional reference block based on a template matching used in a multi-reference mode according to embodiments of the present disclosure.

[0046] Figure 17 is a diagram exemplifying a method of deriving an additional reference block based on an intra template matching prediction used in a multi-reference mode according to embodiments of the present disclosure.

[0047] Figure 18 is a diagram exemplifying a method of deriving an additional reference block based on a template-based intra mode derivation (TIMD) used in a multi-reference mode according to embodiments of the present disclosure.

[0048] Figure 19 is a diagram exemplifying a method of calculating an error (or cost) for deriving an additional reference block in a multi-reference mode according to embodiments of the present disclosure.

[0049] Figure 20 A decoding method according to an embodiment of the disclosure is exemplified.

[0050] Figure 21 An encoding method according to an embodiment of the disclosure is exemplified.

[0051] Figure 22 FIG. 1 is a diagram showing a content streaming system to which an embodiment of the disclosure is applicable. DETAILED DESCRIPTION

[0052] Because the disclosure can be changed variously and has several embodiments, a specific embodiment will be illustrated in the accompanying drawings and described in detail in the detailed description. However, it is not intended to limit the disclosure to a specific embodiment, and it should be understood to include all changes, equivalents, and alternatives included in the spirit and technical scope of the disclosure. While each drawing is described, like reference numerals are used for like components.

[0053] Terms such as first, second, and the like can be used to describe various components, but the components should not be limited by the terms. The terms are used only to distinguish one component from other components. For example, a first component can be referred to as a second component, and similarly, a second component can also be referred to as a first component without departing from the scope of the disclosure. The terms and / or combinations of any one or more of the related statements items are included.

[0054] When a component is referred to as being "connected" or "linked" to another component, it should be understood that it can be directly connected or linked to the other component, but another component can also be present in the middle. On the other hand, when a component is referred to as being "directly connected" or "directly linked" to another component, it should be understood that there is no other component in the middle.

[0055] The terms used in this application are only used to describe the specific embodiments and are not intended to limit the disclosure. Unless the context clearly indicates otherwise, a singular expression includes a plural expression. In this application, it should be understood that terms such as "include" or "have" are intended to designate the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, but do not exclude the possibility of existence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof in advance.

[0056] The present disclosure relates to video / image coding. For example, the methods / embodiments disclosed herein can be applied to methods disclosed in Versatile Video Coding (VVC) standard. In addition, the methods / embodiments disclosed herein can be applied to methods disclosed in Essential Video Coding (EVC) standard, AOMedia Video 1 (AV1) standard, second generation Audio Video Coding standard (AVS2), or next generation video / image coding standards (e.g., H.267 or H.268, etc.).

[0057] This specification presents various embodiments of video / image coding, and unless otherwise stated, the embodiments can be combined with each other to be performed.

[0058] Here, a video can refer to a set of a series of images over time. A picture generally refers to a unit representing one image within a specific time period, and a slice / tile is a unit forming a part of a picture in coding. A slice / tile can include at least one coding tree unit (CTU). One picture can consist of at least one slice / tile. A tile is a rectangular region consisting of a plurality of CTUs within a specific tile column and a specific tile row of one picture. A tile column is a rectangular region of CTUs having the same height as a picture and a width assigned by syntax requirements of a picture parameter set. A tile row is a rectangular region of CTUs having a height assigned by a picture parameter set and a width the same as a width of a picture. CTUs within one tile can be consecutively arranged according to a CTU raster scan, and tiles within one picture can be consecutively arranged according to a raster scan of tiles. One slice can include an integer number of complete tiles or an integer number of consecutive complete CTU rows within tiles of one picture that can be exclusively included in a single NAL unit. Meanwhile, one picture can be divided into at least two sub-pictures. A sub-picture can be a rectangular region of at least one slice within a picture.

[0059] A pixel, a pel, or a picture element can refer to the smallest unit constituting one picture (or image). In addition, a "sample" can be used as a term corresponding to a pixel. A sample can generally represent a pixel or a pixel value, and can represent only a pixel / pixel value of a luminance component or only a pixel / pixel value of a chrominance component.

[0060] A unit can represent a basic unit of image processing. A unit can include at least one of a specific region of a picture and information related to the corresponding region. One unit can include one luminance block and two chrominance (e.g., cb, cr) blocks. In some cases, a unit can be used interchangeably with terms such as a block or a region. In general, an MxN block can include a set (or an array) of transform coefficients or samples (or sample arrays) consisting of M columns and N rows.

[0061] Here, "A or B" can mean "only A", "only B", or "both A and B". In other words, here, "A or B" can be interpreted as "A and / or B". For example, here, "A, B, or C" can mean "only A", "only B", "only C", or "any combination of A, B, and C".

[0062] As used herein, a slash ( / ) or a comma can mean "and / or". For example, "A / B" can mean "A and / or B". Thus, "A / B" can mean "only A", "only B", or "both A and B". For example, "A, B, C" can mean "A, B, or C".

[0063] Here, "at least one of A and B" can mean "only A", "only B", or "both A and B". Also, expressions such as "at least one of A or B" or "at least one of A and / or B" can be interpreted in the same manner as "at least one of A and B" herein.

[0064] Also, here, "at least one of A, B, and C" can mean "only A", "only B", "only C", or "any combination of A, B, and C". Also, "at least one of A, B, or C" or "at least one of A, B, and / or C" can mean "at least one of A, B, and C".

[0065] Also, brackets used herein can mean "for example". Specifically, when indicated as "prediction (intra prediction)", "intra prediction" can be proposed as an example of "prediction". In other words, "prediction" here is not limited to "intra prediction", and "intra prediction" can be proposed as an example of "prediction". Also, even when indicated as "prediction (i.e., intra prediction)", "intra prediction" can be proposed as an example of "prediction".

[0066] Here, technical features described in the individual drawings can be implemented individually or simultaneously.

[0067] Figure 1 A video / image encoding system according to the disclosure is illustrated.

[0068] Reference Figure 1 The video / image encoding system can include a first device (a source device) and a second device (a sink device).

[0069] The source device can transmit the encoded video / image information or data in the form of a file or stream to the receiving device through a digital storage medium or a network. The source device can include a video source, an encoding device, and a transmission unit. The receiving device can include a reception unit, a decoding device, and a renderer. The encoding device can be referred to as a video / image encoding device, and the decoding device can be referred to as a video / image decoding device. The transmitter can be included in the encoding device. The receiver can be included in the decoding device. The renderer can include a display unit, and the display unit can be composed of a separate device or an external component.

[0070] The video source can acquire video / images through a process of capturing, synthesizing, or generating video / images. The video source can include a device that captures video / images and a device that generates video / images. The device that captures video / images can include at least one camera, a video / image archive including previously captured video / images, or the like. The device that generates video / images can include a computer, a tablet, a smartphone, or the like, and can (electronically) generate video / images. For example, a virtual video / image can be generated through a computer or the like, and in this case, the process of capturing video / images can be replaced by a process of generating related data.

[0071] The encoding device can encode the input video / images. The encoding device can perform a series of processes such as prediction, transformation, quantization, or the like for compression and encoding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.

[0072] The transmission unit can transmit the encoded video / image information or data output in the form of a bitstream to the reception unit of the receiving device in the form of a file or stream through a digital storage medium or a network. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, or the like. The transmission unit can include an element for generating a media file through a predetermined file format and can include an element for transmission through a broadcast / communication network. The reception unit can receive / extract the bitstream and transmit it to the decoding device.

[0073] The decoding device can decode the video / images by performing a series of processes such as dequantization, inverse transformation, prediction, or the like corresponding to the operations of the encoding device.

[0074] The renderer can render the decoded video / images. The rendered video / images can be displayed through a display unit.

[0075] Figure 2 A rough block diagram of an encoding device to which embodiments of the disclosure can be applied and which performs encoding of a video / image signal is illustrated.

[0076] ReferenceFigure 2 The encoding apparatus 200 can be composed of an image partitioner 210, a predictor 220, a residue processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 can include an inter-predictor 221 and an intra-predictor 222. The residue processor 230 can include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residue processor 230 can further include a subtractor 231. The adder 250 can be referred to as a reconstructor or a reconstructed block generator. According to an embodiment, the above-described image partitioner 210, predictor 220, residue processor 230, entropy encoder 240, adder 250, and filter 260 can be configured by at least one hardware component (e.g., an encoder chipset or a processor). In addition, the memory 270 can include a decoded picture buffer (DPB) and can be configured by a digital storage medium. The hardware component can further include the memory 270 as an internal / external component.

[0077] The image partitioner 210 can partition an input image (or picture, frame) input to the encoding apparatus 200 into at least one processing unit. As an example, the processing unit can be referred to as a coding unit (CU). In this case, the coding unit can be recursively partitioned from a coding tree unit (CTU) or a largest coding unit (LCU) according to a quad-tree binary-tree ternary (QTBTTT) structure.

[0078] For example, one coding unit can be partitioned into a plurality of coding units having a deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure can be applied first, and the binary-tree structure and / or the ternary structure can be applied later. Alternatively, the binary-tree structure can be applied before the quad-tree structure. The encoding process according to this specification can be performed based on a final coding unit that is no longer partitioned. In this case, based on the coding efficiency according to the characteristics of the image, etc., the largest coding unit can be directly used as the final coding unit, or if necessary, the coding unit can be recursively partitioned into coding units of a deeper depth, and the coding unit having the optimal size can be used as the final coding unit. Here, the encoding process can include processes such as prediction, transformation, and reconstruction, which will be described later.

[0079] As another example, the processing unit can further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit can be divided or partitioned from the above-described final coding unit, respectively. The prediction unit can be a unit of sample prediction, and the transform unit can be a unit for deriving a transform coefficient and / or a unit for deriving a residual signal from the transform coefficient.

[0080] In some cases, the term of unit can be used interchangeably with terms such as block or area. In general, an MxN block can represent a set of transform coefficients or samples consisting of M columns and N rows. The samples can represent pixels or pixel values in general, and can represent only pixels / pixel values of a luma component or only pixels / pixel values of a chroma component. The samples can be used as a term to correspond to pixels or samples for one picture (or image).

[0081] The encoding apparatus 200 can subtract a prediction signal (prediction block, prediction sample array) output from the inter-predictor 221 or the intra-predictor 222 from an input image signal (original block, original sample array) to generate a residual signal (residual signal, residual sample array), and the generated residual signal is transmitted to the transformer 232. In this case, the unit of subtracting the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) within the encoding apparatus 200 can be referred to as a subtractor 231.

[0082] The predictor 220 can perform prediction on a block to be processed (hereinafter, referred to as a current block), and generate a predicted block including predicted samples for the current block. The predictor 220 can determine whether to apply intra-prediction or inter-prediction in units of the current block or CU. The predictor 220 can generate various information on prediction, such as prediction mode information, etc., and transmit it to the entropy encoder 240, as described later in the description of each prediction mode. The information on prediction can be encoded in the entropy encoder 240 and output in the form of a bitstream.

[0083] The intra-predictor 222 can predict the current block by referring to samples within the current picture. Depending on the prediction mode, the referred samples can be positioned in the vicinity of the current block or can be positioned at a distance away from the current block. In intra-prediction, the prediction mode can include at least one non-directional mode and a plurality of directional modes. The non-directional mode can include at least one of a DC mode or a planar mode. Depending on the level of detail of the prediction direction, the directional mode can include 33 directional modes or 65 directional modes. However, this is only an example, and more or less directional modes can be used according to the configuration. The intra-predictor 222 can determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring block.

[0084] The inter predictor 221 can derive a prediction block for a current block based on a reference block (a reference sample array) specified by a motion vector on a reference picture. In this case, to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in a block, sub-block, or sample unit based on the correlation of the motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). For inter prediction, the neighboring blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block can be the same or different. The temporal neighboring block can be referred to as a collocated reference block, a collocated CU (colCU), etc., and the reference picture including the temporal neighboring block can be referred to as a collocated picture (colPic). For example, the inter predictor 221 can configure a motion information candidate list based on the neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or the reference picture index of the current block. The inter prediction can be performed based on various prediction modes, and for example, for a skip mode and a merge mode, the inter predictor 221 can use the motion information of the neighboring blocks as the motion information of the current block. For the skip mode, unlike the merge mode, a residual signal can not be transmitted. For a motion vector prediction (MVP) mode, the motion vectors of surrounding blocks are used as a motion vector predictor, and a motion vector difference is signaled to indicate the motion vector of the current block.

[0085] The predictor 220 can generate a prediction signal based on various prediction methods described later. For example, the predictor can not only apply intra prediction or inter prediction to predict one block, but also simultaneously apply intra prediction and inter prediction. It can be referred to as a combined inter and intra prediction (CIIP) mode. In addition, the predictor can be based on an intra block copy (IBC) prediction mode or can be based on a palette mode for prediction for a block. The IBC prediction mode or the palette mode can be used for content image / video encoding of games, etc., such as screen content coding (SCC), etc. The IBC basically performs prediction within a current picture, but it can be performed similarly to inter prediction in that it derives a reference block within the current picture. In other words, the IBC can use at least one of the inter prediction techniques described herein. The palette mode can be considered as an example of intra coding or intra prediction. When the palette mode is applied, sample values within a picture can be signaled based on information about a palette table and a palette index. The prediction signal generated by the predictor 220 can be used to generate a reconstructed signal or a residual signal.

[0086] The transformer 232 can generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique can include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loève transform (KLT), a graph-based transform (GBT), or a conditional non-linear transform (CNT). Here, the GBT refers to a transform obtained from a graph when relationship information between pixels is expressed as the graph. The CNT refers to a transform obtained based on a prediction signal generated by using all previously reconstructed pixels. In addition, the transform process can be applied to square pixel blocks of the same size or can be applied to non-square blocks of variable sizes.

[0087] The quantizer 233 can quantize the transform coefficients and transmit them to the entropy encoder 240, and the entropy encoder 240 can encode and output information about the quantized transform coefficients (about the quantized transform coefficients) as a bitstream. The information about the quantized transform coefficients can be referred to as residual information. The quantizer 233 can rearrange the quantized transform coefficients in a block form into a one-dimensional vector form based on a coefficient scan order, and can generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form.

[0088] The entropy encoder 240 can perform various encoding methods such as exponential Golomb, context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), or the like. The entropy encoder 240 can encode information necessary for video / video image reconstruction (for example, values of syntax elements, or the like) other than the transform coefficients quantized together or individually.

[0089] Encoded information (e.g., encoded video / image information) can be transmitted or stored in a form of a bitstream in units of network abstraction layer (NAL) units. The video / image information can further include information on various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information can further include general constraint information. Here, information and / or syntax elements transmitted / signaled from the encoding apparatus to the decoding apparatus can be included in the video / image information. The video / image information can be encoded through the above-described encoding process and included in the bitstream. The bitstream can be transmitted through a network or can be stored in a digital storage medium. Here, the network can include a broadcasting network and / or a communication network, etc., and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmission unit (not shown) for transmission and / or a storage unit (not shown) for storing a signal output from the entropy encoder 240 can be configured as an internal / external element of the encoding apparatus 200, or the transmission unit can also be included in the entropy encoder 240.

[0090] The quantized transform coefficients output from the quantizer 233 can be used to generate a prediction signal. For example, a dequantization and inverse transform can be applied to the quantized transform coefficients through the dequantizer 234 and the inverse transformer 235 to reconstruct a residual signal (a residual block or residual samples). The adder 250 can add the reconstructed residual signal to the prediction signal output from the inter-predictor 221 or the intra-predictor 222 to generate a reconstructed signal (a reconstructed picture, a reconstructed block, a reconstructed sample array). When there is no residual of a block to be processed, such as when a skip mode is applied, a prediction block can be used as a reconstructed block. The adder 250 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-prediction of a next block to be processed within the current picture, and can also be used for inter-prediction of a next picture through filtering described later. Meanwhile, luma mapping with chroma scaling (LMCS) can be applied in a picture encoding and / or reconstructing process.

[0091] The filter 260 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 260 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and can store the modified reconstructed picture in the memory 270, particularly in the DPB of the memory 270. The various filtering methods can include a deblocking filter, a sample adaptive offset, an adaptive loop filter, a bilateral filter, etc. The filter 260 can generate and transmit various information on filtering to the entropy encoder 240. The information on filtering can be encoded in the entropy encoder 240 and output in the form of a bitstream.

[0092] The modified reconstructed picture transmitted to the memory 270 can be used as a reference picture in the inter-predictor 221. When inter-prediction is applied thereto, the encoding apparatus can avoid prediction mismatch in the encoding apparatus 200 and the decoding apparatus, and can also improve coding efficiency.

[0093] The DPB of the memory 270 can store the modified reconstructed picture to be used as a reference picture in the inter-predictor 221. The memory 270 can store motion information of a block from which motion information in the current picture is derived (or encoded) and / or motion information of a block in the pre-reconstructed picture. The stored motion information can be transmitted to the inter-predictor 221 to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory 270 can store reconstructed samples of a reconstructed block in the current picture and transmit them to the intra-predictor 222.

[0094] Figure 3 A rough block diagram of a decoding apparatus to which embodiments of the present disclosure can be applied and which performs decoding of a video / image signal is illustrated.

[0095] Reference Figure 3 The decoding apparatus 300 can be configured by including an entropy decoder 310, a residue processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 can include an inter-predictor 331 and an intra-predictor 332. The residue processor 320 can include a dequantizer 321 and an inverse transformer 321.

[0096] According to embodiments, the above-described entropy decoder 310, residue processor 320, predictor 330, adder 340, and filter 350 can be configured by one hardware component (e.g., a decoder chipset or a processor). In addition, the memory 360 can include a decoded picture buffer (DPB) and can be configured by a digital storage medium. The hardware component can further include the memory 360 as an internal / external component.

[0097] When a bitstream including video / image information is input, the decoding apparatus 300 can reconstruct an image in response to a process of processing video / image information in the Figure 2 The decoding apparatus 300 can reconstruct an image in response to a process of processing video / image information in the encoding apparatus of FIG. 1. For example, the decoding apparatus 300 can derive a unit / block based on related information of block partitioning obtained from the bitstream. The decoding apparatus 300 can perform decoding by using a processing unit applied in the encoding apparatus. Accordingly, the decoded processing unit can be an encoding unit, and the encoding unit can be partitioned from a coding tree unit or a largest coding unit according to a quad tree structure, a binary tree structure, and / or a ternary tree structure. At least one transform unit can be derived from the encoding unit. Also, the reconstructed image signal decoded and output by the decoding apparatus 300 can be played back by a playback device.

[0098] Decoding device 300 can receive data in bitstream form from... Figure 2 The signal output by the encoding device and the received signal can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can parse the bitstream to derive information (e.g., video / image information) necessary for image reconstruction (or picture reconstruction). The video / image information may further include information about various parameter sets such as adaptive parameter sets (APS), picture parameter sets (PPS), sequence parameter sets (SPS), or video parameter sets (VPS). In addition, the video / image information may further include general constraint information. The decoding device can further decode the picture based on the information about the parameter sets and / or the general constraint information. The signaled / received information and / or the syntax elements described later herein can be decoded and obtained from the bitstream through the decoding process. For example, the entropy decoder 310 can decode the information in the bitstream based on encoding methods such as exponential Golomb coding, CAVLC, CABAC, etc., and output the values ​​of the syntax elements necessary for image reconstruction and the quantized values ​​of the transform coefficients of the residuals. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element from the bitstream, determine a context model using information about the syntax element to be decoded, decoding information of surrounding blocks and the block to be decoded, or information about symbols / bins decoded in the previous step, perform arithmetic decoding on the bins by predicting the occurrence probability of the bins based on the determined context model, and generate symbols corresponding to the value of each syntax element. In this case, after determining the context model, the CABAC entropy decoding method can update the context model by using information about the decoded symbols / bins for the context model used for the next symbol / bin. Among the information decoded in the entropy decoder 310, information about prediction is provided to the predictors (inter-frame predictor 332 and intra-frame predictor 331), and the residual values ​​of entropy decoding performed on them in the entropy decoder 310, i.e., the quantized transform coefficients and related parameter information, can be input to the residual processor 320. The residual processor 320 can derive the residual signal (residual block, residual sample, residual sample array). In addition, information about filtering among the information decoded in the entropy decoder 310 can be provided to the filter 350. Meanwhile, the receiving unit (not shown) that receives the signal output from the encoding device can be further configured as an internal / external element of the decoding device 300 or the receiving unit can be a component of the entropy decoder 310.

[0099] Meanwhile, a decoding apparatus according to this specification can be referred to as a video / image / picture decoding apparatus, and the decoding apparatus can be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder can include the entropy decoder 310, and the sample decoder can include at least one of the dequantizer 321, the inverse transformer 322, the adder 340, the filter 350, the memory 360, the inter predictor 332, and the intra predictor 331.

[0100] The dequantizer 321 can dequantize the quantized transform coefficients and output the transform coefficients. The dequantizer 321 can rearrange the quantized transform coefficients into a two-dimensional block form. In this case, the rearrangement can be performed based on a coefficient scanning order performed in the encoding apparatus. The dequantizer 321 can perform dequantization on the quantized transform coefficients by using a quantization parameter (e.g., quantization step length information) and obtain the transform coefficients.

[0101] The inverse transformer 322 inversely transforms the transform coefficients to obtain a residual signal (a residual block, a residual sample array).

[0102] The predictor 320 can perform prediction on the current block and generate a prediction block including prediction samples for the current block. The predictor 320 can determine whether to apply intra prediction or inter prediction to the current block based on information about prediction output from the entropy decoder 310, and determine a specific intra / inter prediction mode.

[0103] The predictor 320 can generate a prediction signal based on various prediction methods described later. For example, the predictor 320 can not only apply intra prediction or inter prediction to predict one block, but also simultaneously apply intra prediction and inter prediction. It can be referred to as a combined inter and intra prediction (CIIP) mode. In addition, the predictor can be based on an intra block copy (IBC) prediction mode or can be based on a palette mode for prediction of a block. The IBC prediction mode or the palette mode can be used for content image / video encoding of games, etc., such as screen content coding (SCC), etc. The IBC basically performs prediction within a current picture, but it can be performed similarly to inter prediction in that it derives a reference block within the current picture. In other words, the IBC can use at least one of the inter prediction techniques described herein. The palette mode can be considered as an example of intra coding or intra prediction. When the palette mode is applied, information about a palette table and a palette index can be included in the video / image information and signaled.

[0104] The intra predictor 331 can predict the current block by referring to samples within the current picture. Depending on the prediction mode, the referred samples can be located in the vicinity of the current block or can be located at a distance away from the current block. In intra prediction, the prediction mode can include at least one non-directional mode and a plurality of directional modes. The intra predictor 331 can determine the prediction mode applied to the current block by using the prediction mode applied to a neighboring block.

[0105] The inter predictor 332 can derive a prediction block for the current block based on a reference block (a reference sample array) specified by a motion vector on a reference picture. In this case, to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of a block, a sub-block, or a sample based on the correlation of the motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). For inter prediction, the neighboring blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. The inter predictor 332 can configure a motion information candidate list based on the neighboring blocks, and derive a motion vector and / or a reference picture index of the current block based on received candidate selection information, for example. The inter prediction can be performed based on various prediction modes, and information about the prediction can include information indicating an inter prediction mode for the current block.

[0106] The adder 340 can add the obtained residual signal to a prediction signal (a prediction block, a prediction sample array) output from the predictor (including the inter predictor 332 and / or the intra predictor 331) to generate a reconstructed signal (a reconstructed picture, a reconstructed block, a reconstructed sample array). When there is no residual of a block to be processed, as when a skip mode is applied, the prediction block can be used as the reconstructed block.

[0107] The adder 340 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra prediction of a next block to be processed in the current picture, can be output through filtering described later, or can be used for inter prediction of a next picture. Meanwhile, luma mapping with chroma scaling (LMCS) can be applied in the picture decoding process.

[0108] The filter 350 can improve subjective / objective picture quality by applying filtering to the reconstructed signal. For example, the filter 350 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and transmit the modified reconstructed picture to the memory 360, specifically, the DPB of the memory 360. The various filtering methods can include deblocking filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, etc.

[0109] The (modified) reconstructed pictures stored in the DPB of the memory 360 can be used as reference pictures in the inter prediction unit 332. The memory 360 can store motion information of a block from which the motion information in the current picture is derived (or decoded) and / or the motion information of a block in the pre-reconstructed picture. The stored motion information can be sent to the inter predictor 260 to be used as the motion information of a spatial neighboring block or the motion information of a temporal neighboring block. The memory 360 can store the reconstructed samples of a reconstructed block in the current picture and send them to the intra predictor 331.

[0110] Here, the embodiments described in the filter 260, the inter predictor 221 and the intra predictor 222 of the encoding apparatus 200 can also be applied to the filter 350, the inter predictor 332 and the intra predictor 331 of the decoding apparatus 300, respectively, equally or correspondingly.

[0111] Intra prediction

[0112] The prediction unit of the encoding / decoding apparatus can derive reference samples from neighboring reference samples of the current block according to an intra prediction mode of the current block, and generate prediction samples of the current block based on the reference samples.

[0113] For example, the prediction unit of the encoding / decoding apparatus can derive (i) a prediction sample based on an average or an interpolation of neighboring reference samples of the current block, and (ii) a prediction sample based on a reference sample existing in the neighboring reference samples of the current block in a certain (prediction) direction relative to the prediction sample. The case of (i) can be referred to as a non-directional mode or a non-angular mode, and the case of (ii) can be referred to as a directional mode or an angular mode.

[0114] In addition, the prediction unit of the encoding / decoding apparatus can generate a prediction sample by interpolating a first neighboring sample with a second neighboring sample located in a direction opposite to a prediction direction of the intra prediction mode of the current block, based on the prediction sample of the current block in the neighboring reference samples. The above case can be referred to as linear interpolation intra prediction (LIP).

[0115] In addition, the prediction unit of the encoding / decoding apparatus can derive a temporary prediction sample for the current block based on filtered neighboring reference samples, and derive a prediction sample for the current block by weighting the temporary prediction sample with at least one reference sample derived according to the intra prediction mode among existing neighboring reference samples (that is, unfiltered neighboring reference samples). The above case can be referred to as position-dependent intra prediction (PDPC).

[0116] In addition, the prediction unit of the encoding / decoding device can select a reference sample line having the highest prediction accuracy from among multiple neighboring reference sample lines of the current block, and derive the prediction sample using a reference sample in the corresponding line that is located in the prediction direction. In this case, the prediction unit of the encoding / decoding device can perform intra prediction encoding by indicating (signaling) the used reference sample line to the decoding device. The above case can be referred to as multi-reference line (MRL) intra prediction or MRL-based intra prediction.

[0117] In addition, the prediction unit of the encoding / decoding device can divide the current block into vertical or horizontal sub-partitions to perform intra prediction based on the same intra prediction mode, but derive and use neighboring reference samples in units of the sub-partitions. That is, in this case, the intra prediction mode for the current block can be equally applied to the sub-partitions, but in some cases, the intra prediction performance can be improved by deriving and using peripheral reference samples in units of the sub-partitions. This prediction method can be referred to as intra sub-partition (ISP) or ISP-based intra prediction.

[0118] In addition, when the prediction direction is with respect to a prediction sample point between neighboring reference samples, that is, when the prediction direction points to a fractional sample position, the value of the prediction sample can be derived by interpolation of multiple reference samples located around the corresponding prediction direction (around the corresponding fractional sample position).

[0119] The above-described intra prediction methods can be referred to as intra prediction types to be distinguished from intra prediction modes. The intra prediction types can be referred to by various terms such as intra prediction techniques or additional intra prediction modes. For example, the intra prediction types (or additional intra prediction modes, etc.) can include at least one of the above-described LIP, PDPC, MRL, and ISP.

[0120] Information on the intra prediction types can be encoded by the encoding device, included in the bitstream, and signaled to the decoding device. The intra prediction type information can be implemented in various forms such as flag information indicating whether each intra prediction type is applicable or index information indicating one of multiple intra prediction types.

[0121] A most probable mode (MPM) list for deriving the above-described intra prediction modes can be differently configured according to the intra prediction type. Alternatively, the MPM list can be configured universally regardless of the intra prediction type.

[0122] Further, in decoder-side intra mode derivation (DIMD), the intra prediction is derived as a weighted average between a planar predictor and two derived directional predictors. Two angular modes are selected from a gradient histogram (HoG) computed from neighboring pixels in the current block. Once the two modes are selected, the corresponding predictors and the planar predictor are derived, respectively. Thereafter, a weighted average between the blocks is used as the final predictor. The corresponding amplitudes of the gradient histograms (HoG) for each mode are used to determine the weights.

[0123] The derived intra mode is included in the base list of most probable intra modes (MPM), so the DIMD process is performed before constructing the MPM list. The base derived intra mode of the DIMD block is stored with the block and used to construct the MPM list for the neighboring blocks.

[0124] Further, for each intra prediction mode of the MPM, the sum of absolute transformed differences (SATD) between the prediction samples of the template and the reconstructed samples is computed. The first two intra prediction modes with the smallest SATD are selected as the TIMD modes. The two TIMD modes are weighted and the weighted intra prediction is used to encode the current CU. Position dependent intra prediction combination (PDPC) is included in the derivation of the TIMD modes.

[0125] The cost of the two selected modes is compared with a threshold and the two cost factors satisfying [Equation 1] are applied in the test: [Equation 1] costMode2<2 costMode1.

[0126] If the condition in [Equation 1] is true, fusion is applied; otherwise, only mode 1 is used.

[0127] The mode weights are calculated according to the SATD cost as shown in [Equation 2]: [Equation 2] weight1 = costMode2 / (costMode1 + costMode2) weight2 = 1 - weight1 Figure 4 An example of a video / image encoding method based on intra template matching prediction to which embodiments of the present disclosure can be applied is exemplified.

[0128] Intra template matching prediction (IntraTMP) is a special intra prediction mode that copies the best prediction block, where the current template matches an L-shaped template from the reconstructed part of the current frame. The encoder searches for the most similar template to the current template in the reconstructed part of the current frame within a predetermined search range and uses that block as the prediction block. Next, the encoder signals the use of the mode and the same prediction operation is performed at the decoder side.

[0129] A prediction signal is generated by matching the L-shaped causal neighbors of the current block with other blocks in a predefined search region of Figure 4 R1: the current CTU R2: the top-left CTU R3: the top CTU R4: the left CTU The sum of absolute difference (SAD) is used as the cost function.

[0130] Within each region, the decoder searches for the template with the smallest SAD compared to the current template and uses that block as the prediction block.

[0131] The size of each region (SearchRange_w, SearchRange_h) is set proportionally to the block size (BlkW, BlkH) to ensure a fixed number of SAD comparisons per pixel. That is, this is the same as Equation 3.

[0132] [Equation 3]

[0133] SearchRange_w = a BlkW

[0134] SearchRange_h = a BlkH

[0135] Here, "a" is a constant that controls the gain / complexity trade-off. In practice, "a" can be equal to 5.

[0136] By downsampling the search positions in all search regions within the template by 2, the number of template matching calculations can be reduced to ¼. After finding the best match, a refinement process is performed. The refinement is performed by a second template matching search centered at the best match with a reduced range. The reduced range is defined as min(BlkW, BlkH) / 2.

[0137] The intra template matching tool is enabled for CUs with width and height of 64 or smaller. The maximum coding unit size for intra template matching is configurable.

[0138] When DIMD is not used in the current coding unit, the intra template matching prediction mode is signaled at the coding unit level via a dedicated flag.

[0139] ​Block vectors (BVs) derived from Intra Template Matching Prediction (IntraTMP) can be used for Intra Block Copy (IBC). Stored IntraTMP block vectors (BVs) of neighboring blocks are used as spatial block vector (BV) candidates when constructing an IBC candidate list together with IBC block vectors (BVs).

[0140] IntraTMP block vectors are stored in the IBC block vector buffer. A current IBC block can use both the IBC block vector (BV) and the IntraTMP block vector (BV) of a neighboring block as block vector (BV) candidates for the IBC block vector (BV) candidate list.

[0141] Intra Block Copy (IBC)

[0142] It is known that Intra Block Copy (IBC) significantly improves the coding efficiency of screen content. Since IBC mode is implemented as a block-level coding mode, the encoder performs block matching (BM) to find the best block vector (or motion vector) for each CU. Here, the block vector is used to indicate the displacement from the current block to a reference block that has already been reconstructed within the current picture. The luma block vector of an IBC-coded coding unit has integer precision. The chroma block vector is also rounded to integer precision. When combined with AMVR, the IBC mode can switch between 1-pixel motion vector precision and 4-pixel motion vector precision. An IBC-coded coding unit is considered as a third prediction mode different from the intra prediction mode or the inter prediction mode. The IBC mode can be applied to a coding unit whose width and height are both less than 64 luma samples.

[0143] At the encoder side, hash-based motion estimation is performed for IBC. The encoder performs an RD check for blocks whose width or height is not larger than 16 luma samples. When not in merge mode, the block vector search is first performed using hash-based search. When the hash search does not return a valid candidate, a block matching based local search is performed.

[0144] In the hash-based search, the hash key match (32-bit CRC) between the current block and the reference block is extended to all allowed block sizes. The hash key calculation for each position in the current picture is based on 4x4 sub-blocks. For larger current blocks, the hash key is determined to match the hash key of the reference block when the hash keys of all 4x4 sub-blocks match the hash keys at the corresponding reference positions. When multiple reference blocks are found whose hash keys match the hash key of the current block, the block vector cost of each matching reference is calculated and the block with the smallest cost is selected.

[0145] In the block matching search, the search range is set to include both the previous CTU and the current CTU.

[0146] IBC mode is signaled with a flag at the coding unit level, and it can be signaled as IBC AMVP mode or IBC skip / merge mode as follows: - IBC skip / merge mode: merge candidate index is used to indicate the block vector among the list of neighboring candidate IBC coded blocks for predicting the current block. The merge list consists of spatial, HMVP and paired candidates.

[0147] - IBC AMVP mode: block vector difference is coded in the same way as motion vector difference. Block vector prediction method uses two candidates as predictors, one from the left neighboring block and the other from the top neighboring block (at IBC coding). When either of the two neighboring blocks is not available, the base block vector is used as predictor. A flag indicating the block vector predictor index is signaled.

[0148] Inter prediction

[0149] Further, when inter prediction is applied, a prediction unit of an encoding / decoding apparatus can perform inter prediction on a block-by-block basis to derive prediction samples. Inter prediction can mean prediction derived in a manner according to data elements (e.g., sample values or motion information) of a picture other than a current picture. When inter prediction is applied to a current block, a prediction block (prediction sample array) for the current block can be derived based on a reference block (reference sample array) specified by a motion vector on a reference picture indicated by a reference picture index.

[0150] In this case, in order to reduce the amount of motion information transmitted in the inter prediction mode, motion information of a current block can be predicted in units of a block, a sub-block, or a sample based on a correlation between motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can also include inter prediction type (L0 prediction, L1 prediction, Bi prediction, etc.) information. When inter prediction is applied, the neighboring blocks can include spatial neighboring blocks within a current picture and temporal neighboring blocks in a reference picture.

[0151] The reference picture including the reference block and the reference picture including the temporal neighboring block can be the same or different. The temporal neighboring block can be referred to as a collocated reference block, a colCU, etc., and the reference picture including the temporal neighboring block can be referred to as a collocated picture (colPic). For example, a motion information candidate list can be constructed based on blocks neighboring the current block, and a flag or index information indicating which candidate to select (use) to derive a motion vector and / or a reference picture index of the current block can be signaled.

[0152] Inter prediction can be performed based on various prediction modes. For example, in the case of a skip mode and a merge mode, motion information of a current block can be the same as that of a selected neighboring block. In the case of the skip mode, unlike the merge mode, a residual signal can not be transmitted. In the case of a motion vector prediction (MVP) mode, a motion vector of a selected neighboring block can be used as a motion vector predictor, and a motion vector difference can be signaled. In this case, a motion vector of the current block can be derived using a sum of the motion vector predictor and the motion vector difference.

[0153] Motion information can include L0 motion information and / or L1 motion information according to an inter prediction type (e.g., L0 prediction, L1 prediction, bi prediction). A motion vector in an L0 direction can be referred to as an L0 motion vector or MVL0, and a motion vector in an L1 direction can be referred to as an L1 motion vector or MVL1. Prediction based on an L0 motion vector can be referred to as L0 prediction, prediction based on an L1 motion vector can be referred to as L1 prediction, and prediction based on both an L0 motion vector and an L1 motion vector can be referred to as bi prediction. Here, the L0 motion vector can denote a motion vector associated with a reference picture list L0 (L0), and the L1 motion vector can denote a motion vector associated with a reference picture list L1 (L1). The reference picture list L0 can include pictures preceding a current picture in an output order as reference pictures, and the reference picture list L1 can include pictures succeeding the current picture in the output order. A preceding picture can be referred to as a forward (reference) picture, and a succeeding picture can be referred to as a backward (reference) picture.

[0154] The reference picture list L0 can further include pictures succeeding the current picture in a display order as reference pictures. In this case, within the reference picture list L0, preceding pictures can be indexed first, and thereafter succeeding pictures can be indexed. The reference picture list L1 can further include pictures preceding the current picture in the display order as reference pictures. In this case, within the reference picture list L1, succeeding pictures can be indexed first, and thereafter preceding pictures can be indexed. Here, the output order can correspond to a picture order count (POC) order.

[0155] A video / image encoding process based on inter prediction can include, for example, the following.

[0156] Figure 5 Examples of a video / image encoding method based on inter prediction to which embodiments of the disclosure can be applied are illustrated.

[0157] The encoding device can perform inter prediction for the current block to generate prediction samples (S400). The encoding device can derive an inter prediction mode and motion information for the current block, and generate prediction samples for the current block. Here, the inter prediction mode determination process, the motion information derivation process, and the prediction sample generation process can be performed simultaneously, or one process can precede another process. For example, the inter prediction unit of the encoding device can include a prediction mode determination unit, a motion information derivation unit, and a prediction sample derivation unit. The prediction mode determination unit can determine a prediction mode for the current block, the motion information derivation unit can derive motion information for the current block, and the prediction sample derivation unit can derive prediction samples for the current block.

[0158] For example, the inter prediction unit of the encoding device can search for a block similar to the current block within a certain region (search region) of a reference picture through motion estimation, and derive a reference block having a minimum difference or less than or equal to a certain reference from the current block. Based on this, a reference picture index indicating a reference picture in which the reference block is located can be derived, and a motion vector can be derived based on a position difference between the reference block and the current block. The encoding device can determine a mode suitable for the current block among various prediction modes. The encoding device can compare RD costs for the various prediction modes and determine a best prediction mode for the current block.

[0159] For example, when the skip mode or the merge mode is applied to the current block, the encoding device can construct a merge candidate list as will be described below, and derive a reference block having a minimum difference or less than a certain threshold from the current block among reference blocks indicated by merge candidates included in the merge candidate list. In this case, a merge candidate associated with the derived reference block is selected, and merge index information indicating the selected merge candidate can be generated and signaled to the decoding device. The motion information of the selected merge candidate can be used to derive the motion information of the current block.

[0160] As another example, when (A)MVP mode is applied to the current block, the encoding apparatus can construct a list of (A)MVP candidates, which will be described later, and use a motion vector of an MVP candidate selected from among the (A)MVP candidates included in the (A)MVP candidate list as a motion vector predictor (MVP) candidate of the current block. In this case, for example, a motion vector indicating a reference block derived through the above-described motion estimation can be used as a motion vector of the current block, and among the MVP candidates, an MVP candidate having a motion vector that is the smallest in difference from the motion vector of the current block can be the selected MVP candidate. A motion vector difference (MVD) that is a difference obtained by subtracting the MVP from the motion vector of the current block can be derived. In this case, information about the MVD can be signaled to the decoding apparatus. Also, when the (A)MVP mode is applied, a reference picture index value can be configured as reference picture index information and signaled to the decoding apparatus separately.

[0161] The encoding apparatus can derive residual samples based on the prediction samples (S410). The encoding apparatus can derive the residual samples by comparing the prediction samples with original samples of the current block.

[0162] The encoding apparatus can encode image information including prediction information and residual information (S420). The encoding apparatus can output the encoded image information in the form of a bitstream. The prediction information is information related to a prediction process, and can include prediction mode information (e.g., a skip flag, a merge flag, or a merge index) and / or motion information. The motion information can include candidate selection information (e.g., a merge index, an MVP flag, or an MVP index) for deriving a motion vector. Also, the motion information can include information about the above-described MVD and / or reference picture index information. Also, the motion information can include information indicating whether L0 prediction, L1 prediction, or bi prediction is applied. The residual information is information about residual samples. The residual information can include information about quantized transform coefficients for the residual samples.

[0163] The output bitstream can be stored in a (digital) storage medium and transmitted to the decoding apparatus, or can be transmitted to the decoding apparatus via a network.

[0164] Also, as described above, the encoding apparatus can generate a reconstructed picture (including reconstructed samples and reconstructed blocks) based on the reference samples and the residual samples. This is to derive the same prediction result as that performed by the decoding apparatus, thereby improving encoding efficiency. Accordingly, the encoding apparatus can store the reconstructed picture (or the reconstructed samples, the reconstructed blocks) in a memory and use the reconstructed picture as a reference picture for inter prediction. As described above, an in-loop filtering process or the like can be further applied to the reconstructed picture.

[0165] A video / image decoding process based on inter prediction can include, for example, the following.

[0166] Figure 6 Examples of a video / image decoding method based on inter prediction to which embodiments of the present disclosure can be applied are illustrated.

[0167] Referring to Figure 6 The decoding device can perform operations corresponding to those performed by the encoding device. The decoding device can perform prediction for a current block and derive prediction samples based on received prediction information.

[0168] In particular, the decoding device can determine a prediction mode for the current block based on the received prediction information (S500). The decoding device can determine which inter prediction mode is applied to the current block based on the prediction mode information within the prediction information.

[0169] For example, the decoding device can determine whether to apply the merge mode or the (A)MVP mode to the current block based on the merge flag. Alternatively, one of various inter prediction mode candidates can be selected based on the mode index. The inter prediction mode candidates can include the skip mode, the merge mode, and / or the (A)MVP mode, or can include various inter prediction modes described below.

[0170] The decoding device can derive the motion information of the current block based on the determined inter prediction mode (S510). For example, when the skip mode or the merge mode is applied to the current block, the decoding device can construct a merge candidate list described below and select one merge candidate from among the merge candidates included in the merge candidate list. The selection can be performed based on the above-described selection information (merge index). The motion information of the selected merge candidate can be used to derive the motion information of the current block. The motion information of the selected merge candidate can be used as the motion information of the current block.

[0171] As another example, when the (A)MVP mode is applied to the current block, the decoding device can construct an (A)MVP candidate list described below and use the motion vector of the MVP candidate selected from among the MVP candidates included in the (A)MVP candidate list as the MVP of the current block. The above-described selection can be performed based on the above-described selection information (MVP flag or MVP index). In this case, the MVD of the current block can be derived based on the MVP information, and the motion vector of the current block can be derived based on the MVP and the MVD of the current block. Further, the reference picture index of the current block can be derived based on the reference picture index information. The picture indicated by the reference picture index within the reference picture list for the current block can be derived as a reference picture referred to for inter prediction of the current block.

[0172] Furthermore, as will be described below, the motion information of the current block can be derived without constructing a candidate list. In this case, the motion information of the current block can be derived according to the process described in the prediction modes described below. In this case, the candidate list construction described above can be omitted.

[0173] The decoding device can generate the prediction samples for the current block based on the motion information of the current block (S520). In this case, the reference picture is derived based on the reference picture index of the current block, and the prediction samples for the current block can be derived using the samples of the reference block pointed to on the reference picture by the motion vector of the current block. In this case, as will be described below, depending on the case, a prediction sample filtering process can be further performed for all or some of the prediction samples of the current block.

[0174] For example, the inter prediction unit of the decoding device can include a prediction mode determination unit, a motion information derivation unit, and a prediction sample derivation unit. The prediction mode for the current block can be determined based on the prediction mode information received from the prediction mode determination unit, the motion information of the current block such as the motion vector and / or the reference picture index can be derived based on the information about the motion information received from the motion information derivation unit, and the prediction samples for the current block can be derived from the prediction sample derivation unit.

[0175] The decoding device generates the residual samples for the current block based on the received residual information (S530). The decoding device can generate the reconstructed samples for the current block based on the prediction samples and the residual samples, and generate the reconstructed picture based on the reconstructed samples for the current block (S540). As described above, in-loop filtering processes and the like can be further applied to the reconstructed picture.

[0176] Figure 7 An exemplary inter prediction process to which embodiments of the present disclosure can be applied is exemplarily illustrated. Figure 8 is a diagram exemplarily illustrating spatial candidates used in an inter prediction process to which embodiments of the present disclosure can be applied. Figure 9 is a diagram exemplarily illustrating affine motion prediction used in an inter prediction process to which embodiments of the present disclosure can be applied.

[0177] Reference Figure 7 As described above, the inter prediction process can include an inter prediction mode determination operation, a motion information derivation operation based on the determined prediction mode, and a prediction performance (prediction sample generation) operation based on the derived motion information. As described above, the inter prediction process can be performed by an encoding device and a decoding device. In this document, the term “coding device” can include an encoding device and / or a decoding device.

[0178] Referring to Figure 7 The encoding apparatus determines an inter prediction mode for the current block (S600). Various inter prediction modes can be used to predict the current block within a picture. For example, various modes such as merge mode, skip mode, motion vector prediction (MVP) mode, affine mode, subblock merge mode, and merge mode with motion vector difference (MMVD) mode can be used. Decoder-side motion vector refinement (DMVR) mode, adaptive motion vector resolution (AMVR) mode, bi-prediction with CU-level weights (BCW), bi-directional optical flow (BDOF), etc. can be used as additional modes, or instead of these modes. Also, according to an embodiment of the disclosure, the above-described inter prediction modes can include a multiple hypothesis prediction (MHP) mode. The multiple hypothesis prediction mode is a method of performing prediction by weighting prediction blocks generated based on additional motion information for a bi-predicted (or bi-predicted) block. The multiple hypothesis prediction mode will be described in detail below.

[0179] In the disclosure, the affine mode can also be referred to as an affine motion prediction mode. Also, the MVP mode can also be referred to as an advanced motion vector prediction (AMVP) mode. In the disclosure, motion information candidates derived from some modes and / or some modes can be included as motion information related candidates for other modes. For example, an HMVP candidate can be added as a merge candidate for a merge / skip mode, or as an MVP candidate for an AMVP mode. When the HMVP candidate is used as a motion information candidate for a merge mode or a skip mode, the HMVP candidate can be referred to as an HMVP merge candidate.

[0180] Prediction mode information indicating an inter prediction mode for the current block can be signaled from the encoding apparatus to the decoding apparatus. The prediction mode information can be included in a bitstream and received by the decoding apparatus. The prediction mode information can include index information indicating one of multiple candidate modes. Alternatively, the inter prediction mode can be indicated by hierarchical signaling of flag information.

[0181] In this case, the prediction mode information can include one or more flags. For example, a skip flag can be signaled to indicate whether a skip mode is applied. When the skip mode is not applied, a merge flag can be signaled to indicate whether a merge mode is applied. When the merge mode is not applied, an MVP mode can be applied. Alternatively, an additional flag can be signaled for additional differentiation. The affine mode can be signaled as an independent mode or as a mode according to the merge mode or the MVP mode. For example, the affine mode can include an affine merge mode and an affine MVP mode.

[0182] The encoding apparatus can derive the motion information for the current block (S610). The motion information can be derived based on the inter prediction mode. The encoding apparatus can perform inter prediction using the motion information of the current block. The encoding apparatus can derive the best motion information for the current block through a motion estimation process.

[0183] For example, the encoding apparatus can use the original block within the original picture for the current block to search for a similar reference block having high correlation in a predetermined search range within a reference picture in fractional pixel units, thereby deriving the motion information. The block similarity can be derived according to the difference in the phase-based sample value. For example, the block similarity can be calculated based on the SAD between the current block (or a template of the current block) and the reference block (or a template of the reference block). In this case, the motion information can be derived based on the reference block having the smallest SAD within the search range. The derived motion information can be signaled to the decoding apparatus using various methods based on the inter prediction mode.

[0184] The encoding apparatus can perform inter prediction based on the motion information for the current block (S620). The encoding apparatus can generate predicted samples for the current block based on the motion information. The current block including the predicted samples can be referred to as a predicted block.

[0185] When the merge mode is applied, the motion information of the current predicted block is not directly transmitted, but the motion information of the current predicted block is derived using the motion information of the neighboring predicted block. Accordingly, the motion information of the current predicted block can be indicated by transmitting flag information indicating that the merge mode is used and a merge index indicating which neighboring predicted block is used. The merge mode can be referred to as a regular merge mode.

[0186] To perform the merge mode, the encoder searches for a merge candidate block for deriving the motion information of the current predicted block. For example, up to five merge candidate blocks can be used, but the number of merge candidate blocks is not limited thereto. The maximum number of merge candidate blocks can be transmitted in a slice header or a tile group header, but is not limited thereto. After finding the merge candidate blocks, the encoder can generate a merge candidate list and select a merge candidate block having the lowest cost as a final merge candidate block.

[0187] For example, the above merge candidate list can utilize five merge candidate blocks. For example, four spatial merge candidates and one temporal merge candidate can be utilized. Specifically, in the case of the spatial merge candidate, Figure 8 The blocks illustrated in FIG. 6A can be used as the spatial merge candidate. Hereinafter, the spatial merge candidate or the spatial motion vector prediction (MVP) candidate described below can be referred to as a spatial motion vector prediction (SMVP), and the temporal merge candidate or the temporal motion vector prediction (MVP) candidate described below can be referred to as a temporal motion vector prediction (TMVP).

[0188] Motion vector prediction (MVP) mode can be referred to as advanced motion vector prediction (AMVP) mode. When the MVP mode is applied, a MVP candidate list can be generated using motion vectors of reconstructed spatial neighboring blocks (e.g., of a neighboring block of Figure 8 a reconstructed spatial neighboring block and / or a motion vector corresponding to a temporal neighboring block (or a Col block). In other words, the motion vector of the reconstructed spatial neighboring block and / or the motion vector corresponding to the temporal neighboring block can be used as a MVP candidate. When the pair-wise prediction is applied, a MVP candidate list for deriving L0 motion information and a MVP candidate list for deriving L1 motion information can be generated and used, respectively.

[0189] The above-described prediction information (or prediction-related information) can include selection information (e.g., a MVP flag or a MVP index) indicating a best MVP candidate selected from among the MVP candidates included in the list. The prediction unit can select the MVP for the current block from among the MVP candidates included in the motion vector candidate list using the selection information.

[0190] The prediction unit of the encoding apparatus can obtain a MVD between the motion vector of the current block and the MVP, encode the MVD, and output the encoded MVD in the form of a bitstream. Specifically, the MVD can be obtained by subtracting the MVP from the motion vector of the current block.

[0191] In this case, the prediction unit of the decoding apparatus can obtain the MVD included in the prediction-related information and derive the motion vector of the current block by adding the MVD to the MVP. The prediction unit of the decoding apparatus can obtain or derive a reference picture index indicating a reference picture, etc., from the prediction-related information.

[0192] Existing video coding systems use only a single motion vector (i.e., a translational motion model) to represent the motion of a coding block. However, although the translational motion model can represent the best motion at the block level, it does not necessarily correspond to the best motion at each pixel. Determining the best motion vector at the pixel level can improve coding efficiency.

[0193] To this end, in the affine mode, affine motion prediction using an affine motion model is applied. The affine motion prediction can use two, three, or four motion vectors to represent the motion vector at each pixel level of a block.

[0194] As shown in Figure 9 , the affine motion prediction can determine a motion vector at a pixel position within a block using two or more control point motion vectors (CPMVs). This set of motion vectors is referred to as an affine motion vector field (MVF).

[0195] A combination of inter prediction and intra prediction can be applied to the current block. An additional flag (e.g., ciip_flag) can be signaled to indicate whether a combined inter / intra prediction (CIIP) mode is applied to the current coding unit. For example, the additional flag is signaled to indicate whether the CIIP mode is applied to the current coding unit when the coding unit is coded in merge mode and the coding unit includes 64 or more luma samples (i.e., the coding unit width multiplied by the coding unit height is greater than or equal to 64), and both the coding unit width and the coding unit height are less than 128 luma samples. The CIIP prediction combines an inter prediction signal and an intra prediction signal. The inter prediction signal P_inter for the CIIP mode is derived using the same inter prediction process that applies to regular merge mode, and the intra prediction signal P_intra is derived according to regular intra prediction process together with the planar mode. Thereafter, the intra prediction signal and the inter prediction signal are combined using a weighted average, and the weight is calculated based on the coding modes of the top and left neighboring blocks as follows: A. islntraTop is set to 1 when the top neighbor is available and is intra coded; otherwise, islntraTop is set to 0; B. islntraLeft is set to 1 when the left neighbor is available and is intra coded; otherwise, islntraLeft is set to 0; C. wt is set to 3 when (islntraLeft + islntraLeft) is 2; D. Otherwise, wt is set to 2 when (islntraLeft + islntraLeft) is 1; E. Otherwise, wt is set to 1.

[0196] The CIIP prediction is constructed as shown in Equation 4: [Equation 4]

[0197] Figure 10 is an illustration of template matching (TM) to which embodiments of the present disclosure can be applied.

[0198] TM is a decoder-side motion vector derivation method that refines the motion information of the current coding unit by finding the closest match between a template (i.e., the top and / or left neighboring blocks of the current coding unit) in the current picture and a block (i.e., having the same size as the template) in the reference picture. As shown in Figure 10 , a better motion vector is searched around the initial motion of the current coding unit within a search range of [-8, +8] pixels. The search step is determined by the AMVR mode, and TM can be cascaded with the bi-directional matching process in merge mode.

[0199] In AMVP mode, MVP candidates are determined based on TM error. The template that achieves the smallest difference between the current block template and the reference block template is selected. TM is then performed only for a certain MVP candidate to refine the motion vector. TM uses an iterative diamond search (starting from full-pel MVD precision (4-pel in 4-pel AMVR mode)) within a search range of [-8, +8] pixels to refine the certain MVP candidate.

[0200] According to AMVR mode, a cross search is performed with full-pel MVD precision (4-pel in 4-pel AMVR mode), followed by a half-pel and a quarter-pel search in turn to further refine the AMVP candidate. The search process ensures that the MVP candidate remains the same MVD precision as indicated in AMVR mode after the TM process. The search process terminates when the difference between the previous minimum cost in iteration and the current minimum cost is less than or equal to a threshold of the area of the block.

[0201] [Table 1]

[0202] In merge mode, a similar search method is applied to the merge candidate indicated by the merge index. As shown in Table 1, TM can be performed up to 1 / 8-pel MVD precision, or can skip half-pel MVD precision, depending on the use of an alternative interpolation filter based on the merge motion information (used when AMVR is in half-pel mode). In addition, when TM mode is enabled, TM can operate as a standalone process or as an additional motion enhancement process between a block-based bi-directional matching (BM) scheme and a sub-block-based BM scheme, depending on whether BM is enabled or disabled.

[0203] In multi-hypothesis inter prediction (MHP) mode, one or more additional motion- compensated prediction signals are signaled in addition to the existing bi-prediction signal. The overall prediction signal is thus obtained by sample-wise weighted addition. Using the bi-prediction signal p_bi and the first additional prediction signal h_3, the resulting prediction signal p_3 is obtained as shown in Equation 5: [Equation 5]

[0204] The weight a can be specified with the new syntax element add_hyp_weight_idx according to the mapping in Table 2: [Table 2]

[0205] One or more additional prediction signals can be used as shown in Equation 6. The resulting overall prediction signal is accumulated iteratively for each additional prediction signal.

[0206] [Formula 6]

[0207] Thus, the overall prediction signal is obtained as the last p_n (i.e., p_n with the largest index n). For example, up to two additional prediction signals can be used (e.g., n is limited to 2).

[0208] The motion parameters of each additional prediction hypothesis can be explicitly signaled by specifying a reference index, a motion vector prediction index, and an MVD, or implicitly signaled by specifying a merge index. A separate multi-hypothesis merge flag can distinguish between these two signaling modes.

[0209] In the inter AMVP mode, MHP can be applied only when Bi-prediction with CU-level weight (BCW) weights is not equal in bi-prediction mode.

[0210] Furthermore, a combination of MHP and BDOF is possible, but BDOF can be applied only to the bi-predicted part of the prediction signal (i.e., the first two hypotheses).

[0211] A prediction block for the current block can be derived based on motion information derived according to an inter prediction mode. The prediction block can comprise predicted samples (a prediction sample array) of the current block. When the motion vector of the current block points to a fractional sample unit, an interpolation process can be performed by which a predicted sample for the current block can be derived based on a fractional sample reference sample within the reference picture.

[0212] When affine inter prediction is applied to the current block, the predicted samples can be generated based on sample / sub-block level motion vectors (MVs).

[0213] When bi-prediction is applied, the predicted samples derived based on L0 prediction (that is, using a reference picture in the reference picture list L0 and MVL0) and the predicted samples derived based on L1 prediction (that is, using a reference picture in the reference picture list L1 and MVL1) can be combined by a (phase dependent) weighted sum or weighted average (bi-prediction with CU-level weights, BCW), and the resulting predicted samples can be used as the predicted samples of the current block. When bi-prediction is applied and the reference picture used for L0 prediction and the reference picture used for L1 prediction are located in different temporal directions with respect to the current picture, this can be referred to as bi-prediction.

[0214] The reconstructed samples and the reconstructed picture can be generated based on the derived predicted samples, and then processes such as in-loop filtering can be performed.

[0215] The MHP mode will be described in detail below. As described above, the MHP mode is a prediction method using an additional prediction block (or predictor) other than a basic prediction block. The MHP mode can be selectively used as one of the various inter prediction modes described above. It should be understood that the MHP mode according to the embodiments of the disclosure is not limited to these names. In this specification, the MHP mode can also be referred to as a multi-reference mode, a multi-reference prediction, a multi-reference prediction mode, a multi-reference block mode, an MHP mode, a multi-hypothesis inter prediction mode, an inter-inter combination prediction mode, a combination inter prediction mode, a combination prediction mode, a multi-inter prediction mode, a multi-prediction mode, an additional reference prediction mode, an additional reference mode, a multi-reference block, etc.

[0216] Embodiments

[0217] Figure 11 A prediction method performed by a decoding apparatus according to an embodiment of the disclosure is exemplified. Figure 11 The steps or operations shown in the above-described embodiments are not essential components of the prediction method, and at least some of the steps or operations shown in the above-described embodiments can be omitted Figure 11 The steps or operations shown in the above-described embodiments are not essential components of the prediction method, and at least some of the steps or operations shown in the above-described embodiments can be omitted

[0218] Reference Figure 11 The decoding apparatus can generate a basic prediction block (or reference block) by performing uni-prediction or bi-prediction (S700). When the multi-reference mode is applied, the decoding apparatus can generate and combine an additional reference block other than a basic reference block generated (or derived) through uni-prediction or bi-prediction. For example, the basic reference block can be a block referred to (or used) to generate (or derive) the basic prediction block. For example, the basic reference block can include an L0 reference block and / or an L1 reference block. For example, the existing prediction block can be a block generated (or derived) by combining or weightedly summing the basic reference blocks. For example, the basic prediction block can refer to a block obtained by weightedly summing an L0 prediction block and an L1 prediction block. In the disclosure, the basic prediction block can be referred to as a basic block, an initial prediction block, an initial block, a temporary prediction block, a temporary block, a regular prediction block, a regular block, etc. In the present embodiment, the basic prediction block is mainly described as a block obtained by weightedly summing an L0 prediction block and an L1 prediction block, but the disclosure is not limited thereto.

[0219] In other words, in an embodiment, the image decoding apparatus according to the disclosure can derive a first reference block and a second reference block of a current block by performing bi-prediction, and generate a basic prediction block by weightedly summing the first reference block and the second reference block. Alternatively, in an embodiment, the image decoding apparatus according to the disclosure can derive a third reference block of a current block by performing uni-prediction, thereby generating a basic prediction block. In other words, the basic prediction block can be derived by weightedly summing a plurality of reference blocks, or can be derived using a single reference block.

[0220] The decoding apparatus can derive (or generate) the additional reference blocks (or additional prediction blocks) based on the multi-reference mode (S710). The decoding apparatus can derive the additional reference blocks in addition to the base prediction block, and combine (or weight-sum) the derived additional reference blocks with the base prediction block.

[0221] In an embodiment, when the multi-reference mode is applied, the decoding device can derive and combine up to a pre-defined number of additional reference blocks. In other words, the decoding apparatus can combine (or weight-sum) the base prediction block with less than or equal to the pre-defined number of additional reference blocks. For example, the pre-defined number can be 2. Alternatively, the pre-defined number can be one of 1, 2, 3, or 4. The pre-defined number can be referred to as a maximum number in the multi-reference mode.

[0222] Further, when combining the additional multi-reference blocks, the additional multi-reference blocks can be weight-summed with the base prediction block sequentially. For example, when up to two additional reference blocks are generated, the base prediction block and a first additional reference block can be weight-summed to generate a prediction block, and the generated prediction block and a second additional reference block can be weight-summed to generate a final prediction block. The prediction block generated by weight-summing the base prediction block and the first additional reference block can be referred to as an intermediate prediction block.

[0223] Alternatively, when combining the additional multi-reference blocks, the base prediction block and the generated additional multi-reference blocks can be weight-summed together. That is, after the additional multi-reference blocks are generated, a weight can be applied to each of the additional multi-reference blocks and the base prediction block (or the L0 reference block and the L1 reference block) and then weight-summed together.

[0224] Further, in an embodiment, the decoding apparatus can determine whether to apply the multi-reference mode. In this case, an operation of determining whether to apply the multi-reference mode can be added before operation S710. For example, whether to apply the multi-reference mode can be explicitly signaled or implicitly derived (or determined) by the decoding apparatus.

[0225] Further, as an embodiment, it can be signaled from the encoding apparatus to the decoding apparatus whether or not to apply the multi-reference mode. For example, a multi-reference mode flag indicating whether or not to apply the multi-reference mode can be signaled from the encoding apparatus to the decoding apparatus. In this case, a condition for signaling / parsing the multi-reference mode flag can be predefined. The signaling / parsing condition of the multi-reference mode flag can be a condition for availability of the multi-reference mode. When the signaling / parsing condition is satisfied, the decoding apparatus can parse the multi-reference mode flag from the bitstream. Alternatively, as an embodiment, the decoding apparatus can derive whether or not to apply the multi-reference mode based on predefined encoding information. As an example, whether or not the multi-reference mode is applicable can be defined in the same manner as the availability condition (or the signaling / parsing condition) of the multi-reference mode described below.

[0226] Further, as an embodiment, the decoding apparatus can obtain multi-reference mode information (also referred to as multi-reference mode prediction information) to generate the additional reference block. For example, the multi-reference mode information can include weight information and / or prediction information. Based on the prediction information, the reference block according to the multi-reference mode (i.e., the additional reference block) can be derived, and the additional reference block derived based on the weight information can be weighted-summed with the base prediction block (or the intermediate prediction block). Further, as an example, the multi-reference mode information can further include a multi-reference mode flag indicating whether or not to apply the multi-reference mode.

[0227] For example, the prediction information can include mode information for deriving the additional reference block and motion information according to the mode. The mode information can be inter mode information indicating whether the mode is a merge mode or an AMVP (or inter) mode. For example, the mode information can be a merge flag. That is, the merge mode or the AMVP (or inter) mode can be used to derive the additional reference block, and a flag syntax element indicating whether to use the merge mode or the AMVP (inter) mode can be signaled. Alternatively, a predefined mode among the merge mode or the AMVP (or inter) mode can be used to derive the additional reference block. Alternatively, the merge mode or the AMVP (or inter) mode can be selected based on predefined encoding information.

[0228] For example, when the merge mode is used to derive the additional reference block, the prediction information can include a merge index. The merge index can specify a merge candidate within a merge candidate list. When the AMVP (or inter) mode is used to derive the additional reference block, the prediction information can include an MVP flag, a reference index, and MVD information. The MVP flag can specify a candidate within an MVP candidate list.

[0229] The decoding apparatus can generate a final prediction block by weighted summing the base prediction block and the additional reference blocks (S720). As described above, the number of additional reference blocks can be less than or equal to a predefined number. For example, when the number of additional reference blocks is 2, the final prediction block can be a weighted sum of the base prediction block and the two additional reference blocks. In an embodiment, weight information for the weighted sum can be signaled or derived.

[0230] As described above, when combining the additional multi-reference blocks, the additional multi-reference blocks can be sequentially weighted-summed with the base prediction block, or the base prediction block and the generated additional multi-reference blocks can be weighted-summed together.

[0231] In general, an inter prediction process supports single or bi-directional prediction. However, when the inter prediction process includes more than one prediction block, the inter prediction process can be considered as MHP (or multi-reference prediction). In other words, the multi-reference prediction utilizes multiple reference blocks (or prediction blocks) for prediction, and the following signaling or derivation methods can be considered.

[0232] In an embodiment, the same or similar merge index as the merge mode can be used to signal information about the additional reference blocks. Alternatively, the same or similar reference index, MVP flag (or index), MVD, etc. as the AMVP (or inter) mode can be used to signal information about the additional reference blocks. Alternatively, the motion information about the additional reference blocks can be derived by inheriting from the already-decoded neighboring blocks. In an embodiment, whether to signal information about the additional reference blocks can be determined according to the number of additional reference blocks used. For example, when there is one additional reference block, information about the additional reference block can be signaled, and when there are two additional reference blocks, all or some of the information about the additional reference blocks can not be signaled but can be derived at the decoder side.

[0233] Based on the multi-reference mode according to the embodiment of the disclosure, by signaling / deriving the motion information of the prediction blocks as well as the weight information, blocks having different characteristics from the conventional prediction blocks can be generated, and the prediction accuracy can be improved by utilizing various reference blocks for prediction.

[0234] Figure 12 FIG. 1 is a diagram illustrating reference blocks used in a multi-reference mode according to an embodiment of the disclosure.

[0235] Referring to Figure 12 , the diagram illustrates a case where prediction is performed using multi-reference blocks (prediction blocks) when the multi-reference mode is applied. That is, in Figure 12 , the reference blocks P0 and P1 denote base reference blocks (or regular reference blocks), and the reference blocks P2 and P3 denote additional reference blocks.

[0236] Figure 13 is a diagram illustrating a method of signaling information about an additional reference block used in a multi-reference mode according to an embodiment of the disclosure. Figure 13 The operations shown in Figure 13 at least some of the operations shown in

[0237] When there can be an additional reference block and information for each additional reference block is signaled and parsed, the motion information of the added prediction block can be signaled and parsed in Figure 13 the order shown. Here, MaxNum can indicate the maximum number of additional reference blocks.

[0238] Referring to Figure 13 , a loop can be performed until the number of additional reference blocks reaches the maximum number. When the number of additional reference blocks is less than or equal to the maximum number, mhp_flag can be parsed (signaled). mhp_flag denotes a syntax element indicating whether to use an additional reference block. When an additional reference block is used, mhp_mrg can be parsed. mhp_mrg denotes a syntax element indicating whether to use a merge mode or an AMVP (or inter) mode to derive an additional reference block. According to the value of mhp_mrg, when the merge mode is applied, a merge index and a weight index can be signaled, and when the AMVP (or inter) mode is applied, a reference index, an MVP index, MVD data, and a weight index can be signaled.

[0239] As shown in Figure 13 , when the maximum number of additional reference blocks that can be generated is MaxNum, the presence of MHP-related syntax can be determined through mhp_flag. For example, when MaxNum is 2, mhp_flag can have the values shown in Table 3 below, and the presence of a first additional block and a second additional block can be determined based on the values.

[0240] [Table 3]

[0241] Referring to Table 3, when mhp_flag is "1", an additional reference block can be present. The additional reference block can be identified as an MHP_MERGE mode or an MHP_AMVP (or MHP_INTER) mode through mhp_mrg. In the MHP_MERGE mode, a merge index and a weight index can be signaled. In the MHP_AMVP (or MHP_INTER) mode, a reference index, an MVP index (or flag), MVD data, and a weight index can be signaled.

[0242] The syntax names described in this disclosure are examples, and the names may vary. Furthermore, in this disclosure, the modes for additional reference blocks when they exist are described as MHP_MERGE and MHP_AMVP (or MHP_INTER), which may differ from the merging mode and AMVP (or inter-frame) mode representing motion information for regular reference blocks. Specifically, MVP candidate lists for regular reference blocks and MVP candidate lists for additional reference blocks can be constructed independently. Additionally, the MHP_MERGE and MHP_AMVP (or MHP_INTER) modes for additional reference blocks may include weight index information.

[0243] Furthermore, in the implementation, in addition to the method of notifying the application of multiple reference blocks via the aforementioned signal, multiple reference blocks can also be applied through derivation. Specifically, since the merging mode uses motion information inherited from the decoded neighboring / non-neighboring blocks, the information for multiple reference blocks can also use motion information inherited from neighboring / non-neighboring blocks. For example, when the current block is in merging mode and the neighboring / non-neighboring blocks used to obtain motion information include MHP information, this information can be inherited and used to generate additional reference blocks for the current block.

[0244] In the implementation, as described above, multi-reference block information can be obtained through signal notification or derivation process, and both methods can be applied. For example, when configuring N multi-reference blocks, if there are M multi-reference blocks (M<=N) obtained through derivation process, information for the NM multi-reference blocks can be obtained through signal notification.

[0245] Additional reference blocks obtained through signal notification or derivation can be weighted and summed to generate the final prediction block as follows. As an example, when multiple reference blocks are applied, the final prediction block can be calculated as shown in Equation 7 below. Equation 7 assumes the existence of P0 and P1 as regular reference blocks, and further assumes the existence of P2 and P3.

[0246] [Formula 7]

[0247] Step 1:

[0248] Step 2:

[0249] Step 3:

[0250] In Equation 7, W0 and W1 represent the weights applied to the additional reference blocks P2 and P3, respectively. In the first operation of Equation 7, the basic prediction block can be generated by a weighted sum of the regular reference blocks. In the second and third operations of Equation 7, a weighted summation of the additional reference blocks can be performed. As mentioned above, the basic prediction block can be either the regular reference block before the weighted summation or the prediction block after the weighted summation.

[0251] Although the multiple reference blocks are described as being added to the inter prediction process, this is not limited to the inter prediction process. The multiple reference blocks can be included in the intra prediction process and the IBC process. In other words, the prediction mode for the base prediction block is not limited to the inter mode, and it can also be the intra mode or the IBC mode.

[0252] When the additional reference blocks are included in various base modes (prediction modes for base blocks), the base blocks (regular blocks) can include inter blocks (inter prediction blocks), intra blocks (intra prediction blocks), and IBC blocks, and the prediction modes of the additional reference blocks (hereinafter referred to as "additional prediction modes") can include not only the MHP AMVP (or MHP INTER) mode or the MHP MERGE mode, but also the MHP INTRA mode and the MHP IBC mode. In other words, the prediction modes of the additional reference blocks (additional prediction modes) are not limited to the inter mode, and the intra mode or the IBC mode is also possible. In an embodiment, the additional prediction modes can be the inter mode (MHP INTER mode), the intra mode (MHP INTRA mode), or the IBC mode (MHP IBC mode).

[0253] In an embodiment, information (e.g., a flag or an index) indicating each of the additional prediction modes can be signaled to determine the prediction mode for the additional reference blocks. In addition, the information indicating each of the additional prediction modes can be obtained at the decoder level by inheritance other than signaling. Furthermore, the prediction mode for the multiple reference blocks can be derived and generated at the decoder level when certain conditions are met, other than the signaling or transmission method.

[0254] Embodiment 1

[0255] The prediction mode of the additional reference block (additional prediction mode) can include the MHP INTRA mode, the MHP IBC mode, or the MHP INTER mode. When the prediction mode for the base prediction block (i.e., the base prediction mode (or base mode)) is the intra mode, the IBC mode, or the inter mode, the prediction mode for the additional reference block (i.e., the additional prediction mode) can exist in various combinations.

[0256] In an embodiment, the prediction modes for the additional reference blocks (including the MHP INTRA mode, the MHP IBC mode, and the MHP INTER mode) can exist in a limited number according to the type of the base mode, and can be determined based on a combination with the allowed base mode.

[0257] In an embodiment, the condition for combining the prediction mode for the base prediction block (base prediction mode, MHP mode) and the prediction mode for the additional reference block (additional prediction mode) can be the prediction mode for the base prediction block.

[0258] The condition for combining the base mode and the additional prediction mode can be a prediction tool applied to the base block in the base inter mode. The prediction tool of the base inter mode can include a tool related to a method of generating a prediction block, such as a base affine mode (e.g., affine AMVP and affine merge), a base AMVP mode, a base merge mode, an MMVD mode, a CIIP mode, and a GPM mode.

[0259] The condition for combining the base mode and the additional prediction mode can be a mode (e.g., planar mode, DC mode, directional mode, etc.) applied to the base block in the base intra mode. In other words, whether to allow MHP_INTRA can be determined based on the mode for the base intra mode. For example, when the base intra mode is planar or DC, the MHP_INTRA mode can be allowed. Further, the allowed mode in the MHP_INTRA mode can be determined based on the base intra mode.

[0260] The condition for combining the base mode and the additional prediction mode can be the number of prediction blocks (or base reference blocks) used in the base block. Whether to allow the MHP mode can be determined based on whether the base inter mode uses uni-prediction or bi-prediction, or whether the base intra mode uses a mode using multiple reference blocks such as decoder-side intra mode derivation (DIMD) or template-based intra mode derivation (TIMD).

[0261] The condition for combining the base mode and the additional prediction mode can be the size of the current block, the width / height of the current block, or the block ratio. In an embodiment, when the width or height of the current block is equal to a reference value, the combination between any one of the base inter mode, the base intra mode, or the base IBC mode and at least one of the MHP_INTER mode, the MHP_INTRA mode, or the MHP_IBC mode is allowed, or when the width or height of the current block is greater than the reference value, the combination between any one of the base inter mode, the base intra mode, or the base IBC mode and at least one of the MHP_INTER mode, the MHP_INTRA mode, or the MHP_IBC mode is allowed, or when the width or height of the current block is less than the reference value, the combination between any one of the base inter mode, the base intra mode, or the base IBC mode and at least one of the MHP_INTER mode, the MHP_INTRA mode, or the MHP_IBC mode is allowed.

[0262] The condition for the combination between the base mode and the additional prediction mode can be a quantization parameter (QP). In an embodiment, when the QP is equal to a threshold (e.g., 32), the combination between any one of the base inter mode, the base intra mode, or the base IBC mode and at least one of the MHP INTER mode, the MHP INTRA mode, or the MHP IBC mode is allowed, or when the QP is greater than the threshold, the combination between any one of the base inter mode, the base intra mode, or the base IBC mode and at least one of the MHP INTER mode, the MHP INTRA mode, or the MHP IBC mode is allowed, or when the QP is less than the threshold, the combination between any one of the base inter mode, the base intra mode, or the base IBC mode and at least one of the MHP INTER mode, the MHP INTRA mode, or the MHP IBC mode is allowed.

[0263] The specific condition for the combination between the base mode and the additional prediction mode can be a temporal layer. For example, when the temporal layer is equal to a threshold (e.g., 4), the combination between any one of the base inter mode, the base intra mode, or the base IBC mode and at least one of the MHP INTER mode, the MHP INTRA mode, or the MHP IBC mode is allowed, or when the temporal layer is greater than the threshold, the combination between any one of the base inter mode, the base intra mode, or the base IBC mode and at least one of the MHP INTER mode, the MHP INTRA mode, or the MHP IBC mode is allowed, or when the temporal layer is less than the threshold, the combination between any one of the base inter mode, the base intra mode, or the base IBC mode and at least one of the MHP INTER mode, the MHP INTRA mode, or the MHP IBC mode is allowed.

[0264] Figure 14 FIG. 1 is a diagram illustrating a combination between a base block mode and a mode for an additional reference block according to an embodiment of the present disclosure.

[0265] In an embodiment, when there is an additional reference block, in addition to the inter mode (AMVP mode, merge mode), the intra mode and the IBC mode can also be supported. When the intra mode as the additional reference block other than the base block is referred to as the MHP INTRA mode, the IBC mode is referred to as the MHP IBC mode, and the inter mode is referred to as the MHP INTER mode, the combination of the base block and the additional reference block in each of the inter mode, the intra mode, and the IBC mode can be as shown in Figure 14

[0266] As shown in Figure 14 ​As shown, for each of the inter mode, intra mode, and IBC mode, there can be a MHP INTRA mode, a MHP IBC mode, and a MHP INTER mode. This is an example, and the mode of additional reference blocks allowed for each prediction mode can be as follows: A. Additional reference blocks in MHP INTRA mode can be allowed.

[0267] A-1. A limited number of additional reference blocks in MHP INTRA mode can be included. For example, one additional reference block in MHP INTRA mode can be allowed.

[0268] A-2. Additional reference blocks in MHP INTRA mode can be allowed when certain conditions are met.

[0269] B. Additional reference blocks in MHP-IBC mode can be allowed.

[0270] B-1 A limited number of additional reference blocks in MHP-IBC mode can be included. For example, one additional reference block in MHP-IBC mode can be allowed.

[0271] B-2 Additional reference blocks in MHP-IBC mode can be allowed when certain conditions are met.

[0272] B-3 Since IBC mode is divided into AMVP mode and merge mode, MHP-IBC can also be divided into MHP-IBC AMVP mode and MHP-IBC MERGE mode. Alternatively, both MHP-IBC AMVP mode and MHP-IBC MERGE mode can be allowed, or only a specific mode between MHP-IBC AMVP mode and MHP-IBC MERGE mode can be allowed. For example, only MHP-IBC MERGE mode can be allowed.

[0273] C Additional reference blocks in MHP INTER mode can be allowed.

[0274] C-1 A limited number of additional reference blocks in MHP INTER mode can be included. For example, one MHP INTER mode can be allowed.

[0275] C-2 Additional reference blocks in MHP INTRA mode can be allowed when certain conditions are met.

[0276] C-3 Since the base inter mode is divided into the AMVP mode and the merge mode, the MHP INTER mode for the additional reference block can also have both the MHP AMVP mode and the MHP MERGE mode. Alternatively, both the MHP AMVP mode and the MHP MERGE mode for the additional reference block can be allowed, or only one of the MHP AMVP mode and the MHP MERGE mode for the additional reference block can be allowed. For example, only the MHP MERGE mode can be allowed.

[0277] Various combinations of modes can be generated using the above-described methods. In an embodiment, when the base block mode is the intra mode, a combination of one additional reference block of the MHP INTRA mode and one additional reference block of the MHP IBC mode can be allowed. In an embodiment, when the base block mode is the IBC mode, a combination of one additional reference block of the MHP INTER mode and additional multiple reference blocks of the MHP IBC mode can be allowed.

[0278] Figure 15 FIG. 1 is a diagram illustrating a method of signaling information about an additional reference block used in a multi-reference mode according to an embodiment of the disclosure. Figure 15 The operations shown in FIG. 1 are not essential operations of an embodiment of the disclosure, and at least some of the operations shown in FIG. 1 can be omitted. Figure 15 The operations shown in FIG. 1 are not essential operations of an embodiment of the disclosure, and at least some of the operations shown in FIG. 1 can be omitted.

[0279] Specifically, Figure 15 Examples in which the MHP INTRA mode, the MHP IBC mode, and the MHP INTER mode for the additional reference block are allowed in a specific mode for the base block are illustrated. In an embodiment, when the additional reference block exists and a specific condition is satisfied, information (e.g., a flag or an index) indicating whether the MHP INTER mode is allowed can be signaled (by the encoder) / parsed (by the decoder). For example, a flag mhp_intra_flag indicating whether the MHP INTRA mode is allowed can be signaled / parsed. The name of the flag mhp_intra_flag is not limited thereto, and a flag having various names can be used to indicate whether the MHP INTRA mode is allowed. For example, when the flag mhp_intra_flag is "1", the additional reference block is classified as the MHP INTRA mode, and prediction information for an intra block that is the additional reference block can be signaled / parsed when it exists.

[0280] When the specific condition is not satisfied or the flag mhp_intra_flag is "0", information (e.g., a flag or an index) indicating whether the MHP_IBC mode is allowed can be signaled / parsed. For example, a flag mhp_ibc_flag indicating whether the MHP_IBC mode is allowed can be signaled / parsed. The name of the flag mhp_ibc_flag is not limited thereto, and a flag having various names can be used to indicate whether the MHP_IBC mode is allowed. For example, when the flag mhp_ibc_flag is "1", the additional reference block is identified as the MHP_IBC mode, and prediction information for the IBC block which is the additional reference block can be signaled / parsed. For example, as shown in Figure 15 , prediction information for the MHP_IBC_MERGE block can be signaled / parsed.

[0281] Finally, when the specific condition is not satisfied or the flag mhp_ibc_flag is "0", the additional reference block is identified as the MHP_INTER mode, and prediction information for the inter block which is the additional reference block can be signaled / parsed. For example, as shown in Figure 15 , prediction information for the MHP_MERGE block can be signaled / parsed.

[0282] Figure 15 The methods exemplified in the above are examples, and various combinations are possible according to the methods listed above. As an example, a flag mhp_intra_flag indicating whether the MHP_INTER mode is allowed can be signaled / parsed, when the flag mhp_intra_flag is "1", prediction information for the MHP_INTRA block can be signaled / parsed, when the flag mhp_intra_flag is "0", a flag mhp_inter_flag indicating whether the MHP_INTER mode is allowed can be signaled / parsed, when the flag mhp_inter_flag is "1", prediction information for the MHP_INTER block can be signaled / parsed, and when the flag mhp_inter_flag is "0", prediction information for the MHP_IBC block can be signaled / parsed.

[0283] As an example, a flag mhp ibc flag indicating whether the MHP IB C mode is allowed can be signaled / parsed, and when the flag mhp ibc flag is "1", prediction information for MHP IB C blocks can be signaled / parsed, and when the flag mhp ibc flag is "0", a flag mhp intra flag indicating whether the MHP INTRA mode is allowed can be signaled / parsed, and when the flag mhp intra flag is "1", prediction information for MHP INTRA blocks can be signaled / parsed, and when the flag mhp intra flag is "0", prediction information for MHP INTER blocks can be signaled / parsed. As an example, a flag mhp ibc flag indicating whether the MHP IB C mode is allowed can be signaled / parsed, and when the flag mhp ibc flag is "1", prediction information for MHP IB C blocks can be signaled / parsed, and when the flag mhp ibc flag is "0", a flag mhp inter flag indicating whether the MHP INTER mode is allowed can be signaled / parsed, and when the flag mhp inter flag is "1", prediction information for MHP INTER blocks can be signaled / parsed, and when the flag mhp inter flag is "0", prediction information for MHP INTRA blocks can be signaled / parsed.

[0284] As an example, a flag mhp_inter_flag indicating whether or not the MHP INTER mode is allowed can be signaled / parsed, and when the flag mhp_inter_flag is "1", prediction information for the MHP INTER block can be signaled / parsed, and when the flag mhp_inter_flag is "0", a flag mhp_intra_flag indicating whether or not the MHP INTRA mode is allowed can be signaled / parsed, and when the flag mhp_intra_flag is "1", prediction information for the MHP INTRA block can be signaled / parsed, and when the flag mhp_intra_flag is "0", prediction information for the MHP IBC block can be signaled / parsed. In an embodiment, a flag mhp_inter_flag indicating whether or not the MHP INTER mode is allowed can be signaled / parsed. When the flag mhp_inter_flag is "1", prediction information for the MHP INTER block can be signaled / parsed. When the flag mhp_inter_flag is "0", a flag mhp_ibc_flag indicating whether or not the MHP IBC mode is allowed can be signaled / parsed. When the flag mhp_ibc_flag is "1", prediction information for the MHP IBC block can be signaled / parsed. When the flag mhp_ibc_flag is "0", prediction information for the MHP INTRA block can be signaled / parsed.

[0285] The specific condition allowing the combination of each mode (e.g., condition 1 and condition 2 as shown in FIG. 6A) can include the following, and can include one or more conditions. Figure 16

[0286] A. The specific condition can be related to the prediction mode for the base block. That is, whether or not the specific condition is satisfied can be determined based on whether the prediction mode for the base block is an intra mode, an IBC mode, or an inter mode. For example, when the prediction mode for the base block is an intra mode or an IBC mode, the MHP INTER mode for the additional reference block can be allowed. As another example, when the prediction mode is inter, the MHP IBC can be allowed.

[0287] ​A-1. The specific condition can be related to the prediction tool in the inter mode applied to the base block. In this case, the prediction tool can be related to the generation method of the prediction block, such as affine mode (affine AMVP and affine merge), AMVP mode, base merge mode, MMVD mode, CIIP mode, or GPM mode. For example, when the base block is in AMVP mode, MHP INTRA mode for the additional reference block can not be allowed. As another example, when the base block is in CIIP mode or GPM mode, MHP INTRA mode and / or MHP IBC mode for the additional reference block can not be allowed.

[0288] A-2 The specific condition can be the prediction mode for the intra mode. That is, whether MHP INTRA for the additional reference block is allowed can be determined based on the prediction mode for the base block that is the intra mode. For example, when the base block is the intra mode and has a planar mode or a DC mode, MHP INTRA mode for the additional reference block can be allowed. As another example, when the base block has an angular mode with clear directional (e.g., horizontal mode, vertical mode, or its neighboring directional mode), MHP INTRA mode for the additional reference block can not be allowed. In addition, the allowable prediction mode (e.g., planar mode, DC mode, or angular mode) among MHP INTRA modes for the additional reference block can be determined based on the prediction mode for the base block. For example, when the base block is the intra mode and has an angular mode, the non-directional mode (e.g., planar mode and DC mode) among MHP INTRA modes for the additional reference block can be allowed.

[0289] A-3 The specific condition can be the number of prediction blocks (base reference blocks) used in the base block.

[0290] A-3-1 When the base block is in the inter mode with single prediction, the additional reference block in MHP INTRA mode, MHP INTER mode, or MHP IBC mode can be allowed.

[0291] A-3-2 When the base block is in the inter mode with bi-directional prediction, a limited number of additional reference blocks (additional reference blocks in MHP INTRA mode, MHP INTER mode, or MHP IBC mode) can be allowed. For example, one MHP INTRA can be allowed.

[0292] A-3-3 When the base block includes multiple intra blocks, the additional reference blocks can not be allowed (in MHP INTRA mode, MHP INTER mode, or MHP IBC mode). For example, when the base block includes multiple intra blocks, this can be because the base block uses DIMD or TIMD to derive multiple prediction modes, or because the base block includes different intra modes in a multi-partition region such as SGPM.

[0293] A-3-4 When the base block includes multiple IBC blocks, the additional reference blocks can not be allowed (MHP INTRA mode, MHP INTER mode, or MHP IBC mode). For example, when the base block includes multiple IBC blocks, this can be because the base block is a block with two or more block vectors, or because the base block is an IBC block that supports bi-prediction.

[0294] A-4 The specific condition can be a size of the base block, a width / height ratio of the base block, or a block ratio. For example, when the size (width x height) of the base block is less than 1024, the MHP INTRA mode or the MHP IBC mode can be allowed. As another example, when the block size (width x height) is less than 64, the MHP IBC or MHP INTER mode can be allowed.

[0295] A-5 The specific condition can be a quantization parameter (QP) of the base block. For example, when the QP of the base block is less than a threshold value, the MHP INTRA mode or the MHP IBC mode for the additional reference blocks can be allowed. In this case, the threshold value can be 32. In an embodiment, when the QP is equal to the threshold value (e.g., 32), a combination between any one of the base inter mode, the base intra mode, or the base IBC mode and at least one of the MHP INTER mode, the MHP INTRA mode, or the MHP IBC mode can be allowed, or when the QP is greater than the threshold value, a combination between any one of the base inter mode, the base intra mode, or the base IBC mode and at least one of the MHP INTER mode, the MHP INTRA mode, or the MHP IBC mode can be allowed, or when the QP is less than the threshold value, a combination between any one of the base inter mode, the base intra mode, or the base IBC mode and at least one of the MHP INTER mode, the MHP INTRA mode, or the MHP IBC mode can be allowed.

[0296] A-6 The certain condition can be the temporal layer. For example, the MHP INTRA mode or the MHP IB C mode for the additional reference block can be allowed when the temporal layer is less than a threshold. In this case, the threshold of the temporal layer can be 4. For example, when the temporal layer is equal to the threshold (e.g., 4), the combination between any one of the basic inter mode, the basic intra mode, or the basic IBC mode and at least one of the MHP INTER mode, the MHP INTRA mode, or the MHP IB C mode is allowed, or when the temporal layer is greater than the threshold, the combination between any one of the basic inter mode, the basic intra mode, or the basic IBC mode and at least one of the MHP INTER mode, the MHP INTRA mode, or the MHP IB C mode is allowed, or when the temporal layer is less than the threshold, the combination between any one of the basic inter mode, the basic intra mode, or the basic IBC mode and at least one of the MHP INTER mode, the MHP INTRA mode, or the MHP IB C mode is allowed.

[0297] This limited allowance of the additional reference block can help reduce the signaling overhead. The proposed method can be utilized by considering the trade-off between the encoder / decoder complexity and the performance improvement. In addition, the proposed method can utilize only the mode that obtains information by signaling, only the mode that obtains information by inheritance and derivation, or a combination of the mode that obtains information by signaling and the mode that obtains information by inheritance and derivation.

[0298] Embodiment 2

[0299] The additional reference block can be generated using the signaled information or the information derived through the inheritance process. Since various combinations of the additional reference block are possible, the additional reference block inherited from the neighboring block can improve the compression performance when it satisfies the following constraints. The following constraints can be used for this purpose.

[0300] The inheritance is possible when the neighboring / non-neighboring neighboring block and the current block have the same basic mode. In an embodiment, when the basic prediction mode for the current block is the basic inter mode, the information for the additional reference block can be inherited from the neighboring block having the basic inter mode. In an embodiment, when the basic prediction mode for the current block is the basic intra mode, the information about the additional reference block can be inherited from the neighboring block having the basic intra mode. In another embodiment, when the basic prediction mode for the current block is the basic IBC mode, the information about the additional reference block can be inherited from the neighboring block having the basic IBC mode.

[0301] The inheritance of additional reference block patterns can be determined based on the pattern priority determined according to the basic pattern for the current block. In another implementation, when adjacent blocks include additional reference blocks of various patterns, the order and number of additional reference blocks (or additional prediction patterns) inherited from adjacent blocks can be determined based on the pattern priority determined according to the basic pattern for the current block.

[0302] The inheritance of additional prediction modes can be determined based on each allowed combination. For example, when the base mode is an inter-frame mode and only MHP_INTRA and MHP_IBC modes are allowed as additional prediction modes, information about the intra-frame mode and IBC mode in the additional reference blocks included in the adjacent blocks can be inherited only.

[0303] The inheritance of additional prediction modes can be determined based on the prediction mode (planar mode, DC mode, or directional mode) of the base intra mode. For example, when the base mode for the current block is intra mode and it has planar mode or DC mode, information about the MHP_INTRA mode for adjacent blocks may not be inherited.

[0304] In addition to signaling information about additional prediction patterns, methods for inheriting information from neighboring / non-neighboring blocks are also described.

[0305] Merging patterns borrow motion information directly from neighboring or non-neighboring blocks without signaling the motion information. In multi-reference patterns, information about additional prediction patterns included in neighboring or non-neighboring blocks can be derived using the same or similar methods as other information.

[0306] As described above, additional reference blocks can be generated by signaling information about additional prediction modes (or additional reference blocks) by distinguishing between intra-frame, IBC (AMVP, merge), and inter-frame (AMVP, merge) modes (corresponding to the number of allowed additional reference blocks). Furthermore, additional reference blocks for the current block can be generated by distinguishing between additional prediction modes (or additional reference blocks) and inherited information about them.

[0307] The inheritance of additional prediction patterns can be determined based on the following criteria: A. Basic Mode Types When the basic modes for the current block and neighboring / non-neighboring blocks are the same, the additional prediction modes for neighboring blocks can be inherited by the current block. Specifically, when the current block is in inter-frame mode, it can inherit multi-reference block information (information about the additional prediction modes) from neighboring blocks in inter-frame mode. Similarly, when the current block is in IBC mode or intra-frame mode, it can inherit multi-reference block information (information about the additional prediction modes) from neighboring blocks in IBC mode or intra-frame mode, respectively.

[0308] B. Pattern-Specific Priority

[0309] Information about additional prediction modes can be inherited based on predefined priorities for the base mode of the current block. Table 4 illustrates examples of the priority of additional reference blocks' modes (additional prediction modes) based on the base mode (base prediction mode) of the current block. The types and number of allowed additional prediction modes, as well as their priorities, can vary depending on each base mode.

[0310] [Table 4]

[0311] In the implementation according to Table 4, when the basic mode for the current block is inter-frame mode and the modes of the additional reference blocks of the adjacent blocks with inter-frame mode as the basic mode are configured in the order of MHP_IBC mode, MHP_INTRA mode and MHP_INTER mode, the modes of the multiple reference blocks for the current block can be configured in the order of MHP_INTER mode, MHP_IBC mode and MHP_INTRA mode.

[0312] In another implementation, when the modes of the additional reference blocks for adjacent blocks are configured in the order of MHP_IBC mode 1, MHP_INTRA mode, and MHP_IBC mode 2, and the IBC mode has a higher priority, information about the additional prediction mode can be inherited to the current block in the order of MHP_IBC mode 1, MHP_IBC mode 2, and MHP_INTRA mode.

[0313] C. The permissible number of additional predictive patterns inherited from the basic pattern.

[0314] Information about additional prediction modes can be inherited based on the permissible number of additional prediction modes for each base mode. In an implementation, only one prediction mode can be inherited for each prediction mode of each additional reference block. For example, when a neighboring block includes an additional reference block and additional reference blocks in MHP_INTER mode 1, MHP_INTER mode 2, and an MHP_IBC mode, MHP_INTER mode 1 and MHP_IBC mode can be inherited as additional prediction modes for the current block. A neighboring block including MHP_INTER mode 1, MHP_INTER mode 2, and an MHP_IBC mode can have inherited MHP_INTER mode 1, MHP_IBC mode, and MHP_INTER mode 2. In other words, the permissible number of inheritances for each additional reference block (additional prediction mode) can be applied separately from the permissible number of additional prediction modes for signaling.

[0315] D. Permissible combinations with additional prediction models based on the basic model

[0316] The inheritance of additional prediction modes can be determined based on the permissible combinations of additional prediction modes according to the basic mode. In an implementation, when only combinations of MHP_INTER and MHP_IBC modes are allowed for the basic inter-frame mode, information regarding only MHP_INTER and MHP_IBC modes from the information of additional reference blocks included in adjacent blocks can be inherited into the current block. This is an example, and the possible combinations of additional prediction modes according to the basic mode and the number of permissible inheritances according to the basic mode within the possible combinations can be determined. Furthermore, the permissible combinations of inheritance for additional reference blocks (additional prediction modes) and the permissible combinations of additional prediction modes for signaling can be independent.

[0317] E. Prediction mode for basic intra-block

[0318] The inheritance of additional prediction modes can be determined based on the prediction mode for the basic intra-mode. For example, when the basic mode for the current block is intra-mode and planar or DC mode, information about the MHP_INTRA mode for adjacent blocks may not be inherited. Another example is that when the basic mode for the current block is intra-mode and directional mode, information about the MHP_INTER and MHP_IBC modes for adjacent blocks may not be inherited. These are just examples, and various combinations are allowed depending on the prediction mode for the intra-mode of the current block and the type and prediction mode of the additional reference blocks for adjacent blocks.

[0319] Implementation Method 3

[0320] Additional reference blocks can be generated using information signaled or derived through an inheritance process. Additionally, additional reference blocks in MHP_INTER, MHP_IBC, and / or MHP_INTRA modes can be derived based on TM.

[0321] For example, additional reference blocks in MHP_IBC mode and / or MHP_INTER mode can be generated based on TM. TM can be applied in particular to merge modes (MHP_IBC mode or MHP_INTER mode).

[0322] Additional reference blocks in MHP_INTER and / or MHP_IBC modes, derived based on TM, can initially search for integer pixel units based on the basic motion vector and / or basic block vector. Subsequently, additional searches can be performed for fractional pixel units based on the location with the minimum error (or cost). Then, inter-frame mode blocks and / or IBC mode blocks with refined motion vectors and / or refined block vectors at the location with the minimum error (or cost) can be used as additional reference blocks.

[0323] The basic motion vector can be motion information indicated by the MVP index or BVP index in the merge pattern. The search process (initial search and / or additional search) can be performed within a predefined search range.

[0324] Error (or cost) can refer to various types of error or cost, such as the sum of absolute errors, the mean of absolute errors, the sum of squared errors, and the mean of squared errors. Error (or cost) is calculated based on the difference between the template region of the current block and the template region of the predicted block, and can be calculated as the sum of absolute differences (SA(T)D), the mean-minus-supplemented-absolute-differences (MR-SA(T)D), the sum of squared errors (SSE), or the mean squared error (MSE). Depending on the specific region and location within the search range, the error (or cost) can be multiplied by different weights.

[0325] Additional reference blocks in MHP_INTRA mode can be generated based on TM. In MHP_INTRA mode, the error (or cost) between the reconstructed samples in the template region of the current block and the predicted samples generated from the reference samples in the template region is calculated, and the mode with the minimum error (or cost) can be selected.

[0326] The prediction pattern used to calculate errors (or costs) can be determined within the MPM configured for the MHP_INTRA mode. The MPM can include planar, horizontal, vertical, etc., as a basis, and can include one or more patterns.

[0327] For the aforementioned MHP_INTRA pattern block, the following procedure can be used to determine the prediction pattern for the MHP_INTRA pattern: calculate the horizontal and vertical gradients by applying a Sobel filter to the template region of the current block, obtain the gradient histogram (HoG) based on the horizontal and vertical gradients, derive the angle pattern based on the gradient histogram, and use the derived angle pattern as the prediction pattern for the MHP_INTRA pattern.

[0328] Depending on the allowed number, HoG-based prediction patterns for the MHP_INTRA pattern can include one or more patterns.

[0329] The following section describes a specific method for deriving an additional reference block (or additional prediction pattern) without signaling / resolving information about the additional reference block (or additional prediction pattern).

[0330] Figure 16 This diagram illustrates a method for deriving additional reference blocks based on TM in multiple reference modes according to an embodiment of the present disclosure. Specifically, it illustrates a method for deriving additional reference blocks in MHP_INTER mode and / or MHP_IBC mode based on TM.

[0331] When the base mode for the current block is inter-frame mode, the process described in this embodiment can be used. Specifically, when the base mode for the current block is merge mode, the process described in this embodiment can be applied. However, this is not limited to this, and the process described in this embodiment can also be applied when the base mode for the current block is AMVP mode.

[0332] Specifically, such as Figure 16 As shown, TM-based motion vector refinement (MV refinement) can be performed within the search range using the motion vectors of the block indicated by the motion vector predictor (MVP) as the base motion vector (baseMV). Using the improved motion vectors, inter-frame prediction blocks can be generated as additional reference blocks. The motion vector refinement (MV refinement) process can be performed in the following order.

[0333] A. Based on the base motion vector (baseMV), the position with the minimum error (or cost) is determined by searching integer pixel units.

[0334] B. If possible, perform a fractional pixel search based on the location with error (or cost) to determine the location with the minimum error (or cost).

[0335] C. The location with the minimum error (or cost) is designated as the refined motion vector (refinedMV), and the block with the refined motion vector (refinedMV) is designated as the block in MHP_INTER mode, thus generating the prediction block. When the base motion vector (baseMV) and the refined motion vector (refinedMV) are the same, the process of generating additional reference blocks can be omitted.

[0336] When the base mode for the current block is IBC mode, the process described in this embodiment can be used. Specifically, when the base mode for the current block is merge mode, the process described in this embodiment can be applied. However, this is not limited to this, and the process described in this embodiment can also be applied when the base mode for the current block is AMVP mode.

[0337] Specifically, such asFigure 17 As shown, motion information (block vectors) of the blocks indexed by the Block Vector Predictor (BVP) can be used as base block vectors (baseBV) to perform TM-based block vector refinement (BV refinement) within the search range. Using the refined block vectors, IBC prediction blocks can be generated as additional reference blocks. The BV refinement process can be performed in the following order: A. Based on the basic block vector (baseBV), the location with the minimum error (or cost) is determined by searching integer pixel units.

[0338] B. If possible, perform a search on a fractional pixel basis to determine the location with the minimum error (or cost) based on the location with error (or cost).

[0339] C. Predictive blocks are generated by designating the location with the minimum error (or cost) as the refined block vector (refinedBV) and designating the block with the refined block vector (refinedBV) as the block in MHP_IBC mode. In this case, the process of generating additional predictive blocks can be omitted when the base block vector (baseBV) and the refined block vector (refinedBV) are the same.

[0340] The search process described above can be performed within the search range, and weights can be applied to the error (or cost) based on each search location or region within the search range during the search process. In an implementation, the selectivity of a specific region can be increased through calculations such as in [Equation 8].

[0341] [Formula 8]

[0342] For example, cost can represent error (or cost), and λ can be assigned a smaller value the closer it is to the position indicated by the base motion vector (baseMV) or base block vector (baseBV).

[0343] Figure 18 This is a diagram illustrating a method for deriving an additional reference block based on intra-frame template matching prediction used in a multi-reference mode, according to an embodiment of the present disclosure. Figure 19 This is a diagram illustrating a method for deriving an additional reference block based on TIMD used in a multi-reference mode, according to an embodiment of the present disclosure. Figure 17 This is a diagram illustrating a method for calculating the error (or cost) of an additional reference block in a multi-reference mode according to an embodiment of the present disclosure.

[0344] The MHP_IBC block can be derived using a derivation method based on IntraTMP (Intra-Template Matching Prediction). In other words, as... Figure 18As shown, through multi-reference matching prediction (IntraTMP), the error (or cost) between the template region of the encoded / decoded neighboring block and the template region of the block pointed to by the block vector of the neighboring block can be used, and the position with the minimum error (or cost) can be used as the block vector of the MHP_IBC block for the current block.

[0345] When the basic mode for the current block is intra-frame mode, the process described in this embodiment can be used.

[0346] Under certain conditions, an intra-predictive block can be generated as an additional reference block, and the MHP_INTRA block can be derived using the following method.

[0347] A. A prediction pattern for the MHP_INTRA pattern can be utilized, similar to the TIMD method described above. Specifically, such as... Figure 19 As shown, the error (or cost) between the reconstructed samples within the template region of the current block and the predicted samples generated from the reference samples within the template region is calculated, and the mode with the minimum error (or cost) can be selected. The prediction mode used to calculate the error (or cost) can be determined within the MPM configured for the MHP_INTRA mode. The MPM can include a planar mode, a horizontal mode, a vertical mode as a base, and one or more modes.

[0348] B. Prediction patterns for the MHP_INTRA pattern can be utilized similarly to the DIMD method described above. In other words, by applying a Sobel filter within the template region of the current block, the horizontal and vertical gradients are computed, and the HoG can be obtained based on these gradients. Angle patterns are derived based on the gradient histogram, and these derived angle patterns can be used as prediction patterns for the MHP_INTRA pattern. Depending on the allowed number, the HoG-based prediction patterns for the MHP_INTRA pattern can include one or more patterns.

[0349] The error (or cost) used in deriving blocks in MHP_INTER mode and / or MHP_IBC mode can be calculated as follows. For example, as... Error (cost) As shown, based on the error between the neighboring samples (cA, cL) of the current block and the neighboring samples (pA, pL) of the reference block pointed to by the basic motion vector (baseMV) or the basic block vector baseBV, when the error is greater than a threshold, a block in MHP_INTER mode and / or MHP_IBC mode can be generated.

[0350] Specifically, the error SA between the upper neighbor sample (cA) of the current block and the upper neighbor sample (pA) of the predicted block can be compared with a threshold, or the error SL between the left neighbor sample (cL) of the current block and the left neighbor sample (pL) of the predicted block can be compared with a threshold. Additionally, both the error SA associated with the upper neighbor sample and the error SL associated with the left neighbor sample can be compared with a threshold. For ease of explanation, the error for the top-left neighbor sample is not described. However, samples at these locations can be considered when calculating inter-sample errors using upper neighbor samples, left neighbor samples, or top-left neighbor samples.

[0351] The threshold can be set using Equations 9 and 10.

[0352] [Formula 9]

[0353] Threshold = Width Template height / λ

[0354] [Formula 10]

[0355] Threshold = Template width Height / λ

[0356] For example, λ can be set to a value that satisfies λ>1.

[0357] Here, the following methods can be used to calculate the error (or cost) between samples (or sample values), and combinations of the following methods can be applied.

[0358] A. The error (or cost) can be calculated by considering the sum of absolute differences (SAD) between samples. In other words, the error (or cost) can be calculated as in Equation 11. In [Equation 11], c and p represent the neighboring samples of the current block and the predicted block, respectively, and x represents the position of the neighboring sample. Furthermore, this equation can be applied to N samples in the upper region, the left region, or the upper left region.

[0359] [Equation 11]

[0360] Error (cost) =

[0361] B. The error (or cost) can be calculated by considering the mean-sum-absolute-differences (MR-SAD) among the samples. In other words, the error (or cost) can be calculated as shown in [Equation 12]. In [Equation 12], c and p represent the neighboring samples of the current block and the predicted block, respectively, and x represents the position of the neighboring sample. Furthermore, this equation can be applied to N samples in the upper region, the left region, or the upper left region.

[0362] [Equation 12]

[0363] Error (cost) =

[0364] C. The error (or cost) can be calculated by considering the SSE between samples. In other words, the error (or cost) can be calculated as shown in [Equation 13]. In [Equation 13], c and p represent the neighboring samples of the current block and the predicted block, respectively, and x represents the position of the neighboring sample. Furthermore, this can be applied to N samples in the upper region, the left region, or the upper left region.

[0365] [Equation 13]

[0366] Error (cost) =

[0367] D. The error (or cost) can be calculated by considering the MSE between samples. In other words, this can be calculated as shown in [Equation 14]. In [Equation 14], c and p represent the neighboring samples of the current block and the predicted block, respectively, and x represents the position of the neighboring sample. Furthermore, this can be applied to N samples in the upper region, the left region, or the upper left region.

[0368] [Formula 14]

[0369] Figure 20 =

[0370] E. The method for calculating the error (or cost) for each block can vary depending on the characteristics of the block. For example, the error can be calculated by considering the MR-SAD of a block with LIC applied. As another example, the error can be calculated by considering the MR-SAD of a block with BCW applied. As yet another example, the cost can be calculated by considering the SAD of a block with DMVR and BDOF applied.

[0371] Figure 20 A decoding method according to an embodiment of the present disclosure is illustrated. Figure 20 The steps or operations shown are not essential components of the prediction method and can be omitted. Figure 20 At least some of the steps or operations shown.

[0372] Reference Figure 21 The decoding device can obtain information about the basic prediction (S800).

[0373] For example, a basic prediction can be a series of processes used to generate basic prediction samples. A basic prediction can be any of intra-frame prediction, inter-frame prediction, or inter-block prediction.

[0374] The decoding processor of the decoding device can parse information about the basic predictions for the current block from the bitstream, inherit that information from neighboring blocks, or derive that information using specific tools.

[0375] Information about the basic prediction may include index information specifying the basic prediction and / or generation information used to generate the basic prediction block. For example, a decoding processor can parse the index information of the basic prediction and the generation information of the basic prediction block from the bitstream. As another example, a decoding processor can parse the index information of the basic prediction and inherit the generation information of the basic prediction block from adjacent blocks based on the index information. As yet another example, a decoding processor can parse the index information of the basic prediction and use tools to deduce the generation information of the basic prediction block based on the index information.

[0376] The decoding device can generate basic prediction blocks (S810).

[0377] For example, the decoding processor can generate basic prediction blocks based on information about the basic predictions. The decoding processor can generate basic prediction blocks based on the generation information of parsed basic prediction blocks or based on the generation information of inherited basic prediction blocks. Furthermore, the decoding processor can generate basic prediction blocks based on the generation information of deduced basic prediction blocks.

[0378] The decoding device can obtain information about additional predictions (S820).

[0379] For example, the decoding processor of the decoding device can obtain information about additional predictions. In a broad sense, additional prediction can refer to the prediction process other than the basic prediction, and in a narrow sense, it can refer to a series of processes used to generate additional prediction samples.

[0380] Information about additional predictions may include additional prediction availability information indicating whether additional predictions should be performed, prediction mode index information indicating the additional prediction mode, and additional block generation information used to generate additional prediction blocks.

[0381] The decoding processor can obtain prediction mode index information indicating the additional prediction mode based on the additional prediction availability information, and can obtain block generation information for generating additional prediction blocks based on the prediction mode index information. The decoding processor can parse the additional prediction availability information, prediction mode index information, and additional block generation information from the bitstream.

[0382] The decoding processor can parse additional block generation information for generating additional prediction blocks based on satisfying predetermined conditions. The predetermined conditions may include at least one of the following: (i) whether additional prediction modes are allowed for the basic prediction mode, (ii) whether additional prediction modes are allowed for the tools applied in the basic prediction mode, (iii) whether additional prediction modes are allowed for the width or height of the current block, (iv) whether additional prediction modes are allowed based on the quantization parameters of the current block, and (v) whether additional prediction modes are allowed based on the temporal layer of the current block.

[0383] The decoding processor can inherit block generation information used to generate additional prediction blocks based on additional prediction availability information. For example, the decoding processor can inherit additional block generation information for neighboring blocks that have the same basic prediction pattern as the current block. The decoding processor can derive additional prediction patterns for neighboring blocks that have the same basic prediction pattern as the current block, and inherit additional block generation information for the additional prediction patterns of neighboring blocks selected based on the priority of inheritance for additional prediction patterns. The decoding processor can derive additional prediction patterns for neighboring blocks that have the same basic prediction pattern as the current block, and inherit additional block generation information for the additional prediction patterns of neighboring blocks allowed based on the basic prediction pattern of the current block.

[0384] The decoding processor can derive block generation information for generating additional prediction blocks based on additional prediction availability information. The decoding processor can derive additional prediction performance information for performing additional predictions using the TM method. The decoding processor can derive additional prediction performance information for performing additional predictions using at least one of the TIMD or DIMD methods.

[0385] Furthermore, additional prediction can be one or more of intra-frame prediction, inter-frame prediction, or IBC. In other words, multiple additional predictions can be applied. For example, a first additional prediction, a second additional prediction, a third additional prediction, ... and an nth additional prediction can be applied. The decoding processor can obtain information about the first additional prediction and information about the second additional prediction.

[0386] The decoding device can generate additional prediction blocks (S830).

[0387] For example, the decoding processor can generate additional prediction blocks based on information about additional predictions. The decoding processor can generate additional prediction blocks based on the generation information of parsed additional prediction blocks, or based on the generation information of inherited additional prediction blocks. Furthermore, the decoding processor can generate additional prediction blocks based on the generation information of deduced additional prediction blocks.

[0388] First additional prediction, second additional prediction, third additional prediction, ... and nth additional prediction can be applied, and the decoding processor can generate first additional prediction blocks and second additional prediction blocks.

[0389] The decoding device can generate a final prediction block for the current block based on the merging of the basic prediction block and the additional prediction block (S840).

[0390] For example, a decoding processor can generate a final prediction block for the current block by weighted summation or weighted average of the basic prediction block and the additional prediction blocks.

[0391] First additional prediction, second additional prediction, third additional prediction, ... and nth additional prediction can be applied, and the decoding processor can generate the final prediction block for the current block by weighted summation or weighted average of the basic prediction block, the first additional prediction block and the second additional prediction block.

[0392] Figure 21 An example of an encoding method according to an embodiment of the present disclosure is illustrated. Figure 21 The steps or operations shown are not essential components of the prediction method and can be omitted. Figure 22 At least some of the steps or operations shown.

[0393] The encoding device can encode information about the basic prediction used to generate the basic prediction block (S900).

[0394] For example, the encoding processor of the encoding device can encode information about the basic predictions used to generate the basic prediction block by performing the basic predictions.

[0395] The encoding device can encode information about additional predictions used to generate additional prediction blocks (S910).

[0396] For example, the encoding processor can encode information about additional predictions used to generate additional prediction blocks by performing additional predictions.

[0397] The encoding device can encode the merging information used to generate the final predicted block for the current block (S920).

[0398] For example, the encoding processor can encode the merging information used to generate the final prediction block for the current block based on the merging of the basic prediction block and the additional prediction blocks.

[0399] Figure 22 An example of a content streaming system to which embodiments of the present disclosure can be applied is shown.

[0400] Reference ​The content streaming system using the embodiments of this disclosure may mainly include an encoding server, a streaming server, a web server, media storage, user equipment, and multimedia input devices.

[0401] An encoding server generates a bitstream by compressing content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data, and then sends it to a streaming server. As another example, when multimedia input devices such as smartphones, cameras, and camcorders generate bitstreams directly, the encoding server can be omitted.

[0402] Bitstreams can be generated by applying the encoding method or bitstream generation method of the embodiments of this disclosure, and the streaming server can temporarily store the bitstream during the sending or receiving of the bitstream.

[0403] A streaming server sends multimedia data to a user's device via a web server based on the user's request, and the web server acts as a medium to notify the user what services are available. When a user requests a service from the web server, the web server delivers it to the streaming server, and the streaming server sends the multimedia data to the user. In this scenario, the content streaming system may include a separate control server, which in this case controls the commands / responses between each device in the content streaming system.

[0404] A streaming server can receive content from media storage and / or encoding servers. For example, when receiving content from an encoding server, the content can be received in real time. In this case, to provide a smooth streaming service, the streaming server can store bitstreams for specific time periods.

[0405] Examples of user equipment may include mobile phones, smartphones, laptops, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, tablet PCs, tablet computers, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head-mounted displays (HMDs), digital TVs, desktop computers, digital signage, etc.).

[0406] In a content streaming system, each server can be operated as a distributed server, and in this case, data received from each server can be distributed and processed.

[0407] The claims set forth herein can be combined in various ways. For example, the technical features of the method claims of this disclosure can be combined and implemented as a device, and the technical features of the device claims of this disclosure can be combined and implemented as a method. Furthermore, the technical features of the method claims and the device claims of this disclosure can be combined and implemented as a device, and the technical features of the method claims and the device claims of this disclosure can be combined and implemented as a method.

[0408] Industrial applicability

[0409] The embodiments disclosed herein can be used to encode or decode images.

Claims

1. A decoding method, the method comprising the following steps: Obtain information about routine forecasts; A regular prediction block is generated by performing the regular prediction based on the information about the regular prediction; Obtain information about additional forecasts; as well as An additional prediction block is generated by performing the additional prediction based on the information about the additional prediction; A final prediction block for the current block is generated based on a weighted average of the regular prediction block and the additional prediction block. The conventional prediction includes any one of intra-frame prediction, inter-frame prediction, or intra-block copy (IBC), and The additional prediction includes at least one of the intra-frame prediction, the inter-frame prediction, or the IBC.

2. The method according to claim 1, wherein, The steps to obtain the information about the additional predictions include: Obtain additional prediction availability information indicating whether to perform the additional prediction. Based on the additional prediction availability information, prediction mode index information indicating the additional prediction mode is obtained, and Based on the prediction mode index information, additional block generation information is obtained for generating the additional prediction blocks. The additional block generation information includes at least one of intra-frame block generation information, inter-frame block generation information, or IBC block generation information.

3. The method according to claim 2, wherein, The steps of obtaining the additional block generation information used to generate the additional prediction block include: Based on satisfying preset conditions, the additional block generation information used to generate the additional prediction block is obtained. The preset conditions include at least one of the following: For the standard prediction model, is the additional prediction model allowed? For tools applied to the regular prediction pattern, is the additional prediction pattern allowed? For the width or height of the current block, is the additional prediction mode allowed? For the quantization parameter QP of the current block, is the additional prediction mode allowed, or For the time layer of the current block, is the additional prediction mode allowed? 4. The method according to claim 1, wherein, The steps to obtain the information about the additional predictions include: Obtain additional prediction availability information indicating whether to perform the additional prediction. Based on the additional prediction availability information, the additional block generation information used to generate the additional prediction block is inherited, and The additional block generation information includes at least one of intra-frame block generation information, inter-frame block generation information, or IBC block generation information.

5. The method according to claim 4, wherein, The step of inheriting the additional block generation information used to generate the additional prediction block includes: Inheriting additional block generation information from adjacent blocks that have the same regular prediction pattern as the regular prediction pattern for the current block.

6. The method according to claim 4, wherein, The step of inheriting the additional block generation information used to generate the additional prediction block includes: Derive additional prediction patterns for adjacent blocks that have the same regular prediction pattern as the regular prediction pattern for the current block, and The additional block generation information is inherited from the additional prediction pattern of the adjacent blocks selected based on the inheritance priority of the additional prediction pattern.

7. The method according to claim 4, wherein, The step of inheriting the additional block generation information used to generate the additional prediction block includes: Derive additional prediction patterns for adjacent blocks that have the same regular prediction pattern as the regular prediction pattern for the current block, and The additional block generation information is inherited from the additional prediction patterns of the adjacent blocks, which are allowed based on the regular prediction pattern for the current block.

8. The method according to claim 1, wherein, The steps to obtain the information about the additional predictions include: Obtain additional prediction availability information indicating whether to perform the additional prediction. Based on the additional prediction availability information, additional block generation information for generating the additional prediction block is derived, and The additional block generation information includes at least one of intra-frame block generation information, inter-frame block generation information, or IBC block generation information.

9. The method according to claim 8, wherein, The steps of deriving the additional prediction execution information used to perform the additional prediction include: The additional prediction execution information is derived using a template matching method.

10. The method according to claim 8, wherein, The steps of deriving the additional prediction execution information used to perform the additional prediction include: The additional prediction execution information is derived using at least one of a template-based intra-frame mode derivation (TIMD) method or a directional intra-frame mode derivation (DIMD) method.

11. The method according to claim 1, wherein, The information regarding the additional forecasts includes information regarding the first additional forecast and information regarding the second additional forecast. The additional prediction block includes a first additional prediction block and a second additional prediction block, and The step of generating the final prediction block includes: generating the final prediction block based on a weighted average of the regular prediction block, the first additional prediction block, and the second additional prediction block.

12. A method of encoding, the method comprising the following steps: Information about the regular predictions used to generate a regular prediction block by performing regular predictions is encoded. Information about the additional predictions used to generate the additional prediction block by performing additional predictions is encoded; as well as Information about the weighted average used to generate the final prediction block for the current block based on the weighted average of the regular prediction block and the additional prediction block is encoded. The conventional prediction includes any one of intra-frame prediction, inter-frame prediction, or intra-block copy (IBC), and The additional prediction includes at least one of the intra-frame prediction, the inter-frame prediction, or the IBC.

13. A method for sending image information, the method comprising the following steps: Information about the regular predictions used to generate a regular prediction block by performing regular predictions is encoded. Information about the additional predictions used to generate the additional prediction block by performing additional predictions is encoded; Information about the weighted average used to generate a final prediction block for the current block based on the weighted average of the regular prediction block and the additional prediction block is encoded. as well as The image information is transmitted, including encoded information about the conventional prediction, encoded information about the additional prediction, and encoded information about the weighted average. The conventional prediction includes any one of intra-frame prediction, inter-frame prediction, or intra-block copy (IBC), and The additional prediction includes at least one of the intra-frame prediction, the inter-frame prediction, or the IBC.