Image encoding / decoding method, recording medium storing bitstream thereon, and method for transmitting bitstream

By generating a predicted block of the current block and a weighted sum of multiple reference blocks, combined with filter strength and inverse transform techniques, the problem of low efficiency in encoding and decoding high-resolution and high-quality images is solved, achieving a more efficient encoding and decoding process, and providing a solution for bitstream storage and transmission.

CN122029804APending Publication Date: 2026-05-12LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
LG ELECTRONICS INC
Filing Date
2024-09-13
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing image compression technologies are inefficient in encoding and decoding high-resolution and high-quality images, especially in the case of multiple reference blocks, and lack effective storage and transmission methods.

Method used

By generating a prediction block for the current block, deriving multiple reference blocks, and generating a final prediction block through a weighted sum of the regular prediction block and multiple reference blocks, the encoding and decoding efficiency is improved by combining filter strength and inverse transform techniques.

Benefits of technology

It improves the efficiency of image encoding and decoding, especially in the case of multiple reference blocks, enhances prediction performance, and provides an efficient method for bitstream storage and transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122029804A_ABST
    Figure CN122029804A_ABST
Patent Text Reader

Abstract

An image encoding / decoding method, a bit stream transmission method, and a computer-readable recording medium storing a bit stream are provided. An image decoding method according to the present disclosure comprises the steps of: generating a prediction block of a current block on the basis of a prediction mode of the current block; generating a reconstructed sample of the current block based on the prediction block; and performing filtering on the reconstructed sample, in which the step of generating the prediction block comprises the steps of: deriving a basic prediction block of the current block; deriving a plurality of reference blocks of the current block; and generating a final prediction block of the current block by weighted summation of the basic prediction block and the multiple reference blocks, and determining a filter strength for filtering based on prediction information related to the current block, the prediction block, or the multiple reference blocks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to an image encoding / decoding method, a recording medium having a bit stream stored thereon, and a method for transmitting the bit stream, and more specifically, to an encoding / decoding process associated with multiple reference blocks. Background Technology

[0002] Recently, the demand for high-resolution and high-quality images, such as HD (high-definition) and UHD (ultra-high-definition) images, has been increasing in various application fields, and therefore, efficient image compression technology is being discussed.

[0003] There are various techniques, such as inter-frame prediction techniques that use video compression to predict the pixel values ​​included in the current image from images before or after the current image, intra-frame prediction techniques that use pixel information in the current image to predict the pixel values ​​included in the current image, and entropy coding techniques that assign short symbols to values ​​that occur frequently and long symbols to values ​​that occur infrequently. These image compression techniques can be used to effectively compress image data and transmit or store it.

[0004] Therefore, efficient image compression techniques are needed to effectively transmit, store, and reproduce information about high-resolution and high-quality images. Summary of the Invention

[0005] [Technical Issues]

[0006] This disclosure aims to provide an image encoding method and apparatus with improved encoding / decoding efficiency.

[0007] In addition to the base reference block exported during the prediction process, this disclosure also aims to improve prediction performance when the reference block is exported.

[0008] This disclosure also aims to improve the efficiency of the encoding / decoding process when multiple reference blocks are present.

[0009] This disclosure also aims to provide a non-transitory computer-readable recording medium for storing bitstreams generated using the image encoding method according to this disclosure.

[0010] This disclosure also aims to provide a method for transmitting a bitstream generated using an image encoding method according to this disclosure.

[0011] The technical objectives to be achieved by this disclosure are not limited to those described above, and other technical objectives not described herein will be clearly understood by those skilled in the art to which this disclosure pertains based on the following description.

[0012] [Technical Solution]

[0013] According to embodiments of this disclosure, an image decoding method executed by a decoding device includes: generating a prediction block of the current block based on a prediction mode of the current block; generating a reconstruction sample of the current block based on the prediction block; and performing filtering on the reconstruction sample, wherein generating the prediction block includes: deriving a regular prediction block of the current block; deriving a plurality of reference blocks of the current block; and generating a final prediction block of the current block by a weighted sum of the regular prediction block and the plurality of reference blocks, wherein the filter strength for filtering is determined based on prediction information of the current block, the prediction block, or the plurality of reference blocks.

[0014] According to embodiments of this disclosure, a method for decoding an image performed by a decoding device includes: generating a prediction block for a current block based on prediction information obtained from a bitstream; deriving transform coefficients by performing dequantization based on residual information obtained from the bitstream; deriving a residual block by performing an inverse transform of the transform coefficients; and generating a reconstruction sample based on the prediction block and the residual block, wherein generating the prediction block includes generating a regular prediction block for the current block; deriving a plurality of reference blocks for the current block; and deriving a final prediction block for the current block by a weighted sum of the regular prediction block and the plurality of reference blocks, wherein determining whether to apply a subblock transform (SBT) to the current block is based on the prediction patterns of the plurality of reference blocks.

[0015] According to embodiments of this disclosure, a method for decoding an image performed by a decoding device includes: generating a prediction block for a current block based on prediction information obtained from a bitstream; deriving transform coefficients by performing dequantization based on residual information obtained from the bitstream; deriving a residual block by performing an inverse transform of the transform coefficients; and generating reconstructed samples based on the prediction block and the residual block, wherein generating the prediction block includes: generating a regular prediction block for the current block; deriving a plurality of reference blocks for the current block; and deriving a final prediction block for the current block by a weighted sum of the regular prediction block and the plurality of reference blocks, wherein a transform kernel for the inverse transform is determined based on a prediction mode associated with at least one of the regular prediction block and the plurality of reference blocks.

[0016] According to embodiments of this disclosure, an image encoding method executed by an encoding device includes: generating a prediction block of the current block based on a prediction mode of the current block; generating a reconstruction sample of the current block based on the prediction block; and performing filtering on the reconstruction sample, wherein generating the prediction block includes: deriving a regular prediction block of the current block; deriving a plurality of reference blocks of the current block; and generating a final prediction block of the current block by a weighted sum of the regular prediction block and the plurality of reference blocks, wherein the filter strength for filtering is determined based on prediction information of the current block, the prediction block, or the plurality of reference blocks.

[0017] According to embodiments of this disclosure, a computer-readable digital storage medium is provided for storing a bitstream generated using an image encoding method or apparatus.

[0018] According to embodiments of this disclosure, a method for transmitting image data includes transmitting a bitstream generated using an image encoding method or apparatus.

[0019] The features briefly outlined above are merely illustrative aspects of the detailed description of this disclosure that follows, and do not limit the scope of this disclosure.

[0020] [Beneficial Effects]

[0021] According to this disclosure, an image encoding / decoding method and apparatus with improved encoding / decoding efficiency can be provided.

[0022] According to this disclosure, in addition to the basic reference block exported during the prediction process, prediction performance can be improved when exporting the reference block.

[0023] According to this disclosure, the efficiency of the encoding / decoding process can also be improved when multiple reference blocks are used for the current block.

[0024] According to this disclosure, a non-transitory computer-readable recording medium for storing bitstreams generated using the image encoding method according to this disclosure can also be provided.

[0025] According to this disclosure, a non-transitory computer-readable recording medium can also be provided for storing a bitstream received and decoded by an image decoding apparatus according to this disclosure and used for image reconstruction.

[0026] According to this disclosure, a method for transmitting a bitstream generated using an image encoding method can also be provided.

[0027] The effects of this disclosure are not limited to those described above, and other effects not yet described will be clearly understood by those skilled in the art to which this disclosure pertains based on the following description. Attached Figure Description

[0028] Figure 1 A video / image encoding system according to this disclosure is shown.

[0029] Figure 2 A rough block diagram of an encoding apparatus that can be applied to embodiments of the present disclosure and perform encoding of video / image signals is shown.

[0030] Figure 3 A rough block diagram of a decoding apparatus that can be applied to embodiments of the present disclosure and perform decoding of video / image signals is shown.

[0031] Figure 4 This is a diagram illustrating the Low Frequency Inseparable Transform (LFNST) process.

[0032] Figure 5This is a flowchart illustrating an inter-frame prediction-based video / image coding method applicable to embodiments of the present disclosure.

[0033] Figure 6 This is a flowchart illustrating a video / image decoding method based on inter-frame prediction applicable to embodiments of the present disclosure.

[0034] Figure 7 This is a flowchart illustrating an inter-frame prediction process to which embodiments of the present disclosure can be applied.

[0035] Figure 8 This is a flowchart illustrating an inter-frame prediction method performed by a decoding device according to an embodiment of the present disclosure.

[0036] Figure 9 This is a diagram illustrating an example of a reference block used in a multi-reference block mode according to an embodiment of the present disclosure.

[0037] Figure 10 This is a flowchart illustrating a method for sending / parsing motion information of added prediction blocks according to an embodiment of the present disclosure when sending / parsing motion information of multiple reference blocks using signals.

[0038] Figure 11 An example of a geometric partitioning pattern (GPM) partitioning grouped by the same corner is shown.

[0039] Figure 12 An example of generating mixed weights w0 using GPM is shown.

[0040] Figure 13 This is a diagram illustrating the top and left neighbor blocks used in the combined inter-frame and intra-frame prediction (CIIP) weight derivation.

[0041] Figure 14 This is a set of diagrams illustrating the partitioning method used in corner mode.

[0042] Figure 15 This is a set of diagrams illustrating the available intra-prediction mode (IPM) candidates in a GPM with inter-frame and intra-frame prediction.

[0043] Figure 16 This is a diagram illustrating the coding unit (CU) / prediction unit (PU) boundaries and sub-blocks at sub-PUs in an Advanced Temporal Motion Vector Prediction (ATMVP) pattern.

[0044] Figure 17 This is a flowchart illustrating a method for resolving a CU according to an embodiment, depending on whether multiple reference blocks are applied and whether the resolution is based on their transmission transformation and residual signal information.

[0045] Figure 18This is a flowchart illustrating a method for resolving a CU based on whether a multiple reference block is applied and whether it meets specific conditions, according to an embodiment.

[0046] Figure 19 This is a flowchart illustrating a method for resolving a CU based on whether certain conditions are met when applying multiple reference blocks, according to an embodiment.

[0047] Figure 20 This is a flowchart illustrating a method for determining boundary strength (BS) when multiple reference blocks are applied according to an embodiment.

[0048] Figure 21 This is a diagram illustrating examples of template regions for conventional reference blocks and template regions for multiple reference blocks according to an embodiment.

[0049] Figure 22 Examples of content streaming systems to which embodiments of the present disclosure may be applied are shown.

[0050] These are examples of embodiments of the present disclosure that can be applied to content streaming systems. Detailed Implementation

[0051] Because this disclosure can be modified in various ways and has several embodiments, specific embodiments will be illustrated in the accompanying drawings and described in detail in the detailed description. However, this disclosure is not intended to be limited to the specific embodiments and should be understood to include all variations, equivalents, and substitutions included within the spirit and scope of this disclosure. Similar reference numerals are used for similar components in the description of each drawing.

[0052] Terms such as "first," "second," etc., may be used to describe various components, but components should not be limited by these terms. These terms are used only to distinguish one component from other components. For example, without departing from the scope of this disclosure, a first component may be referred to as a second component, and similarly, a second component may be referred to as a first component. Terms and / or combinations of any one or more related statement items are included.

[0053] When a component is described as "connected" or "linked" to another component, it should be understood that it can be directly connected or linked to another component, but there may also be another component in between. On the other hand, when a component is described as "directly connected" or "directly linked" to another component, it should be understood that there is no other component in between.

[0054] The terminology used in this application is for describing particular embodiments only and is not intended to limit this disclosure. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this application, it should be understood that terms such as “comprising” or “having” are intended to designate the presence of features, numbers, steps, operations, components, portions, or combinations thereof described in the specification, but do not preclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, portions, or combinations thereof.

[0055] This disclosure relates to video / image coding. For example, the methods / exercises disclosed herein can be applied to methods disclosed in the Universal Video Coding (VVC) standard. Additionally, the methods / exercises disclosed herein can be applied to methods disclosed in the Basic Video Coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the second-generation Audio Video Coding (AVS2) standard, or next-generation video / image coding standards (e.g., H.267 or H.268).

[0056] This specification presents various embodiments of video / image encoding, and unless otherwise stated, these embodiments may be combined with each other to perform the task.

[0057] Here, video can refer to a collection of images over time. An image generally refers to a unit representing an image within a specific time period, and a slice / tile is a unit that forms part of an image during encoding. A slice / tile can include at least one Code Tree Unit (CTU). An image can consist of at least one slice / tile. A tile is a rectangular area consisting of multiple CTUs within a specific tile column and a specific tile row of an image. A tile column is a rectangular area of ​​CTUs with the same height as the image and a width assigned by the syntax requirements of the image parameter set. A tile row is a rectangular area of ​​CTUs with the same height assigned by the image parameter set and a width equal to the width of the image. CTUs within a tile can be arranged consecutively according to a CTU raster scan, and tiles within an image can be arranged consecutively according to a tile raster scan. A slice can include an integer number of complete tiles or an integer number of consecutive complete CTU rows that can be exclusively included within a single NAL unit of an image. Simultaneously, an image can be divided into at least two sub-images. A sub-image can be a rectangular area of ​​at least one slice within an image.

[0058] A pixel, cell, or pixel unit can refer to the smallest unit that makes up a picture (or image). Additionally, "sample" can be used as the term corresponding to a pixel. A sample can typically represent a pixel or pixel value, and can represent only the pixel / pixel value of the luminance component, or only the pixel / pixel value of the chrominance component.

[0059] A unit can represent a basic unit of image processing. A unit may include a specific region of an image and at least one of the information associated with that region. A unit may include a luminance block and two chrominance (e.g., cb, cr) blocks. In some cases, units may be used interchangeably with terms such as block or region. In general, an MxN block may include a set (or array) of transform coefficients or samples (or sample arrays) consisting of M columns and N rows.

[0060] Here, "A or B" can refer to "A only", "B only", or "both A and B". In other words, "A or B" can be interpreted as "A and / or B". For example, "A, B or C" can refer to "A only", "B only", "C only", or "any combination of A, B and C".

[0061] The forward slash ( / ) or comma used in this article can refer to "and / or". For example, "A / B" can refer to "A and / or B". Therefore, "A / B" can refer to "A only", "B only", or "both A and B". For example, "A, B, C" can refer to "A, B, or C".

[0062] Here, "at least one of A and B" can refer to "only A", "only B" or "both A and B". Furthermore, expressions such as "at least one of A or B" or "at least one of A and / or B" can be interpreted in the same way as "at least one of A and B".

[0063] Additionally, here, "at least one of A, B, and C" can refer to "A only", "B only", "C only" or "any combination of A, B, and C". Furthermore, "at least one of A, B, or C" or "at least one of A, B, and / or C" can refer to "at least one of A, B, and C".

[0064] Additionally, the parentheses used in this document can refer to "for example". Specifically, when the indication is "prediction (intra-frame prediction)", "intra-frame prediction" can be cited as an example of "prediction". In other words, "prediction" here is not limited to "intra-frame prediction", and "intra-frame prediction" can be cited as an example of "prediction". Furthermore, even when the indication is "prediction (i.e., intra-frame prediction)", "intra-frame prediction" can be cited as an example of "prediction".

[0065] Here, a technical feature described individually in a single figure can be implemented individually or simultaneously.

[0066] Figure 1 A video / image encoding system according to this disclosure is shown.

[0067] refer to Figure 1A video / image encoding system may include a first device (source device) and a second device (receiving device).

[0068] A source device can transmit encoded video / image information or data to a receiving device in the form of a file or stream via digital storage media or a network. The source device may include a video source, an encoding device, and a transmitting unit. The receiving device may include a receiving unit, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, and the display unit may consist of a separate device or external components.

[0069] A video source can acquire video / images through the process of capturing, compositing, or generating video / images. A video source can include devices for capturing video / images and devices for generating video / images. Devices for capturing video / images can include at least one camera, video / image archives containing previously captured video / images, etc. Devices for generating video / images can include computers, tablets, smartphones, etc., and can generate video / images (electronically). For example, virtual video / images can be generated by computers, etc., and in this case, the process of capturing video / images can be replaced by the process of generating related data.

[0070] Encoding devices can encode input video / images. They can perform a series of processes such as prediction, transformation, and quantization for compression and encoding efficiency. The encoded data (encoded video / image information) can be output as a bitstream.

[0071] The transmitting unit can send encoded video / image information or data, output as a bitstream, to the receiving unit of the receiving device in the form of a file or stream, via digital storage media or a network. Digital storage media can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit can include elements for generating media files according to a predetermined file format and may include elements for transmission over a broadcast / communication network. The receiving unit can receive / extract the bitstream and send it to a decoding device.

[0072] Decoding devices can decode video / images by performing a series of processes, such as dequantization, inverse transform, and prediction, that correspond to the operations of encoding devices.

[0073] The renderer can render decoded video / images. The rendered video / images can be displayed through a display unit.

[0074] Figure 2A rough block diagram of an encoding apparatus that can be applied to embodiments of the present disclosure and perform encoding of video / image signals is shown.

[0075] refer to Figure 2 The encoding device 200 may consist of an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstruction block generator. According to an embodiment, the image partitioner 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 may be configured by at least one hardware component (e.g., an encoder chipset or processor). Additionally, the memory 270 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory 270 as an internal / external component.

[0076] Image partitioner 210 can partition an input image (or picture, frame) input to encoding device 200 into at least one processing unit. As an example, a processing unit can be called a coding unit (CU). In this case, coding units can be recursively partitioned from coding tree units (CTUs) or maximum coding units (LCUs) according to a quadtree-binary-tritree (QTBTTT) structure.

[0077] For example, a coding unit can be partitioned into multiple coding units of greater depth based on a quadtree structure, a binary tree structure, and / or a ternary structure. In this case, for example, a quadtree structure can be applied first, and a binary tree structure and / or a ternary structure can be applied later. Alternatively, a binary tree structure can be applied before the quadtree structure. The coding process according to this specification can be performed based on the final coding unit that is no longer partitioned. In this case, based on coding efficiency, etc., according to image characteristics, the largest coding unit can be directly used as the final coding unit, or if necessary, the coding unit can be recursively partitioned into deeper coding units, and the coding unit with the optimal size can be used as the final coding unit. Here, the coding process can include processes such as prediction, transformation, and reconstruction, as described later.

[0078] As another example, the processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may be partitioned or divided from the aforementioned final coding unit, respectively. The prediction unit may be a unit for predicting samples, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving residual signals from transform coefficients.

[0079] In some cases, a unit can be used interchangeably with terms such as block or region. Generally, an MxN block can represent a set of transform coefficients or samples consisting of M columns and N rows. Samples can typically represent pixels or pixel values, and can represent only the pixel / pixel value of the luminance component, or only the pixel / pixel value of the chrominance component. Samples can be used as a term to correspond a picture (or image) to pixels or cells.

[0080] Encoding device 200 can subtract the prediction signal (prediction block, prediction sample array) output from inter-frame predictor 221 or intra-frame predictor 222 from the input image signal (original block, original sample array) to generate a residual signal (residual signal, residual sample array), and the generated residual signal is sent to converter 232. In this case, the unit in encoding device 200 that subtracts the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) can be called subtractor 231.

[0081] Predictor 220 can perform prediction on the block to be processed (hereinafter referred to as the current block) and generate a block of predictions including prediction samples for the current block. Predictor 220 can determine whether to apply intra-frame prediction or inter-frame prediction on a block or CU basis. Predictor 220 can generate various information about the prediction, such as prediction mode information, and send it to entropy encoder 240, as described later in the description of each prediction mode. The information about the prediction can be encoded in entropy encoder 240 and output as a bitstream.

[0082] Intra-predictor 222 can predict the current block by referencing samples within the current image. Depending on the prediction mode, the referenced samples can be located near the current block or at a distance from it. In intra-prediction, the prediction mode can include at least one non-directional mode and multiple directional modes. The non-directional mode can include at least one of a DC mode or a planar mode. Depending on the level of detail of the prediction direction, the directional modes can include 33 or 65 directional modes. However, this is just an example, and more or fewer directional modes can be used depending on the configuration. Intra-predictor 222 can determine the prediction mode applied to the current block by using prediction modes applied to neighboring blocks.

[0083] Inter-frame predictor 221 can derive a prediction block for the current block based on a reference block (reference sample array) specified by motion vectors on a reference image. In this case, to reduce the amount of motion information transmitted in the inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation between motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference image indices. Motion information may further include inter-frame prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). For inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current image and temporally neighboring blocks existing in the reference image. The reference image including the reference block and the reference image including the temporally neighboring block may be the same or different. The temporally neighboring block may be referred to as a juxtaposed reference block, a juxtaposed CU (colCU), etc., and the reference image including the temporally neighboring block may be referred to as a juxtaposed image (colPic). For example, inter-frame predictor 221 can configure a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or reference image index of the current block. Inter-frame prediction can be performed based on various prediction modes, and for example, for skip mode and merge mode, the inter-frame predictor 221 can use the motion information of neighboring blocks as the motion information of the current block. For skip mode, unlike merge mode, residual signals may not be sent. For motion vector prediction (MVP) mode, the motion vectors of surrounding blocks are used as motion vector predictors, and the motion vector difference is signaled to indicate the motion vector of the current block.

[0084] Predictor 220 can generate a prediction signal based on various prediction methods described later. For example, the predictor can not only apply intra-frame prediction or inter-frame prediction to predict a block, but also apply both intra-frame prediction and inter-frame prediction simultaneously. This can be referred to as a combined intra-frame and inter-frame prediction (CIIP) mode. Alternatively, the predictor can be based on an intra-block copy (IBC) prediction mode or a palette mode for prediction against a block. The IBC prediction mode or palette mode can be used for content image / video coding such as screen content coding (SCC) in games, etc. IBC essentially performs prediction within the current image, but it can be performed similarly to inter-frame prediction because it derives a reference block within the current image. In other words, IBC can use at least one of the inter-frame prediction techniques described herein. A palette mode can be considered an example of intra-frame coding or intra-frame prediction. When a palette mode is applied, sample values ​​within the image can be signaled based on information about the palette table and palette index. The prediction signal generated by predictor 220 can be used to generate a reconstructed signal or a residual signal.

[0085] Transformer 232 can generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Karhunen-Loève Transform (KLT), Graphical Based Transform (GBT), or Conditional Nonlinear Transform (CNT). Here, GBT refers to the transform obtained from a graphic when the relationship information between pixels is expressed as a graphic. CNT refers to the transform obtained based on generating a prediction signal using all previously reconstructed pixels. Furthermore, the transform process can be applied to square pixel blocks of the same size or to non-square blocks of variable size.

[0086] Quantizer 233 can quantize the transform coefficients and send them to entropy encoder 240, which can encode the quantized signal (information about the quantized transform coefficients) and output it as a bitstream. The information about the quantized transform coefficients can be referred to as residual information. Quantizer 233 can rearrange the quantized transform coefficients in block form into a one-dimensional vector form based on the coefficient scan order, and can generate information about the quantized transform coefficients based on the one-dimensional vector form of the quantized transform coefficients.

[0087] The entropy encoder 240 can perform various encoding methods, such as exponential Columbus coding, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy encoder 240 can encode information necessary for video / video image reconstruction (e.g., values ​​of syntax elements, etc.) in addition to the transform coefficients quantized together or individually.

[0088] Encoded information (e.g., encoded video / image information) can be transmitted or stored in bitstream form at the network abstraction layer (NAL) unit level. The video / image information may further include information about various parameter sets such as adaptive parameter sets (APS), picture parameter sets (PPS), sequence parameter sets (SPS), or video parameter sets (VPS). Additionally, the video / image information may further include general constraint information. Here, information transmitted from the encoding device / signaled to the decoding device and / or syntax elements can be included in the video / image information. The video / image information can be encoded by the above-described encoding process and included in the bitstream. The bitstream can be transmitted over a network or stored in a digital storage medium. Here, the network may include broadcast networks and / or communication networks, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. Transmission units (not shown) for transmission and / or storage units (not shown) for storing signals output from the entropy encoder 240 can be configured as internal / external components of the encoding device 200, or the transmission unit may also be included in the entropy encoder 240.

[0089] The quantized transform coefficients output from quantizer 233 can be used to generate a prediction signal. For example, the residual signal (residual block or residual sample) can be reconstructed by applying dequantization and inverse transform to the quantized transform coefficients using dequantizer 234 and inverse transformer 235. Adder 250 can add the reconstructed residual signal to the prediction signal output from inter-frame predictor 221 or intra-frame predictor 222 to generate a reconstructed signal (reconstructed image, reconstructed block, reconstructed sample array). When there is no residual for the block to be processed, such as when a skip mode is applied, the prediction block can be used as a reconstructed block. Adder 250 can be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed within the current image, and can also be used for inter-frame prediction of the next image by filtering, which will be described later. Meanwhile, a luminance mapping with chroma scaling (LMCS) can be applied during image encoding and / or reconstruction.

[0090] Filter 260 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 260 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image, and the modified reconstructed image can be stored in memory 270, specifically in the DPB of memory 270. Various filtering methods can include deblocking filtering, sample adaptive shifting, adaptive loop filtering, bilateral filtering, etc. Filter 260 can generate various information about the filtering and send it to entropy encoder 240. The information about the filtering can be encoded in entropy encoder 240 and output as a bitstream.

[0091] The modified reconstructed image sent to memory 270 can be used as a reference image in inter-frame predictor 221. When inter-frame prediction is applied through it, the encoding device can avoid prediction mismatch in encoding device 200 and decoding device, and can also improve encoding efficiency.

[0092] The DPB of memory 270 can store modified reconstructed images for use as reference images in inter-frame predictor 221. Memory 270 can store motion information of blocks from which motion information in the current image is derived (or encoded) and / or motion information of blocks in the pre-reconstructed image. The stored motion information can be sent to inter-frame predictor 221 to be used as motion information for spatially or temporally neighboring blocks. Memory 270 can store reconstructed samples of reconstructed blocks in the current image and send them to intra-frame predictor 222.

[0093] Figure 3 A rough block diagram of a decoding apparatus that can be applied to embodiments of the present disclosure and perform decoding of video / image signals is shown.

[0094] refer to Figure 3 The decoding device 300 can be configured to include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-frame predictor 331 and an intra-frame predictor 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 321.

[0095] According to an embodiment, the entropy decoder 310, residual processor 320, predictor 330, adder 340, and filter 350 described above can be configured by a single hardware component (e.g., a decoder chipset or processor). Additionally, the memory 360 may include a decoded image buffer (DPB) and can be configured by a digital storage medium. The hardware component may further include the memory 360 as an internal / external component.

[0096] When the input includes a bitstream containing video / image information, the decoding device 300 can respond to... Figure 2 The process of processing video / image information in an encoding device reconstructs an image. For example, decoding device 300 can derive units / blocks based on relevant information about block partitions obtained from the bitstream. Decoding device 300 can perform decoding by using processing units applied in the encoding device. Therefore, the decoding processing unit can be an encoding unit, and the encoding unit can be partitioned from encoding tree units or maximum encoding units according to a quadtree structure, binary tree structure, and / or ternary tree structure. At least one transform unit can be derived from the encoding unit. Furthermore, the reconstructed image signal decoded and output by decoding device 300 can be played back by a playback device.

[0097] Decoding device 300 can receive data in bitstream form from... Figure 2 The signal output by the encoding device and the received signal can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can parse the bitstream to derive information (e.g., video / image information) necessary for image reconstruction (or picture reconstruction). The video / image information may further include information about various parameter sets such as adaptive parameter sets (APS), picture parameter sets (PPS), sequence parameter sets (SPS), or video parameter sets (VPS). In addition, the video / image information may further include general constraint information. The decoding device can further decode the picture based on the information about the parameter sets and / or the general constraint information. The information sent / received by the signal and / or the syntax elements described later herein can be decoded and obtained from the bitstream through the decoding process. For example, the entropy decoder 310 can decode the information in the bitstream based on encoding methods such as exponential Golomb coding, CAVLC, CABAC, etc., and output the values ​​of the syntax elements necessary for image reconstruction and the quantized values ​​of the transform coefficients of the residuals. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element from the bitstream, determine a context model using information about the syntax element to be decoded, decoding information of surrounding blocks and the block to be decoded, or information about symbols / bins decoded in the previous step, perform arithmetic decoding on the bins by predicting the occurrence probability of the bins based on the determined context model, and generate symbols corresponding to the value of each syntax element. In this case, after determining the context model, the CABAC entropy decoding method can update the context model by using information about the decoded symbols / bins for the context model used for the next symbol / bin. Among the information decoded in the entropy decoder 310, information about prediction is provided to the predictors (inter-frame predictor 332 and intra-frame predictor 331), and the residual values ​​of entropy decoding performed on them in the entropy decoder 310, i.e., the quantized transform coefficients and related parameter information, can be input to the residual processor 320. The residual processor 320 can derive residual signals (residual blocks, residual samples, residual sample arrays). In addition, information about filtering among the information decoded in the entropy decoder 310 can be provided to the filter 350. Meanwhile, the receiving unit (not shown) that receives the signal output from the encoding device can be further configured as an internal / external element of the decoding device 300 or the receiving unit can be a component of the entropy decoder 310.

[0098] Furthermore, the decoding device according to this specification can be referred to as a video / image / picture decoding device, and the decoding device can be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include an entropy decoder 310, and the sample decoder may include at least one of a dequantizer 321, an inverse transformer 322, an adder 340, a filter 350, a memory 360, an inter-frame predictor 332, and an intra-frame predictor 331.

[0099] Dequantizer 321 can dequantize the quantized transform coefficients and output the transform coefficients. Dequantizer 321 can rearrange the quantized transform coefficients into a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scan order performed in the encoding device. Dequantizer 321 can obtain the transform coefficients by performing dequantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information).

[0100] The inverse transformer 322 performs an inverse transformation on the transformation coefficients to obtain the residual signal (residual block, residual sample array).

[0101] Predictor 320 can perform prediction on the current block and generate a prediction block including prediction samples for the current block. Predictor 320 can determine whether to apply intra-frame prediction or inter-frame prediction to the current block based on the prediction information output from entropy decoder 310, and determine a specific intra-frame / inter-frame prediction mode.

[0102] Predictor 320 can generate prediction signals based on various prediction methods described later. For example, predictor 320 can not only apply intra-frame prediction or inter-frame prediction to predict a block, but also apply intra-frame prediction and inter-frame prediction simultaneously. This can be referred to as a combined intra-frame and inter-frame prediction (CIIP) mode. Alternatively, the predictor can be based on an intra-block copy (IBC) prediction mode or a palette mode for block prediction. The IBC prediction mode or palette mode can be used for content image / video coding such as screen content coding (SCC) in games, etc. IBC essentially performs prediction within the current frame, but it can be performed similarly to inter-frame prediction because it derives a reference block within the current frame. In other words, IBC can use at least one of the inter-frame prediction techniques described herein. Palette mode can be considered an example of intra-frame coding or intra-frame prediction. When a palette mode is applied, information about the palette table and palette index can be included in the video / image information and transmitted as a signal.

[0103] Intra-predictor 331 can predict the current block by referencing samples within the current image. Depending on the prediction mode, the referenced samples can be located near the current block or at a certain distance away from the current block. In intra-prediction, the prediction mode can include at least one non-directional mode and multiple directional modes. Intra-predictor 331 can determine the prediction mode applied to the current block by using prediction modes applied to neighboring blocks.

[0104] Inter-frame predictor 332 can derive a prediction block for the current block based on a reference block (reference sample array) specified by motion vectors on a reference image. In this case, to reduce the amount of motion information transmitted in the inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation between motion information of neighboring blocks and the current block. Motion information may include motion vectors and a reference image index. Motion information may further include inter-frame prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). For inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current image and temporally neighboring blocks existing in the reference image. For example, inter-frame predictor 332 can configure a motion information candidate list based on neighboring blocks and derive the motion vector and / or reference image index of the current block based on the received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and information about the prediction may include information indicating the inter-frame prediction mode used for the current block.

[0105] Adder 340 can add the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor (including inter-frame predictor 332 and / or intra-frame predictor 331) to generate a reconstruction signal (reconstructed image, reconstruction block, reconstruction sample array). When there is no residual for the block to be processed, such as when a skip mode is applied, the prediction block can be used as the reconstruction block.

[0106] Adder 340 can be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current image, output by filtering as described later, or it can be used for inter-frame prediction of the next image. Meanwhile, a luminance map with chroma scaling (LMCS) can be applied during image decoding.

[0107] Filter 350 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 350 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image, and send the modified reconstructed image to memory 360, specifically the DPB of memory 360. Various filtering methods can include deblocking filtering, adaptive sampling offset, adaptive loop filtering, bilateral filtering, etc.

[0108] The (modified) reconstructed image stored in the DPB of memory 360 can be used as a reference image in inter-frame prediction unit 332. Memory 360 can store motion information of blocks derived (or decoded) from motion information in its current image and / or motion information of blocks in the pre-reconstructed image. The stored motion information can be sent to inter-frame predictor 260 as motion information for spatially or temporally neighboring blocks. Memory 360 can store reconstructed samples of reconstructed blocks in the current image and send them to intra-frame predictor 331.

[0109] Here, the embodiments described in the filter 260, inter-frame predictor 221 and intra-frame predictor 222 of the encoding device 200 can also be applied equally or correspondingly to the filter 350, inter-frame predictor 332 and intra-frame predictor 331 of the decoding device 300.

[0110] The transformation / inverse transformation will be described in detail below.

[0111] Encoding device 200 can derive residual blocks (residual samples) based on blocks predicted via intra / inter / intra-block copying (IBC) prediction, and derive quantized transform coefficients by applying transform and quantization to the derived residual samples. Information about the quantized transform coefficients (residual information) can be included in the residual coding syntax, encoded, and then output as a bitstream. Decoding device 300 can obtain information about the quantized transform coefficients (residual information) from the bitstream and derive the quantized transform coefficients by decoding this information. Decoding device 300 can derive residual samples based on the quantized transform coefficients through dequantization / inverse transform. As mentioned above, at least one of quantization / dequantization and / or transform / inverse transform can be omitted. When transform / inverse transform is omitted, transform coefficients can also be referred to as "coefficients" or "residual coefficients," or, for consistency of expression, they can still be referred to as "transform coefficients." Whether transform / inverse transform is omitted can be indicated by a signal based on transform_skip_flag.

[0112] Transform / inverse transforms can be performed based on transform kernels. For example, according to this document, a multiple transform selection (MTS) scheme can be applied. In this case, some from a set of multiple transform kernels can be selected and applied to the current block. Transform kernels can be referred to by various terms, such as "transformation matrix," "transformation type," etc. For example, a set of transform kernels can be a combination of vertical transform kernels (vertical transform kernels) and horizontal transform kernels (horizontal transform kernels).

[0113] For example, MTS index information (or syntax element “tu_mts_idx”) can be generated / encoded by encoding device 200 to indicate one of the transform kernel sets and sent to decoding device 300 by signal. For example, the transform kernel set based on the values ​​of the MTS index information can be derived, as shown in Table 1 below.

[0114] [Table 1]

[0115] Table 1 shows the descriptions of trTypeHor and trTypeVer based on tu_mtx_idx[x][y].

[0116] The transformation kernel set can be determined based on, for example, cu_sbt_horizontal_flag and cu_sbt_pos_flag.

[0117] A value of 1 for `cu_sbt_horizontal_flag` indicates that the current CU is horizontally divided into two TUs. A value of 0 for `cu_sbt_horizontal_flag[x0][y0]` indicates that the current CU is vertically divided into two TUs. A value of 1 for `cu_sbt_pos_flag` indicates that `tu_cbf_luma`, `tu_cbf_cb`, and `tu_cbf_cr` of the first TU in the current CU do not exist in the bitstream. A value of 0 for `cu_sbt_pos_flag` indicates that `tu_cbf_luma`, `tu_cbf_cb`, and `tu_cbf_cr` of the second TU in the current CU do not exist in the bitstream.

[0118] [Table 2]

[0119] Table 2 shows the descriptions of trTypeHor and trTypeVer based on cu_sbt_horizontal_flag and cu_sbt_pos_flag.

[0120] The transform kernel set can be determined based on, for example, the intra-prediction mode (IPM) of the current block.

[0121] [Table 3]

[0122] In the tables (Tables 1 to 3), trTypeHor can indicate the horizontal transform kernel, and trTypeVer can indicate the vertical transform kernel. Here, a trTypeHor / trTypeVer value of 0 indicates DCT2, a trTypeHor / trTypeVer value of 1 indicates DCT7, and a trTypeHor / trTypeVer value of 2 indicates DCT8. However, these are illustrative, and other values ​​can be mapped to other DCTs / DSTs according to predetermined rules.

[0123] Table 4 below shows examples of the basis functions for the aforementioned DCT2, DCT8, and DST7.

[0124] [Table 4]

[0125] In this document, an MTP-based transform can be applied as the primary transform, and a secondary transform can be further applied. The secondary transform can be applied only to the coefficients in the upper left w×h region of the coefficient block to which the primary transform has been applied, and can be called a “reduced secondary transform (RST)”. For example, w and / or h can be 4 or 8. In the case of the transform, the primary and secondary transforms can be applied sequentially to the residual block, and in the case of the inverse transform, the inverse secondary transform and the inverse primary transform can be applied sequentially to the transform coefficients. The secondary transform (RST transform) can be called a “low-frequency coefficient transform (LFCT)” or a “low-frequency inseparable transform (LFNST)”. The inverse secondary transform can be called an “inverse LFCT” or an “inverse LFNST”.

[0126] Figure 4 This is a diagram illustrating the LFNST process.

[0127] like Figure 4 As shown, LFNST is referred to as RST applied between the forward primary transform and quantization (in the encoder) and between dequantization and the inverse primary transform (in the decoder). In LFNST, either a 4x4 or 8x8 non-separable transform can be applied depending on the block size. For example, a 4x4 LFNST can be applied to a small block (i.e., min(width, height) < 8), and an 8x8 LFNST can be applied to a large block (i.e., min(width, height) > 4).

[0128] The application of the inseparable transformation used in LFNST can be described as follows. To apply 4x4 LFNST, the 4x4 input block X can be represented as a vector as shown in Equation 1 below.

[0129] [Formula 1]

[0130] It can be used To compute the inseparable transformation. Here, It can be a vector of transformation coefficients, and it can also be a 16x16 transformation matrix. (16×1 coefficient vector) The blocks can be reconfigured later using the scan order (horizontal, vertical, or diagonal) of the corresponding blocks. Coefficients with smaller indices can be placed in a 4x4 coefficient block along with smaller scan indices.

[0131] Transform / inverse transform can be performed within a CU or TU unit. In other words, the transform / inverse transform can be applied to residual samples in a CU or a TU. The CU size can be the same as the TU size, or multiple TUs can exist within a CU region. The CU size typically indicates the size of the luma component (sample) coded block (CB). The TU size typically indicates the size of the luma component (sample) TB. The chroma component (sample) CB or TB size can be derived based on the luma component (sample) CB or TB size according to the component ratio of the color format (chroma format) (e.g., 4:4:4, 4:2:2, 4:2:0, etc.). The TU size can be derived based on maxTbSize. For example, when the CU size is larger than maxTbSize, multiple maxTbSize TUs (TBs) can be derived from the CU, and the transform / inverse transform can be performed within the TU (TB) units. maxTbSize can be considered when determining whether to apply various types of intra-frame prediction (such as intra-fragment prediction (ISP)). Information about maxTbSize can be predetermined, or it can be generated and encoded by the encoding device 200 and sent to the decoding device 300 via a signal.

[0132] Figure 5 This is a flowchart illustrating an inter-frame prediction-based video / image coding method applicable to embodiments of the present disclosure.

[0133] The encoding device 200 can perform inter-frame prediction on the current block (S500). The encoding device 200 can derive the inter-frame prediction mode and motion information of the current block, and generate prediction samples for the current block. Here, the inter-frame prediction mode determination process, the motion information derivation process, and the prediction sample generation process can be performed simultaneously, or any one of these processes can be performed before the others. For example, the inter-frame predictor 221 of the encoding device 200 can include a prediction mode determiner, a motion information deriver, and a prediction sample deriver. The prediction mode determiner can determine the prediction mode of the current block, the motion information deriver can derive the motion information of the current block, and the prediction sample deriver can derive the prediction samples of the current block.

[0134] The inter-frame predictor 221 of the encoding device 200 can search for blocks similar to the current block within a specific region (search region) of the reference image through motion estimation, and derive the reference block with the smallest difference from the current block, or a certain threshold or smaller. The inter-frame predictor 221 can derive a reference image index indicating the existence of the reference block based on the reference block, and derive a motion vector based on the positional difference between the reference block and the current block. The encoding device 200 can determine the prediction mode to be applied to the current block among various prediction modes. The encoding device 200 can compare the rate-distortion (RD) costs of various prediction modes and determine the optimal prediction mode for the current block.

[0135] When a merging mode is applied to the current block, the encoding device 200 can generate a merging candidate list, which will be described below, and derive the reference block from the current block whose inter-sample difference (i.e., the sum of absolute differences (SAD) or the sum of absolute transform differences (SATD)) is the smallest or less than a certain threshold among the reference blocks indicated by the merging candidates included in the merging candidate list. In this case, a merging candidate associated with the derived reference block can be selected, and merging index information indicating the selected merging candidate can be generated and sent as a signal to the decoding device 300. The encoding device 200 can use the motion information of the selected merging candidate to derive the motion information of the current block.

[0136] When the skip mode is applied to the current block, the motion information of the current block can be derived in the same way as when the merge mode is applied. However, when the skip mode is applied, the residual signal of the corresponding block is omitted, and the predicted samples can be directly used as reconstructed samples.

[0137] As another example, when the Advanced (A)MVP mode is applied to the current block, the encoding device 200 can generate an (A)MVP list, which will be described below, and use the MVP candidate selected from the MVP candidates included in the (A)MVP candidate list as the MVP of the current block. In this case, the motion vector of the reference block obtained by the aforementioned motion estimation can be used as the motion vector of the current block. The encoding device 200 can determine the MVP candidate with the motion vector that has the smallest difference from the motion vector of the current block as the selected MVP candidate. The encoding device 200 can derive the Motion Vector Difference (MVD), which is the difference obtained by subtracting the MVP from the motion vector of the current block. Information about the MVD can be sent to the decoding device 300 by signal. Moreover, when the (A)MVP mode is applied, the reference image index value can constitute reference image index information, which is sent to the separate decoding device 300 by signal.

[0138] The encoding device 200 can derive residual samples based on the predicted samples (S510). The encoding device 200 can derive residual samples by comparing the original samples and the predicted samples of the current block.

[0139] The encoding device 200 can encode image information including prediction information and residual information (S520). The encoding device 200 can output the encoded image information in the form of a bitstream. The prediction information may be information related to the prediction process and may include prediction mode information (e.g., skip flag, merge flag, merge index, etc.) and / or motion information. The motion information may contain candidate selection information (e.g., merge index, MVP flag, or MVP index), which is information used to derive motion vectors. In addition, the motion information may contain the aforementioned MVD information and / or reference image index information. Furthermore, the motion information may contain information indicating whether L0 prediction, L1 prediction, or bidirectional prediction is applied. The residual information is information about the residual samples and may include information about the quantized transform coefficients of the residual samples.

[0140] The output bitstream can be stored in (digital) storage medium and sent to the decoding device 300, or it can be sent to the decoding device 300 via a network.

[0141] Simultaneously, as described above, the encoding device 200 can generate a reconstructed image (including reconstructed samples and reconstructed blocks) based on reference samples and residual samples. This aims to derive the same prediction result from the encoding device as the prediction result obtained by the decoding device 300, and can improve encoding efficiency. Therefore, the encoding device 200 can store the reconstructed image (or reconstructed samples or reconstructed blocks) in memory and use it as a reference image for inter-frame prediction. Loop filtering processes, etc., can be further applied to the reconstructed image as described above.

[0142] Figure 6 This is a flowchart illustrating a video / image decoding method based on inter-frame prediction applicable to embodiments of the present disclosure.

[0143] refer to Figure 6 The decoding device 300 can perform operations corresponding to those performed by the encoding device 200. The decoding device 300 can predict the current block based on the received prediction information and derive prediction samples.

[0144] Specifically, the decoding device 300 can determine the prediction mode of the current block based on the received prediction information (S600). The decoding device 300 can determine which inter-frame prediction mode is applied to the current block based on the prediction mode information in the prediction information.

[0145] For example, decoding device 300 may determine whether to apply a merge mode or an (A)MVP mode to the current block based on a merge flag. Alternatively, decoding device 300 may select one of a variety of inter-frame prediction mode candidates based on a mode index. Inter-frame prediction mode candidates may include skip mode, merge mode, and / or (A)MVP mode, or may include a variety of inter-frame prediction modes described below.

[0146] The decoding device 300 can derive motion information for the current block based on the determined inter-frame prediction mode (S610). For example, when a skip mode or merge mode is applied to the current block, the decoding device 300 can generate a merge candidate list (described below) and select one of the merge candidates included in the merge candidate list. This selection can be based on the aforementioned selection information (merge index). The decoding device 300 can use the motion information of the selected merge candidate to derive motion information for the current block. The motion information of the selected merge candidate can be used as motion information for the current block.

[0147] As another example, when the (A)MVP mode is applied to the current block, the decoding device 300 can generate an (A)MVP candidate list, which will be described below, and use the motion vector of one of the selected MVP candidates included in the (A)MVP candidate list as the MVP of the current block. This selection can be performed based on the aforementioned selection information (MVP flag or MVP index). In this case, the decoding device 300 can derive the MVD of the current block based on the MVD information, and derive the motion vector of the current block based on the MVP and MVD of the current block. Furthermore, the decoding device 300 can derive the reference image index of the current block based on the reference image index information. The image indicated by the reference image index in the reference image list of the current block can be exported as a reference image, which is referenced for inter-frame prediction of the current block.

[0148] Furthermore, as will be described below, the motion information of the current block can be exported without generating a candidate list. In this case, the motion information of the current block can be exported according to the process that begins with the prediction mode described below. Then, the aforementioned candidate list generation can be omitted.

[0149] The decoding device 300 can generate a prediction sample for the current block based on the motion information of the current block (S620). In this case, the decoding device 300 can derive a reference image based on the reference image index of the current block, and derive the prediction sample for the current block using samples of the reference block indicated by the motion vector of the current block within the reference image. In this case, a prediction sample filtering process for all or some of the prediction samples of the current block can be further performed, which will be described below.

[0150] For example, the inter-frame predictor 332 of the decoding device 300 may include a prediction mode determiner, a motion information deriver, and a prediction sample deriver. The prediction mode determiner can determine the prediction mode of the current block based on the received prediction mode information. The motion information deriver can derive the motion information (motion vector, reference image index, etc.) of the current block based on the received motion information. The prediction sample deriver can derive the prediction samples of the current block.

[0151] The decoding device 300 derives residual samples of the current block based on the received residual information (S630). The decoding device 300 can generate reconstructed samples of the current block based on the predicted samples and residual samples, and generate a reconstructed image based on the reconstructed samples (S640). Subsequently, loop filtering and other processes can be further applied to the reconstructed image as described above.

[0152] Figure 7 This is a flowchart illustrating an inter-frame prediction process to which embodiments of the present disclosure can be applied.

[0153] refer to Figure 7 The inter-frame prediction process may include operations such as determining an inter-frame prediction mode, deriving motion information based on the determined prediction mode, and performing prediction (generating prediction samples) based on the derived motion information. The inter-frame prediction process may be performed by the encoding device 200 and the decoding device 300 as described above. In this disclosure, the encoding device may include encoding device 200 and / or decoding device 300.

[0154] refer to Figure 7 The encoding device determines the inter-frame prediction mode for the current block (S700). Various inter-frame prediction modes can be used to predict the current block within the image. For example, various modes can be used, including merge mode, skip mode, MVP mode, affine mode, sub-block merge mode, merge mode with MVD (MMVD), etc. Decoder-side motion vector refinement (DMVR) mode, adaptive motion vector resolution (AMVR) mode, bidirectional prediction with CU-level weights (BCW), bidirectional optical flow (BDOF), etc., can be used as additional modes. Furthermore, according to embodiments of this disclosure, the aforementioned inter-frame prediction modes may include a multiple hypothesis prediction (MHP) mode. The MHP mode represents a method of performing prediction by calculating a weighted sum of additional prediction blocks generated based on additional motion information about unidirectional or bidirectional prediction (or prediction) blocks. This MHP mode will be described in detail below.

[0155] In this application, the affine pattern can also be referred to as an "affine motion prediction pattern." Furthermore, the MVP pattern can be referred to as an "Advanced Motion Vector Prediction (AMVP) pattern." This disclosure may include several patterns and / or motion information candidates derived from some patterns as one of the candidates related to motion information from other patterns. For example, a history-based motion vector prediction (HMVP) candidate can be added as a merge candidate for merging / skipping patterns or as an MVP candidate for AMVP patterns. When an HMVP candidate is used as a motion information candidate for merging or skipping patterns, the HMVP candidate can be referred to as an "HMVP merge candidate."

[0156] When prediction mode information indicating the inter-frame prediction mode of the current block can be signaled from the encoding device 200 to the decoding device 300, the prediction mode information can be included in the bitstream and sent to the decoding device 300. The prediction mode information may include index information indicating one of a plurality of candidate modes. Alternatively, the inter-frame prediction mode can be indicated by hierarchical signaling of flag information.

[0157] In this case, the prediction mode information may include one or more flags. For example, the encoding device 200 may signal a skip flag to indicate whether a skip mode is applied, signal a merge flag to indicate whether a merge mode is applied when a skip mode is not applied, and further signal a flag indicating the application of the MVP mode or a flag for an additional mode when the merge mode is not applied. The affine mode may be signaled as an independent mode or as a mode dependent on the merge mode, MVP mode, etc. For example, the affine mode may include an affine merge mode and an affine MVP mode.

[0158] The encoding device can derive motion information for the current block (S710). Motion information can be derived based on inter-frame prediction modes. The encoding device can then use the motion information of the current block to perform inter-frame prediction.

[0159] The encoding device 200 can derive the optimal motion information for the current block through a motion estimation process.

[0160] For example, encoding device 200 can search a defined search range within a reference image for similar reference blocks that are highly correlated with the original block in the original image of the current block in the fractional pixel unit, and derive motion information from the reference blocks. Block similarity can be derived based on the difference in sample values ​​based on phase. For example, block similarity can be calculated based on the SAD between the current block (or a template of the current block) and the reference block (or a template of the reference block). In this case, encoding device 200 can derive motion information based on the reference block with the minimum SAD within the search area. The derived motion information can be transmitted as a signal to decoding device 300 based on an inter-frame prediction mode according to various methods.

[0161] The encoding device can perform inter-frame prediction based on the motion information of the current block (S720). The encoding device can generate prediction samples for the current block based on the motion information. The current block containing the prediction samples can be called a "prediction block".

[0162] MHP will be described in detail below. In this document, MHP may also be referred to as "Multi-reference Mode," "Multi-reference Prediction," "Multi-reference Prediction Mode," "Multi-reference Block Mode," "MHP Mode," "Multi-hypothesis Inter-frame Prediction Mode," "Inter-frame Combined Prediction Mode," "Combined Inter-frame Prediction Mode," "Combined Prediction Mode," "Multi-frame Prediction Mode," "Multi-prediction Mode," "Additional Reference Prediction Mode," "Additional Reference Mode," "Multi-reference Block," etc. Furthermore, MHP mode can be a mode that uses (or applies, is available, exists, or allows) additional reference blocks in addition to the basic reference block.

[0163] Figure 8 This is a flowchart illustrating an inter-frame prediction method performed by a decoding device according to an embodiment of the present disclosure.

[0164] refer to Figure 8 The decoding device 300 can generate a basic prediction block for the current block (S800). In other words, the decoding device 300 can generate a basic prediction block by performing one-way or two-way prediction. When multiple reference blocks exist, the decoding device 300 can derive additional reference blocks besides the prediction blocks generated (or derived) by one-way or two-way prediction, and combine the additional reference blocks with the basic prediction block. As an example, the basic prediction block can be an L0 prediction block or an L1 prediction block, or it can be a block obtained by calculating the weighted sum of the L0 and L1 prediction blocks. In this embodiment, the case where the basic prediction block is obtained by calculating the weighted sum of the L0 and L1 prediction blocks will be mainly described, but this disclosure is not limited to this. The basic prediction block can be a reference block before calculating the weighted sum. In this application, the basic prediction block can be referred to as an "initial prediction block", "temporary prediction block", "reference prediction block", "regular prediction block", etc., and the basic reference block can be referred to as an "initial reference block", "temporary reference block", "regular reference block", etc.

[0165] According to an embodiment, the decoding apparatus 300 can derive a first reference block and a second reference block of the current block by performing bidirectional prediction, and generate a basic prediction block by calculating a weighted sum of the first and second reference blocks. Alternatively, the decoding apparatus 300 according to this disclosure can generate a basic prediction block by performing unidirectional prediction to derive a third reference block of the current block. In other words, the basic prediction block can be derived by calculating a weighted sum of multiple reference blocks or from a single reference block.

[0166] According to an embodiment, weights can be used for a weighted sum of L0 and L1 prediction blocks. Weights are used for weighted prediction and can be collectively referred to as bidirectional prediction with CU-based weights (BCW) and CU-based weights (CW). Weights can be derived from a weight candidate list. The weight candidate list can include multiple weight candidates and can be predefined in the encoding device 200 and the decoding device 300.

[0167] Weight candidates can be a set of weights (i.e., a first weight and a second weight) indicating the weights applied to each bidirectional prediction block, or they can be weights applied to prediction blocks corresponding to either of the two directions. When only weights applied to prediction blocks corresponding to one direction are derived from the weight candidate list, weights applied to prediction blocks corresponding to the other direction can be derived based on the weights derived from the weight candidate list. For example, weights applied to predictions corresponding to the other direction can be derived by subtracting the weights derived from the weight candidate list from a predetermined value.

[0168] According to an embodiment, a weight index indicating the weights used for weighted prediction of the current block within the weight candidate list can be derived. In this disclosure, the weight index may be referred to as "bcw_idx" or "bcw index". The weight index can be derived by a decoding device or transmitted by a signal from an encoding device. When derived by a decoding device, the weight index can be derived as the weight index of a specific merge candidate within the merge candidate list. As an example, a specific merge candidate can be specified within the merge candidate list using the merge index.

[0169] The decoding device 300 can export (or generate) an additional reference block for the current block (S810). In other words, the decoding device 300 can export a multi-reference block for the current block.

[0170] The decoding device 300 can derive additional reference blocks in addition to the basic prediction blocks, and combine the basic reference blocks and the derived additional reference blocks (or calculate their weighted sum).

[0171] According to an embodiment, when multiple reference blocks are applied (or allowed or present), the decoding device 300 can derive up to a predetermined number of additional reference blocks and combine the additional reference blocks with the basic reference block. In other words, the decoding device 300 can combine (or calculate their weighted sum) a predefined number or fewer of additional reference blocks with the basic reference block. As an example, the predefined number can be 2. As another example, the predefined number can be one of 1, 2, 3, and 4. As yet another example, the predefined number can be referred to as the "maximum number of additional reference blocks".

[0172] When multiple additional reference blocks are combined with a base reference block, a weighted sum can be used to sequentially add the additional reference blocks to the base prediction block. For example, when generating up to two additional reference blocks, a prediction block can be generated by calculating a weighted sum of the base prediction block and the first additional reference block, and a final prediction block can be generated by calculating a weighted sum of the generated prediction block and the second additional reference block. The prediction block generated by calculating a weighted sum of the base prediction block and the first additional reference block can be called an "intermediate prediction block".

[0173] When multiple additional reference blocks are combined with a base reference block, a weighted sum can be used to add the generated additional reference blocks to the base prediction block simultaneously. In other words, after generating multiple additional reference blocks, weights can be applied to each of the multiple additional reference blocks and the base prediction block (or the L0 prediction block and the L1 prediction block), and the results can be summed simultaneously.

[0174] According to various embodiments, the decoding device 300 can determine whether various reference blocks are applied. In this case, an operation to determine whether multiple reference blocks are applied can be added before operating S810. Here, the decoding device 300 can explicitly send a signal or implicitly derive (or determine) whether multiple reference blocks are applied.

[0175] According to various embodiments, the encoding device 200 can signal to the decoding device 300 whether a multi-reference block is applied. For example, the encoding device 200 can signal to the decoding device 300 a flag indicating whether a multi-reference block is applied. Here, conditions for signaling / parsing the flag can be predefined. The conditions for signaling / parsing the flag can be multi-reference block availability conditions. When the signaling / parsing conditions are met, the decoding device 300 can parse the flag from the bitstream.

[0176] According to various embodiments, the decoding device 300 can deduce whether to apply a multi-reference block based on predefined encoding information.

[0177] According to various embodiments, the decoding apparatus 300 can acquire multi-reference block information (also referred to as "multi-reference block prediction information") to generate additional reference blocks. The multi-reference block information may include weighting information and / or prediction information. Additional reference blocks can be derived based on the prediction information, and the derived additional reference blocks can be added to a base prediction block (or intermediate prediction block) using weighted summation based on the weighting information. Furthermore, the multi-reference block information may also include flags indicating whether multi-reference blocks are applied.

[0178] Prediction information may include prediction mode information for deriving additional reference blocks and motion information based on the prediction mode. The prediction mode information may be inter-frame prediction mode information indicating a merge mode or AMVP mode. For example, the prediction mode information may also include a merge flag. In other words, a merge mode or AMVP mode can be used for deriving multiple reference blocks, and flag syntax elements indicating this can be signaled.

[0179] Predefined prediction modes from the merge mode and AMVP mode can be used for multi-reference block derivation. A predefined prediction mode can be selected from the merge mode and AMVP mode based on predefined encoding information.

[0180] When the merge mode is used to derive a multi-reference block, the prediction information may include a merge index. The merge index can specify merge candidates from the merge candidate list. When the AMVP mode is used to derive a multi-reference block, the prediction information may include MVP flags, reference indices, and / or MVD information. The MVP flags can specify candidates from the MVP candidate list.

[0181] The decoding device 300 can generate a final prediction block by calculating a weighted sum of a basic prediction block and additional reference blocks (S820). As described above, the number of additional reference blocks can be a predefined number or less. For example, when the number of additional reference blocks is 2, the final prediction block can be a block obtained by calculating a weighted sum of the basic prediction block and the two additional reference blocks. According to an embodiment, the weight information of the weighted sum can be transmitted or derived by a signal.

[0182] As described above, when multiple additional reference blocks are combined with a base reference block, the additional reference blocks can be added to the base prediction block sequentially or simultaneously using a weighted sum.

[0183] Generally, inter-frame prediction processes support unidirectional or bidirectional prediction. However, when more prediction blocks are involved, a block can be considered a multi-reference block. In this disclosure, reference blocks used for unidirectional / bidirectional prediction in video codecs (such as High Efficiency Video Coding (HEVC), Universal Video Coding (VVC), etc.) are described as regular (basic) reference blocks, and other additional reference blocks are described as additional reference blocks, multi-reference blocks, etc.

[0184] According to an embodiment, information about the additional reference block can be signaled as a merging index in the same or similar manner as in the merging mode. Alternatively, reference indexes, MVP flags (or indexes), MVD, etc., can be used to signal information about the additional reference block in the same or similar manner as in the AMVP mode. Alternatively, information about the additional reference block can be inherited from decoded neighboring blocks to derive motion information. According to an embodiment, whether to signal information about the additional reference blocks can be determined based on the number of additional reference blocks used. For example, when one additional reference block exists, information about the additional reference block can be signaled, and when two additional reference blocks exist, all or some of the information about the additional reference blocks may not be signaled, but can be derived on the decoder side.

[0185] According to embodiments of this disclosure, in-screen prediction blocks and IBC prediction blocks are included as multiple prediction blocks by signaling / deriving additional motion information and weight information of the prediction blocks, in addition to orientation-specific motion information, during the inter-frame prediction process. Therefore, compression performance can be improved by generating various final prediction blocks.

[0186] Figure 9 This is a diagram illustrating an example of a reference block used in a multi-reference block mode according to an embodiment of the present disclosure.

[0187] When multiple reference blocks exist, it can be as follows: Figure 9 Each reference block is shown. Figure 9 In this context, reference blocks P0 and P1 are regular reference blocks, while reference blocks P2 and P3 are multi-reference blocks (i.e., multiple additional reference blocks).

[0188] For ease of description, Figure 9 The example shown illustrates the case with two reference blocks in each prediction direction, but the number of reference blocks is not limited to this. In other words, the number of reference blocks in each prediction direction can vary, and the motion compensation order can also change. The following description of multiple reference blocks is based on... Figure 9 .

[0189] Figure 10 This is a flowchart illustrating a method for sending / parsing motion information of added prediction blocks according to an embodiment of the present disclosure when sending / parsing motion information of multiple reference blocks using signals. Figure 10 The operations described herein can be performed on the encoding device 200 and / or the decoding device 300.

[0190] When multiple reference blocks exist and information about the blocks is sent / resolved using signals, it can be done by... Figure 9 The sequence shown uses signals to send / parse motion information for the added prediction blocks. Here, MaxNum can be the maximum number of additional reference blocks.

[0191] refer to Figure 10 , Figure 10 The operation can be repeated in a loop until the number of additional reference blocks reaches the maximum. When the number of additional reference blocks is the maximum or less, the mhp_flag can be resolved. The mhp_flag can be a flag used to indicate whether multiple reference blocks exist.

[0192] When multiple reference blocks exist (i.e., when mhp_flag is "1"), mhp_mrg can be resolved. mhp_mrg can indicate either the merging mode (MH-MERGE) or the inter-frame mode (MH-AMVP) for multiple reference blocks. mhp_mrg enables the determination of either the merging mode or the inter-frame (AMVP) mode for multiple reference blocks.

[0193] In merge mode (i.e., when mhp_mrg is "1"), the merge index and weight index can be sent using signals. In inter-frame MVP (AMVP) mode, the reference index, MVP index, MVP data, and weight index can be sent using signals.

[0194] like Figure 10 As shown, when the maximum number of additional reference blocks that can be generated is MaxNum, the existence of multiple reference blocks can be determined from mhp_flag. For example, when MaxNum is 2, mhp_flag can have the values ​​shown in Table 5 below. Depending on the value, it can be determined whether a first additional block (i.e., the first additional reference block) and a second additional block (i.e., the second additional reference block) exist.

[0195] [Table 5]

[0196] Referring to Table 5, when mhp_flag is "1", additional reference blocks can exist. mhp_mrg enables the determination of either the multi-reference block merging mode (MH-MERGE) or the multi-reference block inter-frame mode (MH-AMVP). In MH-MERGE mode, the merging index and weight index can be signaled. In MH-AMVP mode, the reference index, MVP index (or flag), MVD data, and weight index can be signaled.

[0197] The syntax names such as mhp_flag and mhp_mrg described in this disclosure are illustrative and the names can vary. In this disclosure, the prediction modes for the case of multiple reference blocks are described as MH-MERGE and MH-AMVP, but these prediction modes can be distinguished from the merging modes and inter-frame modes that represent motion information for regular reference blocks. Specifically, MVP candidate lists for regular reference blocks and MVP candidate lists for additional reference blocks can be generated separately. Furthermore, the MH-MERGE mode and MG-AMVP mode for additional reference blocks may include weight index information.

[0198] In addition to the aforementioned methods for applying multiple reference blocks via signaling, derivation can be used to apply multiple reference blocks. Specifically, since the merging mode employs motion information transmitted (inherited) from encoded / decoded adjacent and non-adjacent blocks, the motion information received from encoded / decoded adjacent and non-adjacent blocks can also be used as information for multiple reference blocks. For example, when the current block is in merging mode and the adjacent and non-adjacent blocks used to obtain motion information include MHP information, the corresponding information can be received and used to generate additional reference blocks for the current block.

[0199] As mentioned above, multi-reference block information can be obtained through signaling or an export procedure, and both methods can be applied. Specifically, when the number of multi-reference blocks obtained through the export procedure is less than MaxNum, multi-reference blocks can be added via signaling. For example, when N multi-reference blocks are generated and M blocks are obtained through the export procedure (M≤N), up to N multi-reference blocks can be obtained via signaling.

[0200] Weighted summation can be used to add additional reference blocks obtained through signaling or derivation to the base reference block, enabling the generation of the final prediction block. As an example, when there are two additional reference blocks P2 and P3 and the weights obtained through signaling or derivation of each additional reference block are assumed to be w0 and w1, the process of generating the final prediction block using the weights can be represented as shown in Equation 2 below.

[0201] [Equation 2]

[0202] In Equation 2, w0 and w1 are the weights applied to the additional reference blocks P2 and P3, respectively. In the first step 1 of Equation 2, the basic prediction block P' can be generated using the basic reference blocks P0 and P1. In the second step 2 of Equation 2, the weighted sum of the basic prediction block P' and the additional reference block P2 can be calculated. In the third step 3 of Equation 2, the weighted sum of the prediction block (intermediate prediction block) P” obtained through weighted summation and the additional reference block P3 can be calculated to generate the final prediction block P.

[0203] This describes the addition of multiple reference blocks during the inter-frame prediction process; however, the addition of multiple reference blocks is not limited to this process and can be included in intra-frame prediction and IBC processes. In various prediction modes, when multiple reference blocks are included, a regular block (i.e., a regular reference block) can collectively refer to inter-frame blocks, intra-frame blocks, and IBC blocks in addition to the multiple reference blocks, and the multiple reference blocks include not only inter-frame modes or merge modes but also intra-frame modes and IBC modes. To determine the mode of the multiple reference blocks, flags indicating each mode can be included, and these flags can be obtained via transport (or inheritance) other than signaling. Alternatively, multiple reference blocks can be derived and generated on the decoder side when specific conditions other than signaling or transport methods are met.

[0204] However, when multiple reference blocks exist or intra-frame modes exist within multiple reference blocks, signaling and decoding processes for transformation, filtering, correlation syntax, etc., can be performed under various conditions.

[0205] Therefore, in embodiments of this disclosure, signaling and decoding processes performed based on the presence of multiple reference blocks or the presence of intra-frame modes in the multiple reference blocks will be described.

[0206] Before proceeding, the following will describe the various tools related to signaling and decoding processes in the presence of intra-frame mode or multiple reference blocks.

[0207] IBC

[0208] IBC is a tool for extending HEVC for Screen Content Coding (SCC). It is well-known for significantly improving the coding efficiency of screen content material. Because IBC mode is implemented as a block-level coding mode, block matching (BM) can be performed on the encoder to find the optimal block vector (BV) (or motion vector) for each CU. Here, the BV can be used to represent the displacement from the current block to a reference block that has already been reconstructed within the current block. The luma BV of an IBC-coded CU can have an integer resolution (or precision). The chroma BV can be rounded to an integer resolution. When combined with AMVR, IBC mode can switch between 1-pixel motion vector resolution and 4-pixel motion vector resolution. IBC-coded CUs can be processed as a third prediction mode in addition to intra-frame or inter-frame prediction modes. IBC mode can be applied to CUs with a width and height of 64 luma samples or less.

[0209] On the encoder side, hash-based motion estimation can be performed for IBC. The encoder can perform BD checks on blocks with a width and height no greater than 16 luminance samples. In non-merging mode, a BV search can be performed first using a hash-based search. When the hash search does not return valid candidates, a BM-based local search can be performed.

[0210] In hash-based searches, hash key matching (32-bit Cyclic Redundancy Check (CRC)) between the current block and the reference block can be extended to all allowed block sizes. Hash key calculation for each location in the current image can be based on 4x4 sub-blocks. For a current block with a larger size, a hash key matching the reference block's hash key can be determined when all hash keys of all 4x4 sub-blocks match the hash key at the corresponding reference location.

[0211] When a hash key of multiple reference blocks is found to match the hash key of the current block, the BV cost of the matching reference blocks can be calculated, and the lowest cost can be selected.

[0212] In block matching searches, the search range can be set to cover both previous and current CTUs. At the CU level, the IBC mode can be signaled along with a flag and can be signaled as either IBV AMVP mode or IBC skip / merge mode.

[0213] IBC Skip / Merge Mode: The merge candidate index can be used to indicate which BVs from the list of neighboring candidate IBC coding blocks are used to predict the current block. The merge list can include spatial candidates, HMVP candidates, and paired candidates.

[0214] IBC AMVP mode: BV difference can be encoded in the same way as MVD. The BV prediction method can use two candidates as predictors (in IBC encoding), one from the left neighbor and one from the top neighbor. When either neighbor is unavailable, a default block can be used as the predictor. A flag can be sent to indicate the BV predictor index.

[0215] Geometric Partitioning (GPM)

[0216] In VVC, GPM can be supported for inter-frame prediction. GPM is a type of merging mode that can be signaled using CU-level flags. Other merging modes can include regular merging mode, MMVD mode, CIIP mode, and sub-block merging mode. For each possible CU size w×h=2 m ×2 n It can support a total of 64 partitions. Here, m, n∈{3, ..., 6} and 8x64 and 64x8 are excluded.

[0217] When using this mode, the CU can be divided into two partitions by geometrically positioned straight lines. The position of the partition lines can be mathematically derived from the angles of a specific partition and offset parameters. Each part of the geometric partition of the CU can be used for inter-frame prediction using its own motion. Only a single prediction is allowed for each partition; that is, each part has one motion vector and one reference index. A single prediction motion constraint is applied to ensure that, similar to regular bidirectional prediction, only two motion-compensated predictions are needed per CU.

[0218] Figure 11 An example of GPM partitions grouped by the same corner is shown.

[0219] Usable Figure 11 The process illustrated in the figure derives the unidirectional predicted motion for each partition.

[0220] When GPM is used for the current CU, a geometric partition index indicating the partition pattern (angles and offsets) and two merge indices (one per partition) can be additionally signaled. The maximum number of GPM candidates can be explicitly signaled in SPS to specify the syntax binaryization of the GPM merge index. After predicting each part of the geometric partition, a mixing process with adaptive weights can be used to adjust the sample values ​​according to the geometric partition edges, as described below for mixing along the geometric partition edges. This is the prediction signal for the entire CU, and the transformation and quantization processes can be applied to the entire CU as in other prediction modes. Finally, the motion field of the CU predicted using GPM can be stored in a motion field storage device for the geometric partition pattern described below.

[0221] Blending along the edges of geometric partitions

[0222] After using each part of its own motion prediction geometry partition, a blend can be applied to the two prediction signals to derive samples around the edges of the geometry partition. The blend weights for each location in the CU can be derived based on the distance between the individual location and the partition edge.

[0223] The distance from the location (x, y) to the partition edge can be derived as shown in Equation 3 below.

[0224] [Formula 3]

[0225] Here, i and j are the indices of the angle and offset of the geometric partition, which depend on the geometric partition index sent by the signal. ρ x,j and ρ y,j The sign depends on the corner index i.

[0226] As shown in Equation 4 below, the weights of each part of the geometric partition can be derived.

[0227] [Formula 4]

[0228] partIdx is based on corner indexes. Figure 12 An example of generating mixed weights w0 using GPM is shown. An example of weights w0 can be illustrated as follows: Figure 12 As shown.

[0229] Sports field storage for GPM

[0230] Mv1 from the first part of the geometric partition, Mv2 from the second part of the geometric partition, and the combination Mv of Mv1 and Mv2 can be stored in the motion field of the GPU-encoded CU.

[0231] As shown in Equation 5 below, the type of motion vector stored at each individual location in the sports field is determined.

[0232] [Formula 5]

[0233] Here, motionIdx equals d(4x+2, 4y+2), and partIdx is based on the corner index i.

[0234] When sType equals 0 or 1, Mv0 or Mv1 can be stored in the corresponding motion field. Otherwise, when sType equals 2, a combination Mv of Mv1 and Mv2 can be stored. The combined Mv can be generated through the following process.

[0235] When Mv1 and Mv2 come from different lists of reference images (one from L0 and the other from L1), Mv1 and Mv2 can be easily combined to generate bidirectional predicted motion vectors.

[0236] Otherwise, when Mv1 and Mv2 are in the same list, only the unidirectional predicted motion Mv2 can be stored.

[0237] CIIP

[0238] CIIP can be applied to the current block. Additional flags (e.g., ciip_flag) can be signaled to indicate whether CIIP mode is applied to the current CU. For example, when the CU is encoded in merge mode, if the CU contains 64 luma samples (i.e., if the product of the CU width and CU height is 64 or greater) and both the CU width and CU height are less than 128 luma samples, then an additional flag can be signaled to indicate whether CIIP mode is applied to the current CU. As its name suggests, CIIP prediction involves combining inter-frame prediction signals and intra-frame prediction signals. The inter-frame prediction signal P_inter for CIIP mode can be derived using the inter-frame prediction process applied to regular merge mode, and the intra-frame prediction signal P_intra can be derived using the regular intra-frame prediction process with planar mode. A weighted average can then be used to combine the intra-frame and inter-frame prediction signals. Figure 13 This is a diagram illustrating the top and left neighbor blocks used in CIIP weight derivation. Here, it can be seen from... Figure 13 The encoding patterns of the top and left neighboring blocks shown are calculated as follows.

[0239] When the upper neighbor is available and intra-frame coding is enabled, isIntraTop is set to 1; otherwise, isIntraTop is set to 0. When the left neighbor is available and intra-coded, isIntraLeft is set to 1; otherwise, isIntraLeft is set to 0. When (isIntraTop+isIntraLeft) equals 2, wt is set to 3; Otherwise, when (isIntraTop+isIntraLeft) equals 1, wt is set to 2; Otherwise, set wt to 1.

[0240] CIIP predictions can be configured as shown in Equation 6 below.

[0241] [Formula 6]

[0242] The combination includes CIIP with template-based intra-frame mode export (TIMD) and template matching (TM) merging.

[0243] In CIIP mode, prediction samples can be generated by assigning weights to the inter-frame prediction signal that uses CIIP-TM to merge candidate predictions and the intra-frame prediction signal that uses TIMD predictions derived in intra-frame prediction mode. This method can be applied only to CBs comprising regions of 1024 pixels or smaller.

[0244] TIMD can be used to export IPMs from CIIP. Specifically, you can select the IPM with the lowest SATD value from the TIMD pattern list and map it to one of 67 regular IPMs.

[0245] Additionally, when the exported IPM is in corner mode, modifications can be made to the weights (wIntra, wInter) used for the two tests.

[0246] Figure 14 This is a set of diagrams illustrating the partitioning method used in corner mode.

[0247] In near-horizontal mode (2 ≤ corner mode index < 34), the current block is vertically partitioned, such as... Figure 14 As shown in A. In near-vertical mode (34 ≤ angular mode index < 66), as... Figure 14 The current block of the horizontal partition shown in B.

[0248] The different sub-blocks (wIntra, wInter) can be shown in Table 6 below.

[0249] [Table 6]

[0250] CIIP-TM can be used to generate a CIIP-TM merge candidate list for the CIIP-TM pattern. Merge candidates can be improved using TM. Similar to regular merge candidates, CIIP-TM merge candidates can also be rearranged through Adaptive Reordering of Merge Candidates (ARMC). The maximum number of CIIP-TM merge candidates can be 2.

[0251] GPM including inter-frame and intra-frame prediction

[0252] In a GPM with inter-frame prediction and intra-frame prediction, the final prediction samples are generated by assigning weights to the inter-frame prediction samples and intra-frame prediction samples of each GPM-separated region. Inter-frame prediction samples are derived using the inter-frame GPM, while intra-frame prediction samples can be derived from the IPM candidate list and from the index sent by the encoder with signals. The IPM candidate list size can be predefined as 3.

[0253] Figure 15 This is a set of diagrams illustrating the available IPM candidates in a GPM with inter-frame and intra-frame prediction.

[0254] like Figure 15 A to Figure 15 As shown in D, the available IPM candidates can be parallel angle mode (parallel mode), perpendicular angle mode (vertical mode), and planar mode relative to the GPM block boundary. Furthermore, as... Figure 15 As shown, GPM with inter-frame prediction and intra-frame prediction can be limited to reduce signaling overhead for IPM and prevent an increase in the size of the intra-frame prediction circuitry in the hardware decoder. Furthermore, to further improve coding performance, direct motion vectors from the GPU mixing region and the IPM storage device can be introduced.

[0255] In decoder-side intra-mode export (DIMD) and neighbor-based IMP export, parallel modes can be registered first. Therefore, when no identical IMP candidates exist in the list, up to two IMP candidates can be registered via DIMD and / or derived from neighboring blocks.

[0256] In the nearest neighbor pattern derivation, up to five locations exist for available neighbor blocks, but as shown in Table 7, these locations may be constrained by the boundary angles of GPM blocks already used in GPM with TM (GPM-TM).

[0257] Table 7 shows the locations of available neighboring blocks derived for IPM candidates based on the GPM block boundary angles. A and L can represent the top and left sides of the predicted block.

[0258] [Table 7]

[0259] GPM intra-frame can be combined with GPM-MMVD. For further improvement in coding performance, TIMD can be used as an IPM candidate within a GPM frame. Parallel modes can be registered first, followed by TIMD, DIMD, and IPM candidates for neighboring blocks.

[0260] MHP

[0261] In multi-hypothesis inter-frame prediction mode, one or more additional motion-compensated prediction signals can be transmitted in addition to the existing bidirectional prediction signal, and the resulting overall prediction signal can be obtained by weighted superposition of samples. The bidirectional prediction signal p bi The first additional prediction signal / hypothesis h3 can be used to obtain the result prediction signal p3, as shown in Equation 7 below.

[0262] [Formula 7]

[0263] The weighting factor α can be specified by add_hyp_weight_idx, which is a new syntax element according to Table 8 below.

[0264] [Table 8]

[0265] Similarly, one or more additional prediction signals can be used. The resulting overall prediction signal can be repeatedly accumulated with each additional signal, as shown in Equation 8 below.

[0266] [Formula 8]

[0267] The resulting overall prediction signal can be obtained as the final p n (That is, p with the largest index n) n In this exploratory experiment (EE), up to two additional prediction signals can be used (i.e., n is limited to 2).

[0268] The motion parameters for each additional prediction hypothesis can be explicitly signaled by specifying a reference index, predicted motion vector values, and MVD, or implicitly signaled by specifying a merging index. A multi-hypothesis merging flag distinguishes between these two signaling modes.

[0269] MHP can be applied to the AMVP inter-mode only when different weights of BCW are selected in the bidirectional prediction mode.

[0270] MHP and BDOF can be combined, but BDOF can only be applied to the bidirectional signal portion of the predicted signal (i.e., the first general assumption).

[0271] Overlapping Block Motion Compensation (OBMC)

[0272] OBMC can be performed on all motion-compensated block boundaries within the right and bottom boundaries of the CU. Therefore, this can be applied to both luma and chroma components. To handle CU / sub-block boundaries in a uniform manner, OBMC can be performed on all MC block boundaries at the sub-block level.

[0273] Figure 16 This is a diagram illustrating sub-blocks at the CU / PU boundaries and sub-PUs in the Advanced Temporal Motion Vector Prediction (ATMVP) mode. Here, the size of the sub-block can be set to 4x4, as shown below. Figure 16 As described in the text.

[0274] When OBMC is applied to the current sub-block, in addition to the current motion vector, the prediction blocks for the current sub-block can be derived using the motion vectors of four connected neighboring sub-blocks that are available and different from the current motion vector. Weights can be assigned to these multiple prediction blocks based on multiple motion vectors to generate the final prediction signal.

[0275] like Figure 16 A and Figure 16 As shown in B, the predicted block based on the motion vectors of neighboring sub-blocks is P. N (Here, N is the index of the top, bottom, left, and right neighboring sub-blocks), and the predicted block based on the motion vector of the current block is P. C When P N and P C When they belong to the same PU (because they contain the same motion information), the OBMC does not need to be from the P. N Execute. Otherwise, P can be... N Add all pixels to P C The same pixels. In other words, P can be... N Add four rows / columns to P C Weighting factors {1 / 4, 1 / 8, 1 / 16, 1 / 32} are used for P. N The weighting factors {3 / 4, 7 / 8} can be used for P. C For P generated based on the motion vectors of vertical (horizontal) neighboring sub-blocks N , exists in relation to P N Pixels in the same row (column) can be added to P as equal weighting factors. C .

[0276] Localized lighting compensation (LIC)

[0277] LIC is an inter-frame prediction technique used to model the local illumination variation between the current block and the predicted block as a function of the variation between the current block template and the reference block template. The parameters of this function can be a scaling factor α and an offset β, which form a linear equation, i.e., α... p[x]+β is used to compensate for illumination variations, where p[x] is a reference sample indicated by the motion vector at position x on the reference image. When surround motion compensation is enabled, surround offset can be considered to clip the motion vector. Since α and β can be derived based on the current block template and the reference block template, no signaling overhead is required for them, except for signaling the LIC flag for AMVP mode to indicate the use of LIC.

[0278] LIC can be used for inter-frame CUs with the following features.

[0279] Intra-frame neighbor samples can be used in LIC parameter export.

[0280] LIC can be disabled for blocks with fewer than 32 luminance samples.

[0281] For both non-subblock mode and affine mode, LIC parameter export can be performed based on the template sample corresponding to the current CU, rather than the partial template block sample corresponding to the first top-left 16x16 cell.

[0282] Samples of the reference block template can be generated using block motion vectors and motion compensation without rounding them to integer pixel resolution (or precision).

[0283] For bidirectional inter-frame prediction CU, two sets of LIC parameters can be derived separately for L0 and L1 predicted samples. An iterative method can be applied to derive the L0 and L1 parameters. Specifically, the L0 LIC parameters can first be derived by minimizing the difference between the L0 template prediction T0 and the template T, and the samples in T can be updated by subtracting the corresponding samples in T0. Subsequently, the L1 parameters can be calculated to minimize the difference between the L1 template prediction T1 and the updated template. Finally, the L0 parameters can be readjusted in the same manner.

[0284] The signaling and decoding process for each tool is described below when intra-frame modes exist as multiple reference blocks. In this disclosure, the presence of multiple reference blocks is described as the presence of at least one additional reference block in addition to the regular reference block (or regular prediction block), and multiple reference blocks are described as multiple additional reference blocks.

[0285] According to an embodiment, when an intra-frame mode exists as a multiple reference block, the signaling conditions used to determine whether transform-related flags (i.e., cu_coded_flag) are included and residual signals can be changed. When a specific condition is met, the cu_coded_flag can be signaled, and when the specific condition is not met, the cu_coded_flag value can be inferred as "1" or "0". This document has described that the cu_coded_flag can be signaled, but the expression is not limited to this and can be "parsed", "derived", etc. For example, the expression can be "signaled" on the encoding side, and the expression can be "parsed", "derived", etc. on the decoding side. In this disclosure, "parsed" can mean "derived", "obtained", etc.

[0286] When the current block is in inter-frame mode or IBC mode and multiple reference blocks exist, the signaling condition for cu_coded_flag can be the number of MH-intra-prediction modes. Here, the number of MH-intra-prediction modes can be the number of intra-prediction modes present in the multiple reference blocks. When the number of MH-intra-prediction modes is 0, cu_coded_flag can be sent using a signal, and when the number of MH-intra-prediction modes is 1 or greater, it can be inferred that cu_coded_flag has a value of 1.

[0287] When the current block is in inter-frame mode or IBC mode and multiple reference blocks exist, the signaling condition for cu_coded_flag can be the block size. When the block size meets a specific condition, cu_coded_flag can be a signal; otherwise, cu_coded_flag can be inferred to be "1".

[0288] If the current block is in intra-frame mode and multiple reference blocks exist, the cu_coded_flag value can be deduced as "1" if certain conditions are met, and the cu_coded_flag can be sent by signal. Here, the specific condition can be the number of MH-iner or MH-IBC blocks. When the number of MH-iner or MH-IBC blocks is greater than 0, the cu_coded_flag can be sent by signal.

[0289] When the current block is in intra-frame mode and multiple reference blocks exist, the signaling condition for cu_coded_flag can be the block size.

[0290] According to embodiments, when an intra-frame mode exists as a multiple reference block, the boundary strength (BS) determination method in the deblocking filter process can be modified and applied. When each of the current block or adjacent blocks includes one or more MH-intra-frame blocks, the BS can be assigned a value of 2. According to this embodiment, the BS can be assigned a value from 0 to 2. According to various embodiments, in intra-frame mode or CIIP mode, the BS can have a value of 2 or can be extended to the range of 0 to 3. For example, the BS can be assigned a value of 3 in intra-frame mode and a value of 2 in combined mode. Here, the combined mode can be a CIIP mode. According to various embodiments, the combined mode can represent a mode where the regular block is an inter-frame or IBC mode and MH-intra-frame is included as a multiple reference block. According to various embodiments, the combined mode can represent a mode where the regular block is an intra-frame mode and MH-inter-frame or MH-IBC is included as a multiple reference block.

[0291] According to an embodiment, when an intra-frame mode exists as a multiple reference block, the sub-block transform (SBT) signaling conditions can be modified and applied. For example, when an MH-intra-frame prediction block exists, SBT information can be transmitted without signaling. Here, the number of MH-intra-frame blocks can include only the MH-intra-frame blocks transmitted with signaling.

[0292] According to an embodiment, when LIC is applied to the current block and multiple reference blocks exist, multiple reference blocks that are signaled may not be allowed. Specifically, intra-frame blocks may not be allowed as multiple reference blocks.

[0293] When applying LIC to the current block, LIC parameters can be derived from the multi-reference block and applied. Here, additional LIC flag signaling is not required, and the weight information for the multi-reference block can be omitted. When weight information is omitted, default weight values ​​can be used for a weighted sum of the multi-reference block and the regular block. Here, the default weights can have variations depending on the mode of the multi-reference block.

[0294] When applying a LIC to the current block, the application of the LIC to multiple reference blocks can be determined based on the difference between the LIC parameters of the current block and the LIC parameters of multiple reference blocks. During the process of deriving the LIC parameters, the template region used to calculate the LIC parameters for a regular block may differ from the template region used to calculate the LIC parameters for multiple reference blocks. In the presence of multiple reference blocks, if the signaling information for the multiple reference blocks precedes the LIC, the LIC of the current block can be restricted. Whether to apply the LIC can be determined only within the MH-frame.

[0295] According to an embodiment, when OBMC is applied to the current block and multiple reference blocks exist, multiple reference blocks may not be allowed to be transmitted by signal. Specifically, intra-frame blocks may not be allowed as multiple reference blocks.

[0296] When OBMC is applied to the current block, the weight information for the multi-reference blocks may not be included. When weight information is not included, the default weight values ​​can be used for the weighted sum of the multi-reference blocks and the regular blocks. Here, the default weights may have varying values ​​depending on the mode of the multi-reference blocks.

[0297] In the presence of multiple reference blocks, if the signaling information of the multiple reference blocks precedes the OBMC, the OBMC of the current block can be restricted. Whether to apply the OBMC can be applied only within the MH-frame.

[0298] In the presence of multiple reference blocks, OBMC at the CU boundary can be restricted, while OBMC executed in sub-block units can be maintained.

[0299] According to the embodiments, when the intra-frame mode exists as a multi-reference block, the LFNST can be considered as follows. When the regular block is an inter-frame mode or IBC mode and a multi-reference block exists, the following can be applied to determine whether to select (or apply) the LFNST core to effectively account for the characteristics of the residual signal.

[0300] Specifically, when all multi-reference blocks are intra-mode or the multi-reference blocks include one or more intra-modes, a representative mode can be determined based on the directionality of each intra-mode, and an LFNST kernel based on that mode can be used. When the directionality of each intra-mode is inconsistent, an LFNST kernel corresponding to the intra-mode of the first multi-reference block can be used. Alternatively, when the directionality of the modes is inconsistent, LFNST may not be applied, or a planar mode kernel may be used.

[0301] In addition, when the multi-reference block is in inter-frame mode or IBC mode, LFNST may not be applied, or a planar mode kernel may be used.

[0302] Additionally, when a multi-reference block includes an intra-frame mode and an inter-frame / IBC mode, the LFNST core can be selected based on the given intra-frame mode.

[0303] The aforementioned method can be applied similarly even when the regular block is in intra-frame mode and multiple reference blocks exist. When all multiple reference blocks are in intra-frame mode or the multiple reference blocks include one or more intra-frame modes, an LFNST kernel based on the intra-frame mode of the regular block can be used.

[0304] When the intra-frame mode of a regular block is similar to the orientation (horizontal, vertical, or non-angular) of each intra-frame mode, the intra-frame mode representing that orientation can be determined, and the LFNST kernel can be used based on that mode. When the orientation of each intra-frame mode is inconsistent, an LFNST kernel based on the intra-frame mode of the regular block can be used, LFNST can be omitted, or a kernel based on the planar mode can be used.

[0305] When all multi-reference blocks are in intra-frame or IBC mode, an LFNST core in intra-frame mode based on regular blocks can be used, LFNST can be omitted, or a planar mode core can be used.

[0306] When a multi-reference block includes an intra-frame mode and an inter-frame / ICB mode, an LFNST kernel based on the intra-frame mode of the regular block can be used. Alternatively, when the intra-frame mode of the regular block and the intra-frame mode of the multi-reference block have similar orientations, an intra-frame mode representing that orientation can be determined, and an LFNST kernel can be used based on that mode. Otherwise, LFNST may not be applied, or a planar mode kernel can be used.

[0307] Alternatively, when the intra-frame mode of a regular block and the intra-frame mode of a multi-reference block have similar orientations, a kernel based on the intra-frame mode of the regular block can be used; otherwise, a kernel based on the planar mode can be used.

[0308] The following describes the interaction with other tools when the intra-frame (MH-intra-frame) mode exists as a multi-reference block. In particular, the presence of the MH-intra-frame mode can represent a relatively large amount of residual signal, and therefore the existence of an intra-frame mode for interaction can be considered.

[0309] In skip mode, the motion information signaling and derivation methods used in merge mode are employed, but residual signals are not transmitted. As described above, unlike inter-frame mode, many residual signals are included in intra-frame mode. Therefore, intra-frame mode, as a multi-reference block, can also be applied to blocks with relatively large residual signals.

[0310] Figure 17 This is a flowchart illustrating a method for resolving a CU based on an embodiment, depending on whether a multi-reference block is applied and whether the transformed and residual signal information is transmitted based on it.

[0311] Figure 17 This can show the correlation between whether multiple reference blocks are applied during the process of parsing CU (coding_unit()) and whether it is based on its transmission transformation and residual signal information.

[0312] refer to Figure 17 The process of resolving the CU can begin with resolving the prediction mode. In other words, the type of prediction mode can be determined. In this case, it can be determined (or decided) whether the prediction mode is an intra-frame mode or an inter-frame mode.

[0313] In intra-frame mode, multiple reference blocks may not be applied (allowed) to the corresponding block. In other words, the corresponding block may not have multiple reference blocks. Due to the characteristics of intra-frame mode with many residual signals, the transform-related flag (i.e., cu_coded_flag information) is not sent by signal, but can always be inferred to be a value of 1. Here, a cu_coded_flag with a value of 1 can indicate the presence of transform information and residual signal information.

[0314] Otherwise, when the corresponding block is in skip mode instead of intra mode, multiple reference blocks should not be used. Because skip mode does not send any residual signals, the cu_coded_flag information is not sent, but it can always be inferred to be 0.

[0315] Otherwise, when the corresponding block is neither in intra-frame mode nor skip mode, the `general_merge_flag` can be used to determine whether the corresponding block is in inter-frame (AMVP) mode or merge mode. In the case of merge mode, multiple reference blocks can be applied, and these multiple reference blocks can include both intra-frame mode and inter-frame (AMVP and merge) mode. Since whether the corresponding block is in skip mode or merge mode is determined based on the presence of a residual signal, merge mode indicates the presence of a residual signal. Therefore, instead of signaling `cu_coded_flag`, it can be inferred to be a value of 1.

[0316] Finally, when the corresponding block is determined to be in inter-frame (AMVP) mode via `general_merge_flag`, multiple reference blocks can exist, and these multiple reference blocks can contain both intra-frame and inter-frame (AMVP and merge) modes. Furthermore, multiple reference blocks can have signaled modes and inherited modes, but the type and number of inherited modes are unknown during syntax parsing. Therefore, when the number of multiple reference signals sent via signaled mode is `numMHP`, if `numMHP` is greater than 0, it is determined that the corresponding block has a large amount of information to be signaled due to many residual signals, and `cu_coded_flag` is not sent via signaled mode, but can be inferred to be a value of 1. When `numMHP` is 0, `cu_coded_flag` can be sent via signaled mode using either a value of 0 or 1.

[0317] More specifically, see reference Figure 17 This allows us to determine the type of prediction mode for the current block, and when the prediction mode is intra-frame mode, multi-reference blocks can be omitted (multi-reference blocks are not applied). Here, cu_coded_flag can be inferred to be the value 1.

[0318] When the prediction mode is not an intra-frame mode, a flag indicating whether the prediction mode is a skip mode can be parsed to determine whether the prediction mode is a skip mode.

[0319] When the prediction mode is not skip mode, the flag indicating whether the prediction mode is inter-frame mode or merge mode can be parsed to determine whether the prediction mode is inter-frame mode or merge mode.

[0320] When the prediction mode is merge mode, multiple reference blocks can be applied to the corresponding blocks (application of multiple reference blocks). Here, multiple reference blocks can include intra-frame mode and inter-frame mode (AMVP and merge). Since skip mode and merge mode are distinguished by the presence or absence of residual signals, in the case of merge mode, cu_coded_flag is not sent, but due to the presence of residual signals, cu_coded_flag can be inferred to be a value of 1.

[0321] When the prediction mode is inter-frame (AMVP) mode, the corresponding block can have multiple reference blocks. Here, multiple reference blocks can include intra-frame mode and inter-frame (AMVP and merge) mode. Furthermore, multiple reference blocks can have signaled transmission mode and inheritance mode. However, since the number of inheritance mode types is unknown during syntax parsing, when the number of signaled multiple reference blocks (numMHP) is greater than 0, the corresponding block can be identified as having a large amount of information to be signaled, attributable to many residual signals. In this case, cu_coded_flag is not signaled, but can be inferred to be a value of 1. When numMHP is 0, cu_coded_flag can be signaled as a value of 0 or 1.

[0322] The following describes a method for parsing cu_coded_flag under specific parsing conditions when applying multiple reference blocks.

[0323] Figure 18 This is a flowchart illustrating a method for resolving a CU based on whether a multiple reference block is applied and whether it meets specific conditions, according to an embodiment.

[0324] As described above, prediction modes can be categorized into intra-frame, skip, merge, and inter-frame (AMVP) modes, and the support (or application) of multiple reference blocks can be determined based on each mode. Furthermore, for blocks considered to have many residual signals, the cu_coded_flag is not signaled, but can be inferred as having a value of 1. For blocks considered to have few residual signals, the cu_coded_flag is not signaled, but can be inferred as having a value of 0.

[0325] When certain conditions are met, the `cu_coded_flag` can be sent using a signal; when these conditions are not met, the `cu_coded_flag` can be inferred to be 1. For example, when the number of intra-prediction modes in a multi-reference block sent using a signal is `numMHIntra`, `numMHIntra` can be considered to determine the value of `cu_coded_flag`. In other words, Figure 18 The specific condition can be "if (num MHIntra == 0)". When the condition is met, cu_coded_flag can be sent with a signal, and when the condition is not met, cu_coded_flag is not sent with a signal, but it can be inferred to be the value 1.

[0326] As another example, a specific condition could be "if (width × height > 1024)". In other words, when the block size (width × height) meets the condition, the cu_coded_flag can be signaled, and when the block size does not meet the condition, the cu_coded_flag can be inferred to be a value of 1. This is because it can be assumed that blocks with smaller sizes have more residual signals (when no similar blocks exist, the block continues to split). The two conditions mentioned above can be considered together, and in addition to the block size, the aspect ratio of the block can also be considered.

[0327] More specifically, see reference Figure 18 This allows us to determine the prediction mode type of the current block. In this case, we can determine whether the prediction mode is intra-frame mode or inter-frame mode.

[0328] When the prediction mode is intra-frame mode, multi-reference blocks may not be applied (or allowed) to blocks. In other words, multi-reference blocks may not exist in a block. Because intra-frame mode has many residual signals due to its characteristics, the transform-related flag (i.e., cu_coded_flag information) is not signaled, but can always be inferred to be 1.

[0329] When the prediction mode is not intra-frame mode, the flag cu_skip_flag indicating whether the prediction mode is skip mode can be parsed to determine whether the prediction mode is skip mode.

[0330] When the prediction mode is not skip mode, the general_merge_flag indicating whether the prediction mode is inter-frame mode or merge mode can be parsed to determine whether the prediction mode is inter-frame mode or merge mode.

[0331] When the prediction mode is merge mode, multiple reference blocks can be applied to the block. Here, multiple reference blocks can include intra-frame mode and inter-frame mode (AMVP and merge).

[0332] Since skip mode and merge mode are distinguished by the presence or absence of the residual signal, in the case of merge mode, cu_coded_flag is not sent, but due to the presence of the residual signal, it can be inferred to be a value of 1.

[0333] When the prediction mode is inter-frame (AMVP) mode, a block can have multiple reference blocks. Here, multiple reference blocks can include intra-frame mode and inter-frame (AMVP and merge) mode.

[0334] When certain conditions are met, the `cu_coded_flag` can be signaled, and when these conditions are not met, the `cu_coded_flag` can be inferred to be the value 1. As an example, consider using the number of intra-prediction modes `numMHIntra` in the signaled multi-reference block to determine the value of `cu_coded_flag`. In this case, the specific conditions could include the condition that the number of intra-prediction modes is 0 (if (num MHIntra == 0)). When the specific conditions are met (i.e., when the number of intra-prediction modes is 0 (num MHIntra == 0)), the `cu_coded_flag` can be signaled, and when the specific conditions are not met (i.e., when the number of intra-prediction modes is not 0), the `cu_coded_flag` is not signaled, but can be inferred to be the value 1.

[0335] As another example, the specific condition could be the block size. In this case, the specific condition could include a block size (width × height) greater than a threshold size (1024) (if (width × height) > 1024). When the specific condition is met (i.e., when the block size is greater than 1024), the cu_coded_flag can be signaled, and when the specific condition is not met (i.e., the block size is 1024 or smaller), the cu_coded_flag can be inferred to be the value 1.

[0336] Since smaller block sizes involve more residual signals, the aforementioned specific conditions can be considered together, including the block aspect ratio in addition to the block size.

[0337] The above method can be applied to IBC mode in the same way. In other words, even when applying IBC mode to regular blocks, MH-intra-frame, MH-inter-frame, and MH-IBC blocks can be applied. Moreover, the aforementioned cu_coded_flag parsing / inference process can be applied depending on whether intra-frame mode is applied as a multi-reference block.

[0338] The following describes the cu_coded_flag parsing / inference method based on specific conditions when the prediction mode of the current block is intra-frame mode.

[0339] Figure 19 This is a flowchart illustrating a method for resolving a CU based on whether certain conditions are met when applying multiple reference blocks, according to an embodiment.

[0340] When the prediction mode of the current block (CU) is intra-frame mode, MH-intra-frame, MH-inter-frame, and MH-IBC blocks can be applied as multi-reference blocks. When applying multi-reference blocks, if certain conditions are met, cu_coded_flag is not sent as a signal, but is inferred to be 1. This is an existing rule and can be changed to send cu_coded_flag as a signal.

[0341] When the current block is an intra-frame block, a specific condition can be the number of multi-reference blocks transmitted by signaling. Even in the case of intra-frame blocks, cu_coded_flag can have a value of 0 when the residual signal from the multi-reference blocks is sufficiently reduced. Therefore, when the number of blocks with inter-frame or IBC modes among the multi-reference blocks transmitted by signaling is defined as numMHInter, this number can be considered in determining the cu_coded_flag value. In other words, Figure 19 The specific condition can be "if (numMHIntra>0)". When the condition is met, cu_coded_flag can be sent with a signal, and when the condition is not met, cu_coded_flag is not sent with a signal, but it can be inferred to be the value 1.

[0342] As another example, a specific condition could be "if (width × height > 1024)". In other words, when the block size (width × height) meets the condition, the cu_coded_flag can be signaled, and when the block size does not meet the condition, the cu_coded_flag can be inferred to be 1. This is because it can be assumed that blocks with smaller sizes have more residual signals. The two conditions mentioned above can be considered together, and in addition to the block size, the aspect ratio of the block can also be considered.

[0343] More specifically, see reference Figure 19When the current block is in intra-frame mode and multi-reference blocks are applied, `cu_coded_flag` can be resolved if a specific condition is met, and if the specific condition is not met, `cu_coded_flag` can be inferred to be 1. In this case, the specific condition can be the number of multi-reference blocks sent by signaling. Even when the current block is an intra-frame block, `cu_coded_flag` can have a value of 0 if the residual signal from the multi-reference blocks is sufficiently reduced. Therefore, the value of `cu_coded_flag` can be determined by the number of blocks with inter-frame or IBC mode among the signaled multi-reference blocks, `numMHInter`. In this case, the specific condition can include the condition that the number of signaled multi-reference blocks is greater than 0 (if (numMHIntra>0)). When the specific condition is met (i.e., when the number of signaled multi-reference blocks is greater than 0), `cu_coded_flag` can be sent by signaling, and when the specific condition is not met (i.e., when the number of signaled multi-reference blocks is 0), `cu_coded_flag` is not sent by signaling, but can be inferred to be 1.

[0344] As another example, the specific condition could be the block size. In this case, the specific condition could include a block size (width × height) greater than a threshold size (1024) (if (width × height) > 1024). When the specific condition is met (i.e., when the block size is greater than 1024), the cu_coded_flag can be signaled, and when the specific condition is not met (i.e., the block size is 1024 or smaller), the cu_coded_flag can be inferred to be the value 1.

[0345] Since smaller block sizes involve more residual signals, the aforementioned specific conditions can be considered together, and the block aspect ratio can also be considered in addition to the block size.

[0346] The following describes the interaction with other tools when multiple reference blocks are present. In particular, the MH-intra-frame mode can be considered as a condition used to determine the BS during the deblocking filter process.

[0347] The deblocking filter process may include: determining the length of the filter to be applied to the CU, TU, and sub-block boundaries; determining the strength of the filter to be applied to each boundary; determining the final filter length and coefficients by determining the variation of each sample and whether each sample is an actual edge; and applying the filter. The determination of the filter strength in this process will be described below.

[0348] Figure 20 This is a flowchart illustrating a method for determining a BS when multiple reference blocks are applied according to an embodiment.

[0349] When the current block is Q and the neighboring block is P, the BS (Block Scale) is first determined based on the prediction mode type of each block. Then, factors such as TB (Block Tolerance) edges, the presence of transform coefficients, and differences in motion vectors are considered to ultimately determine the BS value. Since intra-frame modes involve a relatively large amount of residual signal, high discontinuities between blocks are generally expected, resulting in high BS values. Similarly, CIIP (Concurrent Interpretive Injection) modes also include intra-frame prediction blocks and are therefore determined to have block characteristics similar to those of intra-frame modes, leading to high BS values.

[0350] In this disclosure, a multi-reference block containing an MH-intra-frame block can be identified as a block expected to have additional compression effects due to intra-frame prediction blocks because it contains many residual signals. Therefore, a block containing an MH-intra-frame block can be identified as an intra-frame block and thus assigned a large BS value.

[0351] BS can be allocated 2 when the current block Q and its neighboring block P are in intra-frame mode or CIIP mode, or when numMHIntra > 0. As defined above, numMHIntra can represent the number of intra-frame mode blocks in a multi-reference block transmitted by signaling. Moreover, as a modified definition, numMHIntra can represent the number of MH-intra blocks that the corresponding block has. This can include not only MH-intra modes transmitted by signaling, but also inherited MH-intra modes.

[0352] Furthermore, BS can be subdivided and defined as values ​​from 0 to 3. In this case, BS can be set to 3 in intra-frame mode and to 2 in combined mode. The combination method can be one of the following: CIIP mode GPU In-Frame Mode When the inter-frame / IBC mode is a regular block, the mode is obtained by combining multiple reference blocks. When the inter-frame / IBC mode is a regular block, including the MH-intra-frame mode. Here, MH-intraframe can represent the number of MH-intraframe modes transmitted by signaling, including inherited MH-intraframe modes.

[0353] When the intra-frame mode is a regular block, the mode is obtained by combining multiple reference blocks.

[0354] When the intra-frame mode is a regular block, including MH-inter-frame and MH-IBC modes.

[0355] Here, MH-inter-frame and MH-IBC can be the number of multi-reference blocks transmitted by signaling, including inherited modes.

[0356] More specifically, see reference Figure 20It determines whether to apply block-based incremental pulse code modulation (BDPCM) to the current block Q and the adjacent block P.

[0357] When BDPCM is not applied to the current block Q and the adjacent block P, it can be determined whether the current block Q and the adjacent block P are in intra-frame mode or CIIP mode. When the current block Q and the adjacent block P are in intra-frame mode or CIIP mode, BS can be set to 2 (BS=2).

[0358] When BDPCM is applied to the current block Q and the adjacent block P, BS can be set to 0 (BS=0).

[0359] When the current block Q and the adjacent block P are not in intra-frame mode or CIIP mode, it can be determined whether the current block Q and the adjacent block P are edges of the TB and whether the current block Q and the adjacent block P have transform coefficients. When the current block Q and the adjacent block P are edges of the TB and have transform coefficients, BS can be set to 1 (BS=1); otherwise, it can be determined whether the current block Q and the adjacent block P are different modes. When the current block Q and the adjacent block P are different modes, BS can be set to 1 (BS=1); otherwise, it can be determined whether the current block Q and the adjacent block P are in IBC mode and have different BV. When the current block Q and the adjacent block P are in IBC mode and have different BV, BS can be set to 1 (BS=1); otherwise, it can be determined whether the current block Q and the adjacent block P are in inter-frame mode and have different reference images or different MVs. When the current block Q and the adjacent block P are in inter-frame mode (INTER) and have different reference images or different MVs, BS can be set to 1 (BS=1); otherwise, it can be determined whether the current block Q and the adjacent block P are in inter-frame mode and whether the different MVs are greater than 1 / 2 luminance sample. When the current block Q and the adjacent block P are in inter-frame mode and the different MV is greater than 1 / 2 luminance sample, BS can be set to 1 (BS=1); otherwise, BS can be set to 0 (BS=0).

[0360] The following describes the interaction with other tools when multiple reference blocks are present. In particular, the multi-reference condition can be considered a condition for SBT application.

[0361] As shown in Table 9 below, the syntax table in Table 9 illustrates the SBT portion of the syntax structure of the coding unit (). When the current block is not in intra-frame mode, the transform can be applied in the sub-block unit. As shown in the syntax table in Table 9 below, when the current block is a CIIP with intra-frame characteristics and GPU intra-frame mode, SBT may not be applied. In GPU intra-frame mode, corner-based prediction blocks are partitioned, but one of the two prediction blocks includes an intra-frame prediction block.

[0362] In this disclosure, the MH-intra-prediction block also has intra-frame characteristics. Therefore, the presence of the MH-intra-frame block can be used to determine whether SBT should be applied.

[0363] [Table 9]

[0364] More specifically, SBT can be omitted when numMHIntra > 0 is satisfied. In other words, SBT-related syntax information can be included when numMHIntra == 0. As defined above, numMHIntra can represent the number of intra-mode prediction blocks in a multi-reference block transmitted as a signal.

[0365] The following describes the operation for multi-reference blocks when LIC is applied to the current block. LIC is a method for generating prediction blocks using parameters α and β derived from the pixel differences between the current block and the templates adjacent to each reference block. In other words, when the reference sample is p[x], the prediction sample for the current block can be derived as α. p[x]+β.

[0366] Applying LIC to the current block indicates a local change between the current block and the reference block in the reference image. In this case, due to the LIC properties based on template-derived parameters, the final predicted block can be generated by considering not only the predicted block in the reference image but also the changes in neighboring samples. LIC can be applied to both IBC mode and inter-frame mode.

[0367] The following describes how to modify the application of multiple reference blocks when applying LIC.

[0368] When applying LIC to the current block, multiple reference blocks (MPLCs) may not be allowed to be signaled. Specifically, intra-frame blocks may not be allowed as MPLCs.

[0369] When applying a LIC to the current block, inheriting multiple reference blocks may not be allowed. Specifically, intra-blocks may not be allowed as inherited multiple reference blocks.

[0370] When LIC is applied to the current block, LIC parameters can be derived and applied to multi-reference blocks. As shown in Equation 9 below, when LIC is applied to a regular block p0[x], the regular block to which LIC is applied can be defined as p0'[x], and when LIC is applied to a multi-reference block p2[x], the multi-reference block to which LIC is applied can be defined as p2'[x]. These can be used to derive the final prediction block p[x]. In other words, the method for deriving the final prediction block can be presented as shown in Equation 9 below.

[0371] [Formula 9]

[0372] In Equation 9, for ease of description, the process of deriving the final prediction block has been described as an example for the case of a unidirectional prediction block. However, the present disclosure is not limited thereto, and the same / similar method can also be applied to a bidirectional prediction block.

[0373] When applying LIC to the current block, LIC parameters can be derived and applied to multiple reference blocks. Here, additional LIC flag signaling may not be required, and the weight information of the multiple reference blocks may not be included. When the weight information of the multiple reference blocks is not included, a default value can be used as the weight for the weighted sum of the regular (prediction) block and the multiple reference blocks. Here, the default value can be equal weights, or different default values can be set for the mode types of the multiple reference blocks. As an example, when the multiple reference block is an inter-frame mode, the default value can be changed and applied as equal weights (16 / 32), and when the multiple reference block is an intra-frame mode, the default value can be changed and applied as equal weights (14 / 32).

[0374] When applying LIC to the current block, the candidate weight information of the intra-frame block used as the multiple reference block can vary. Since the LIC parameters are derived using a template, the LIC parameters can partially reflect information about adjacent blocks. As an example, the weight information with a definition of {1 / 4, -1 / 8} for the multiple reference block can be changed and defined as {1 / 16, 1 / 32}, etc. for use.

[0375] In the case of applying LIC to the current block, if the difference between the parameter α (α0) of the regular block and the LIC parameter α (α2) of the multiple reference block is less than the threshold T0, the multiple reference block can be applied. In particular, this can be applied to an inherited multiple reference block. When the inherited multiple reference block is in an inter-frame or IBC mode, under specific conditions, the multiple reference block may not be applied to the current block. The specific conditions may include the case where the absolute value of the difference between the parameter α (α0) of the regular block and the LIC parameter α (α2) of the multiple reference block is less than the threshold T0 (|α0 - α2| < T0). This can also be applied to the parameter β in the same / similar manner. When the difference between the parameters β of the two blocks is less than the threshold T1, the multiple reference block can be applied. In this case, the specific conditions may include the case where the absolute value of the difference between the parameter β (β0) of the regular block and the LIC parameter β (β2) of the multiple reference block is less than the threshold T1 (|β0 - β2| < T1).

[0376] When LIC is applied to the current block, the weights of intra-blocks used as multi-reference blocks can be set larger. Due to the application of LIC, neighboring samples of the current block can be more accurate than those of inter-frame / IBC prediction blocks. In this case, when the MH-intra-frame weight of a block without LIC is 16 / 32, the MH-intra-frame weight of a block with LIC can be set to a larger value, such as 20 / 32. This can be modified and is limited to cases where the intra-block is used as an inherited multi-reference block, etc.

[0377] When LIC is applied to the current block, the template region of the regular reference block p0 can differ from the template region of the multi-reference block p2. This will reference... Figure 21 Describe it.

[0378] Figure 21 This is a diagram illustrating examples of template regions for conventional reference blocks and template regions for multiple reference blocks according to an embodiment.

[0379] refer to Figure 21 When the template region used to derive parameter p0 is limited to the upper side of the regular block p0[x], the template region used to derive parameter p2 can be modified and limited to the left side of the multi-reference block p2[x], etc.

[0380] In the presence of multiple reference blocks, if the signaling information of the multiple reference blocks precedes the LIC, the LIC of the current block can be restricted. This can be modified and applied depending on the type of multiple reference blocks. As an example, when an MH-frame is included in a multiple reference block, the LIC of the current block can be restricted.

[0381] The following describes the operation of OBMC for multi-reference blocks when applied to the current block. OBMC is a method for removing discontinuities caused by motion compensation using motion vectors from each predicted block from block edges. It can effectively remove discontinuities between blocks when the block to which OBMC is applied has a different motion vector than its neighbors due to the weighted sum of predicted blocks obtained using motion vectors from neighboring blocks and predicted blocks obtained using motion vectors from the current block. OBMC can be applied not only to inter-frame modes but also to IBC modes.

[0382] This disclosure describes a method for modifying the application of multiple reference blocks when the OBMC is applied to the current block.

[0383] When OBMC is applied to the current block, multi-reference blocks that can be signaled may not be allowed. Specifically, intra-frame blocks may not be allowed as multi-reference blocks. Since intra-frame blocks use prediction patterns derived from neighboring samples, discontinuities between blocks can be relatively eliminated.

[0384] When OBMC is applied to the current block, inherited multi-reference blocks may not be allowed. Specifically, intra-frame blocks may not be allowed as multi-reference blocks. Since intra-frame blocks use prediction patterns derived from neighboring samples, discontinuities between blocks can be relatively eliminated.

[0385] When OBMC is applied to the current block, it may not include weight information for multi-reference blocks. When multi-reference block weight information is not included, default weight values ​​can be used as a weighted sum of regular blocks (i.e., regular prediction blocks) and multi-reference blocks. Here, the default value can be equal weights, or it can be set differently depending on the mode type of the multi-reference blocks. As an example, when the multi-reference blocks are in inter-frame mode, the default value can be changed and applied as equal weights (16 / 32), and when the multi-reference blocks are in intra-frame mode, the default value can be changed and applied as equal weights (14 / 32).

[0386] When OBMC is applied to the current block, the candidate weight information for the intra-block used as a multi-reference block can be changed. The application of OBMC can eliminate discontinuities with neighboring samples to some extent. As an example, the weight information of the multi-reference block, defined as {1 / 4, -1 / 8}, can be changed to {1 / 16, 1 / 32}, etc., for use.

[0387] In the presence of multiple reference blocks, if the signaling information of the multiple reference blocks precedes the OBMC, the OBMC of the current block can be restricted. Alternatively, when multiple reference blocks exist and include intra-frame blocks, this can be modified and applied to restrict the OBMC.

[0388] In the presence of multiple reference blocks, if the signaling information of the multiple reference blocks precedes the OBMC, the OBMC execution can be restricted to the CU of the current block, while the OBMC execution can be maintained at the sub-block level.

[0389] The following describes the method for selecting the LFNST core when LFNST is applicable to multiple reference blocks. In the case of multiple reference blocks, one or more intra-frame modes can exist in the reference blocks, and multiple reference blocks can be considered simultaneously with intra-frame modes. However, in some cases, the characteristics of the residual signal may differ from the characteristics of the intra-frame modes defined in LFNST. Therefore, when the regular block (i.e., the regular reference block) is an inter-frame or IBC mode and multiple reference blocks exist, the characteristics of the residual signal can be effectively considered by applying whether to select (or apply) the LFNST core, as described below.

[0390] All multi-reference blocks are in intra-mode or include one or more intra-mode cases: When the orientations of intra-frame modes are similar, the intra-frame mode representing the intra-frame mode can be determined, and the LFNST kernel can be used based on that mode. When the orientations of intra-frame modes are inconsistent, the LFNST kernel corresponding to the intra-frame mode of the first multi-reference block can be used. Alternatively, when the orientations of the modes are inconsistent, LFNST may not be applied, or a kernel for a planar mode may be used.

[0391] All cases where multiple reference blocks are in inter-frame / IBC mode: LFNST can be omitted, or a planar kernel can be used.

[0392] The case where multiple reference blocks include one intra-frame mode and one inter-frame / IBC mode: The LFNST core can be selected based on a given intra-frame mode.

[0393] When the regular block is in intra-frame mode and multiple reference blocks exist, the characteristics of the residual signal can be effectively considered by applying whether to select (or apply) the LFNST core, as described below.

[0394] All multi-reference blocks are in intra-mode or include one or more intra-mode cases: The LFNST core can be used in the intra-frame mode based on regular blocks.

[0395] Alternatively, when the intra-mode of the regular block and each intra-mode have similar orientations, an intra-mode representing the intra-mode can be determined, and an LFNST kernel can be used based on that mode. Alternatively, when the orientations of the intra-modes are inconsistent, an LFNST kernel based on the intra-mode of the regular block can be used, LFNST can be omitted, or a kernel based on a planar mode can be used.

[0396] All cases where multiple reference blocks are in inter-frame / IBC mode: The LFNST core can be used in the intra-frame mode based on regular blocks.

[0397] Alternatively, LFNST can be omitted, or a planar kernel can be used.

[0398] The case where multiple reference blocks include one intra-frame mode and one inter-frame / IBC mode: The LFNST core can be used according to the intra-frame mode of regular blocks.

[0399] Alternatively, when the intra-mode of a regular block and the intra-mode of a multi-reference block have similar orientations, an intra-mode representing the intra-mode can be determined, and an LFNST kernel can be used based on that mode. Otherwise, LFNST can be omitted, or a planar mode kernel can be used.

[0400] Alternatively, when the intra-frame mode of a regular block and the intra-frame mode of a multi-reference block have similar orientations, a kernel based on the intra-frame mode of the regular block can be used; otherwise, a kernel based on the planar mode can be used.

[0401] The following describes in further detail: 1) a method for determining whether to parse or infer transform-related flags; 2) a method for determining the filter strength; and 3) a method for determining whether to apply SBT when an intra-frame block exits as the current block and / or a multiple reference block. Furthermore, the following describes in further detail: 4) a method for determining whether multiple reference blocks are allowed when LIC is applied to the current block; 5) a method for determining whether multiple reference blocks are allowed when OBMC is applied to the current block; and 6) a method for determining the LFNST core when LFNST is applied to multiple reference blocks. The methods described below can be performed by the encoding device 200 and / or the decoding device 300.

[0402] Example 1: When an intra-frame block exists as the current block and / or multiple reference blocks, determine whether to resolve or infer the transform. Methods of related marking

[0403] When the current block is in inter-frame (AMVP) mode and multiple reference blocks exist, the determination of whether to parse or infer transform-related flags can depend on whether a first condition is met. The transform-related flag indicates whether the current block includes transform and residual information and can be a `cu_coded_flag`. The first condition is a condition used to determine whether to parse or infer transform-related flags and can include conditions based on the number of multiple reference blocks (`numMHP`), the number of reference blocks in the intra-frame mode among the multiple reference blocks transmitted with signals (`numMHIntra`), or the block size (i.e., the CU size). For example, the first condition can include the case where `numMHP` is 0 (if (num MHP == 0)), the case where `numMHIntra` is 0 (if (num MHIntra == 0)), and / or the case where the block size (width × height) is greater than a specific value (if (width × height) > 1024). The specific value can be, but is not limited to, 1024 and can be a value used to determine blocks with few residual signals. According to various embodiments, the first condition can also include a condition based on the block's aspect ratio.

[0404] When the first condition is met (i.e., numMHP is 0), cu_coded_flag can be parsed, and when the first condition is not met (i.e., numMHP is greater than 0), cu_coded_flag can be inferred to have a value of 1. In this case, there can be multi-reference blocks sent with signals and inherited multi-reference blocks. However, the type and number of multi-reference blocks are unknown during the syntax parsing process. Therefore, when the number of multi-reference blocks sent with signals is greater than 0, the current block can be determined to be a block with a large amount of information to be sent with signals due to many residual signals. When the current block is determined to be a block with many residual signals, cu_coded_flag can be inferred to have a value of 1 and not sent with signals, and when the number of multi-reference blocks sent with signals is 0, cu_coded_flag can be parsed to have a value of 0 or 1.

[0405] As another example, when the first condition is met (i.e., numMHIntra is 0), the current block is determined to be a block with few residual signals, and therefore cu_coded_flag can be resolved. When the first condition is not met (i.e., numMHIntra is greater than 0), the current block is determined to be a block with many residual signals, and therefore cu_coded_flag can be inferred to be a value of 1. In other words, cu_coded_flag is inferred to be a value of 1 and is not sent with a signal.

[0406] As another example, cu_coded_flag can be parsed when the first condition is met (i.e., the width × height value is greater than 1024), and when the first condition is not met (i.e., the width × height value is 1024 or less), cu_coded_flag can be inferred to be 1.

[0407] According to various embodiments, conditions included in the first condition can be considered together. For example, the if (numMHP == 0) condition, the if (numMHIntra == 0) condition, and / or the if (width × height) > 1024) condition can be considered together, and it can be determined whether to parse or infer the transformation-related flags depending on whether these conditions are met.

[0408] According to various embodiments, when the current block is in merge mode and multiple reference blocks exist, cu_coded_flag can be inferred to be a value of 1 and not sent with a signal. When the current block is in merge mode, residual signals are included, and therefore cu_coded_flag can be inferred to be a value of 1 and not sent with a signal.

[0409] According to various embodiments, when the current block is in skip mode, the residual signal is not included, and therefore cu_coded_flag can be inferred to be a value of 0 and is not sent with a signal.

[0410] According to various embodiments, when the current block is in intra-frame mode, there are no multiple reference blocks, and therefore cu_coded_flag can be inferred to be a value of 1 and is not sent with a signal.

[0411] According to various embodiments, the following describes a method for resolving or inferring transformation-related flags when the current block is an intra-frame block if multiple reference blocks exist.

[0412] When the current block is an intra-frame block, if multiple reference blocks exist, it can be determined whether to parse or infer the transform-related flag under a predetermined second condition. In other words, according to the basic method, when the current block is an intra-frame block, cu_coded_flag is inferred to be a value of 1 and is not transmitted with a signal, while according to various embodiments, it can be determined whether to parse cu_coded_flag depending on whether a second condition is met. Here, the second condition includes a condition based on the number of blocks numMHInter, which is the inter-frame / IBC mode or block size among the multiple reference blocks transmitted with a signal. For example, the second condition may include the case where numMHInter is greater than 0 (if (num MHInter>0)) and / or the case where the block size (width × height) is greater than a specific value (e.g., 1024) (if (width × height)>1024). In this case, even when the current block is in intra-frame mode, if the multiple reference blocks include blocks in inter-frame mode or IBC mode, it can be determined that the residual signal is sufficiently reduced, and therefore whether to parse or infer cu_coded_flag under this second condition. According to various embodiments, the second condition may also include a condition based on the block aspect ratio.

[0413] When the second condition is met (i.e., numMHIntra is greater than 0), cu_coded_flag can be parsed, and when the second condition is not met (i.e., numMHIntra is 0), cu_coded_flag can be inferred to be 1 without sending a signal.

[0414] As another example, `cu_coded_flag` can be parsed when the second condition is met (i.e., the width × height value is greater than 1024), and when the second condition is not met (i.e., the width × height value is 1024 or less), `cu_coded_flag` can be inferred as a value of 1 without signaling. In this case, since a smaller block size may be considered to contain more residual signals, `cu_coded_flag` may be inferred as a value of 1 without signaling when the block size is considered to contain many residual signals.

[0415] According to various embodiments, the two conditions included in the second condition can be considered together. For example, the condition if (numMHInter>0) and the condition if (width×height)>1024) can be considered together to determine whether to parse or infer cu_coded_flag.

[0416] Example 2: A method for determining filter intensity based on multi-reference block intra-frame mode

[0417] When multiple reference blocks exist, whether the multiple reference blocks are in intra-frame mode (MH-intra-frame) can be regarded as a condition for determining the filtering strength during the deblocking filtering process.

[0418] During the filtering process using a deblocking filter, the BS value of the edge can be determined first based on the prediction modes of the current block and each adjacent block. Then, by considering whether the block edge is a transform block edge, whether there are transform coefficients, the difference in motion vectors, etc., the BS value is finally determined.

[0419] When the multi-reference block is in intra-frame mode, the current block and / or neighboring blocks include numerous residual signals. Therefore, the current block can be identified as the block expected to have additional compression effects. In this case, the current block, which contains multi-reference blocks and / or neighboring blocks in intra-frame mode, can be identified as an intra-frame block, and the BS value can be determined (or set or assigned) based on this.

[0420] According to embodiments of this disclosure, the BS value can be determined based on prediction information related to the current block, the predicted block, or multiple reference blocks. For example, the prediction information may include information about the prediction mode applied to the current block, information about the generation method (or tool) used to generate the predicted block, and / or information about the prediction mode applied to multiple reference blocks. For example, the prediction information may include information indicating whether a CIIP mode is applied to the current block.

[0421] The BS value can be set to a first value when certain conditions are met. These conditions may include the current block and / or adjacent blocks being in intra-frame mode or CIIP mode, or the number of intra-frame mode blocks (numMHIntra) in a multi-reference block being greater than 0. Here, the first value is the largest value within a specific range, such as 2 or 3, but not limited to these. As an example, when the BS value has a range of 0 to 2, the maximum value can be 2 and the minimum value can be 0. As another example, when the BS value has a range of 0 to 3, the maximum value can be 3 and the minimum value can be 0. According to various embodiments, in addition to the foregoing definition, numMHIntra can represent the number of intra-frame mode (reference) blocks in a multi-reference block transmitted by signaling or the number of intra-frame mode (reference) blocks in a inherited multi-reference block.

[0422] When the BS value is in the range of 0 to 3, if the current block and / or adjacent blocks are in intra-frame mode, the BS value can be set to 3, and if the current block and / or adjacent blocks are in intra-frame mode or the predicted block is in combination mode, the BS value can be set to 2. Here, combination mode can be CIIP mode, GPU intra-frame mode, a mode that combines multiple reference blocks with a regular block (i.e., regular reference block) when the regular block is in inter-frame / IBC mode, a mode that includes blocks in intra-frame mode (MH-intra) among multiple reference blocks that are signaled or inherited when the regular block is in inter-frame / IBC mode, a mode that combines multiple reference blocks with a regular block when the regular block is in intra-frame mode, and / or a mode that includes blocks in inter-frame mode (MH-inter) and IBC mode (MH-IBC) among multiple reference blocks that are signaled or inherited when the regular block is in intra-frame mode.

[0423] According to various embodiments, it can be determined whether the current block and / or adjacent blocks are in BDPCM mode, and when the current block and / or adjacent blocks are in BDPCM mode, the BS value can be set to 0. Otherwise, it can be determined based on whether the current block and / or adjacent blocks are in intra-frame mode or whether the predicted block is in CIIP mode.

[0424] The BS value can be set to 2 when the current block and / or adjacent blocks are in intra-frame mode or the predicted block is in CIIP mode. Otherwise, it can be determined based on whether the block edge is a transform block edge and whether the current block and / or adjacent blocks have transform coefficients.

[0425] The BS value can be set to 1 when the block edge is a transform block edge and the current block and / or adjacent blocks have transform coefficients. Otherwise, it can be determined whether the current block and / or adjacent blocks are different modes.

[0426] The BS value can be set to 1 when the current block and / or adjacent blocks are in different modes. Otherwise, it can be determined whether the current block and / or adjacent blocks are in IBC mode and have different BVs.

[0427] The BS value can be set to 1 when the current block and / or adjacent blocks are in IBC mode and have different BVs. Otherwise, it can be determined whether the current block and / or adjacent blocks are in inter-frame mode and have different reference images or different motion vectors.

[0428] The BS value can be set to 1 when the current block and / or adjacent blocks are in inter-frame mode and have different reference images or different motion vectors. Otherwise, the BS value can be determined based on whether the current block and / or adjacent blocks are in inter-frame mode and whether the different motion vectors are greater than 1 / 2 luma sample. The BS value can be set to 1 when the current block and / or adjacent blocks are in inter-frame mode and the different motion vectors are greater than 1 / 2 luma sample. Otherwise, the BS value can be set to 0.

[0429] Example 3: Whether to apply SBT depends on whether there is a multi-reference block intra-frame mode as an intra-frame mode. Method

[0430] SBT can be applied (or enabled) to the current block when it is not in intra-frame mode. Furthermore, since CIIP mode and GPU intra-frame mode include intra-frame mode features, SBT may not be applied.

[0431] When a block as an intra-frame mode (MH-intra) exists within a multi-reference block, intra-frame mode features are included. Therefore, in this disclosure, the application of SBT can be determined based on the presence of an MH-intra block.

[0432] In other words, the application of SBT can be determined under specific conditions used to determine the presence of an intra-MG block. Here, the specific conditions can be based on the number of intra-mode blocks, numMHIntra, which are multiple reference blocks. For example, the specific conditions could include the case where numMHIntra is greater than 0 (if (num MHIntra>0)).

[0433] When certain conditions are met (i.e., numMHIntra is greater than 0), SBT may not be applied, and the relevant syntax information of SBT may not be included in the CU syntax structure, as shown in Table 9.

[0434] SBT can be applied when a specific condition is not met (i.e., numMHIntra is 0), and the relevant syntax information of SBT can be included in the CU syntax structure in Table 9.

[0435] Example 4: Method for determining whether multiple reference blocks are allowed when LIC is applied to the current block

[0436] Applying LIC to the current block indicates a local illumination change between the reference block in the reference image and the current block. Because LIC uses a scaling factor and an offset to generate the prediction block—parameters derived from the pixel differences between the current block and the template adjacent to each reference block—it can consider not only the illumination changes of the reference block in the reference image but also the illumination changes of adjacent blocks to generate the final prediction block.

[0437] In other words, when applying LIC to the current block, the process of generating the final predicted block may include applying LIC to a regular block (i.e., a regular reference block) P0[x] to generate a regular block P0'[x] with LIC applied, applying LIC to an additional reference block P2[x] to generate an additional reference block P2'[x] with LIC applied, and calculating a weighted sum of P0'[x] and P2'[x] to generate the final predicted block P[x]. Here, a one-way predicted block (e.g., ...) is applied. Figure 8The cases of P0 and P2 are described as examples, but the same / similar methods can also be applied to cases where bidirectional prediction blocks are applied.

[0438] When LIC is applied to the current block, the weights used to calculate the weighted sum of P0'[x] and P2'[x] can be default values. These default values ​​can be (but are not limited to) equal weights and can contain different weight values ​​depending on the mode type of the multi-reference block. As an example, when the multi-reference block is in inter-frame mode (which can indicate both MH merge mode and MH AMVP mode), the default values ​​can have equal weights (16 / 32), and when the multi-reference block is in intra-frame mode, the default values ​​can be changed and applied to equal weights (14 / 32).

[0439] According to various embodiments, when LIC is applied to the current block, LIC parameters can be derived using a multi-reference block. Here, additional LIC flag signaling is not required, and the weight information of the multi-reference block (as described in the aforementioned encoded information) may not be included. When the weight information of the multi-reference block is not included, default values ​​can be used as the weights of the weighted sum of the regular block and the multi-reference block in the same / similar manner as described above. Signaling information for the multi-reference block may exist (e.g., MVP index, weight index, etc. in MH merging mode), and when the LIC flag of the current block is true, LIC and the weighted sum are partially complementary to each other. Therefore, the amount of information transmitted by signaling can be reduced by omitting information such as the weight index of the multi-reference block.

[0440] According to various embodiments, when a LIC is applied to the current block, the weight information of the intra-block used as a multi-reference block can include different candidate values ​​compared to the weight information of the intra-block used as a multi-reference block when no LIC is applied. Since the LIC parameters are derived using the template region of a regular reference block, they partially reflect information from neighboring blocks (illumination, resolution, etc.). Therefore, for example, the weight information of the multi-reference block, defined as {1 / 4, -1 / 8}, can be changed and defined as {1 / 16, 1 / 32}, etc., for use.

[0441] According to various embodiments, in the case where LIC is applied to a current block, if the difference between the scaling factor value α0 of a regular block and the scaling factor value α2 of a multi-reference block is less than a first threshold T0, the multi-reference block may be applied. In particular, this may be applied to an inherited multi-reference block. When the inherited multi-reference block is in an inter-frame or IBC mode, under a specific condition (|α0 - α2| < T0), the multi-reference block may or may not be applied to the current block. When a specific condition is satisfied (i.e., |α0 - α2| is less than T0), the inherited multi-reference block may be applied to the current block. When the specific condition is not satisfied (i.e., |α0 - α2| is T0 or greater), the inherited multi-reference block may not be applied to the current block. This may be applied in the same / similar manner to the offset β. In other words, when the difference between the offset value β0 of a regular block and the offset value β2 of a multi-reference block is less than a second threshold T1, the multi-reference block may be applied. The inherited multi-reference block may or may not be applied to the current block under a specific condition (|β0 - β2| < T1). When a specific condition is satisfied (i.e., |β0 - β2| is less than T1), the inherited multi-reference block may be applied to the current block. When the specific condition is not satisfied (i.e., |β0 - β2| is T1 or greater), the inherited multi-reference block may not be applied to the current block.

[0442] According to various embodiments, when LIC is applied to a current block, the weight of an intra-block used as a multi-reference block may be set to be greater than the weight of an intra-block used as a multi-reference block when LIC is not applied. Due to the application of LIC, the neighboring samples (or neighboring blocks) of the current block may be more accurate than the neighboring blocks of an inter-frame / IBC block. For this reason, the weight of the MH_intra-block to which LIC is applied may be set to be greater than the weight of the MH_intra-block to which LIC is not applied. For example, when the weight of the MH_intra-block to which LIC is not applied is 16 / 32, the weight of the MH_intra-block to which LIC is applied may be set to 20 / 32. This may be modified and is limited to cases such as when an intra-block is used as an inherited multi-reference block.

[0443] According to various embodiments, when LIC is applied to a current block, the template region of a regular reference block p0 for deriving LIC parameters may be different from the template region of a multi-reference block p2. For example, as Figure 21 shown, the template region for deriving the LIC parameters of p0 may be limited to the upper template region of p0. In addition, the template region for deriving the LIC parameters of p1 may be limited to the left template region of p2.

[0444] According to various embodiments, when LIC is applied to a current block, it may not be allowed (to apply) the multi-reference block signaled by the current block. Specifically, it may not be allowed for an intra-block to be a multi-reference block. In other words, for the current block to which LIC is applied, there may be no intra-block among the multi-reference blocks.

[0445] According to various embodiments, when a LIC is applied to the current block, multiple reference blocks inherited by the current block may not be allowed. Specifically, intra blocks may not be allowed as inherited multiple reference blocks. In other words, for the current block to which the LIC is applied, there may not be any intra blocks among the inherited multiple reference blocks.

[0446] Depending on the various embodiments where multiple reference blocks exist, the LIC of the current block can be restricted if the signaling / parsing information of the multiple reference blocks precedes the LIC. This can vary depending on the type of multiple reference blocks. As an example, when an intra-frame block exists as a multiple reference block, the LIC of the current block can be restricted.

[0447] Example 5: Method for determining whether multiple reference blocks are allowed when OBMC is applied to the current block.

[0448] As mentioned above, when applying OBMC, discontinuities at block edges (boundaries) can be removed by calculating a weighted sum of the predicted blocks obtained using the motion vectors of neighboring blocks and the predicted blocks obtained using the motion vector of the current block. Here, discontinuities between blocks can be effectively removed when the motion vectors have different motion vector values.

[0449] When OBMC is applied to the current block, multiple reference blocks that may not be allowed to be signaled for the current block may be excluded. Specifically, intra-frame blocks may not be allowed as multiple reference blocks. In other words, for the current block to which OBMC is applied, intra-frame blocks may not exist among the multiple reference blocks. Since intra-frame blocks use prediction patterns derived from neighboring samples, discontinuities between blocks can be relatively eliminated.

[0450] According to various embodiments, when OBMC is applied to the current block, multiple reference blocks inherited by the current block may not be allowed. Specifically, intra blocks may not be allowed as inherited multiple reference blocks. In other words, for the current block to which OBMC is applied, there may not be any intra blocks among the inherited multiple reference blocks.

[0451] According to various embodiments, when OBMC is applied to the current block, the weight information of the multi-reference block may not be included in the aforementioned coding information. When the aforementioned coding information does not include weight information, the weighted sum of the regular block (i.e., the regular prediction block) and the multi-reference block can use default weight values. The default value may be, but is not limited to, equal weights, or may include different default values ​​depending on the mode type of the multi-reference block. As an example, when the multi-reference block is in inter-frame mode, the default value may have an equal weight of 16 / 32, and when the multi-reference block is in intra-frame mode, the default value may be changed and applied to an equal weight such as 14 / 32.

[0452] According to various embodiments, when OBMC is applied to the current block, the weight information of the intra-block used as a multi-reference block can include different candidate values. This is because the application of OBMC can, to some extent, eliminate discontinuities with adjacent blocks. Therefore, the weight information of the multi-reference block, defined as {1 / 4, -1 / 8}, can be changed and defined as {1 / 16, 1 / 32}, etc., for use.

[0453] According to various embodiments, in the presence of multiple reference blocks, if the signaling information of the multiple reference blocks precedes the OBMC, the IBMC of the current block can be restricted. According to various embodiments, this can be modified and applied to restrict the OBMC when an intra-frame block exists as a multiple reference block.

[0454] According to various embodiments, in the presence of multiple reference blocks, if the signaling information of the multiple reference blocks precedes the OBMC, the OBMC executed on a CU-by-CU basis of the current block can be restricted, while the OBMC executed on a sub-block-by-sub-block basis can be maintained.

[0455] Example 6: Method for determining the LFNST core when the Secondary Transform (LFNST) is applicable to multiple reference blocks.

[0456] When multiple reference blocks exist, one or more reference blocks that are intra-frame modes may exist within the multiple reference blocks. As an example, all multiple reference blocks may be intra-frame modes, or one or more reference blocks that are intra-frame modes may exist. As another example, none of the multiple reference blocks may be intra-frame modes. As another example, a regular block (i.e., a regular reference block) may be inter-frame or IBC modes, and one or more multiple reference blocks that are intra-frame modes may exist. As another example, a regular block may be intra-frame mode, and one or more multiple reference blocks that are intra-frame modes may exist. In this case, the method for determining the transform kernel used for the inverse transform will be described in detail.

[0457] Where all multi-reference blocks are intra-mode or include one or more intra-mode blocks: A representative intra-frame mode can be determined based on the directionality of each intra-frame mode, and the LFNST kernel based on the representative intra-frame mode can be used as a transform kernel. As an example, a representative intra-frame mode can be determined within the intra-frame modes based on the matching of the directionality of each intra-frame mode, and the LFNST kernel based on the representative intra-frame mode can be used as a transform kernel.

[0458] Any intra-mode can be selected based on the directionality of each intra-mode, and the LFNST kernel based on the selected intra-mode can be used as a transform kernel. As an example, when the directionality of each intra-mode is inconsistent, the LFNST kernel based on the intra-mode of any one of the multiple reference blocks can be used as a transform kernel. For example, any one reference block can be the first reference block among the multiple reference blocks.

[0459] LFNST can be applied without considering the directionality of each intra-frame mode, or a planar mode kernel can be used as the transform kernel. As an example, when the directionality of modes is inconsistent, LFNST may not be applied, or a planar mode kernel may be used as the transform kernel.

[0460] 2) Case where all multi-reference blocks are in inter-frame / IBC mode: LFNST can be omitted, or a planar kernel can be used as a transform kernel.

[0461] 3) The case where multiple reference blocks include one intra-frame mode and one inter-frame / IBC mode: The LFNST kernel based on intra-frame mode can be used as a transform kernel.

[0462] 4) The case where the regular block is in intra mode and all multiple reference blocks are in intra mode or include one or more intra modes: The LFNST kernel, based on the intra-frame mode of regular blocks, can be used as a transform kernel.

[0463] A representative intra-frame mode can be determined based on the directionality of the intra-frame modes in a regular block and each intra-frame mode, and the LFNST kernel based on the representative intra-frame mode can be used as a transform kernel. As an example, by matching the directionality of the intra-frame modes in a regular block with the directionality of each intra-frame mode, a representative intra-frame mode can be determined within the intra-frame modes, and the LFNST kernel based on the representative intra-frame mode can be used as a transform kernel.

[0464] The LFNST kernel based on regular block intra modes can be used as the transform kernel for both the directionality of regular block intra modes and each intra mode. As an example, the LFNST kernel based on regular block intra modes can be used as a transform kernel when the directionality of regular block intra modes and each intra mode is inconsistent.

[0465] Based on the directionality of the intra-mode of the regular block and each intra-mode, LFNST may not be applied, or a planar mode kernel may be used as the transform kernel. As an example, when the directionality of the intra-mode of the regular block and each intra-mode is inconsistent, LFNST may not be applied, or a planar mode kernel may be used as the transform kernel.

[0466] 5) The case where the regular block is in intra-frame mode, and multiple reference blocks are in inter-frame / IBC mode: The LFNST kernel, based on the intra-frame mode of regular blocks, can be used as a transform kernel.

[0467] LFNST can be omitted, or a planar kernel can be used as a transform kernel.

[0468] 6) The case where the regular block is in intra-frame mode, and the multiple reference block includes one intra-frame mode and one inter-frame / IBC mode: The LFNST kernel, based on the intra-frame mode of regular blocks, can be used as a transform kernel.

[0469] Matching the directionality of intra-modes based on regular blocks with the directionality of each intra-mode allows for the determination of representative intra-modes within the intra-modes, and the LFNST kernel based on the representative intra-modes can be used as a transform kernel.

[0470] When the directionality of the intra-mode of a regular block is similar to that of each intra-mode, the LFNST kernel based on the intra-mode of the regular block can be used as a transform kernel.

[0471] When the intra-mode of a regular block and the directionality of each intra-mode are inconsistent, LFNST may not be applied, or the kernel of the planar mode may be used as the transform kernel.

[0472] In the above embodiments, the method is described as a series of steps or blocks based on the flowchart; however, the embodiments are not limited to the order of the steps. Some steps may occur in a different order or simultaneously with other steps. Furthermore, those skilled in the art will understand that the steps shown in the flowchart are not exclusive and may include other steps or one or more steps of the flowchart may be deleted without affecting the scope of the embodiments described herein.

[0473] The methods according to the above embodiments of the present disclosure can be implemented in software, and the encoding and / or decoding devices according to the present disclosure can be included in a device for performing image processing, such as a television (TV), a computer, a smartphone, a set-top box, a display device, etc.

[0474] When the embodiments of this document are implemented in software, the methods described above can be implemented using modules (processes, functions, etc.) that perform the functions described above. Modules can be stored in memory and executed by a processor. Memory can be internal or external to the processor and connected to the processor via various known means. The processor may include an application-specific integrated circuit (ASIC), another chipset, logic circuitry, and / or data processing devices. Memory may include read-only memory (ROM), random access memory (RAM), flash memory, memory cards, storage media, and / or other storage devices. In other words, the embodiments described in this disclosure can be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units shown in each figure can be implemented and executed on a computer, processor, microprocessor, controller, or chip. In this case, information for implementation (e.g., information about instructions) or algorithms can be stored in a digital storage medium.

[0475] Furthermore, the decoding and encoding devices used in the embodiments of this disclosure can be included in multimedia broadcast transceivers, mobile communication terminals, home theater video equipment, digital cinema video equipment, surveillance cameras, video chat devices, real-time communication devices such as video communication devices, mobile streaming media devices, storage media, cameras, video-on-demand (VoD) service providers, over-the-air (OTT) video devices, internet streaming media service providers, three-dimensional (3D) video devices, virtual reality (VR) devices, augmented reality (AR) devices, video telephony devices, vehicle terminals (e.g., vehicle (including autonomous vehicles) terminals, aircraft terminals, ship terminals, etc.), medical video equipment, etc., and can be used to process video signals or data signals. For example, OTT video devices can include game consoles, Blu-ray players, internet access TVs, home theater systems, smartphones, tablet PCs, digital video recorders (DVRs), etc.

[0476] Furthermore, the processing methods applied in the embodiments of this specification can be generated in the form of a program executed by a computer and stored in a computer-readable recording medium. Multimedia data having the data structure according to the embodiments of this specification can also be stored in a computer-readable recording medium. Computer-readable recording media include various storage devices and distributed storage devices. Computer-readable recording media can include, for example, Blu-ray discs (BD), Universal Serial Bus (USB) devices, ROMs, programmable ROMs (PROMs), erasable programmable ROMs (EPROMs), RAM, optical disc (CD)-ROMs, magnetic tapes, floppy disks, and optical data storage devices. Additionally, computer-readable recording media include media implemented in the form of a carrier (e.g., transmission via the Internet). Bitstreams generated by encoding methods can be stored in computer-readable recording media or transmitted via wired or wireless communication networks.

[0477] The embodiments of this specification can be implemented as computer program products of program code, and the program code can be executed on a computer according to the embodiments of this specification. The program code can be stored on a computer-readable medium.

[0478] Figure 22 Examples of content streaming systems to which embodiments of the present disclosure may be applied are shown.

[0479] refer to Figure 22 The content streaming system using embodiments of this disclosure may mainly include an encoding server, a streaming server, a web server, media storage, user equipment, and multimedia input devices.

[0480] An encoding server generates a bitstream by compressing content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data, and then sends it to a streaming server. As another example, when multimedia input devices such as smartphones, cameras, and camcorders generate bitstreams directly, the encoding server can be omitted.

[0481] A bitstream can be generated by applying the encoding method or bitstream generation method of the embodiments of this disclosure, and the streaming server can temporarily store the bitstream during the sending or receiving of the bitstream.

[0482] A streaming server sends multimedia data to a user's device via a web server based on the user's request, and the web server acts as a medium to notify the user what services are available. When a user requests a service from the web server, the web server delivers it to the streaming server, and the streaming server sends the multimedia data to the user. In this scenario, the content streaming system may include a separate control server, which in this case controls the commands / responses between each device in the content streaming system.

[0483] A streaming server can receive content from media storage and / or encoding servers. For example, when receiving content from an encoding server, the content can be received in real time. In this case, to provide a smooth streaming service, the streaming server can store bitstreams for specific time periods.

[0484] Examples of user equipment may include mobile phones, smartphones, laptops, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, tablet PCs, tablet computers, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head-mounted displays (HMDs), digital TVs, desktop computers, digital signage, etc.).

[0485] In a content streaming system, each server can be operated as a distributed server, and in this case, data received from each server can be distributed and processed.

[0486] The claims set forth herein can be combined in various ways. For example, the technical features of the method claims of this disclosure can be combined and implemented as a device, and the technical features of the device claims of this disclosure can be combined and implemented as a method. Furthermore, the technical features of the method claims and the technical features of the device claims of this disclosure can be combined and implemented as a device, and the technical features of the method claims and the technical features of the device claims of this disclosure can be combined and implemented as a method.

[0487] [Industrial Applicability]

[0488] The embodiments disclosed herein can be used to encode or decode images.

Claims

1. An image decoding method performed by a decoding device, the image decoding method comprising: Generate a prediction block for the current block based on the prediction pattern of the current block; Based on the predicted block, a reconstruction sample for the current block is generated; as well as Perform filtering on the reconstructed samples. Generating the prediction block includes: Export the regular prediction block of the current block; Export multiple reference blocks of the current block; and The final prediction block for the current block is generated by weighting the regular prediction block and the multiple reference blocks. The filtering strength used for filtering is determined based on the prediction information of the current block, the prediction block, or the plurality of reference blocks.

2. The image decoding method according to claim 1, wherein, The prediction information includes the prediction mode of the current block, the method for generating the prediction block, or the prediction mode of the multiple reference blocks.

3. The image decoding method according to claim 1, wherein, The filtering includes deblocking filtering, and When a predetermined first condition is met, the filter strength is set to a first value, which includes a value of 2 or 3.

4. The image decoding method according to claim 3, wherein, The predetermined first condition includes the current block or a neighboring block adjacent to the current block being in intra-frame mode, the current block or the neighboring block being in a combined inter-frame and intra-frame prediction (CIIP) mode, or the number of blocks in intra-frame mode among the plurality of reference blocks being greater than 0.

5. The image decoding method according to claim 4, wherein, When the current block or the adjacent block is in the intra-frame mode, the filter strength is set to the first value.

6. The image decoding method according to claim 4, wherein, When the current block or the adjacent block is in the CIIP mode, the filter strength is set to the first value.

7. The image decoding method according to claim 4, wherein, When the number of blocks in the intra-frame mode among the plurality of reference blocks is greater than 0, the filter strength is set to the first value.

8. The image decoding method according to claim 7, wherein, The multiple reference blocks located in the current block or the adjacent block are multiple reference blocks that are sent by the signal or inherited.

9. A method for decoding an image performed by a decoding device, the method comprising: The prediction block for the current block is generated based on the prediction information obtained from the bit stream; The transform coefficients are derived by performing dequantization based on the residual information obtained from the bitstream; The residual block is derived by performing an inverse transformation of the transformation coefficients; as well as Reconstructed samples are generated based on the predicted blocks and the residual blocks. The generated prediction blocks include: Generate a regular prediction block for the current block; Export multiple reference blocks of the current block; and The final prediction block for the current block is derived by weighting the regular prediction block and the multiple reference blocks. Specifically, the determination of whether a sub-block transformation (SBT) is applied to the current block is based on the prediction patterns of the multiple reference blocks.

10. The image decoding method according to claim 9, wherein, When the prediction mode of the plurality of reference blocks is not in intra-frame mode, it is determined that the SBT will be applied to the current block.

11. The image decoding method according to claim 9, wherein, When the predetermined second condition is met, it is determined that the SBT will be applied to the current block, and The predetermined second condition includes the case where the number of blocks in intra-frame mode among the plurality of reference blocks is greater than 0.

12. The image decoding method according to claim 11, wherein, When the number of blocks in the intra-frame mode among the plurality of reference blocks is greater than 0, it is determined that the SBT will be applied to the current block.

13. A method for decoding an image performed by a decoding device, the method comprising: The prediction block for the current block is generated based on the prediction information obtained from the bit stream; The transform coefficients are derived by performing dequantization based on the residual information obtained from the bitstream; The residual block is derived by performing an inverse transformation of the transformation coefficients; as well as Reconstructed samples are generated based on the predicted blocks and the residual blocks. The generated prediction blocks include: Generate a regular prediction block for the current block; Export multiple reference blocks of the current block; and The final prediction block for the current block is derived by weighting the regular prediction block and the multiple reference blocks. The transformation kernel used for the inverse transformation is determined based on a prediction mode associated with at least one of the conventional prediction blocks and the plurality of reference blocks.

14. The image decoding method according to claim 13, wherein, When the prediction modes of the plurality of reference blocks are in intra-frame mode, a representative intra-frame mode is determined by considering the directionality of each intra-frame mode, and the secondary transform kernel of the representative intra-frame mode is determined as the transform kernel for the inverse transform.

15. The image decoding method according to claim 13, wherein, When the prediction modes of the plurality of reference blocks are in intra-frame mode, an arbitrary intra-frame mode is selected based on the directionality of each intra-frame mode, and the secondary transform kernel of the selected intra-frame mode is determined as the transform kernel for the inverse transform.

16. The image decoding method according to claim 13, wherein, When the prediction mode of the plurality of reference blocks is not in intra-frame mode, the inverse secondary transform is not applied, or the kernel of the planar mode is determined as the transform kernel for the inverse transform.

17. The image decoding method according to claim 13, wherein, The conventional prediction block is derived based on at least one of the conventional reference blocks, and when the prediction mode of at least one of the conventional prediction blocks is in intra-frame mode, the secondary transform kernel of at least one of the conventional prediction blocks is determined as the transform kernel for the inverse transform.

18. The image decoding method according to claim 17, wherein, When the prediction mode of at least one of the conventional prediction blocks is in intra-frame mode and the prediction mode of the plurality of reference blocks is in intra-frame mode, a representative intra-frame mode is determined by considering the directionality of each intra-frame mode, and the secondary transform kernel of the representative intra-frame mode is determined as the transform kernel for the inverse transform.

19. The image decoding method according to claim 13, wherein, When the prediction mode of at least one of the conventional prediction blocks is in intra-frame mode and the prediction modes of the plurality of reference blocks are in intra-frame mode, the secondary transform kernel of the conventional prediction block is determined as the transform kernel for the inverse transform based on the directionality matching of each intra-frame mode.

20. An image encoding method performed by an encoding device, the image encoding method comprising: Generate a prediction block for the current block based on the prediction pattern of the current block; Based on the predicted block, a reconstruction sample for the current block is generated; as well as Perform filtering on the reconstructed samples. Generating the prediction block includes: Export the regular prediction block of the current block; Export multiple reference blocks of the current block; and The final prediction block for the current block is generated by weighting the regular prediction block and the multiple reference blocks. The filtering strength used for filtering is determined based on the prediction information of the current block, the prediction block, or the plurality of reference blocks.

21. A computer-readable digital storage medium for storing a bitstream generated using an image encoding method, the bitstream being generated by generating a prediction block of the current block based on a prediction pattern of the current block; Based on the predicted block, a reconstruction sample for the current block is generated; And perform filtering on the reconstructed samples, and Generating the prediction block includes: Export the regular prediction block of the current block; Export multiple reference blocks of the current block; and The final prediction block for the current block is generated by weighting the regular prediction block and the multiple reference blocks. The filtering strength used for filtering is determined based on the prediction information of the current block, the prediction block, or the plurality of reference blocks.

22. A method for transmitting image data, the method comprising: A bitstream of the image is generated by generating a prediction block for the current block based on the prediction mode of the current block. Based on the predicted block, a reconstruction sample for the current block is generated; And perform filtering on the reconstructed samples, and Transmitting data including the bit stream, Generating the prediction block includes: Export the regular prediction block of the current block; Export multiple reference blocks of the current block; and The final prediction block for the current block is generated by weighting the regular prediction block and the multiple reference blocks. The filtering strength used for filtering is determined based on the prediction information of the current block, the prediction block, or the plurality of reference blocks.