Method for decoding image information, method for encoding image information, and method for bitstream

The CIIP mode improves image compression by deriving inter-screen and intra-screen reference sample arrays, addressing the high-cost challenge of high-resolution image transmission and storage.

WO2026029642A1PCT designated stage Publication Date: 2026-02-05LG ELECTRONICS INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/011632
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-02
Filing Date
2025-08-04
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

The increasing demand for high-resolution, high-quality images leads to higher transmission and storage costs due to the increased amount of information, necessitating a highly efficient image compression technology.

Method used

Improving prediction performance and coding efficiency through combined inter and intra prediction (CIIP) mode by deriving inter-screen and intra-screen reference sample arrays to enhance the decoding and encoding processes.

Benefits of technology

Enhances prediction performance, coding efficiency, and data transmission efficiency by optimizing the image compression process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025011632_05022026_PF_FP_ABST
    Figure KR2025011632_05022026_PF_FP_ABST
Patent Text Reader

Abstract

A method for decoding image information according to the present disclosure may comprise: acquiring the image information including prediction information; deriving a prediction mode for a current unit in a current picture; deriving an inter-screen reference sample array from a reference picture different from the current picture on the basis of the prediction mode; deriving an intra-screen reference sample array from the current picture on the basis of the prediction mode; deriving a prediction sample array for the current block on the basis of the inter-screen reference sample array and the intra-screen reference sample array; and deriving a reconstructed sample array for the current block on the basis of the prediction sample array.
Need to check novelty before this filing date? Find Prior Art

Description

Method for decoding image information, method for encoding image information, and method for bitstream

[0001] The present disclosure relates to a method for decoding image information, a method for encoding image information, and a method for a bitstream.

[0002] Recently, demand for high-resolution, high-quality images, such as HD (High Definition) and UHD (Ultra High Definition) images, has been increasing across various fields. As image data becomes higher resolution and higher quality, the amount of information transmitted, or bits, increases relative to conventional image data. This increase in information or bits transmitted leads to increased transmission and storage costs.

[0003] Accordingly, a highly efficient image compression technology is required to effectively transmit, store, and play high-resolution, high-quality image information.

[0004] The present disclosure seeks to improve prediction performance by intra-screen prediction in combined inter and intra prediction (CIIP) mode.

[0005] The present disclosure seeks to improve coding efficiency by CIIP mode.

[0006] The present disclosure seeks to improve data transmission efficiency by CIIP mode.

[0007] The technical problems to be achieved in the present disclosure are not limited to the technical problems mentioned above, and other technical problems not mentioned will be clearly understood by a person having ordinary skill in the technical field to which the present disclosure belongs from the description below.

[0008] According to one aspect of the present disclosure, a method for decoding image information may include obtaining the image information including prediction information; deriving a prediction mode for a current unit within a current picture; deriving an inter-screen reference sample array from a reference picture different from the current picture based on the prediction mode; deriving an intra-screen reference sample array within the current picture based on the prediction mode; deriving a prediction sample array for a current block based on the inter-screen reference sample array and the intra-screen reference sample array; and deriving a reconstructed sample array for the current block based on the prediction sample array.

[0009] According to one aspect of the present disclosure, a device for decoding image information includes a memory and at least one processor connected to the memory, wherein the at least one processor obtains the image information including prediction information; derives a prediction mode for a current unit within a current picture; derives an inter-screen reference sample array from a reference picture different from the current picture based on the prediction mode; derives an intra-screen reference sample array within the current picture based on the prediction mode; derives a prediction sample array for a current block based on the inter-screen reference sample array and the intra-screen reference sample array; and derives a reconstructed sample array for the current block based on the prediction sample array.

[0010] In the method or device for decoding the image information, the prediction information includes first flag information and second flag information, the first flag information is obtained based on a value of the second flag information, a general CIIP (Combined inter and intra prediction) mode is derived for the current unit based on a value of the second flag information, and a modified CIIP mode can be derived for the current unit based on a value of the second flag information and a value of the first flag information.

[0011] In the method or device for decoding the image information, the prediction information further includes third flag information, the second flag information is obtained based on the value of the third flag information, and whether a merge mode is derived for the current unit can be determined based on the value of the third flag information.

[0012] In the method or device for decoding the image information, the prediction information further includes fourth flag information, the second flag information is obtained based on the value of the fourth flag information, and a blended mode can be derived for the current unit based on the value of the fourth flag information.

[0013] In the method or device for decoding the above image information, the prediction mode may be determined to be a general combined inter and intra prediction (CIIP) mode based on the fact that the inter-screen prediction sample array is derived based on a motion vector of a spatially adjacent block of the current block, and the prediction mode may be determined to be a modified CIIP mode based on the fact that the inter-screen prediction sample array is not derived based on a motion vector of a spatially adjacent block of the current block.

[0014] In the method or device for decoding the image information, the prediction mode may be determined to be a general CIIP (combined inter and intra prediction) mode based on the number of blocks of the intra prediction mode adjacent to the current block being less than a threshold value, and the prediction mode may be determined to be a modified CIIP mode based on the number of blocks of the intra prediction mode adjacent to the current block being greater than or equal to a threshold value.

[0015] In the method or device for decoding the image information, a first error is derived between a sample value of a template region around the current block derived using directional or non-directional intra prediction and a reconstructed sample value of the template region, a second error is derived between a sample value of the template region derived using intra block copy (IBC) prediction or intra template matching prediction (IntraTMP) and a reconstructed sample value of the template region, and the prediction mode may be determined to be a normal CIIP mode based on the first error being smaller than the second error, and the prediction mode may be determined to be a modified CIIP mode based on the first error being equal to or greater than the second error.

[0016] In the method or device for decoding the image information, the prediction information further includes fifth indexing information, and based on the fifth indexing information, the within-screen reference sample array can be derived based on information about a block vector included in the prediction information, or based on an area with a minimum error with a template area around the current block, or based on an area with a minimum error with the between-screen reference sample array.

[0017] In the method or device for decoding the image information, the prediction information further includes sixth indexing information, and based on the sixth indexing information, the in-screen reference sample array can be derived based on a motion vector candidate list including at least some of a motion vector of a spatial surrounding block of the current block, a motion vector of a temporal surrounding block, or a record-based motion vector, and a block vector candidate list including at least some of a block vector for the current block and a block vector based on a template area around the current block.

[0018] In the method or device for decoding the above image information, the prediction sample array can be derived based on an inter-screen prediction sample array derived using the inter-screen reference sample array and an intra-prediction sample array derived using the intra-screen reference sample array.

[0019] In the method or device for decoding the above image information, the prediction sample array can be derived by weighting the inter-screen prediction sample array and the intra-prediction sample array.

[0020] According to one aspect of the present disclosure, a method for encoding image information may include determining a prediction mode for a current unit of a current picture; deriving an inter-screen reference sample array from a reference picture different from the current picture based on the determined prediction mode; deriving an intra-screen reference sample array within the current picture based on the determined prediction mode; deriving a prediction sample array for a current block based on the inter-screen reference sample array and the intra-screen reference sample array; deriving a residual sample array for the current block based on the prediction sample array; and encoding the image information including prediction information based on the determined prediction mode.

[0021] According to one aspect of the present disclosure, a device for encoding image information includes a memory and at least one processor connected to the memory, wherein the at least one processor determines a prediction mode for a current unit of a current picture; derives an inter-screen reference sample array from a reference picture different from the current picture based on the determined prediction mode; derives an intra-screen reference sample array within the current picture based on the determined prediction mode; derives a prediction sample array for a current block based on the inter-screen reference sample array and the intra-screen reference sample array; derives a residual sample array for the current block based on the prediction sample array; and encodes the image information including prediction information based on the determined prediction mode.

[0022] In the method or device for encoding the image information, the prediction information includes first flag information and second flag information, the first flag information is obtained based on a value of the second flag information, a general CIIP (Combined inter and intra prediction) mode is derived for the current unit based on a value of the second flag information, and a modified CIIP mode is derived for the current unit based on a value of the second flag information and a value of the first flag information.

[0023] In the method or device for encoding the image information, the prediction information further includes third flag information, the second flag information is obtained based on the value of the third flag information, and whether a merge mode is derived for the current unit can be determined based on the value of the third flag information.

[0024] In the method or device for encoding the image information, the prediction information further includes fourth flag information, the second flag information is obtained based on the value of the fourth flag information, and a blended mode can be derived for the current unit based on the value of the fourth flag information.

[0025] In the method or device for encoding the above image information, the prediction mode may be determined to be a general combined inter and intra prediction (CIIP) mode based on the fact that the inter-screen prediction sample array is derived based on a motion vector of a spatially adjacent block of the current block, and the prediction mode may be determined to be a modified CIIP mode based on the fact that the inter-screen prediction sample array is not derived based on a motion vector of a spatially adjacent block of the current block.

[0026] In the method or device for encoding the above image information, the prediction mode may be determined to be a general CIIP (combined inter and intra prediction) mode based on the number of blocks of the intra prediction mode adjacent to the current block being less than a threshold value, and the prediction mode may be determined to be a modified CIIP mode based on the number of blocks of the intra prediction mode adjacent to the current block being greater than or equal to a threshold value.

[0027] In the method or device for encoding the image information, a first error is derived between a sample value of a template region around the current block derived using directional or non-directional intra prediction and a reconstructed sample value of the template region, a second error is derived between a sample value of the template region derived using intra block copy (IBC) prediction or intra template matching prediction (IntraTMP) and a reconstructed sample value of the template region, and the prediction mode may be determined to be a general CIIP mode based on the first error being smaller than the second error, and the prediction mode may be determined to be a modified CIIP mode based on the first error being equal to or greater than the second error.

[0028] In the method or device for encoding the image information, the prediction information further includes fifth indexing information, and based on the fifth indexing information, the within-screen reference sample array can be derived based on information about a block vector included in the prediction information, or derived based on an area with a minimum error with a template area around the current block, or derived based on an area with a minimum error with the between-screen reference sample array.

[0029] In the method or device for encoding the image information, the prediction information further includes sixth indexing information, and based on the sixth indexing information, the in-screen reference sample array can be derived based on a motion vector candidate list including at least some of a motion vector of a spatial surrounding block of the current block, a motion vector of a temporal surrounding block, or a record-based motion vector, and a block vector candidate list including at least some of a block vector for the current block and a block vector based on a template area around the current block.

[0030] In the method or device for encoding the above image information, the prediction sample array can be derived based on an inter-screen prediction sample array derived using the inter-screen reference sample array and an intra-prediction sample array derived using the intra-screen reference sample array.

[0031] In the method or device for encoding the above image information, the prediction sample array can be derived by weighting the inter-screen prediction sample array and the intra-prediction sample array.

[0032] According to one aspect of the present disclosure, a method for a bitstream comprises: generating a bitstream; and transmitting data including the bitstream, wherein generating the bitstream comprises: determining a prediction mode for a current unit of a current picture; deriving an inter-screen reference sample array from a reference picture different from the current picture based on the determined prediction mode; deriving an intra-screen reference sample array within the current picture based on the determined prediction mode; deriving a prediction sample array for a current block based on the inter-screen reference sample array and the intra-screen reference sample array; deriving a residual sample array for the current block based on the prediction sample array; and generating the bitstream including prediction information based on the determined prediction mode.

[0033] According to one aspect of the present disclosure, a device for a bitstream comprises: at least one processor for generating a bitstream; and a transmission unit for transmitting data including the bitstream, wherein the at least one processor determines a prediction mode for a current unit of a current picture; derives an inter-screen reference sample array from a reference picture different from the current picture based on the determined prediction mode; derives an intra-screen reference sample array within the current picture based on the determined prediction mode; derives a prediction sample array for a current block based on the inter-screen reference sample array and the intra-screen reference sample array; derives a residual sample array for the current block based on the prediction sample array; and generates the bitstream including prediction information based on the determined prediction mode.

[0034] According to one aspect of the present disclosure, a computer-readable medium can store data including a bitstream generated based on determining a prediction mode for a current unit of a current picture; deriving an inter-screen reference sample array from a reference picture different from the current picture based on the determined prediction mode; deriving an intra-screen reference sample array within the current picture based on the determined prediction mode; deriving a prediction sample array for a current block based on the inter-screen reference sample array and the intra-screen reference sample array; deriving a residual sample array for the current block based on the prediction sample array; and generating the bitstream including prediction information based on the determined prediction mode.

[0035] The features briefly summarized above regarding the present disclosure are merely exemplary aspects of the detailed description of the present disclosure that follows and do not limit the scope of the present disclosure.

[0036] According to the present disclosure, prediction performance is improved by intra-screen prediction in combined inter and intra prediction (CIIP) mode.

[0037] According to the present disclosure, coding efficiency is improved by the CIIP mode.

[0038] According to the present disclosure, data transmission efficiency is improved by the CIIP mode.

[0039] According to the present disclosure, in CIIP mode

[0040] The effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned will be clearly understood by a person having ordinary skill in the art to which the present disclosure pertains from the description below.

[0041] FIG. 1 is a diagram schematically illustrating a video coding system to which an embodiment according to the present disclosure can be applied.

[0042] FIG. 2 is a schematic diagram of an encoding device to which an embodiment according to the present disclosure can be applied.

[0043] FIG. 3 is a schematic diagram showing a decoding device to which an embodiment according to the present disclosure can be applied.

[0044] FIG. 4 is a diagram showing a search area for intra-template matching prediction in an embodiment according to the present disclosure.

[0045] FIG. 5 is a diagram showing a block vector of intra template matching prediction for intra block copying in an embodiment according to the present disclosure.

[0046] FIG. 6 is a diagram showing padding candidates for replacing zero vectors in an IBC list according to an embodiment of the present disclosure.

[0047] FIG. 7 is a diagram showing an IBC reference area according to a current CU position in an embodiment according to the present disclosure.

[0048] FIG. 8 is a diagram showing a reference area for coding the current CTU in an embodiment according to the present disclosure.

[0049] FIG. 9 is a diagram illustrating a method for decoding image information according to one embodiment of the present disclosure.

[0050] FIG. 10 is a diagram showing the order of signaling or transmitting information regarding a prediction mode according to one embodiment of the present invention.

[0051] FIG. 11 is a diagram showing the order of signaling or transmitting information regarding a prediction mode according to one embodiment of the present invention.

[0052] Fig. 12 is a diagram showing the directional mode and non-directional mode of the intra prediction mode according to one embodiment of the present invention.

[0053] FIG. 13 is a diagram illustrating a method for encoding image information according to one embodiment of the present disclosure.

[0054] FIG. 14 is a diagram exemplifying a content streaming system to which an embodiment according to the present disclosure can be applied.

[0055] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the attached drawings so that those skilled in the art can easily implement the present disclosure. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein.

[0056] In describing embodiments of the present disclosure, detailed descriptions of known configurations or functions will be omitted if they are deemed to obscure the gist of the present disclosure. Furthermore, portions unrelated to the description of the present disclosure in the drawings have been omitted, and similar portions have been designated with similar reference numerals.

[0057] In the present disclosure, when a component is said to be "connected," "coupled," or "connected" to another component, this may include not only a direct connection, but also an indirect connection in which another component exists in between. Furthermore, when a component is said to "include" or "have" another component, unless otherwise specifically stated, this does not exclude the other component, but rather implies that the other component may be included.

[0058] In this disclosure, terms such as first, second, etc. are used solely to distinguish one component from another, and do not limit the order or importance of components unless specifically stated otherwise. Accordingly, within the scope of this disclosure, a first component in one embodiment may be referred to as a second component in another embodiment, and similarly, a second component in one embodiment may be referred to as a first component in another embodiment.

[0059] In this disclosure, distinct components are used to clearly illustrate their respective characteristics, and do not necessarily imply that the components are separated. That is, multiple components may be integrated into a single hardware or software unit, or a single component may be distributed into multiple hardware or software units. Therefore, even if not specifically mentioned, such integrated or distributed embodiments are also included within the scope of this disclosure.

[0060] In the present disclosure, the components described in various embodiments are not necessarily essential components, and some may be optional components. Therefore, embodiments comprising a subset of the components described in one embodiment are also within the scope of the present disclosure. Furthermore, embodiments including other components in addition to the components described in various embodiments are also within the scope of the present disclosure.

[0061] The present disclosure relates to encoding and decoding of video. For example, the methods and embodiments disclosed in this document can be applied to methods disclosed in the versatile video coding (VVC) standard, the essential video coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the second generation of audio video coding standard (AVS2), or the next generation of video / image coding standards (e.g., H.267 or H.268).

[0062] The present disclosure presents various embodiments of video / image coding, and unless otherwise stated, the embodiments may be performed in combination with each other.

[0063] Terms used in this disclosure may have their usual meanings commonly used in the technical field to which this disclosure belongs, unless newly defined in this disclosure.

[0064] In this disclosure, "video" may mean a set of images over time. In this disclosure, "picture" generally means a unit representing one image at a specific time point, and a slice / tile is a unit that constitutes a part of a picture in coding. A slice / tile may include one or more CTUs (coding tree units). A picture may be composed of one or more slices / tiles. A picture may be composed of one or more tile groups. A tile group may include one or more tiles. A brick may represent a rectangular area of ​​CTU rows within a tile in a picture. In this document, tile group and slice may be used interchangeably. For example, in this document, a tile group / tile group header may be referred to as a slice / slice header.

[0065] In the present disclosure, "pixel" or "pel" may refer to the smallest unit that constitutes a picture (or image). Additionally, "sample" may be used as a term corresponding to a pixel. A sample may generally represent a pixel or a pixel value, and may represent only a pixel / pixel value of a luma component or only a pixel / pixel value of a chroma component.

[0066] In the present disclosure, a "unit" may represent a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to the region. One unit may include one luma block and two chroma (e.g., cb, cr) blocks. In some cases, the term "unit" may be used interchangeably with terms such as "block" or "area." In general, an MxN block may include a set (or array) of samples (or sample array) or transform coefficients consisting of M columns and N rows.

[0067] In the present disclosure, the "current block" may mean one of the following: a "current coding block," a "current coding unit," a "block to be encoded," a "block to be decoded," or a "block to be processed." When prediction is performed, the "current block" may mean a "current prediction block" or a "block to be predicted." When transformation (inverse transformation) / quantization (inverse quantization) is performed, the "current block" may mean a "current transformation block" or a "block to be transformed." When filtering is performed, the "current block" may mean a "block to be filtered."

[0068] In the present disclosure, a "current block" may mean a block that includes both a luma component block and a chroma component block, or a "luma block of the current block," unless explicitly described as a chroma block. The luma component block of the current block may be explicitly expressed by including an explicit description of the luma component block, such as "luma block" or "current luma block." Additionally, the chroma component block of the current block may be explicitly expressed by including an explicit description of the chroma component block, such as "chroma block" or "current chroma block."

[0069] In this disclosure, " / " and "," can be interpreted as "and / or". For example, "A / B" and "A, B" can be interpreted as "A and / or B". Additionally, "A / B / C" and "A, B, C" can mean "at least one of A, B, and / or C."

[0070] In this disclosure, "or" may be interpreted as "and / or." For example, "A or B" may mean 1) "A" only, 2) "B" only, or 3) "A and B." Alternatively, "or" in this disclosure may mean "additionally or alternatively."

[0071] FIG. 1 is a schematic diagram illustrating a video / image coding system to which an embodiment according to the present disclosure can be applied.

[0072] Referring to FIG. 1, a video / image coding system may include a first device (source device) and a second device (receiving device). The source device may transmit encoded video / image or data to the receiving device via a digital storage medium or a network in the form of a file or streaming.

[0073] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / video encoding device, and the decoding device may be referred to as a video / video decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, and the display unit may be configured as a separate device or an external component.

[0074] A video source may obtain video / images through a process of capturing, synthesizing, or generating video / images. The video source may include a video / image capture device and / or a video / image generation device. A video / image capture device may include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. A video / image generation device may include, for example, a computer, a tablet, a smartphone, etc., and may (electronically) generate video / images. For example, a virtual video / image may be generated through a computer, etc., in which case the video / image capture process may be replaced by a process of generating related data.

[0075] An encoding device can encode input video / images. The encoding device can perform a series of procedures, such as prediction, transformation, and quantization, to improve compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.

[0076] The transmission unit can transmit encoded video / image information or data output in the form of a bitstream to the receiving unit of a receiving device via a digital storage medium or network in the form of a file or streaming. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit can include an element for generating a media file via a predetermined file format and an element for transmission via a broadcasting / communication network. The receiving unit can receive / extract the bitstream and transmit it to a decoding device.

[0077] The decoding device can decode the video / image by performing a series of procedures such as inverse quantization, inverse transformation, and prediction corresponding to the operation of the encoding device.

[0078] The renderer can render decoded video / images. The rendered video / images can be displayed through the display unit.

[0079] FIG. 2 is a schematic diagram illustrating an encoding device to which an embodiment according to the present disclosure can be applied.

[0080] Referring to FIG. 2, the encoding device (200) may be configured to include an image partitioner (210), a prediction unit (predictor) 220, a residual processor (residual processor) 230, an entropy encoder (entropy encoder) 240, an adder (adder) 250, a filter (filter) 260, and a memory (memory) 270. The prediction unit (220) may include an inter prediction unit (221) and an intra prediction unit (222). The residual processor (230) may include a transformer (transformer) 232, a quantizer (quantizer) 233, a dequantizer (dequantizer) 234, and an inverse transformer (inverse transformer) 235. The residual processing unit (230) may further include a subtractor (231). The addition unit (250) may be called a reconstructor or a recontructed block generator. The image segmentation unit (210), the prediction unit (220), the residual processing unit (230), the entropy encoding unit (240), the addition unit (250), and the filtering unit (260) described above may be configured by one or more hardware components (e.g., an encoder chipset or a processor) depending on the embodiment. In addition, the memory (270) may include a DPB (Decoded Picture Buffer) and may be configured by a digital storage medium. The hardware component may further include the memory (270) as an internal / external component.

[0081] The image segmentation unit (210) can segment an input image (or picture, frame) input to the encoding device (200) into one or more processing units. For example, the processing units may be referred to as coding units (CUs). A coding unit may be recursively segmented into a coding tree unit (CTU) or a largest coding unit (LCU) according to a QTBTTT (Quad-tree binary-tree ternary-tree) structure. For example, one coding unit may be segmented into a plurality of coding units of deeper depth based on a quad-tree structure, a binary tree structure, and / or a ternary tree structure. For example, the quad-tree structure may be applied first, and the binary tree structure and / or the ternary tree structure may be applied later. Alternatively, the binary tree structure may be applied first. A coding procedure according to the present disclosure may be performed based on a final coding unit that is no longer segmented. In this case, based on coding efficiency according to image characteristics, etc., the maximum coding unit may be used as the final coding unit, or, if necessary, the maximum coding unit may be recursively divided into coding units of lower depths, and the coding unit of the optimal size may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration, which will be described later. As another example, the processing unit may further include a prediction unit (PU) or a transformation unit (TU). The prediction unit and the transformation unit may each be divided or partitioned from the final coding unit.The above prediction unit may be a unit of sample prediction, and the above transformation unit may be a unit for deriving a transformation coefficient and / or a unit for deriving a residual signal from a transformation coefficient.

[0082] The term "unit" may be used interchangeably with terms such as "block" or "area" depending on the case. In general, an MxN block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luminance component, or only the pixel / pixel value of the chroma component. A sample can be used as a term corresponding to a pixel or pel in a picture (or image).

[0083] The encoding device (200) can generate a residual signal (residual block, residual sample array) by subtracting a prediction signal (predicted block, prediction sample array) output from an inter prediction unit (221) or an intra prediction unit (222) from an input video signal (original block, original sample array), and the generated residual signal is transmitted to a conversion unit (232). In this case, as illustrated, a unit that subtracts a prediction signal (prediction block, prediction sample array) from an input video signal (original block, original sample array) within the encoding device (200) may be called a subtraction unit (231). The prediction unit (220) can perform prediction on a block to be processed (hereinafter, current block) and generate a predicted block including prediction samples for the current block. The prediction unit (220) can determine whether intra prediction or inter prediction is applied on a current block or CU basis. The prediction unit (220) can generate various information regarding prediction, such as prediction mode information, as described later in the description of each prediction mode, and transmit the information to the entropy encoding unit (240). The information regarding prediction can be encoded in the entropy encoding unit (240) and output in the form of a bitstream.

[0084] The intra prediction unit (222) can predict the current block by referring to samples within the current picture. The referenced samples may be located in the neighborhood of the current block or may be located away from it depending on the prediction mode. In intra prediction, the prediction modes may include multiple non-directional modes and multiple directional modes. The non-directional modes may include, for example, a DC mode and a planar mode. The directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the degree of detail in the prediction direction. However, this is merely an example, and a greater or lesser number of directional prediction modes may be used depending on the settings. The intra prediction unit (222) may also determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.

[0085] The inter prediction unit (221) can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, subblocks, or samples based on the correlation of the motion information between the neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include information on the inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring block can include a spatial neighboring block existing in the current picture and a temporal neighboring block existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different from each other. The temporal neighboring block may be called a collocated reference block, a collocated CU (colCU), etc. A reference picture including the above temporal neighboring blocks may be called a collocated picture (colPic). For example, the inter prediction unit (221) may construct a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction may be performed based on various prediction modes, and for example, in the case of skip mode and merge mode, the inter prediction unit (221) may use the motion information of neighboring blocks as the motion information of the current block. In the case of skip mode, unlike the merge mode, a residual signal may not be transmitted.In the motion vector prediction (MVP) mode, the motion vector of the surrounding blocks is used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.

[0086] The prediction unit (220) can generate a prediction signal based on various prediction methods and / or prediction techniques described below. For example, the prediction unit (220) can apply intra prediction or inter prediction to predict the current block, and can also apply intra prediction and inter prediction simultaneously. A prediction method that simultaneously applies intra prediction and inter prediction to predict the current block may be called combined inter and intra prediction (CIIP). In addition, the prediction unit (220) may be based on an intra block copy (IBC) prediction mode or a palette mode for predicting a block. The IBC prediction mode or palette mode may be used for content image / video coding such as games, such as screen content coding (SCC). IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC may utilize at least one of the inter prediction techniques described in this document. Palette mode can be viewed as an example of intracoding or intraprediction. When applied, palette mode can signal sample values ​​within a picture based on information about the palette table and palette index.

[0087] The prediction signal generated through the prediction unit (220) can be used to generate a restoration signal or a residual signal. The subtraction unit (231) can generate a residual signal (residual block, residual sample array) by subtracting the prediction signal (predicted block, predicted sample array) output from the prediction unit (220) from the input image signal (original block, original sample array). The generated residual signal can be transmitted to the conversion unit (232).

[0088] The transform unit (232) can apply a transform technique to the residual signal to generate transform coefficients. For example, the transform technique can include at least one of a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), a Karhunen-Loeve Transform (KLT), a Graph-Based Transform (GBT), or a Conditionally Non-linear Transform (CNT). Here, GBT refers to a transform obtained from a graph when the relationship information between pixels is expressed as a graph. CNT refers to a transform obtained based on generating a prediction signal using all previously reconstructed pixels. The transform process can be applied to a pixel block having a square equal size, or can be applied to a block of a non-square variable size.

[0089] The quantization unit (233) can quantize the transform coefficients and transmit them to the entropy encoding unit (240). The entropy encoding unit (240) can encode the quantized signal (information about the quantized transform coefficients) and output it as a bitstream. The information about the quantized transform coefficients can be called residual information. The quantization unit (233) can rearrange the quantized transform coefficients in a block form into a one-dimensional vector form based on the coefficient scan order, and can also generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form.

[0090] The entropy encoding unit (240) can perform various encoding methods such as, for example, exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. The entropy encoding unit (190) can also encode, together or separately, information necessary for video / image restoration (e.g., values ​​of syntax elements, etc.) in addition to quantized transform coefficients. The encoded information (e.g., encoded video / image information) can be transmitted or stored in the form of a bitstream in the form of a network abstraction layer (NAL) unit. The video / image information may further include information on various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. The signaled information, transmitted information, and / or syntax elements mentioned in the present disclosure may be included in video / image information. The video / image information may be encoded through the encoding procedure described above and included in the bitstream.

[0091] The above bitstream may be transmitted via a network or stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmission unit (not shown) for transmitting the signal output from the entropy encoding unit (240) and / or a storage unit (not shown) for storing the signal may be provided as an internal / external element of the encoding device (200), or the transmission unit may be provided as a component of the entropy encoding unit (240).

[0092] The quantized transform coefficients output from the quantization unit (233) can be used to generate a residual signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients through the inverse quantization unit (234) and inverse transformation unit (235), a residual signal (residual block or residual samples) can be restored.

[0093] Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture encoding and / or restoration process.

[0094] The addition unit (250) can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter prediction unit (221) or the intra prediction unit (222). When there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as the reconstructed block. The addition unit (250) can be called a reconstructed unit or a reconstructed block generation unit. The generated reconstructed signal can be used for intra prediction of the next block to be processed within the current picture, and can also be used for inter prediction of the next picture after filtering as described below.

[0095] The filtering unit (260) can improve subjective / objective picture quality by applying filtering to the restoration signal. For example, the filtering unit (260) can apply various filtering methods to the restoration picture to generate a modified restoration picture, and store the modified restoration picture in the memory (270), specifically, the DPB of the memory (170). The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit (260) can generate various information regarding filtering and transmit it to the entropy encoding unit (240), as described later in the description of each filtering method. The information regarding filtering can be encoded by the entropy encoding unit (240) and output in the form of a bitstream.

[0096] The modified restored picture transmitted to the memory (270) can be used as a reference picture in the inter prediction unit (221). Through this, when inter prediction is applied, the encoding device (200) can avoid prediction mismatch between the encoding device (200) and the decoding device, and can also improve encoding efficiency.

[0097] The DPB in the memory (270) can store a modified reconstructed picture to be used as a reference picture in the inter prediction unit (221). The memory (270) can store motion information of a block from which motion information in the current picture is derived (or encoded) and / or motion information of blocks in a picture that has already been reconstructed. The stored motion information can be transferred to the inter prediction unit (221) to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory (270) can store reconstructed samples of reconstructed blocks in the current picture and transfer them to the intra prediction unit (222).

[0098] FIG. 3 is a schematic diagram illustrating a decoding device to which an embodiment according to the present disclosure can be applied.

[0099] As illustrated in FIG. 3, the decoding device (300) may be configured to include an entropy decoder (310), a residual processor (320), a predictor (330), an adder (340), a filter (350), and a memory (360). The predictor (330) may include an inter-prediction unit (332) and an intra-prediction unit (331). The residual processor (320) may include a dequantizer (321) and an inverse transformer (321). The entropy decoding unit (310), residual processing unit (320), prediction unit (330), addition unit (340), and filtering unit (350) described above may be configured by a single hardware component (e.g., decoder chipset or processor) depending on the embodiment. In addition, the memory (360) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (360) as an internal / external component.

[0100] When a bitstream including video / image information is input, the decoding device (300) can restore the image by performing a process corresponding to the process performed in the encoding device (200) of FIG. 2. For example, the decoding device (300) can perform decoding using a processing unit applied in the encoding device (200). Therefore, the processing unit for decoding may be, for example, a coding unit. The coding unit may be a coding tree unit or may be obtained by dividing the maximum coding unit according to a quad tree structure, a binary tree structure, and / or a ternary tree structure. In addition, the restored image signal decoded and output by the decoding device (300) can be reproduced through a reproduction device (not shown).

[0101] The decoding device (300) can receive a signal output from the encoding device (200) of FIG. 2 in the form of a bitstream. The received signal can be decoded through the entropy decoding unit (310). For example, the entropy decoding unit (310) can parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. The decoding device (300) can decode a picture further based on the information on the parameter set and / or the general constraint information. The signaling / received information and / or syntax elements described below can be decoded through the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit (310) can decode information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the values ​​of syntax elements required for image restoration and the quantized values ​​of transform coefficients for residuals. More specifically, the CABAC entropy decoding method receives a bin corresponding to each syntax element in the bitstream, determines a context model using information of the syntax element to be decoded and decoding information of the surrounding and decoding target blocks or information of symbols / bins decoded in the previous step, and predicts the occurrence probability of the bin according to the determined context model to perform arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element.At this time, the CABAC entropy decoding method can update the context model using the information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model. Information regarding prediction among the information decoded by the entropy decoding unit (310) is provided to the prediction unit (330), and residual values ​​on which entropy decoding is performed by the entropy decoding unit (310), i.e., quantized transform coefficients and related parameter information, can be input to the residual processing unit (320). The residual processing unit (320) can derive a residual signal (residual block, residual samples, residual sample array). In addition, information regarding filtering among the information decoded by the entropy decoding unit (310) can be provided to the filtering unit (350). Meanwhile, a receiving unit (not shown) that receives a signal output from an encoding device may be further configured as an internal / external element of the decoding device (300), or the receiving unit may be a component of an entropy decoding unit (310). Meanwhile, the decoding device according to the present document may be called a video / video / picture decoding device, and the decoding device may be divided into an information decoder (video / video / picture information decoder) and a sample decoder (video / video / picture sample decoder). The information decoder may include the entropy decoding unit (310), and the sample decoder may include at least one of the inverse quantization unit (321), the inverse transformation unit (322), the addition unit (340), the filtering unit (350), the memory (360), the inter prediction unit (332), and the intra prediction unit (331).

[0102] The inverse quantization unit (321) can inverse quantize the quantized transform coefficients and output the transform coefficients. The inverse quantization unit (321) can rearrange the quantized transform coefficients into a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scanning order performed in the encoding device (200). The inverse quantization unit (321) can perform inverse quantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain transform coefficients.

[0103] In the inverse transform unit (322), the transform coefficients can be inversely transformed to obtain a residual signal (residual block, residual sample array).

[0104] The prediction unit (330) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit can apply intra prediction or inter prediction for prediction of a single block, and can also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). In addition, the prediction unit can be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode can be used for content image / video coding such as games, such as screen content coding (SCC). IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this document. Palette mode can be viewed as an example of intra coding or intra prediction. When palette mode is applied, information about the palette table and palette index may be signaled and included in the video / image information.

[0105] The intra prediction unit (331) can predict the current block by referring to samples within the current picture. The description of the intra prediction unit (222) can be equally applied to the intra prediction unit (331). The referenced samples may be located in the neighborhood of the current block or may be located away from it, depending on the prediction mode. In intra prediction, the prediction modes may include multiple non-directional modes and multiple directional modes. The intra prediction unit (331) can also determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.

[0106] The inter prediction unit (332) can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include information on the inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit (332) can construct a motion information candidate list based on the neighboring blocks, and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction can be performed based on various prediction modes (techniques), and the information about the prediction can include information indicating the mode (technique) of inter prediction for the current block.

[0107] The addition unit (340) can generate a restoration signal (restored picture, restoration block, restoration sample array) by adding the acquired residual signal to the prediction signal (predicted block, prediction sample array) output from the prediction unit (330) (including the inter prediction unit (332) and / or the intra prediction unit (331)). When there is no residual for the block to be processed, such as when the skip mode is applied, the predicted block can be used as the restoration block. The description of the addition unit (250) can be equally applied to the addition unit (340). The addition unit (340) can be called a restoration unit or a restoration block generation unit. The generated restoration signal can be used for intra prediction of the next block to be processed within the current picture, and can also be used for inter prediction of the next picture after going through filtering as described below.

[0108] Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture decoding process.

[0109] The filtering unit (350) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (350) can apply various filtering methods to the restored picture to generate a modified restored picture, and store the modified restored picture in the memory (360), specifically, in the DPB of the memory (360). The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.

[0110] The (modified) reconstructed picture stored in the DPB of the memory (360) can be used as a reference picture in the inter prediction unit (332). The memory (360) can store motion information of a block from which motion information is derived (or decoded) within the current picture and / or motion information of blocks within a picture that has already been reconstructed. The stored motion information can be transmitted to the inter prediction unit (332) to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory (360) can store reconstructed samples of reconstructed blocks within the current picture and transmit them to the intra prediction unit (331).

[0111] In this specification, the embodiments described in the filtering unit (260), the inter prediction unit (221), and the intra prediction unit (222) of the encoding device (200) can be applied to the filtering unit (350), the inter prediction unit (332), and the intra prediction unit (331) of the decoding device (300) in the same or corresponding manner, respectively.

[0112] Intra prediction

[0113] The prediction unit of the encoding device / decoding device can derive a reference sample according to the intra prediction mode of the current block among surrounding reference samples of the current block, and can generate a prediction sample of the current block based on the reference sample.

[0114] For example, (i) a prediction sample may be derived based on an average or interpolation of neighboring reference samples of the current block, and (ii) the prediction sample may be derived based on a reference sample that exists in a specific (prediction) direction with respect to the prediction sample among the neighboring reference samples of the current block. The case of (i) may be called a non-directional mode or a non-angular mode, and the case of (ii) may be called a directional mode or an angular mode. In addition, the prediction sample may be generated through interpolation of the first neighboring sample and the second neighboring sample, which are located in the opposite direction of the prediction direction of the intra prediction mode of the current block with respect to the prediction sample of the current block among the neighboring reference samples. The above-described case may be called linear interpolation intra prediction (LIP). In addition, based on the filtered peripheral reference samples, a temporary prediction sample of the current block may be derived, and a prediction sample of the current block may be derived by weighting at least one reference sample derived according to the intra prediction mode among the existing peripheral reference samples, i.e., unfiltered peripheral reference samples, and the temporary prediction sample. The above-described case may be called Position dependent intra prediction (PDPC). In addition, intra prediction encoding may be performed by selecting a reference sample line with the highest prediction accuracy among the peripheral multiple reference sample lines of the current block, deriving a prediction sample using a reference sample located in the prediction direction in the corresponding line, and instructing (signaling) the used reference sample line to a decoding device. The above-described case may be called multi-reference line intra prediction (MRL) or MRL-based intra prediction.In addition, the current block can be divided into vertical or horizontal subpartitions, and intra prediction can be performed based on the same intra prediction mode, and surrounding reference samples can be derived and used for each subpartition. That is, in this case, the intra prediction mode for the current block is applied equally to the subpartitions, and surrounding reference samples can be derived and used for each subpartition, thereby improving the intra prediction performance in some cases. This prediction method can be called intra sub-partitions (IPS) or IPS-based intra prediction. In addition, when the prediction direction based on the prediction sample points between the surrounding reference samples, that is, when the prediction direction points to the fractional sample position, the value of the prediction sample can also be derived through interpolation of a plurality of reference samples located around the corresponding prediction direction (around the corresponding fractional sample position).

[0115] The intra prediction methods described above may be referred to as intra prediction types, as distinguished from intra prediction modes. The intra prediction types may be referred to by various terms, such as intra prediction techniques or supplementary intra prediction modes. For example, the intra prediction types (or supplementary intra prediction modes, etc.) may include at least one of the above-described LIP, PDPC, MRL, and ISP. Information regarding the intra prediction types may be encoded in an encoding device, included in a bitstream, and signaled to a decoding device. Information regarding the intra prediction types may be implemented in various forms, such as flag information indicating whether each intra prediction type is applied or index information indicating one of multiple intra prediction types.

[0116] The MPM list for deriving the above-described intra prediction mode may be configured differently depending on the intra prediction type. Alternatively, the MPM list may be configured in common regardless of the intra prediction type.

[0117] FIG. 4 is a diagram showing a search area for intra-template matching prediction in an embodiment according to the present disclosure.

[0118] Intra-template matching prediction (IntraTMP) is a special intra prediction mode that copies the optimal prediction block where the L-shaped template matches the current template from the reconstructed portion of the current frame. For a predefined search range, the encoder searches the reconstructed portion of the current frame for the template most similar to the current template and uses that block as the prediction block. The encoder then signals the use of this mode, and the decoder performs the same prediction operation.

[0119] The prediction signal is generated by matching the L-shaped causal neighbors of the current block with other blocks in the predefined search region of Fig. 4. The search region consists of:

[0120] R1: Current CTU

[0121] R2: Top left CTU

[0122] R3: Top of CTU

[0123] R4: Left CTU

[0124] The sum of absolute differences (SAD) is used as the cost function.

[0125] Within each region, the decoder searches for the template with the smallest SAD compared to the current frame and uses that block as the prediction block.

[0126] The size of all regions (SearchRange_w, SearchRange_h) is set proportionally to the block size (BlkW, BlkH) and can have a fixed number of SAD comparisons per pixel.

[0127] SearchRange_w = a * BlkW

[0128] SearchRange_h = a * BlkH

[0129] Here, 'a' is a constant that controls the gain / complexity tradeoff. For example, 'a' could be 5.

[0130] To speed up the template matching process, the search range of all search areas is subsampled by a factor of two. This reduces the number of template matching searches by a factor of four. After finding the optimal matching range, a refinement process is performed. The refinement is performed through a second template matching search centered on the optimal matching range with the reduced range. The reduced range can be defined as min(BlkW, BlkH) / 2.

[0131] The intra-template matching tool can be enabled for CUs with a width and height of 64 or less. The maximum CU size for intra-template matching is configurable.

[0132] Intra template matching prediction mode can be signaled at the coding unit level via a dedicated flag when DIMD is not used in the current coding unit.

[0133] FIG. 5 is a diagram showing a block vector of intra template matching prediction for intra block copying in an embodiment according to the present disclosure.

[0134] Block vectors (BVs) derived from intra-template matching prediction (IntraTMP) are used for intra-block copying (IBC). The stored IntraTMP BVs and IBC BVs of neighboring blocks are used as spatial BV candidates when generating the IBC candidate list.

[0135] The IntraTMP block vector is stored in the IBC block vector buffer, and the current IBC block can use both the IBC BV and the IntraTMP BV of the neighboring block as BV candidates in the IBC BV candidate list, as shown in Fig. 5.

[0136] IntraTMP block vectors are added as spatial candidates to the IBC block vector candidate list.

[0137] Here, intra-block copy (IBC) is well known to significantly improve the coding efficiency of screen content. Since IBC mode is implemented as a block-level coding mode, the encoder performs block matching (BM) to find the optimal block vector (or motion vector) for each CU. The block vector is used to indicate the displacement from the current block to a reference block already reconstructed within the current picture. The luma block vector of an IBC-coded CU has integer precision. The chroma block vector is also rounded to integer precision. When combined with adaptive motion vector resolution (AMVR), IBC mode can switch between 1-pixel and 4-pixel motion vector precision. An IBC-coded CU is processed as a third prediction mode other than intra- or inter-prediction mode. IBC mode is applicable to CUs whose width and height are both 64 luma samples or less.

[0138] On the encoder side, hash-based motion estimation is performed for IBC. The encoder performs RD checks on blocks with a width or height of 16 luma samples or less. In non-merge mode, block vector search is first performed using a hash-based search. If the hash search does not return a valid candidate, a block-matching-based local search is performed.

[0139] In hash-based search, the hash key matching (32-bit CRC) between the current block and reference blocks is extended to all allowed block sizes. The hash key calculation for each location in the current picture is based on 4x4 subblocks. For larger current blocks, a reference block is considered to have a match if the hash keys of all 4x4 subblocks match the hash keys of the corresponding reference locations. If the hash keys of multiple reference blocks match the hash key of the current block, the block vector cost of each matching reference is calculated, and the reference with the lowest cost is selected.

[0140] In block matching searches, the search scope is set to include both the previous and current CTUs.

[0141] At the CU level, IBC mode is signaled by a flag, which can be either IBC AMVP mode or IBC skip / merge mode as follows:

[0142] - IBC Skip / Merge Mode: The merge candidate index is used to indicate the block vector used to predict the current block among the block vectors of neighboring candidate IBC coding blocks. The merge list consists of spatial candidates, HMVP candidates, and pairwise candidates.

[0143] - IBC AMVP mode: Block vector differences are encoded in the same way as motion vector differences. The block vector prediction method uses two candidates as predictors: one is the left neighbor candidate and the other is the upper neighbor candidate (in the case of IBC coding). If one of the two neighbor candidates is unavailable, the base block vector is used as the predictor. A flag indicating the block vector predictor index is signaled.

[0144] FIG. 6 is a diagram showing padding candidates for replacing zero vectors in an IBC list according to an embodiment of the present disclosure.

[0145] The IBC Merge / AMVP list configuration can be modified as follows:

[0146] - Can be inserted into the IBC Merge / AMVP candidate list only if the IBC Merge / AMVP candidate is valid.

[0147] - You can add the top right, bottom left, top left space candidates and one pairwise average candidate to the IBC merge / AMVP candidate list.

[0148] - Template-based adaptive reordering (ARMC-TM) is applied to IBC merge lists.

[0149] The IBC HMVP table size can be increased to 25. After a full pruning process yields up to 20 IBC merge candidates, they are reordered together. After reordering, the top six candidates with the lowest template matching cost are selected as the final candidates in the IBC merge list.

[0150] Zero vector candidates for filling the IBC Merge / AMVP list are replaced with the set of BVP candidates in the IBC reference area. In IBC Merge mode, zero vectors are not valid as block vectors, so they are removed from the IBC candidate list as BVPs.

[0151] Referring to Figure 6, three candidates are located at the nearest corners of the reference area, and three additional candidates are located in the middle of three sub-areas (A, B, C). The coordinates of each sub-area are determined by the width and height of the current block, and the ΔX and ΔY parameters.

[0152] FIG. 7 is a diagram showing an IBC reference area according to a current CU position in an embodiment according to the present disclosure.

[0153] In IBC, template matching can be used in both IBC merge mode and IBC AMVP mode.

[0154] The IBC-TM merge list is modified compared to the list used in the regular IBC merge mode, and candidates are selected based on a pruning method that considers the motion distance between candidates, as in the regular TM merge mode. The 0 motion execution at the end point is replaced with the motion vectors for left (-W, 0), up (0, -H), and top-left (-W, -H), where W is the width of the current CU and H is its height.

[0155] In IBC-TM merge mode, selected candidates are refined using template matching before the RDO or decoding process. IBC-TM merge mode competes with the standard IBC merge mode and is signaled by the TM merge flag.

[0156] In IBC-TM AMVP mode, up to three candidates are selected from the IBC-TM merge list. Each of the three selected candidates is refined using a template matching method and sorted based on the resulting template matching cost. Only the first two candidates are then considered in the motion estimation process as usual.

[0157] Template matching refinement for both IBC-TM merge mode and AMVP mode is very straightforward, as IBC motion vectors are (i) constrained to integers and (ii) constrained within the reference region, as shown in Figure 7. Therefore, in IBC-TM merge mode, all refinements are performed with integer precision, while in IBC-TM AMVP mode, they are performed with integer or 4-pixel precision, depending on the AMVR value. This refinement accesses only samples, without interpolation. In both cases, the refined motion vectors and the templates used at each refinement step must adhere to the constraints of the reference region.

[0158] FIG. 8 is a diagram showing a reference area for coding the current CTU in an embodiment according to the present disclosure.

[0159] The IBC reference region can extend over two CTU rows. Figure 8 shows the reference region for coding CTU(m, n). Specifically, for coding CTU(m, n), the reference region contains CTUs with indices (m-2, n-2)-(W, n-2), (0, n-1)-(W, n-1), (0, n)-(m, n), where W represents the maximum horizontal index within the current tile, slice, or picture. When the CTU size is 256, the reference region is limited to one CTU row. This setting prevents IBC from requiring additional memory when the CTU size is 128 or 256. The per-sample block vector lookup (or local lookup) range is limited horizontally to [-(C << 1), C >> 2] and vertically to [-C, C >> 2] to accommodate the reference region extension. Here, C represents the CTU size.

[0160] Inter prediction

[0161] The prediction unit of the encoding device / decoding device can perform inter prediction on a block-by-block basis to derive prediction samples. Inter prediction can refer to a prediction derived in a manner dependent on data elements (e.g., sample values, or motion information, etc.) of pictures other than the current picture. When inter prediction is applied to the current block, a predicted block (prediction sample array) for the current block can be derived based on a reference block (reference sample array) specified by a motion vector on a reference picture pointed to by a reference picture index. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information of the current block can be predicted on a block, sub-block, or sample basis based on the correlation of the motion information between the surrounding blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include information on the inter prediction type (L0 prediction, L1 prediction, Bi prediction, etc.). When inter prediction is applied, the neighboring blocks may include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring blocks may be called collocated reference blocks, collocated CUs (colCUs), etc., and the reference pictures including the temporal neighboring blocks may be called collocated pictures (colPic). For example, a motion information candidate list may be constructed based on the neighboring blocks of the current block, and a flag or index information indicating which candidate is selected (used) to derive the motion vector and / or reference picture index of the current block may be signaled.Inter prediction can be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the motion information of the current block may be the same as the motion information of the selected neighboring block. In the case of skip mode, unlike the merge mode, a residual signal may not be transmitted. In the case of motion vector prediction (MVP) mode, the motion vector of the selected neighboring block is used as a motion vector predictor, and the motion vector difference may be signaled. In this case, the motion vector of the current block can be derived using the sum of the motion vector predictor and the motion vector difference.

[0162] The above motion information may include L0 motion information and / or L1 motion information depending on the inter prediction type (L0 prediction, L1 prediction, Bi prediction, etc.). A motion vector in the L0 direction may be called an L0 motion vector or MVL0, and a motion vector in the L1 direction may be called an L1 motion vector or MVL1. Prediction based on an L0 motion vector may be called an L0 prediction, prediction based on an L1 motion vector may be called an L1 prediction, and prediction based on both the L0 motion vector and the L1 motion vector may be called a bi-prediction (Bi). Here, an L0 motion vector may represent a motion vector associated with a reference picture list L0 (L0), and an L1 motion vector may represent a motion vector associated with a reference picture list L1 (L1). The reference picture list L0 may include pictures preceding the current picture in output order as reference pictures, and the reference picture list L1 may include pictures succeeding the current picture in output order. The preceding pictures may be called forward (reference) pictures, and the succeeding pictures may be called backward (reference) pictures. The reference picture list L0 may further include pictures succeeding the current picture in output order as reference pictures. In this case, the preceding pictures may be indexed first and the succeeding pictures may be indexed next within the reference picture list L0. The reference picture list L1 may further include pictures preceding the current picture in output order as reference pictures. In this case, the succeeding pictures may be indexed first and the succeeding pictures may be indexed next within the reference picture list 1. Here, the output order may correspond to a POC (picture order count) order.

[0163] Various inter prediction modes can be used to predict the current block in a picture. For example, various modes can be used, such as merge mode, skip mode, motion vector prediction (MVP) mode, affine mode, sub-block merge mode, and merge with MVD (MMVD) mode. Decoder-side motion vector refinement (DMVR) mode, adaptive motion vector resolution (AMVR) mode, bi-prediction with CU-level weight (BCW), and bi-directional optical flow (BDOF) can be used in addition to or instead of these modes. Affine mode may also be called affine motion prediction mode. MVP mode may also be called advanced motion vector prediction (AMVP) mode. In this document, motion information candidates derived by some modes and / or some modes may be included as one of the motion information-related candidates of other modes. For example, an HMVP candidate may be added as a merge candidate in the above merge / skip mode, or may be added as an mvp candidate in the above MVP mode.

[0164] Prediction mode information indicating the inter prediction mode of the current block may be signaled from the encoding device to the decoding device. The prediction mode information may be included in the bitstream and received by the decoding device. The prediction mode information may include index information indicating one of a plurality of candidate modes. Alternatively, the inter prediction mode may be indicated through hierarchical signaling of flag information. In this case, the prediction mode information may include one or more flags. For example, a skip flag may be signaled to indicate whether skip mode is applied, a merge flag may be signaled to indicate whether merge mode is applied if skip mode is not applied, and if merge mode is not applied, MVP mode may be indicated to be applied, or a flag for additional distinction may be further signaled. The affine mode may be signaled as an independent mode, or as a mode dependent on the merge mode or the MVP mode. For example, the affine mode may include an affine merge mode and an affine MVP mode.

[0165] Inter prediction can be performed using motion information of the current block. The encoding device can derive optimal motion information for the current block through a motion estimation procedure. For example, the encoding device can search for similar reference blocks with high correlation within a predetermined search range within the reference picture using the original block within the original picture for the current block, in fractional pixel units, and thereby derive motion information. The similarity between blocks can be derived based on the difference in phase-based sample values. For example, the similarity between blocks can be calculated based on the SAD between the current block (or a template of the current block) and the reference block (or a template of the reference block). In this case, motion information can be derived based on the reference block with the smallest SAD within the search range. The derived motion information can be signaled to the decoding device in various ways based on the inter prediction mode.

[0166] When merge mode is applied, the motion information of the current prediction block is not directly transmitted, but rather the motion information of the surrounding prediction blocks is used to derive the motion information of the current prediction block. Accordingly, the motion information of the current prediction block can be indicated by transmitting flag information indicating that merge mode is used and a merge index indicating which surrounding prediction block was used. The above merge mode may be referred to as regular merge mode.

[0167] To perform merge mode, the encoder must search for merge candidate blocks used to derive motion information of the current prediction block. For example, up to five merge candidate blocks may be used, but the number is not limited thereto. In addition, the maximum number of merge candidate blocks may be transmitted in the slice header or tile group header, but is not limited thereto. After finding the merge candidate blocks, the encoder can generate a merge candidate list, and select the merge candidate block with the lowest cost among them as the final merge candidate block.

[0168] In addition to the merge mode, where implicitly derived motion information is directly used to generate prediction samples for the current CU, there is also a merge mode with motion vector differences (MMVD). Since similar motion information derivation methods are used in both skip mode and merge mode, MMVD can also be applied to skip mode. The MMVD flag (e.g., mmvd_flag) can be specified immediately after the skip and merge flags to indicate whether the CU uses MMVD mode.

[0169] In MMVD, after a merge candidate is selected, it is further refined using the signaled MVD information. When MMVD is applied to the current block (i.e., when mmvd_flag is 1), additional information about MMVD can be specified. The additional information includes a merge candidate flag (e.g., mmvd_merge_flag) that indicates whether to use the first (0) or second (1) candidate in the merge candidate list with the motion vector difference, an index that specifies the motion magnitude (e.g., mmvd_distance_idx), and an index that indicates the motion direction (e.g., mmvd_direction_idx). In MMVD mode, one of the first two candidates in the merge list is selected to use as the motion vector criterion. The merge candidate flag is used to specify which candidate to use.

[0170] The MVP (Motion Vector Prediction) mode may be referred to as the AMVP (advanced motion vector prediction) mode. When the MVP mode is applied, a motion vector predictor (mvp) candidate list may be generated using the motion vectors of the reconstructed spatial neighboring blocks and / or the motion vectors corresponding to the temporal neighboring blocks (or Col blocks). That is, the motion vectors of the reconstructed spatial neighboring blocks and / or the motion vectors corresponding to the temporal neighboring blocks may be used as motion vector predictor candidates. When paired prediction is applied, an mvp candidate list for deriving L0 motion information and an mvp candidate list for deriving L1 motion information may be generated and used separately. The above-described prediction information (or information related to prediction) may include selection information (e.g., an MVP flag or an MVP index) indicating an optimal motion vector predictor candidate selected from among the motion vector predictor candidates included in the above list. At this time, the prediction unit can select the motion vector predictor of the current block from among the motion vector predictor candidates included in the motion vector candidate list using the selection information. The prediction unit of the encoding device can obtain the motion vector difference (MVD) between the motion vector of the current block and the motion vector predictor, and can encode and output it in the form of a bitstream. That is, the MVD can be obtained as a value obtained by subtracting the motion vector predictor from the motion vector of the current block. At this time, the prediction unit of the decoding device can obtain the motion vector difference included in the information about the prediction, and derive the motion vector of the current block by adding the motion vector difference and the motion vector predictor. The prediction unit of the decoding device can obtain or derive a reference picture index indicating a reference picture, etc., from the information about the prediction.

[0171] Existing video coding systems use only one motion vector (using a translaton motion model) to represent the motion of an encoded block. However, although the above method may represent the optimal motion at the block level, it is not the optimal motion for each pixel. If the optimal motion vector can be determined at the pixel level, encoding efficiency can be improved. To this end, this embodiment describes an affine motion prediction method that uses an affine motion model to encode. The affine motion prediction method can represent a motion vector at each pixel level of a block using two, three, or four motion vectors.

[0172] The affine motion model can express four types of motions (translation, scale, rotation, and shear). According to the affine motion prediction method, three types of motions (translation, scale, and rotation) can be expressed by the affine motion model.

[0173] Subblock-based temporal motion vector prediction (SbTMVP) can be used. SbTMVP improves motion vector prediction and merge mode for the current picture's CU by using the motion fields of collocated pictures. The same collocated pictures used in TMVP are also used in SbTVMP. SbTMVP differs from TMVP in two key aspects:

[0174] 1. TMVP predicts actions at the CU level, whereas SbTMVP predicts actions at the sub-CU level.

[0175] 2. While TMVP obtains temporal motion vectors from the collocated block of the collocated picture (the collocated block is the lower-right or center (lower-right center) block relative to the current CU), SbTMVP applies motion shift before obtaining temporal motion information from the collocated picture. In this case, the motion shift is obtained from the motion vector of one of the spatially adjacent blocks of the current CU.

[0176] Geometric partitioning mode (GPM) is supported for inter prediction. Geometric partitioning mode is indicated as a type of merge mode via a CU-level flag, and there are other merge modes such as general merge mode, MMVD mode, CIIP mode, and sub-block merge mode. Geometric partitioning mode supports a total of 64 partitions for each CU size (w×h=2^m×2^n), excluding 8x64 and 64x8.

[0177] In this mode, a CU is divided into two parts by a geometrically positioned straight line. The location of the split line is mathematically derived from the angle and offset parameters of a particular partition. Each part of a geometric partition within a CU is cross-predicted using its own motion. Only a single prediction is allowed for each partition; that is, each part has a single motion vector and a single reference index. The single-prediction motion constraint, similar to conventional dual prediction, requires only two motion-compensated predictions for each CU.

[0178] Combined inter and intra prediction (CIIP) can be applied to the current block. An additional flag (e.g., ciip_flag) can be signaled to indicate whether the combined inter / intra prediction (CIIP) mode is applied to the current CU. For example, when a CU is coded in merge mode, if the CU contains at least 64 luma samples (i.e., the CU width times the CU height is greater than or equal to 64), and both the CU width and the CU height are less than 128 luma samples, an additional flag is signaled to indicate whether the combined inter / intra prediction (CIIP) mode is applied to the current CU. As the name suggests, CIIP prediction combines the inter prediction signal and the intra prediction signal. The inter prediction signal P_inter in CIIP mode is generated using the same inter prediction process as applied in the regular merge mode, and the intra prediction signal P_intra is generated according to the regular intra prediction process and the planar mode. Then, the intra and inter prediction signals are combined using weighted averaging, where the weight values ​​are calculated according to the coding modes of the upper and left adjacent blocks as shown in Equation 1.

[0179] [Mathematical Formula 1]

[0180]

[0181] If the top adjacent block is available and intra-coded, set isIntraTop to 1, otherwise set isIntraTop to 0.

[0182] If the left adjacent block is available and intra-coded, set isIntraLeft to 1, otherwise set isIntraLeft to 0.

[0183] Here, if (isIntraLeft + isIntraLeft) is 2, wt is set to 3.

[0184] Otherwise, if (isIntraLeft + isIntraLeft) is 1, set wt to 2.

[0185] Otherwise, set wt to 1.

[0186] A predicted block for a current block can be derived based on motion information derived according to a prediction mode. The predicted block can include prediction samples (prediction sample array) of the current block. If the motion vector of the current block indicates a fractional sample unit, an interpolation procedure can be performed, through which prediction samples of the current block can be derived based on reference samples of the fractional sample unit within a reference picture. If affine inter prediction is applied to the current block, prediction samples can be generated based on a sample / subblock unit MV. If bi-prediction is applied, prediction samples derived based on L0 prediction (i.e., prediction using a reference picture in a reference picture list L0 and MVL0) and L1 prediction (i.e., prediction using a reference picture in a reference picture list L1 and MVL1) can be used as prediction samples of the current block through a weighted sum or weighted average (according to phase).

[0187] Each embodiment or combination of embodiments of the present disclosure can be applied to inter-prediction and intra-prediction processes, and in particular, can be applied to Combined Inter and Intra Prediction (CIIP). The CIIP mode is a mode that generates a final prediction block through a weighted sum of an inter-prediction block and an intra-prediction block. However, the existing CIIP mode has a disadvantage in that it is difficult to reflect changes in actual pixel values ​​within a block because it generates an intra-prediction block using adjacent samples of the current block. One embodiment includes a method (hereinafter, referred to as a CIIP mode together with the existing CIIP mode or a "modified CIIP mode" to distinguish it from the existing CIIP mode) of searching for a block similar to the current block and using the block as an intra-prediction block in order to increase the accuracy of the intra-prediction block.

[0188] Unlike the conventional CIIP mode that uses surrounding sample values ​​of the current block to derive an on-screen prediction block of the current block, in the modified CIIP mode, the processor searches for a reference block in a given search area within the current picture / slice / tile, and uses the searched reference block to derive an on-screen prediction block of the current block.

[0189] For example, a processor can retrieve a reference flock using the following method:

[0190] - Method 1: The block with the smallest error between the adjacent template area of ​​the current block and the template area within the pre-specified area within the current picture / slice / tile.

[0191] - Method 2: The block with the smallest error between the inter-screen prediction block of the current block and the block samples within the pre-specified area within the current picture / slice / tile.

[0192] - Method 3: A block obtained by constructing a list of block vector predictor (BVP) candidates using the block vector (BV) information included in the adjacent blocks of the current block, and then using the block vector (BV) information in the candidate list pointed to by an index determined by signaling / parsing or a predefined method.

[0193] The method of utilizing the above-mentioned prediction blocks within the screen can be considered as a tool within the general merge mode and as part of the CIIP mode. Furthermore, the method can be used separately as another mode, along with the INTER mode and the MERGE mode.

[0194] The above method can be applied to replace the existing CIIP mode, and whether to select the existing on-screen prediction block derivation method or the proposed on-screen prediction block derivation method can be determined through signaling or a predefined derivation method.

[0195] The above method can operate as a sub-mode of the CIIP mode, and can operate as a sub-mode when considered as another mode together with the INTER mode and the MERGE mode. The sub-mode can include one or more of the methods for deriving intra-screen prediction blocks described in this document.

[0196] FIG. 9 is a diagram illustrating a method for decoding image information according to an embodiment of the present disclosure. FIG. 10 is a diagram illustrating a sequence for signaling or parsing information regarding a prediction mode according to an embodiment of the present disclosure. FIG. 11 is a diagram illustrating a sequence for signaling or parsing information regarding a prediction mode according to an embodiment of the present disclosure. FIG. 12 is a diagram illustrating directional and non-directional modes of an intra prediction mode according to an embodiment of the present disclosure.

[0197] The decoding method (S900) may include the operations described below.

[0198] The terms or names described below (e.g., names of syntax elements or names of variables, etc.) are merely examples, and the technical features of the present disclosure are not limited to the terms or names described below. For example, the image information described below may include various information according to the embodiments described in the present disclosure, and may include information described in at least one of the tables described above.

[0199] The operations described below do not constitute essential components of a decoding method according to an embodiment, and at least some of the operations described below may be omitted. Furthermore, the operations described below do not constitute sufficient components of a decoding method according to an embodiment, and the operations described above may be added. Furthermore, the operations described below form an embodiment together with the operations described above, unless they contradict the operations described above, and do not form a separate embodiment distinct from the operations described above.

[0200] The decoding method (S900) can be executed by a decoding device including a memory and a processor electrically connected to the memory, and can be executed by, for example, a processor.

[0201] The decoding device can obtain image information (S910).

[0202] For example, a processor of a decoding device can obtain image information including prediction information and residual information.

[0203] Video information can take various forms. For example, the video information can be a syntax element or a syntax structure comprising one or more syntax elements. Furthermore, the video information can be a raw byte sequence payload (RBSP) comprising one or more syntax elements or comprising one or more syntax structures. Furthermore, the video information can be a Network Abstraction Layer (NAL) unit comprising one or more RBSPs or a bitstream comprising one or more NALs.

[0204] The prediction information may include information related to the prediction of coding blocks included in each coded picture. For example, the prediction information may include information related to prediction modes, such as intra prediction mode, inter prediction mode, and intra block copy (IBC) prediction mode. For example, the prediction information may further include intra information related to non-directional mode, directional mode, matrix-weighted intra prediction (MIP), multi-reference line (MRL), or intra sub-partitions (ISP) with respect to intra prediction. For example, the prediction information may further include information related to inter modes, such as skip mode, regular MERGE mode, MMVD (Merge with Motion Vector Difference) mode, CIIP (Combined Inter and Intra Prediction) mode, TRIANGULAR mode, SbTMVP (Sbblock-based Temporal Motion Vector Prediction) mode, AFFINE MERGE mode, regular AMVP (Regular Advanced Motion Vector Prediction) mode, SMVD (Symmetric MVD) mode, and AFFINE AMVP mode, in relation to inter prediction. For example, the prediction information may further include information related to block-based delta pulse code modulation (BDPCM), palette, etc., in relation to screen content coding.

[0205] The residual information may include residual samples of coding blocks included in each of the coded pictures and information related to processing of the residual samples. For example, the residual information may include information related to residual samples, information related to quantization parameters (QPs), information related to multiple transform kernel selection (MTS), information related to sub-block transforms (SBTs), information related to low frequency non-separable transforms (LFNSTs), and the like.

[0206] The decoding device can derive a prediction mode for the current unit within the current picture (S920).

[0207] For example, a processor of a decoding device can derive a combined inter and intra prediction (CIIP) mode for a current unit within a current picture based on prediction information included in image information. For example, the prediction information can include information related to a CIIP mode for a coding unit, and the information related to the CIIP mode can include CIIP flag information indicating whether the CIIP mode is applied to the coding unit. Here, the CIIP mode can include a conventional CIIP mode (hereinafter, referred to as a “general CIIP mode”) and a CIIP mode according to an embodiment (hereinafter, referred to as a “CIIP mode” together with the conventional CIIP mode or referred to as a “modified CIIP mode” to distinguish it from the conventional CIIP mode).

[0208] One embodiment includes a signaling method when a mode is applied to generate a final prediction block through a weighted sum between an inter-screen prediction block and an IntraTMP or IBC prediction block in the process of generating a prediction block.

[0209] The prediction mode is divided into INTRA mode, IBC mode, and INTER mode, and each decoding process is performed, but various approaches can be allowed to improve compression performance through combination of heterogeneous modes, such as CIIP mode, which generates a final prediction block by weighted summing of inter-screen prediction blocks and intra-screen prediction blocks. For example, the CIIP mode, the weighted sum method of inter-screen prediction blocks and intra-screen prediction blocks, helps generate prediction blocks that are difficult to obtain with existing single prediction methods. However, the intra-screen prediction block in the above process is a prediction block derived using limited samples adjacent to the current block, and although it can reflect changes in brightness and samples compared to adjacent samples of the current block, it is difficult to reflect texture within the block. Therefore, one embodiment weight-sums the intra-screen prediction block obtained within the current picture / slice / tile with the inter-screen prediction block, thereby maintaining the existing brightness change compensation effect and reflecting more accurate pixel information, thereby increasing compression efficiency. In the weighted sum method of inter-screen prediction blocks and intra-screen prediction blocks proposed in one embodiment, the intra-screen prediction block can be obtained as follows.

[0210] - Method 1: The block with the smallest error between the adjacent template area of ​​the current block and the template area within the pre-specified area within the current picture / slice / tile.

[0211] - Method 2: The block with the smallest error between the inter-screen prediction block of the current block and the block samples within the pre-specified area within the current picture / slice / tile.

[0212] - Method 3: A block obtained by constructing a BVP candidate list using the BV information included in the adjacent blocks of the current block, and then using the BV information in the candidate list pointed to by the index determined by signaling / parsing or a predefined method.

[0213] The processor can derive a conventional CIIP mode for the current block (general CIIP mode) or a CIIP mode according to an embodiment for the current block (modified CIIP mode).

[0214] For example, the prediction information may include flag information or indexing information for indicating application of the CIIP mode, and may also include flag information or indexing information for indicating application of the general CIIP mode or application of the modified CIIP mode. The flag information or indexing information for indicating application of the general CIIP mode or application of the modified CIIP mode may be signaled from the encoding device to the decoding device and parsed by the decoding device.

[0215] One embodiment includes a signaling method for a mode when a method of generating a final prediction block through a weighted sum between inter-screen prediction blocks and intra-screen prediction blocks derived by the above method is set to one mode. According to one embodiment, the mode may be included as one of the merge modes. The general merge mode may be branched into a subblock merge mode, an mmvd mode, a regular merge mode, a GPM mode, a CIIP mode, etc., and the mode described in one embodiment (Proposed mode) may be included as a mode distinct from the above modes. Fig. 10 shows a signaling / parsing structure of each detailed mode within the general merge mode when the proposed mode is included.

[0216] For example, the prediction information may include first flag information (e.g., normal_intra in FIG. 10) and second flag information (e.g., ciip_flag in FIG. 10), and the first flag information may be signaled / parsed based on the second flag information. Here, whether a CIIP mode (including a normal CIIP mode and a modified CIIP mode) is applied to the current unit may be derived based on the value of the second flag information. Whether a modified CIIP mode (a CIIP mode according to an embodiment) is applied to the current unit may be derived based on the value of the second flag information and the value of the first flag information. In addition, whether a normal CIIP mode is applied to the current unit may be derived based on the value of the second flag information and the value of the first flag information.

[0217] Additionally, the prediction information may further include third flag information (e.g., regular_merge_flag in FIG. 10), and whether the merge mode is applied to the current unit may be derived based on the value of the third flag information. Additionally, second flag information may be signaled / parsed to derive whether the CIIP mode is applied based on the value of the third flag information.

[0218] In one embodiment, the mode may be included as another prediction mode. The prediction modes may be divided into INTER mode and MERGE mode, and the proposed method may be divided into another mode as BLENDED mode. The BLENDED mode described in one embodiment is an example, and it is obvious that the name and signaling method may be changed and applied. Fig. 11 shows an example of signaling / parsing the CIIP mode, GPM mode, and the proposed method included in the general merge mode separately from the MERGE mode as BLENDED mode, thereby reducing the signaling bits of the BLENDED mode.

[0219] For example, the prediction information may include first flag information (e.g., normal_intra in FIG. 11) and second flag information (e.g., ciip_flag in FIG. 11), and the first flag information may be obtained based on the second flag information. Here, whether a CIIP mode (including a normal CIIP mode and a modified CIIP mode) is applied to the current unit may be derived based on the second flag information. Whether a modified CIIP mode (a CIIP mode according to an embodiment) is applied to the current unit may be derived based on the value of the second flag information and the value of the first flag information. In addition, whether a normal CIIP mode is applied to the current unit may be derived based on the value of the second flag information and the value of the first flag information.

[0220] Additionally, the prediction information may further include fourth flag information (e.g., blend_flag in FIG. 11), and based on the value of the fourth flag information, whether the blended mode is applied to the current unit may be derived. Additionally, based on the value of the fourth flag information, second flag information may be signaled / parsed to derive whether the CIIP mode is applied.

[0221] According to the two signaling examples above, when ciip_flag is TRUE, it is possible to determine whether to apply the intra-screen prediction mode using the proposed method or the conventional method through the normal_intra flag.

[0222] Additionally, the prediction information may include flag information or indexing information to indicate application of the CIIP mode, and the prediction information may not include flag information or indexing information to indicate application of the general CIIP mode or application of the modified CIIP mode. The processor may derive application of the general CIIP mode or application of the modified CIIP mode based on set conditions.

[0223] One embodiment includes another signaling method for generating a final prediction block by weighting inter-screen prediction blocks and intra-screen prediction blocks derived by another method, when setting a mode. A mode utilizing this method can be applied as a replacement for the existing CIIP mode. In other words, the proposed method can be applied without additional signaling / parsing, such as normal_intra, for the proposed mode.

[0224] In one embodiment, the proposed method can be applied as part of the CIIP mode without additional signaling / parsing. That is, the applicability of the proposed method and the existing method for deriving prediction blocks within a screen can be determined without signaling. The method for deriving prediction blocks within a screen can be determined as follows, and a combination of the methods listed above can also be used.

[0225] - The processor may use the existing intra-screen prediction block derivation method if the inter-screen prediction block is a block obtained using the motion vector of a spatially adjacent block of the current block, and may use the proposed intra-screen prediction block derivation method if not (i.e., if the motion vector of a temporally adjacent block, a non-adjacent block, etc. is used). This is because the similarity between the current block and the adjacent block is high, and thus the similarity of the intra-screen prediction block derived from the adjacent sample can also be determined to be high.

[0226] For example, based on the fact that the inter-prediction sample array is derived based on the motion vectors of spatially adjacent blocks of the current block, the processor can derive the application of a general combined inter and intra prediction (CIIP) mode to the current block. Alternatively, based on the fact that the inter-prediction sample array is not derived based on the motion vectors of spatially adjacent blocks of the current block, the processor can derive the application of a modified CIIP mode to the current block.

[0227] - The processor can make a judgment based on the prediction mode of the spatially adjacent blocks of the current block. More specifically, the processor can use the existing intra-screen prediction block derivation method based on the number of blocks encoded / decoded in INTRA mode if the value is lower than a threshold, and use the intra-screen prediction block derivation method using the proposed method if the value is lower than a threshold.

[0228] For example, based on the number of blocks in the intra prediction mode adjacent to the current block being less than a threshold, the processor can derive the application of the general combined inter and intra prediction (CIIP) mode to the current block. Additionally, based on the number of blocks in the intra prediction mode adjacent to the current block being greater than or equal to a threshold, the processor can derive the application of the modified CIIP mode to the current block.

[0229] - The processor can determine whether to apply the proposed method without signaling using the error value. For example, the processor can generate a prediction sample in an adjacent template area of ​​the current block using the regular intra mode (e.g., the directional and non-directional modes of Fig. 12), which is an existing method for generating a prediction block within a screen, and calculate the error with the restored sample. The processor can calculate the error between the template area at the position indicated by the BV and the restored sample in the adjacent template area of ​​the current block using the proposed method, and generate the prediction block within the screen using the method with the smaller error among the two methods. The error calculation method using the regular intra mode can be limited to the error calculated using a specific mode (e.g., non-directional mode: Planar, DC).

[0230] For example, the processor may derive a first error between sample values ​​of a template region around the current block derived using directional or non-directional intra prediction and reconstructed sample values ​​of the template region, and may derive a second error between sample values ​​of the template region derived using intra block copy (IBC) prediction or intra template matching prediction (IntraTMP) and reconstructed sample values ​​of the template region. The processor may derive application of a normal CIIP mode to the current block based on the first error being less than the second error, and may derive application of a modified CIIP mode to the current block based on the first error being greater than or equal to the second error.

[0231] For example, the prediction information may include flag information or indexing information to indicate application of the CIIP mode, and may also include flag information or indexing information to indicate application of the general CIIP mode or application of the modified CIIP mode. The prediction information may include flag information or indexing information to indicate application of any one of a plurality of methods proposed in the modified CIIP mode. Furthermore, the processor may derive application of any one of the methods for deriving an array of samples within a screen in the modified CIIP mode according to a predetermined condition.

[0232] In a weighted sum method of an inter-screen prediction block and an intra-screen prediction block of one embodiment, the intra-screen prediction block can be obtained as follows.

[0233] - Method 1: The block with the smallest error between the adjacent template area of ​​the current block and the template area within the pre-specified area within the current picture / slice / tile.

[0234] - Method 2: The block with the smallest error between the inter-screen prediction block of the current block and the block samples within the pre-specified area within the current picture / slice / tile.

[0235] - Method 3: A block obtained by constructing a BVP candidate list using the BV information included in the adjacent blocks of the current block, and then using the BV information in the candidate list pointed to by the index determined by signaling / parsing or a predefined method.

[0236] According to one embodiment, only one of the three methods and the existing CIIP in-picture prediction block derivation method may be used, or two or more may be used, depending on predefined conditions or encoding / decoding period agreements. When only one is used, the processor may signal / parse according to the mode signaling method described above or determine the CIIP mode according to the mode derivation method described above. On the other hand, when two or more are used, the three methods may be implicitly distinguished as sub-modes of the CIIP mode described above or sub-modes of the BLENDED mode described above depending on additional conditions, or may be explicitly signaled in the form of a flag or index.

[0237] The priorities for the above three methods can be pre-specified according to the method defined identically in the decoder / encoder, and considering the differences in the signaling and error calculation methods of the intra-picture prediction block derivation methods of the above three methods and the existing CIIP, certain combinations may not be allowed as a sub-mode of the CIIP mode or a sub-mode of the BLENDED mode at the same time. For example, the IBC described in Method 3 calculates the optimal intra-picture prediction block based on the error with the original block at the current block position, whereas the intra-picture prediction block for the weighted sum method proposed in Methods 1 and 2 calculates the optimal intra-picture prediction block based on the error with the reconstructed block or the inter-picture prediction block, and therefore, they may not be used as a sub-mode of the CIIP mode or a sub-mode of the BLENDED mode at the same time.

[0238] For example, the prediction information may further include fifth indexing information. Based on the value of the fifth indexing information, the processor may derive an within-screen reference sample array based on information about a block vector included in the prediction information, or derive an within-screen reference sample array based on an area with a minimum error with respect to a template area surrounding the current block, or derive an within-screen reference sample array based on an area with a minimum error with respect to a screen-to-screen reference sample array.

[0239] For example, the prediction information may further include sixth indexing information. Based on the value of the sixth indexing information, the processor may derive an array of reference samples within the screen based on a motion vector candidate list including at least some of the motion vectors of spatially neighboring blocks of the current block, the motion vectors of temporally neighboring blocks, or the history-based motion vectors, and a block vector candidate list including at least some of the block vectors for the current block and the block vectors based on the template region surrounding the current block.

[0240] In this way, the processor can derive the application of the CIIP mode to the current unit based on the CIIP flag information for the current unit. The processor can derive the application of the modified CIIP mode based on additional flag information or based on a predetermined condition. Furthermore, the processor can derive a method for deriving an array of samples within the screen in the modified CIIP mode based on additional flag information or based on a predetermined condition.

[0241] The decoding device can derive an inter-screen reference sample array from the reference picture (S930).

[0242] For example, a processor of a decoding device can derive an array of inter-screen reference samples based on prediction information contained in the image information.

[0243] The image information may include information related to an inter prediction mode for the current unit. The information related to the inter prediction mode for the current unit may include information related to the inter mode, information related to a reference picture, and information related to a motion vector.

[0244] The processor can obtain information related to an inter-mode, information related to a reference picture, and information related to a motion vector, and derive an inter-screen reference sample array based on the information related to the inter-mode, information related to the reference picture, and information related to the motion vector.

[0245] For example, the processor can generate a list of reference picture candidates and derive a reference picture from among the reference picture candidates based on information related to the reference pictures.

[0246] The processor can generate a list of motion vector candidates based on at least some of motion vectors of spatial neighboring blocks of the current block, motion vectors of temporal neighboring blocks, or historical motion vectors, and can derive a motion vector from among the motion vector candidates based on information related to the motion vectors.

[0247] The processor can derive an inter-reference block based on a reference picture and a motion vector, and can derive an inter-screen reference sample array from the inter-reference block.

[0248] In this way, the processor can derive an array of inter-screen reference samples.

[0249] The decoding device can derive an intra-screen reference sample array within the current picture (S940).

[0250] For example, a processor of a decoding device can derive an array of reference samples within a screen based on prediction information contained in the image information.

[0251] In a weighted sum method of inter-screen prediction blocks and intra-screen prediction blocks of one embodiment, the intra-screen prediction blocks can be derived as follows.

[0252] - Method 1: The block with the smallest error between the adjacent template area of ​​the current block and the template area within the pre-specified area within the current picture / slice / tile.

[0253] - Method 2: The block with the smallest error between the inter-screen prediction block of the current block and the block samples within the pre-specified area within the current picture / slice / tile.

[0254] - Method 3: A block obtained by constructing a BVP candidate list using the BV information included in the adjacent blocks of the current block, and then using the BV information in the candidate list pointed to by the index determined by signaling / parsing or a predefined method.

[0255] The prediction information may include information related to the intra prediction mode for the current unit or information related to the IBC prediction mode. If the prediction information includes information related to the IBC prediction mode for the current unit, the information related to the IBC prediction mode may include information related to the block vector. Additionally, if the prediction information includes information related to the intra prediction mode for the current unit, the information related to the intra prediction mode may include intra mode information.

[0256] For example, when the prediction information of the image information includes information related to the IBC prediction mode for the current unit, the prediction information may include information related to a block vector. The processor may derive an array of reference samples within the screen based on the information about the block vector included in the prediction information. For example, the processor may generate a list of block vector candidates including at least some of the block vectors of the surrounding blocks of the current block and the block vectors based on the template region around the current block, and derive the array of reference samples within the screen based on the block vector derived from the information about the block vector. The IBC prediction mode basically performs prediction within the current picture, but may be performed similarly to the inter prediction mode in that it derives a reference block within the current picture. That is, at least one of the inter prediction techniques may be used in the IBC prediction mode. Here, the block vector is used to indicate a displacement from the current block to a reference block that has already been reconstructed within the current picture.

[0257] For example, if the prediction information includes information related to an intra prediction mode for the current unit, the prediction information may include information related to intra-template matching prediction (IntraTMP). The processor may derive a region with a minimum error with respect to a template region surrounding the current block within a search region, and derive an array of reference samples within the screen based on the derived region. Intra-template matching prediction (IntraTMP) is an intra-prediction mode in which an L-shaped template in a reconstructed portion of the current frame copies an optimal prediction block that matches the current template. For a predefined search range, the processor searches for a template most similar to the current template in the reconstructed portion of the current frame and uses the block as a prediction block.

[0258] For example, an intra-screen prediction method separate from the IBC prediction mode and the intra-prediction mode may be proposed. An intra-screen prediction method using an inter-screen reference block may be proposed. In addition, the prediction information may include information related to prediction separate from the IBC prediction mode and the intra-prediction mode. For example, the prediction information may include information related to prediction using an inter-screen reference block. The processor may derive an inter-screen reference block, i.e., an inter-screen reference sample array, as described above. The processor may derive an area in the search area with a minimum error with respect to the inter-screen reference sample array, and derive an intra-screen reference block, i.e., an intra-screen reference sample array, based on the derived area.

[0259] In this way, the processor can derive an array of reference samples within the screen.

[0260] The decoding device can derive an array of predicted samples for the current block (S950).

[0261] For example, a processor of a decoding device can derive a predicted sample array based on an inter-screen reference sample array and an intra-screen reference sample array.

[0262] The processor can derive an inter-screen prediction sample array, i.e., an inter-screen prediction block, by copying and pasting an inter-screen reference sample array or by performing sample interpolation. Furthermore, the processor can derive an intra-screen prediction sample array, i.e., an intra-screen prediction block, by copying and pasting an intra-screen reference sample array or by performing sample interpolation.

[0263] The processor can derive a prediction sample array, i.e., a prediction block, by weighting or weighting the inter-screen prediction sample array and the intra-screen prediction sample array. The weights for the weighted sum or weighted average can be derived based on at least one of the number of intra-mode blocks located around the current block, the POC difference between the current picture and the reference picture, the temporal identifier (Temporal ID) of the current picture, or a predetermined value.

[0264] For example, a processor of a decoding device can derive an array of predicted samples based on a plurality of intra-screen reference block candidates, a plurality of intra-screen reference block candidates, and a plurality of weight candidates.

[0265] The processor may generate a motion vector candidate list including a plurality of motion vector candidates in relation to inter prediction. For example, the processor may generate the same candidate list as the motion vector candidate list of the merge mode, or the same candidate list as the motion vector candidate list of the AMVP mode. For example, the processor may generate the motion vector candidate list based on at least some of the motion vectors of the spatial neighboring blocks of the current block, the motion vectors of the temporal neighboring blocks, or the history-based motion vectors.

[0266] The processor may generate a block vector candidate list including a plurality of block vector candidates in relation to intra prediction or IBC prediction. For example, the processor may derive block vector candidates for a predicted block within a screen based on intra-template matching prediction or obtain block vector candidates based on IBC prediction, and may generate a block vector candidate list based on the block vector candidates. For example, the processor may generate a block vector candidate list including at least some of the block vectors for a current block and block vectors based on template regions surrounding the current block.

[0267] Additionally, the processor may generate a weight candidate list including multiple weight candidates in relation to the weighted sum between inter-screen prediction block candidates and intra-screen prediction block candidates. The weight candidates may include, but are not limited to, 3:1, 2:2, and 3:1, for example.

[0268] The processor can derive a motion vector and a block vector based on a combination of motion vector candidates included in a motion vector candidate list and block vector candidates included in a block vector candidate list. For example, the processor can derive weighted template samples by weighting samples of a template region around an inter-screen prediction block candidate based on the motion vector candidate and samples of a template region around an intra-screen prediction block candidate based on the block vector candidate, and derive a motion vector and a block vector based on an error between the weighted template samples and samples of the template region of the current block. For example, the processor can derive a motion vector candidate and a block vector candidate having a minimum error between the weighted template samples and samples of the template region of the current block as a motion vector and a block vector.

[0269] The decoding device can derive a restored sample array for the current block (S960).

[0270] For example, a processor of a decoding device can derive a reconstructed sample array based on a predicted sample array and a residual sample array for the current block.

[0271] The processor can generate an array of prediction samples in the manner described above.

[0272] The processor can derive a residual sample array based on residual information included in the image information. For example, the processor can derive quantized transform coefficients based on information related to the residual samples. The processor can derive transform coefficients by performing inverse quantization on the quantized transform coefficients based on information related to quantization parameters. The processor can derive corrected transform coefficients by performing a second transform on the transform coefficients based on information related to a low-frequency non-separable transform. The processor can derive a residual sample array by performing a first transform on the corrected transform coefficients based on information related to multiple transform kernel selection.

[0273] The processor can derive a reconstructed sample array based on a sum between the predicted sample array and the residual sample array.

[0274] As described above, the decoding device can utilize the reconstructed region within the current picture to derive an array of predicted samples within the picture in CIIP mode. For example, the decoding device can derive an within-picture reference block using intra-template matching prediction or IBC mode, and derive an within-picture prediction block, i.e., an array of predicted samples within the picture, using the within-picture reference block, i.e., the array of reference samples within the picture.

[0275] Accordingly, the decoding device can reflect texture within a block, which is difficult to reflect with an array of predicted samples within the screen derived using reference samples adjacent to the current block, using blocks within the current picture. In addition, the decoding device can suppress, minimize, or prevent the problem of data for residual samples increasing due to a large difference between the reconstructed block and the predicted block even when moving away from the top and left of the current block by using a reference block most similar to the reconstructed block within the current picture.

[0276] FIG. 13 is a diagram illustrating a method for encoding image information according to one embodiment of the present disclosure.

[0277] The encoding method (S1300) may include the operations described below.

[0278] The terms or names described below (e.g., names of syntax elements or names of variables, etc.) are merely examples, and the technical features of the present disclosure are not limited to the terms or names described below. For example, the image information described below may include various information according to the embodiments described in the present disclosure, and may include information described in at least one of the tables described above.

[0279] The operations described below do not constitute essential components of an encoding method according to an embodiment, and at least some of the operations described below may be omitted. Furthermore, the operations described below do not constitute sufficient components of an encoding method according to an embodiment, and the operations described above may be added. Furthermore, the operations described below form an embodiment together with the operations described above, unless they contradict the operations described above, and do not form a separate embodiment distinct from the operations described above.

[0280] The encoding method (S1300) can be executed by an encoding device including a memory and a processor electrically connected to the memory, and can be executed by, for example, a processor.

[0281] The encoding device can determine the prediction mode for the current unit within the current picture (S1310).

[0282] For example, a processor of an encoding device may compare Rate Distortion (RD) costs for various prediction modes and determine a combined inter and intra prediction (CIIP) mode for a current unit within a current picture based on the RD cost. Furthermore, the processor may generate prediction information including information related to the CIIP mode.

[0283] Prediction modes can include various modes such as intra prediction mode, inter prediction mode, intra block copy (IBC) prediction mode, etc. With regard to intra prediction mode, there are various intra modes / types including non-directional mode, directional mode, matrix-weighted intra prediction (MIP), multi-reference line (MRL), or intra sub-partitions (ISP). Also, with respect to inter prediction modes, there are various inter modes / types including skip mode, regular MERGE mode, MMVD (Merge with Motion Vector Difference) mode, CIIP (Combined Inter and Intra Prediction) mode, TRIANGULAR mode, SbTMVP (Sbblock-based Temporal Motion Vector Prediction) mode, AFFINE MERGE mode, regular AMVP (Regular Advanced Motion Vector Prediction) mode, SMVD (Symmetric MVD) mode, and AFFINE AMVP mode.

[0284] The processor can compare RD (Rate Distortion) costs for intra modes / types and inter modes / types, and determine the CIIP mode for the current unit within the current picture based on the RD cost. In addition, the processor can generate prediction information including information about the CIIP mode.

[0285] The encoding device can derive an inter-screen reference sample array from a reference picture (S1320).

[0286] For example, a processor of an encoding device may derive an inter-screen reference sample array based on a combined inter- and intra-prediction (CIIP) mode. Furthermore, the processor may generate prediction information for deriving the inter-screen reference sample array based on the combined inter- and intra-prediction (CIIP) mode.

[0287] For example, the processor can compare the RD cost for various inter modes / types and derive the inter mode / type based on the RD cost. The inter prediction modes can include various inter modes / types such as skip mode, regular MERGE mode, MMVD (Merge with Motion Vector Difference) mode, CIIP (Combined Inter and Intra Prediction) mode, TRIANGULAR mode, SbTMVP (Sbblock-based Temporal Motion Vector Prediction) mode, AFFINE MERGE mode, regular AMVP (Regular Advanced Motion Vector Prediction) mode, SMVD (Symmetric MVD) mode, and AFFINE AMVP mode.

[0288] For example, the processor can derive reference pictures and motion vectors based on the inter mode / type. The processor can generate a list of reference picture candidates based on the inter mode / type, and derive a reference picture from among the reference picture candidates based on an RD cost for the reference picture candidates. In addition, the processor can generate a list of motion vector candidates based on the inter mode / type, and derive a motion vector from among the motion vector candidates based on an RD cost for the motion vector candidate list.

[0289] For example, the processor may perform inter prediction based on the inter mode, the reference picture, and the motion vector, and derive an array of inter-screen reference samples. Furthermore, the processor may generate prediction information including information related to the inter mode, information related to the reference picture, and information related to the motion vector, based on the inter mode, the reference picture, and the motion vector.

[0290] The encoding device can derive an intra-screen reference sample array within the current picture (S1330).

[0291] For example, a processor of an encoding device may derive an array of reference samples within a picture based on a combined inter and intra prediction (CIIP) mode. Furthermore, the processor may generate prediction information for deriving the array of reference samples within a picture based on the combined inter and intra prediction (CIIP) mode.

[0292] For example, the processor can compare RD costs for various intra modes / types or IBC prediction modes, and determine a mode / type for deriving an array of within-screen reference samples based on the RD cost. In particular, the processor can compare RD costs for intra modes / types by copying and pasting an within-screen block into the current block or performing sample interpolation, and determine a mode / type for deriving an array of within-screen reference samples based on the RD cost. For example, the processor can compare RD costs for intra-template matching prediction (intraTMP) or IBC prediction modes, and determine a mode / type for deriving an array of within-screen reference samples based on the RD cost.

[0293] In a weighted sum method of inter-screen prediction blocks and intra-screen prediction blocks of one embodiment, the intra-screen prediction blocks can be derived as follows.

[0294] - Method 1: The block with the smallest error between the adjacent template area of ​​the current block and the template area within the pre-specified area within the current picture / slice / tile.

[0295] - Method 2: The block with the smallest error between the inter-screen prediction block of the current block and the block samples within the pre-specified area within the current picture / slice / tile.

[0296] - Method 3: A block obtained by constructing a BVP candidate list using the BV information included in the adjacent blocks of the current block, and then using the BV information in the candidate list pointed to by the index determined by signaling / parsing or a predefined method.

[0297] For example, if the IBC prediction mode is determined in relation to deriving an array of prediction samples within a screen, the processor can derive a block vector based on the IBC prediction mode. The processor can generate a list of block vector candidates based on the IBC prediction mode, and can derive a block vector from among the block vector candidates based on an RD cost for the list of block vector candidates. The IBC prediction mode basically performs prediction within the current picture, but can be performed similarly to the inter prediction mode in that it derives a reference block within the current picture. That is, at least one of the inter prediction techniques can be used in the IBC prediction mode. Here, the block vector is used to indicate a displacement from the current block to a reference block that has already been reconstructed within the current picture.

[0298] The processor can generate prediction information including information related to the IBC prediction mode and information related to the block vector based on the IBC prediction mode and the block vector.

[0299] For example, in relation to deriving an array of predicted samples within a screen, if intra-template matching prediction (intraTMP) is determined, the processor can derive a region within a search region with a minimum error with respect to a template region surrounding the current block based on intraTMP, and derive an array of reference samples within the screen based on the derived region. Intra-template matching prediction (IntraTMP) is an intra-prediction mode in which an L-shaped template in a reconstructed portion of the current frame copies an optimal prediction block that matches the current template. For a predefined search range, the processor searches for a template most similar to the current template in the reconstructed portion of the current frame and uses the block as a prediction block.

[0300] The processor can generate prediction information including information related to the intra-template matching prediction (IntraTMP) based on the intra-template matching prediction (IntraTMP).

[0301] A separate intra-screen prediction method may be proposed, separate from the IBC prediction mode and intra-prediction mode. For example, an intra-screen prediction method using inter-screen reference blocks may be proposed.

[0302] The processor can derive an inter-screen reference block, i.e., an inter-screen reference sample array, as described above. The processor can derive an area within the search region with a minimum error with respect to the inter-screen reference sample array, and derive an intra-screen reference block, i.e., an intra-screen reference sample array, based on the derived area.

[0303] Additionally, if the inter-screen reference sample array includes a plurality of inter-screen reference sample arrays, the processor can derive the representative reference sample array. For example, the processor can derive the representative reference sample array based on at least one of a weighted sum of the plurality of inter-screen reference sample arrays, a picture of count (POC) of the reference picture, a POC difference between the current picture and the reference picture, a sample value of a template region around the plurality of inter-screen reference sample arrays, a quantization parameter of the reference picture, or a temporal identifier (Temporal ID) of the current picture.

[0304] The processor can generate prediction information including information related to an intra-screen prediction method using an inter-screen reference block based on an intra-screen prediction method using an inter-screen reference block.

[0305] In this way, the processor can derive an array of reference samples within the screen.

[0306] The encoding device can derive an array of predicted samples for the current block (S1340).

[0307] For example, a processor of an encoding device may derive a prediction sample array based on an inter-screen reference sample array and an intra-screen reference sample array. Additionally, the processor may generate prediction information for deriving the prediction sample array based on the inter-screen reference sample array and the intra-screen reference sample array.

[0308] The processor can derive an inter-screen prediction sample array, i.e., an inter-screen prediction block, by copying and pasting an inter-screen reference sample array or by performing sample interpolation. Furthermore, the processor can derive an intra-screen prediction sample array, i.e., an intra-screen prediction block, by copying and pasting an intra-screen reference sample array or by performing sample interpolation.

[0309] The processor can derive a prediction sample array, i.e., a prediction block, by weighting or weighting the inter-screen prediction sample array and the intra-screen prediction sample array. The weights for the weighted sum or weighted average can be derived based on at least one of the number of intra-mode blocks located around the current block, the POC difference between the current picture and the reference picture, the temporal identifier (Temporal ID) of the current picture, or a predetermined value.

[0310] The processor may optionally generate prediction information that includes information about the weights.

[0311] For example, a processor of an encoding device can derive an array of prediction samples based on a plurality of intra-screen reference block candidates, a plurality of intra-screen reference block candidates, and a plurality of weight candidates.

[0312] The processor may generate a motion vector candidate list including a plurality of motion vector candidates in relation to inter prediction. For example, the processor may generate the same candidate list as the motion vector candidate list of the merge mode, or the same candidate list as the motion vector candidate list of the AMVP mode. For example, the processor may generate the motion vector candidate list based on at least some of the motion vectors of the spatial neighboring blocks of the current block, the motion vectors of the temporal neighboring blocks, or the history-based motion vectors.

[0313] The processor may generate a block vector candidate list including a plurality of block vector candidates in relation to intra prediction or IBC prediction. For example, the processor may derive block vector candidates for a predicted block within a screen based on intra-template matching prediction or obtain block vector candidates based on IBC prediction, and may generate a block vector candidate list based on the block vector candidates. For example, the processor may generate a block vector candidate list including at least some of the block vectors for a current block and block vectors based on template regions surrounding the current block.

[0314] Additionally, the processor may generate a weight candidate list including multiple weight candidates in relation to the weighted sum between inter-screen prediction block candidates and intra-screen prediction block candidates. The weight candidates may include, but are not limited to, 3:1, 2:2, and 3:1, for example.

[0315] The processor can derive a motion vector and a block vector based on a combination of motion vector candidates included in a motion vector candidate list and block vector candidates included in a block vector candidate list. For example, the processor can derive weighted template samples by weighting samples of a template region around an inter-screen prediction block candidate based on the motion vector candidate and samples of a template region around an intra-screen prediction block candidate based on the block vector candidate, and derive a motion vector and a block vector based on an error between the weighted template samples and samples of the template region of the current block. For example, the processor can derive a motion vector candidate and a block vector candidate having a minimum error between the weighted template samples and samples of the template region of the current block as a motion vector and a block vector.

[0316] The processor can optionally generate prediction information that includes information about the block vector and information about the weights.

[0317] The encoding device can derive a residual sample array for the current block (S1350).

[0318] For example, a processor of an encoding device can derive a residual sample array based on an original sample array and a predicted sample array for the current block.

[0319] The processor can generate an array of prediction samples in the manner described above.

[0320] The processor can derive a residual sample array based on the difference between the original sample array and the predicted sample array. The processor can also generate residual information associated with the residual sample array.

[0321] The processor can derive transform coefficients by performing a first transform on an array of residual samples, and can also generate information related to multiple transform kernel selection. The processor can derive corrected transform coefficients by performing a second transform on the transform coefficients, and can generate information related to a low-frequency inseparable transform. The processor can derive quantized transform coefficients by performing quantization on the corrected transform coefficients, and can generate information related to a quantization parameter. The processor can derive information related to residual samples based on the quantized transform coefficients. Furthermore, the processor can generate residual information based on the information related to multiple transform kernel selection, the information related to the low-frequency inseparable transform, the information related to the quantization parameter, and the information related to the residual samples.

[0322] The encoding device can encode image information (S1660).

[0323] For example, a processor of an encoding device can encode image information including prediction information and residual information.

[0324] Video information can take various forms. For example, the video information can be a syntax element or a syntax structure comprising one or more syntax elements. Furthermore, the video information can be a raw byte sequence payload (RBSP) comprising one or more syntax elements or comprising one or more syntax structures. Furthermore, the video information can be a Network Abstraction Layer (NAL) unit comprising one or more RBSPs or a bitstream comprising one or more NALs.

[0325] For example, a processor of an encoding device may generate prediction information for deriving a combined inter and intra prediction (CIIP) mode for a current unit within a current picture based on prediction information included in image information. For example, the prediction information may include information related to a CIIP mode for a coding unit, and the information related to the CIIP mode may include CIIP flag information indicating whether the CIIP mode is applied to the coding unit.

[0326] Here, the CIIP mode may include a conventional CIIP mode (hereinafter referred to as a “general CIIP mode”) and a CIIP mode according to an embodiment (hereinafter referred to as a “CIIP mode” together with the conventional CIIP mode or a “modified CIIP mode” to distinguish it from the conventional CIIP mode).

[0327] One embodiment includes a signaling method when a mode is applied to generate a final prediction block through a weighted sum between an inter-screen prediction block and an IntraTMP or IBC prediction block in the process of generating a prediction block.

[0328] The prediction mode is divided into INTRA mode, IBC mode, and INTER mode, and each decoding process is performed, but various approaches can be allowed to improve compression performance through combination of heterogeneous modes, such as CIIP mode, which generates a final prediction block by weighted summing of inter-screen prediction blocks and intra-screen prediction blocks. For example, the CIIP mode, the weighted sum method of inter-screen prediction blocks and intra-screen prediction blocks, helps generate prediction blocks that are difficult to obtain with existing single prediction methods. However, the intra-screen prediction block in the above process is a prediction block derived using limited samples adjacent to the current block, and although it can reflect changes in brightness and samples compared to adjacent samples of the current block, it is difficult to reflect texture within the block. Therefore, one embodiment weight-sums the intra-screen prediction block obtained within the current picture / slice / tile with the inter-screen prediction block, thereby maintaining the existing brightness change compensation effect and reflecting more accurate pixel information, thereby increasing compression efficiency. In the weighted sum method of inter-screen prediction blocks and intra-screen prediction blocks proposed in one embodiment, the intra-screen prediction block can be obtained as follows.

[0329] - Method 1: The block with the smallest error between the adjacent template area of ​​the current block and the template area within the pre-specified area within the current picture / slice / tile.

[0330] - Method 2: The block with the smallest error between the inter-screen prediction block of the current block and the block samples within the pre-specified area within the current picture / slice / tile.

[0331] - Method 3: A block obtained by constructing a BVP candidate list using the BV information included in the adjacent blocks of the current block, and then using the BV information in the candidate list pointed to by the index determined by signaling / parsing or a predefined method.

[0332] The processor can generate prediction information to derive a conventional CIIP mode for the current block (general CIIP mode) or to derive a CIIP mode according to one embodiment for the current block (modified CIIP mode).

[0333] For example, the prediction information may include flag information or indexing information for indicating application of the CIIP mode, and may also include flag information or indexing information for indicating application of the general CIIP mode or application of the modified CIIP mode. The flag information or indexing information for indicating application of the general CIIP mode or application of the modified CIIP mode may be signaled from the encoding device to the decoding device and parsed by the decoding device.

[0334] One embodiment includes a signaling method for a mode when a method of generating a final prediction block through a weighted sum between inter-screen prediction blocks and intra-screen prediction blocks derived by the above method is set to one mode. According to one embodiment, the mode may be included as one of the merge modes. The general merge mode may be branched into a subblock merge mode, an mmvd mode, a regular merge mode, a GPM mode, a CIIP mode, etc., and the mode described in one embodiment (Proposed mode) may be included as a mode distinct from the above modes. Fig. 10 shows a signaling / parsing structure of each detailed mode within the general merge mode when the proposed mode is included.

[0335] For example, the prediction information may include first flag information (e.g., normal_intra in FIG. 10) and second flag information (e.g., ciip_flag in FIG. 10), and the first flag information may be signaled / parsed based on the second flag information. Here, whether a CIIP mode (including a normal CIIP mode and a modified CIIP mode) is applied to the current unit may be derived based on the value of the second flag information. Whether a modified CIIP mode (a CIIP mode according to an embodiment) is applied to the current unit may be derived based on the value of the second flag information and the value of the first flag information. In addition, whether a normal CIIP mode is applied to the current unit may be derived based on the value of the second flag information and the value of the first flag information.

[0336] Additionally, the prediction information may further include third flag information (e.g., regular_merge_flag in FIG. 10), and whether the merge mode is applied to the current unit may be derived based on the value of the third flag information. Additionally, second flag information may be signaled / parsed to derive whether the CIIP mode is applied based on the value of the third flag information.

[0337] In one embodiment, the mode may be included as another prediction mode. The prediction modes may be divided into INTER mode and MERGE mode, and the proposed method may be divided into another mode as BLENDED mode. The BLENDED mode described in one embodiment is an example, and it is obvious that the name and signaling method may be changed and applied. Fig. 11 shows an example of signaling / parsing the CIIP mode, GPM mode, and the proposed method included in the general merge mode separately from the MERGE mode as BLENDED mode, thereby reducing the signaling bits of the BLENDED mode.

[0338] For example, the prediction information may include first flag information (e.g., normal_intra in FIG. 11) and second flag information (e.g., ciip_flag in FIG. 11), and the first flag information may be obtained based on the second flag information. Here, whether a CIIP mode (including a normal CIIP mode and a modified CIIP mode) is applied to the current unit may be derived based on the second flag information. Whether a modified CIIP mode (a CIIP mode according to an embodiment) is applied to the current unit may be derived based on the value of the second flag information and the value of the first flag information. In addition, whether a normal CIIP mode is applied to the current unit may be derived based on the value of the second flag information and the value of the first flag information.

[0339] Additionally, the prediction information may further include fourth flag information (e.g., blend_flag in FIG. 11), and based on the value of the fourth flag information, whether the blended mode is applied to the current unit may be derived. Additionally, based on the value of the fourth flag information, second flag information may be signaled / parsed to derive whether the CIIP mode is applied.

[0340] According to the two signaling examples above, when ciip_flag is TRUE, it is possible to determine whether to apply the intra-screen prediction mode using the proposed method or the conventional method through the normal_intra flag.

[0341] Additionally, the prediction information may include flag information or indexing information for indicating application of the CIIP mode, and the prediction information may not include flag information or indexing information for indicating application of the general CIIP mode or application of the modified CIIP mode. The processor may induce the encoding / decoding device to derive application of the general CIIP mode or application of the modified CIIP mode according to a set condition.

[0342] One embodiment includes another signaling method for generating a final prediction block by weighting inter-screen prediction blocks and intra-screen prediction blocks derived by another method, when setting a mode. A mode utilizing this method can be applied as a replacement for the existing CIIP mode. In other words, the proposed method can be applied without additional signaling / parsing, such as normal_intra, for the proposed mode.

[0343] In one embodiment, the proposed method can be applied as part of the CIIP mode without additional signaling / parsing. That is, the applicability of the proposed method and the existing method for deriving prediction blocks within a screen can be determined without signaling. The method for deriving prediction blocks within a screen can be determined as follows, and a combination of the methods listed above can also be used.

[0344] - The encoding / decoding device can use the existing intra-screen prediction block derivation method if the inter-screen prediction block is a block obtained by using the motion vector of a spatially adjacent block of the current block, and can use the proposed intra-screen prediction block derivation method if not (i.e., if the motion vector of a temporally adjacent block, a non-adjacent block, etc. is used). This is because the similarity between the current block and the adjacent block is high, and thus the similarity of the intra-screen prediction block derived from the adjacent sample can also be determined to be high.

[0345] For example, the processor may induce the encoding / decoding device to derive application of a general combined inter and intra prediction (CIIP) mode for the current block based on the fact that the inter-prediction sample array is derived based on motion vectors of spatially adjacent blocks of the current block. Furthermore, the processor may induce the encoding / decoding device to derive application of a modified CIIP mode for the current block based on the fact that the inter-prediction sample array is not derived based on motion vectors of spatially adjacent blocks of the current block.

[0346] - The encoding / decoding device can make a judgment based on the prediction mode of the spatially adjacent blocks of the current block. More specifically, the encoding / decoding device can use the existing intra-screen prediction block derivation method based on the number of blocks encoded / decoded in INTRA mode, if the value is lower than a threshold value, and otherwise use the intra-screen prediction block derivation method using the proposed method.

[0347] For example, the processor may induce the encoding / decoding device to derive application of a general combined inter and intra prediction (CIIP) mode to the current block based on the number of blocks in the intra prediction mode adjacent to the current block being less than a threshold value. Furthermore, the processor may induce the encoding / decoding device to derive application of a modified CIIP mode to the current block based on the number of blocks in the intra prediction mode adjacent to the current block being greater than or equal to the threshold value.

[0348] - The encoding / decoding device can determine whether the proposed method is applied using the error value without signaling. For example, the encoding / decoding device can generate a prediction sample in an adjacent template area of ​​the current block using the regular intra mode (e.g., the directional and non-directional modes of Fig. 12), which is an existing method for generating a prediction block within a screen, and calculate the error with respect to the restored sample. The encoding / decoding device can calculate the error between the template area at the position indicated by the BV and the restored sample in the adjacent template area of ​​the current block using the proposed method, and generate the prediction block within the screen using the method with the smaller error among the two methods. The error calculation method using the regular intra mode can be limited to the error calculated using a specific mode (e.g., non-directional mode: Planar, DC).

[0349] For example, the processor may cause the encoding / decoding device to derive a first error between sample values ​​of a template region around the current block derived using directional or non-directional intra prediction and reconstructed sample values ​​of the template region, and to derive a second error between sample values ​​of the template region derived using intra block copy (IBC) prediction or intra template matching prediction (IntraTMP) and reconstructed sample values ​​of the template region. The processor may cause the encoding / decoding device to derive application of a normal CIIP mode to the current block based on the first error being less than the second error, and to derive application of a modified CIIP mode to the current block based on the first error being greater than or equal to the second error.

[0350] For example, the prediction information may include flag information or indexing information to indicate application of the CIIP mode, and may also include flag information or indexing information to indicate application of the general CIIP mode or application of the modified CIIP mode. The prediction information may include flag information or indexing information to indicate application of any one of a plurality of methods proposed in the modified CIIP mode. In addition, the processor may induce the encoding / decoding device to derive application of any one of the methods for deriving an array of samples within a screen in the modified CIIP mode according to a predetermined condition.

[0351] In a weighted sum method of an inter-screen prediction block and an intra-screen prediction block of one embodiment, the intra-screen prediction block can be obtained as follows.

[0352] - Method 1: The block with the smallest error between the adjacent template area of ​​the current block and the template area within the pre-specified area within the current picture / slice / tile.

[0353] - Method 2: The block with the smallest error between the inter-screen prediction block of the current block and the block samples within the pre-specified area within the current picture / slice / tile.

[0354] - Method 3: A block obtained by constructing a BVP candidate list using the BV information included in the adjacent blocks of the current block, and then using the BV information in the candidate list pointed to by the index determined by signaling / parsing or a predefined method.

[0355] According to one embodiment, only one of the three methods and the existing CIIP in-picture prediction block derivation method may be used, or two or more may be used, depending on predefined conditions or encoding / decoding period agreements. When only one is used, the processor may signal / parse according to the mode signaling method described above or determine the CIIP mode according to the mode derivation method described above. On the other hand, when two or more are used, the three methods may be implicitly distinguished as sub-modes of the CIIP mode described above or sub-modes of the BLENDED mode described above depending on additional conditions, or may be explicitly signaled in the form of a flag or index.

[0356] The priorities for the above three methods can be pre-specified according to the method defined identically in the decoder / encoder, and considering the differences in the signaling and error calculation methods of the intra-picture prediction block derivation methods of the above three methods and the existing CIIP, certain combinations may not be allowed as a sub-mode of the CIIP mode or a sub-mode of the BLENDED mode at the same time. For example, the IBC described in Method 3 calculates the optimal intra-picture prediction block based on the error with the original block at the current block position, whereas the intra-picture prediction block for the weighted sum method proposed in Methods 1 and 2 calculates the optimal intra-picture prediction block based on the error with the reconstructed block or the inter-picture prediction block, and therefore, they may not be used as a sub-mode of the CIIP mode or a sub-mode of the BLENDED mode at the same time.

[0357] For example, the prediction information may further include fifth indexing information. Based on the value of the fifth indexing information, the processor may induce the encoding / decoding device to derive an within-screen reference sample array based on information about a block vector included in the prediction information, or to derive an within-screen reference sample array based on an area with a minimum error with respect to a template area surrounding the current block, or to derive an within-screen reference sample array based on an area with a minimum error with respect to an inter-screen reference sample array.

[0358] For example, the prediction information may further include sixth indexing information. The processor may induce the encoding / decoding device to derive an array of reference samples within the screen based on a list of motion vector candidates including at least some of motion vectors of spatial neighboring blocks of the current block, motion vectors of temporal neighboring blocks, or history-based motion vectors, and a list of block vector candidates including at least some of block vectors for the current block and block vectors based on template regions surrounding the current block, based on a value of the sixth indexing information.

[0359] As described above, the encoding device can utilize the reconstructed region within the current picture to derive an array of predicted samples within the screen in the CIIP mode. For example, the encoding device can derive an within-screen reference block using intra-template matching prediction or IBC mode, and derive an within-screen prediction block, i.e., an array of predicted samples within the screen, using the within-screen reference block, i.e., the array of reference samples within the screen.

[0360] Accordingly, the encoding device can reflect texture within a block, which is difficult to reflect with an array of predicted samples within the screen derived using reference samples adjacent to the current block, using a block within the current picture. In addition, the encoding device can suppress, minimize, or prevent a problem in which data for residual samples increases due to a large difference between the reconstructed block and the predicted block even when moving away from the top and left of the current block, by using a reference block most similar to the reconstructed block within the current picture.

[0361] A bitstream is generated based on image information encoded according to the encoding method (S1300) described above, and the bitstream can be stored in a computer-readable storage medium.

[0362] Additionally, a bitstream is generated based on the image information encoded according to the encoding method (S1300) described above, and the bitstream can be transmitted through a transmission unit and / or a transmission medium.

[0363] FIG. 14 is a diagram exemplifying a content streaming system to which an embodiment according to the present disclosure can be applied.

[0364] As illustrated in FIG. 14, a content streaming system to which an embodiment of the present disclosure is applied may largely include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.

[0365] The encoding server compresses content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data, generates a bitstream, and transmits it to the streaming server. Alternatively, if multimedia input devices such as smartphones, cameras, and camcorders directly generate bitstreams, the encoding server may be omitted.

[0366] The above bitstream can be generated by a video encoding method and / or encoding device to which an embodiment of the present disclosure is applied, and the streaming server can temporarily store the bitstream during the process of transmitting or receiving the bitstream.

[0367] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server can act as an intermediary to inform the user of available services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, and the streaming server can transmit multimedia data to the user. At this time, the content streaming system may include a separate control server, and in this case, the control server may control commands / responses between each device within the content streaming system.

[0368] The streaming server can receive content from a media repository and / or encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.

[0369] Examples of the user devices may include mobile phones, smart phones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, HMDs), digital TVs, desktop computers, digital signage, etc.

[0370] Each server within the above content streaming system can be operated as a distributed server, in which case data received from each server can be processed in a distributed manner.

[0371] The scope of the present disclosure includes software or machine-executable instructions (e.g., operating systems, applications, firmware, programs, etc.) that cause operations according to the methods of various embodiments to be executed on a device or a computer, and a non-transitory computer-readable medium having such software or instructions stored thereon and executable on the device or computer.

[0372] Embodiments according to the present disclosure can be used to encode / decode images.

Claims

1. In a method for decoding video information, Obtaining the image information including prediction information; Derive a prediction mode for the current unit within the current picture; Based on the above prediction mode, an inter-screen reference sample array is derived from the current picture and other reference pictures; Based on the above prediction mode, an intra-screen reference sample array is derived within the current picture; Derive a prediction sample array for the current block based on the inter-screen reference sample array and the intra-screen reference sample array; A method comprising deriving a restoration sample array for the current block based on the predicted sample array.

2. In paragraph 1, The above prediction information includes first flag information and second flag information, The first flag information is obtained based on the value of the second flag information, Based on the value of the second flag information, a general CIIP (Combined inter and intra prediction) mode is derived for the current unit, A method for deriving a modified CIIP mode for the current unit based on the value of the second flag information and the value of the first flag information.

3. In paragraph 2, The above prediction information further includes third flag information, The second flag information is obtained based on the value of the third flag information, A method for determining whether a merge mode is derived for the current unit based on the value of the third flag information.

4. In paragraph 2, The above prediction information further includes fourth flag information, The second flag information is obtained based on the value of the fourth flag information, A method for deriving a blended mode for the current unit based on the value of the fourth flag information.

5. In paragraph 1, Based on the fact that the above inter-screen prediction sample array is derived based on the motion vector of the spatially adjacent block of the current block, the prediction mode is determined to be a general CIIP (combined inter and intra prediction) mode, A method wherein the prediction mode is determined to be a modified CIIP mode based on the fact that the above inter-screen prediction sample array is not derived based on a motion vector of a spatially adjacent block of the current block.

6. In paragraph 1, The prediction mode is determined to be a general CIIP (combined inter and intra prediction) mode based on the number of blocks of intra prediction mode adjacent to the current block being less than a threshold value, A method in which the prediction mode is determined to be a modified CIIP mode based on the number of blocks of intra prediction mode adjacent to the current block being greater than or equal to a threshold value.

7. In paragraph 1, A first error is derived between a sample value of a template region around the current block derived using directional or non-directional intra prediction and a restored sample value of the template region, A second error is derived between a sample value of the template region derived using intra block copy (IBC) prediction or intra template matching prediction (IntraTMP) and a restored sample value of the template region, Based on the fact that the first error is smaller than the second error, the prediction mode is determined to be the general CIIP mode, A method in which the prediction mode is determined to be a modified CIIP mode based on the first error being greater than or equal to the second error.

8. In paragraph 1, The above prediction information further includes fifth indexing information, A method in which, based on the fifth indexing information, the intra-screen reference sample array is derived based on information about a block vector included in the prediction information, or based on an area with a minimum error with a template area around the current block, or based on an area with a minimum error with the inter-screen reference sample array.

9. In paragraph 1, The above prediction information further includes sixth indexing information, A method in which, based on the sixth indexing information, the in-screen reference sample array is derived based on a motion vector candidate list including at least some of the motion vectors of spatial surrounding blocks of the current block, the motion vectors of temporal surrounding blocks, or the record-based motion vectors, and a block vector candidate list including at least some of the block vectors for the current block and the block vectors based on the template area surrounding the current block.

10. In paragraph 1, A method in which the above prediction sample array is derived based on an inter-screen prediction sample array derived using the inter-screen reference sample array and an intra-screen prediction sample array derived using the intra-screen reference sample array.

11. In paragraph 10, A method for deriving the prediction sample array by weighting the inter-screen prediction sample array and the intra-prediction sample array.

12. In a method for encoding image information, Determine the prediction mode for the current unit of the current picture; Based on the above-determined prediction mode, an inter-screen reference sample array is derived from the current picture and other reference pictures; Based on the above-determined prediction mode, an intra-screen reference sample array is derived within the current picture; Derive a prediction sample array for the current block based on the inter-screen reference sample array and the intra-screen reference sample array; Deriving a residual sample array for the current block based on the predicted sample array; A method comprising encoding the image information including prediction information based on the above-determined prediction mode.

13. In paragraph 12, The above prediction information includes first flag information and second flag information, The first flag information is obtained based on the value of the second flag information, Based on the value of the second flag information, a general CIIP (Combined inter and intra prediction) mode is derived for the current unit; A method for deriving a modified CIIP mode for the current unit based on the value of the second flag information and the value of the first flag information.

14. In Article 13 The above prediction information further includes third flag information, The second flag information is obtained based on the value of the third flag information, A method for determining whether a merge mode is derived for the current unit based on the value of the third flag information.

15. In paragraph 13, The above prediction information further includes fourth flag information, The second flag information is obtained based on the value of the fourth flag information, A method for deriving a blended mode for the current unit based on the value of the fourth flag information.

16. In paragraph 12, Based on the fact that the above inter-screen prediction sample array is derived based on the motion vector of the spatially adjacent block of the current block, the prediction mode is determined to be a general CIIP (combined inter and intra prediction) mode, A method wherein the prediction mode is determined to be a modified CIIP mode based on the fact that the above inter-screen prediction sample array is not derived based on a motion vector of a spatially adjacent block of the current block.

17. In paragraph 12, The prediction mode is determined to be a general CIIP (combined inter and intra prediction) mode based on the number of blocks of intra prediction mode adjacent to the current block being less than a threshold value, A method in which the prediction mode is determined to be a modified CIIP mode based on the number of blocks of intra prediction mode adjacent to the current block being greater than or equal to a threshold value.

18. In paragraph 12, A first error is derived between a sample value of a template region around the current block derived using directional or non-directional intra prediction and a restored sample value of the template region, A second error is derived between a sample value of the template region derived using intra block copy (IBC) prediction or intra template matching prediction (IntraTMP) and a restored sample value of the template region, Based on the fact that the first error is smaller than the second error, the prediction mode is determined to be the general CIIP mode, A method in which the prediction mode is determined to be a modified CIIP mode based on the first error being greater than or equal to the second error.

19. In paragraph 12, A method in which the above prediction sample array is derived based on an inter-screen prediction sample array derived using the inter-screen reference sample array and an intra-screen prediction sample array derived using the intra-screen reference sample array.

20. In a method regarding bitstream, Generate a bitstream; Including transmitting data including the above bitstream, Generating the above bitstream is: Determine the prediction mode for the current unit of the current picture; Based on the above-determined prediction mode, an inter-screen reference sample array is derived from the current picture and other reference pictures; Based on the above-determined prediction mode, an intra-screen reference sample array is derived within the current picture; Derive a prediction sample array for the current block based on the inter-screen reference sample array and the intra-screen reference sample array; Deriving a residual sample array for the current block based on the predicted sample array; A method comprising encoding the image information including prediction information based on the above-determined prediction mode.

Citation Information

Patent Citations

  • Method for automatically analyzing micro particles distributed over a large area using artificial intelligence, electron microscope and EDS analyzer

    KR1020230063276A

  • Server and system for providing post-mutual aid service and method thereof

    KR1020260005665A

  • Composition for treatment of cognitive dysfunction

    KR102804780B1

  • Methods and devices for intra block copy and intra template matching

    WO2024108228A1

  • Improvements to intra template matching prediction mode for motion prediction

    WO2024146574A1