Method, apparatus, and recording medium for encoding / decoding image

By configuring a merging candidate list using an intra-frame prediction method and selecting merging candidates to generate prediction blocks, the problem of low efficiency in intra-frame prediction coding in existing technologies is solved, and image quality and resolution are improved, meeting users' needs for high-quality video.

CN121264046APending Publication Date: 2026-01-02ELECTRONICS & TELECOMM RES INST
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202480038339.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-07-03
Filing Date
2024-07-03
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Existing image coding techniques are unable to effectively improve the coding efficiency of intra-frame prediction, resulting in limited improvements in image quality and resolution.

Method used

The intra-frame prediction method is adopted. By configuring a merging candidate list, merging candidates are selected to predict the target block, generating a prediction block, and the bit stream is sent and stored during the video encoding process.

Benefits of technology

It improves the efficiency of image encoding, enhances image quality and resolution, and meets users' demand for high-resolution and high-quality video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121264046A_ABST
    Figure CN121264046A_ABST
Patent Text Reader

Abstract

Disclosed is an image decoding method using intra prediction. The image decoding method includes: configuring at least one candidate list including one or more merge candidates, each merge candidate including intra coding method information, on the basis of a pre-reconstructed block reconstructed prior to a target block; selecting at least one merge candidate for the target block from the at least one candidate list; and generating a prediction block of the target block based on the selected at least one merge candidate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to a method, apparatus, and recording medium for encoding / decoding images. In one aspect, this disclosure relates to an intra-frame prediction method using a merging mode. Background Technology

[0002] With the continuous development of the information and communication industry, video services provided through broadcasting, the Internet, and other means have expanded globally.

[0003] Users demand videos with higher resolution and higher image quality. To meet these needs, image encoding / decoding technologies suitable for such videos are required. Image encoding technologies generate compressed video by compressing the video representing the image into a smaller data size. Image decoding technologies can use the compressed video to generate reconstructed images.

[0004] Various image encoding / decoding techniques exist, such as partitioning, prediction, transform, quantization, filtering, and entropy coding / decoding. By introducing, modifying, enhancing, and combining these techniques, video and images can be compressed, transmitted, and stored more efficiently. Summary of the Invention

[0005] Technical issues One aspect of this disclosure relates to an enhanced intra-frame prediction method for improving image coding efficiency.

[0006] Another aspect of this disclosure relates to encoding or decoding a target block using an intra-frame merging mode.

[0007] Technical solution One aspect of the present invention provides an image decoding method using intra-frame prediction. The image decoding method includes: configuring at least one candidate list comprising one or more merging candidates based on a reconstructed block reconstructed prior to a target block, wherein each merging candidate includes intra-frame coding method information; selecting at least one merging candidate from the at least one candidate list for the target block; and generating a predicted block of the target block by performing intra-frame prediction on the target block based on the selected at least one merging candidate.

[0008] Another aspect of the present invention provides an image coding method using intra-frame prediction. The image coding method includes: configuring at least one candidate list comprising one or more merging candidates based on pre-coded blocks encoded prior to a target block, wherein each merging candidate includes intra-frame coding method information; selecting at least one merging candidate from the at least one candidate list for the target block; and generating a predicted block of the target block based on the selected at least one merging candidate.

[0009] Another aspect of the present invention provides a method for transmitting a bitstream generated by the video encoding method described above.

[0010] Another aspect of the present invention provides a computer-readable recording medium for storing a bitstream generated by the video encoding method described above. Attached Figure Description

[0011] Figure 1 A system for video encoding and decoding according to an embodiment is shown.

[0012] Figure 2 An image partitioning structure according to an embodiment is shown.

[0013] Figure 3 The structure of intra-frame prediction according to an embodiment is shown.

[0014] Figure 4 The structure of inter-frame prediction for explaining inter-frame prediction processing according to an embodiment is shown.

[0015] Figure 5 The order in which spatial candidates are added to the candidate list is shown according to an embodiment.

[0016] Figure 6 Several loop filters are shown based on the example.

[0017] Figure 7 The structure of entropy encoding and entropy decoding is shown based on the example.

[0018] Figure 8 This is an exemplary diagram illustrating partition boundaries in a geometric partitioning pattern according to one embodiment.

[0019] Figure 9 This is an exemplary diagram illustrating partition boundaries, partition offsets, and partition angles in a geometric partitioning pattern according to one embodiment.

[0020] Figure 10 This is an exemplary diagram illustrating a weighted graph of each prediction block along a specific partition boundary in a geometric partitioning pattern according to one embodiment.

[0021] Figure 11 An example of template matching is shown.

[0022] Figure 12 The target template that can be used for the target block is shown.

[0023] Figure 13 This is an exemplary embodiment illustrating the target template for each sub-block in a template matching pattern based on sub-block units.

[0024] Figure 14This is another exemplary embodiment illustrating the target template for each sub-block in a template matching pattern based on sub-block units.

[0025] Figure 15 This is an exemplary diagram illustrating a method for configuring a target template for each sub-block of a target block to which a geometric partitioning pattern has been applied.

[0026] Figure 16 This is another exemplary diagram illustrating a method for configuring a target template for each sub-block of a target block to which a geometric partitioning pattern has been applied.

[0027] Figure 17 This is yet another exemplary diagram illustrating a method for configuring a target template for each sub-block of a target block to which a geometric partitioning pattern has been applied.

[0028] Figure 18 and Figure 19 Various embodiments of subsampling methods for template matching are shown.

[0029] Figures 20 to 25 Each illustrates a search method for template matching according to one embodiment.

[0030] Figure 26 This illustrates a method for configuring a target template in affine mode.

[0031] Figure 27a and Figure 27b A method for configuring a reference template in affine mode is shown according to one embodiment.

[0032] Figure 28 This illustrates a bilateral match according to one embodiment.

[0033] Figure 29 The location of a reconstructed block, according to one embodiment, is shown as being considered for deriving predictive information about a target block.

[0034] Figure 30 A method for configuring a target template in a method for exporting a template-based intra-frame mode is shown according to one embodiment.

[0035] Figure 31 The locations of adjacent or non-adjacent blocks for deriving spatial merging candidates are shown according to some embodiments.

[0036] Figure 32 This shows the chromaticity component blocks when the luminance component blocks and the corresponding chromaticity component blocks in the target CTU have independent block partitioning structures.

[0037] Figure 33 This is a flowchart illustrating an image coding method using an intra-frame merging mode according to one embodiment.

[0038] Figure 34 This is a flowchart illustrating an image decoding method using an intra-frame merging mode according to one embodiment. Detailed Implementation

[0039] Various modifications can be applied to this disclosure. Furthermore, this disclosure can have various embodiments. Specific embodiments will be described with reference to the accompanying drawings and detailed description.

[0040] The specific embodiments are not intended to limit this disclosure to a particular mode of practice, and it should be understood that all changes, equivalents, and alternatives are included as embodiments in this disclosure without departing from the spirit or technical scope of this disclosure.

[0041] These embodiments are described to enable those skilled in the art to readily practice them. It should be noted that the various embodiments differ from one another but are not necessarily mutually exclusive. For example, shapes, structures, and characteristics associated with the embodiments may be applied or implemented in other embodiments without departing from the spirit and scope of this disclosure. Furthermore, it should be understood that the position or arrangement of various components in the embodiments may be changed without departing from the spirit and scope of this disclosure. Therefore, the appended detailed description is not intended to limit the scope of this disclosure, and the scope of the exemplary embodiments is limited only by the appended claims and their equivalents, provided they are properly described.

[0042] A detailed description of the embodiments, which will be described later, can be referenced in the accompanying drawings. The descriptions made in or shown in the drawings are to be considered part of the detailed description of this disclosure. In the drawings, similar reference numerals may be used to designate functions that are identical or similar in various respects. Dependencies between components are not limited to those shown in the drawings.

[0043] In embodiments, singular expressions may include, and may be limited to, plural expressions unless specifically indicated in the context. In other words, in embodiments, expressions such as “at least one” and “one or more” can be replaced by the term “multiple”. Terms such as “ / ”, “and / or”, “at least one of…” and “one or more of…” describing multiple items can refer to 1) one of the multiple items, 2) some of the multiple items, 3) a combination of some of the multiple items, or 4) a combination of the multiple items. Furthermore, plural expressions can be replaced by singular expressions. Here, a plural number can represent an integer of 1, 2, 3, 4, or 5 or greater.

[0044] In embodiments, terms associated with numbers such as "first" and "second" may be used to describe various components. These terms are used only to distinguish one component from another and are not intended to limit the components. For example, a first component may be referred to as a second component without departing from the scope of this disclosure. Similarly, a second component may also be referred to as a first component.

[0045] The fact that the first component sends (or provides) information to the second component can mean that the first component sends the information directly to the second component, or it can mean that the first component sends the information to the second component through a third component. Here, the information received (or acquired) by the second component can be information sent by the first component, or information generated by applying specific processing to information sent by the first component.

[0046] The components in the embodiments may be shown independently to indicate different functional characteristics, and the illustration does not imply that each component corresponds to a separate hardware or software element. That is, for ease of description, the components in the embodiments may be divided and enumerated. Two or more components described in an embodiment may be considered as one component. Furthermore, a component described in an embodiment may be divided into multiple components, which are divided and perform the function of the component individually. Embodiments in which these components are integrated and embodiments in which these components are separate may be included within the scope of this disclosure without departing from its spirit.

[0047] The terminology used in the embodiments is for describing specific embodiments only and is not intended to limit this disclosure. In the embodiments, terms such as "comprising" or "having" are intended to indicate the presence of features, numbers, steps, operations, components, portions, or combinations thereof described in the embodiments. These terms do not exclude the possibility of the presence or addition of other features, numbers, steps, operations, components, portions, or combinations thereof not explicitly described in the embodiments. That is, the description of a component as "comprising" a specific component in the embodiments does not exclude additional components beyond the specific component, and means that additional components may be included within the scope of the embodiments of this disclosure or the technical spirit of this disclosure.

[0048] Some components in the embodiments may be optional components in addition to the basic components of this disclosure used to perform basic functions. Such optional components can be used to improve performance. Embodiments may be implemented as structures that include only the basic components necessary to implement the essential points of the embodiments, excluding optional components. Such structures are included within the scope of the embodiments.

[0049] Embodiments of this disclosure will now be described in detail with reference to the accompanying drawings to enable those skilled in the art to readily practice this disclosure. In the description of this disclosure, detailed descriptions of known functions and configurations that are considered to obscure the main points of the disclosure will be omitted. Throughout the following description of this disclosure, the same reference numerals will be used to denote the same or similar components, and repeated descriptions of the same components will be omitted.

[0050] Substitution of terms in the embodiments In the following description, terms listed in a single line may be used in the embodiments to have the same meaning and may be used interchangeably with each other in the embodiments.

[0051] - "one or more" and "at least one" - "two or more", "multiple", "more than one" and "more" (in embodiments, "one or more" or "at least one" may be further limited to "two or more", "multiple" or "more") - "Information" and "Signal" - "Value", "Predefined Value", "Specific Value", "Threshold", "Baseline Value", and "Reference Value" - "Statistical Values" and "Statistical Values" - "Indicator", "Index", "Flag", and "Information" - "Encoder" and "Encoding Device" - "Decoder" and "Decoding Device" - "Entropy coding", "encoding" and "encoding" - "Entropy decoding", "decoding" and "decoding" - "Encoding / Decoding" and "Encoding and / or Decoding" - "video", "moving image", "image", "picture", "frame", and "screen" - "Reference Screenshot" and "Reference Image" - "Reference Screen List (RPL)" and "Reference Image List" - "Original", "Input", and "Source" - "block", "unit", and "signal" - "Square" and "Square Shape" - "pixel", "sample", and "pel" - "Region", "District", "Part", and "Fragment" - "Partition", "Segment", and "Divide" - "quaternary" and "quaternary" - "luma component", "luma", "luminance component", "luminance", and "Y" - "chroma component", "chroma", "chromaticity", "chromaticity component", "Cb and Cr", "Cb or Cr", "Cb", "Cr", "U and V", "U or V", "U or "V" - "target", "current" (e.g., target block and current block, or target image and current image) - "Neighbor", "Nearby", "Adjacent", "Neighbor / Nearby" (e.g., neighbor block, adjacent block, and neighbor / nearby block) - "Paired" and "COL" - "Reconstruction", "Reconstruction", and "Decoding" - "Reconstructed", "Reconstructed", and "Decoded" - "difference", "error", "residual" - Maximum coding unit (LCU) and coding tree unit (CTU) - "Inter-frame" and "inter" - "Inter-frame prediction", "inter-prediction", and "motion compensation" - "Inter-frame mode", "Inter-frame prediction mode", "Inter-mode", and "Inter-prediction mode" - "Motion Vectors", "Predicted Motion Vectors", and "Advanced Motion Vector Prediction (AMVP)" - "List" and "Candidate List" - "Spatial Candidates" and "Spatial Merge Candidates" - "Time Candidates" and "Time Merged Candidates" - "Predicting Motion Vector Candidates" and "Motion Vector Predictors" - "Forecasting Methods" and "Forecasting Models" - "intra" and "withinframe" - "Intra-prediction" and "Intra-frame prediction" - "Intra-frame mode" and "Intra-frame prediction mode" - "Inverse quantization" and "scaling" - "Quantization Matrix" and "Scaling List" - "Quantization matrix coefficients" and "Matrix coefficients" - "Transformation Coefficient Level", "Quantization Level", "Quantization Coefficient", "Quantization Transformation Coefficient", and "Quantization Transformation Coefficient Level" - "Inverse quantization coefficients" and "Inverse quantization transform coefficients" - "Scan Type" and "Scan Direction" - "Directional Mode", "Angle Mode", "Angled Mode", and "Intra-Frame Prediction Mode" - "Intra-prediction mode (mode number)," "Intra-prediction mode (mode index)," "Intra-prediction mode (mode value)," "Intra-prediction mode (mode angle)," "Intra-prediction mode (mode direction)," "Intra-prediction direction (mode number)," "Intra-prediction direction (mode index)," "Intra-prediction direction (mode value)," and "Intra-prediction direction (mode angle)." - "Merge Mode" and "Motion Merge Mode" - "Geometric Partitioning (GPM)" and "Triangle Partitioning" In addition to the terms exemplified above, terms that have the same meaning based on common knowledge in the art may be used interchangeably with each other in the embodiments.

[0052] The information and range of information values ​​described in the embodiments In embodiments, information may include constants, flags, indexes, variables, encoding parameters, elements, syntax elements, motion information, attributes, entities, objects, data, etc. In other words, the term "information" may be used interchangeably with "data," "flags," "indexes," "variables," "elements," "syntax elements," "motion information," "attributes," or "entities."

[0053] This information can have one of multiple values. The term "nth value" can refer to the nth value among multiple values.

[0054] For example, the first value can represent "0" or logical false. The second value can represent "1" or (logical) true. Alternatively, the first value can represent "1" or (logical) true. The second value can represent "0" or (logical) false.

[0055] A flag can be information with either a value of "0" or "1". In embodiments, the values ​​of flags, i.e., "0" and "1", can be replaced with "1" and "0", respectively. For example, information indicating whether to perform a specific process or whether to apply a specific process can be considered a flag.

[0056] When variables such as i or j are used to indicate rows, columns, or indices, the variables can be integers equal to or greater than 0 and less than or equal to n-1. Optionally, the variables can be integers equal to or greater than 1 and less than or equal to n. Here, n can be the number of rows, the number of columns, or the number of entities indicated by the index.

[0057] Concepts related to encoding and decoding The concepts related to encoding and decoding will be described below. These descriptions can be applied to embodiments.

[0058] Predefined value: A predefined value can refer to a value shared by both the encoding and decoding devices. For example, a predefined value can be restricted to and interpreted as a fixed value. Optionally, a predefined value can be a value shared between the encoding and decoding devices via signaling. Optionally, a predefined value can be a value derived by both the encoding and decoding devices through the same process, such that the encoding and decoding devices have a common value. Optionally, a predefined value can be a common value possessed by both the encoding and decoding devices. The description of predefined values ​​can also be applied to predefined information. In the above description, "information" can be used instead of "value".

[0059] Values ​​derived through the same process in both the encoding and decoding devices can include values ​​derived through the same process for the same value and / or the same information in both the encoding and decoding devices.

[0060] Values ​​derived through the same process in both the encoding and decoding devices can include values ​​derived using the same conditional statements for the same value and / or the same information in both the encoding and decoding devices.

[0061] - Descriptions of predefined values ​​can also be applied to predefined information. In the above description, the term "information" can be used instead of the term "value".

[0062] Availability: The fact that a particular pattern is available for a particular objective means that the pattern selected from the particular patterns is used for the particular objective. Other patterns belonging to the category of a particular pattern may be unavailable patterns. Unavailable patterns may not be used for a particular objective. The description of a particular pattern may also apply to other specific information. In the above description, "information" can be used instead of "pattern".

[0063] Adjacent: The terms "direction" and "second entity" for "first entity" can refer to the "direction" of the first entity and the "second entity" adjacent to a corner / surface of the first entity. For example, the term "top-left block" for "target block" can be the block adjacent to the top-left corner of the target block. Here, "first entity" can be a target cell, target block, or target sample. The term "direction" can refer to one of the directions corresponding to top-left, top, top-right, top-right, left, right, bottom-left, bottom, and bottom-right. The term "second entity" can be a cell, block, or sample. Here, regarding the directions corresponding to top-left, top-right, bottom-left, and bottom-right, the corners of the first entity and the corners of the second entity can be diagonally adjacent to each other. Regarding the directions corresponding to top, left, right, and bottom, a surface of the first entity and a surface of the second entity can be in contact with each other.

[0064] - For example, the block adjacent to the top left of the target block can be the top adjacent block of the block adjacent to the left of the target block. - The block adjacent to the top right of the target block can be the right adjacent block of the block adjacent to the top of the target block. - The block adjacent to the bottom left of the target block can be the bottom adjacent block of the block adjacent to the left of the target block.

[0065] Encoding and decoding: Encoding and decoding can refer to the encoding and / or decoding of images.

[0066] Signals: Signals can refer to information about an image, unit, or block. A specific signal can refer to a specific image, unit, or block.

[0067] Image: An image can refer to one of the frames that make up a video, or it can refer to the video itself. For example, "image encoding and / or decoding" can refer to "video encoding and / or decoding," and it can also refer to "encoding and / or decoding of one of the images that make up a video."

[0068] - An image can refer to the entirety of a picture, or it can refer to a part of a picture, such as a block.

[0069] Target image: The target image can be an encoded target image that serves as the encoding target and / or a decoded target image that serves as the decoding target. Furthermore, the target image can be an input image processed by an encoding device or a reconstructed image processed by a decoding device. The target image can be an image that includes target blocks.

[0070] Sub-screen: A screen can be divided into one or more sub-screens.

[0071] - A sub-screen can be a square or rectangular area within the screen. Each sub-screen can have one or more CTUs.

[0072] A subscreen may include one or more stripes and / or one or more parallel blocks. For example, a subscreen may include one or more stripe rows and one or more stripe columns. Optionally, each subscreen may include one or more parallel block rows and one or more parallel block columns.

[0073] - A sub-picture may include one or more strips that collectively cover a rectangular area within the picture. Therefore, the boundary of each sub-picture can always be the boundary of a strip. Furthermore, the boundary of each vertical sub-picture can always be the boundary of a vertical parallel block.

[0074] Strip: A strip may include one or more parallel blocks in the frame. A strip may include one or more rows of parallel blocks and one or more columns of parallel blocks.

[0075] Parallel Blocks: Parallel blocks can be square or rectangular areas on the screen. A parallel block can have one or more CTUs. The screen can be divided into one or more parallel block rows and one or more parallel block columns.

[0076] CTU: An image can be partitioned into multiple coding tree units (CTUs).

[0077] - A CTU may include a Y-coded tree block (CTB) and at least one of a Cb CTB and a Cr CTB associated with the Y CTB, and may include information about each CTB. This information may include syntax elements.

[0078] - One or more partition types can be used for each CTU partition to form sub-units such as coding units (CUs), prediction units (PUs), and transform units (TUs). One or more partition types can include quadtree (QT) partitions, binary tree (BT) partitions, and ternary tree (TT) partitions. Furthermore, multi-type tree (MTT) partitions can be used for each CTU partition, employing a combination of multiple partition types.

[0079] CTB: CTB can refer to one of Y CTB, Cb CTB, and Cr CTB.

[0080] Unit: A unit can be defined for a specific process in encoding and decoding. A unit can be information about a specific region in an image. An image can be recursively divided into multiple parts to perform specific encoding and decoding processes. A unit can refer to the region to which a specific process is applied and information about that region.

[0081] - Unit type can refer to the specific processing applied to a unit. A specific processing can be applied to a unit based on its unit type. A "specific" unit can be a unit that is named as a "specific" processing in the encoding / decoding process. For example, a unit can be at least one of the following: original unit, CTU, coding unit, prediction unit, residual unit, reconstructed residual unit, transform unit, and reconstruction unit.

[0082] - A cell may include samples having a two-dimensional (2D) form or arrangement. In this respect, the term "cell" may refer to a "block". For example, a block may be at least one of a raw block, a CTB, a coded block (CB), a prediction block (PB), a residual block, a reconstructed residual block, a transform block (TB), and a reconstructed block. For example, a partition of a cell may refer to a partition of the block corresponding to the cell.

[0083] - A unit can include syntax elements. In other words, a block and its syntax elements can be collectively referred to as a "unit".

[0084] - A block can be an M×N sample array. Here, M and N can refer to positive integer values, and a block can typically refer to a sample array in 2D form. The current block can refer to an encoded target block as the target to be encoded in the encoding process and a decoded target block as the target to be decoded. Furthermore, the current block can be at least one of an encoded block, a prediction block, a residual block, a transform block, and a reconstruction block. Blocks can have various sizes and shapes. For example, block shapes can include one or more of quadrilaterals, rectangles, squares, rectangles with a horizontal length different from their vertical length (i.e., elongated rectangles), trapezoids, triangles, right triangles, and pentagons. Here, the horizontal and vertical lengths of a rectangle can be different from each other. Furthermore, block shapes can include other geometric shapes that can be represented in two dimensions. For example, a block shape can be a rectangle or a pentagon defined by excluding right triangle regions from a rectangular region. Here, the right-angled vertex of the right triangle can be one of the vertices of the rectangle. Furthermore, a block shape can be a combination of two or more of the shapes described above. Additionally, a block shape can be the remaining shape obtained by excluding one shape from the aforementioned shapes.

[0085] - In embodiments, the rectangle may be limited to non-square rectangles. In embodiments, when the shape of a particular target is described as a rectangular shape, the description may further imply that the horizontal length and vertical length of the particular target are different from each other.

[0086] - In an embodiment, the block may be limited to at least one of a vertically oriented block and a horizontally oriented block. A vertically oriented block may be a block with a vertical length greater than its horizontal length. A horizontally oriented block may be a block with a horizontal length greater than its vertical length.

[0087] - A unit may include a luminance component block (i.e., a Y block) and two chrominance component blocks (i.e., at least one of a Cb block, a Cr block, or a combination thereof), and may include information about each block. The information may contain syntax elements.

[0088] Information about a cell may include the cell type, cell size, cell depth, cell encoding order, cell decoding order, etc.

[0089] Target unit: A target unit can be an encoding target unit that is to be encoded in encoding and a decoding target unit that is to be decoded. A target unit can be a specific region in a target image to which one or more specific encoding / decoding processes will be applied. A specific type of unit can be generated by applying specific processes to the target unit. Optionally, a target unit can refer to a unit having a specific type for a specific encoding / decoding process.

[0090] Depth: A block can be hierarchically partitioned into multiple sub-blocks while having a depth that depends on the tree structure. The multiple sub-blocks generated from the partitioning operation of a block can be called "partitions".

[0091] - When the blocks that make up an image are represented in a tree structure, the block depth can represent the level of the node corresponding to each block. Optionally, the block depth can indicate the number of partitions applied until the block is determined. As the block is further partitioned, the block depth can increase by 1.

[0092] In a tree structure, the root node can be considered to have the lowest level, and leaf nodes have the highest level. The root node can be the topmost node of the tree structure and can correspond to the unpartitioned initial block. The level of the root node can be 0 or 1. When the root node's level is 0, a node with level 1 can refer to the block determined as the initial block is partitioned once. A node with level n can refer to the block determined when the initial block is partitioned n times. Leaf nodes can be the bottommost nodes of the tree structure. Leaf nodes can be nodes that cannot be further partitioned. The depth of a leaf node can be a predefined maximum depth. For example, the maximum depth can be a positive integer such as 3. The root node can refer to CTU. Leaf nodes can refer to at least one of CU, PU, ​​and TU.

[0093] - Depth can have a type that depends on the partition type. QT depth can represent the depth in a quadtree partition. BT depth can represent the depth in a binary tree partition. TT depth can represent the depth in a ternary tree partition.

[0094] Sample: A sample can be the basic unit that makes up a block. A sample can consist of one or more bits. Bit depth refers to the number of bits that make up a sample. Samples can range from 0 to 2. Bd The value of -1 indicates that this depends on the bit depth.

[0095] PU: PU can represent the basic unit used for prediction-related processing. For example, prediction-related processing may include inter-frame prediction, intra-frame prediction, intra-block copy (IBC) prediction, intra-frame compensation, and motion compensation.

[0096] A prediction unit (PU) can be partitioned into multiple sub-PUs, each with a size smaller than the PU. Each of the multiple sub-PUs can also be a basic unit for prediction-related processing. In other words, prediction unit partitions generated from the partitioning operation of prediction units can also be prediction units.

[0097] TU: A TU can be a basic unit for processing associated with a residual block. Processing associated with a residual block can include at least one of transform, inverse transform, quantization, dequantization, transform coefficient encoding, transform coefficient decoding, entropy encoding, or entropy decoding, or combinations thereof. A TU can be partitioned into multiple sub-transform units, each sub-transform unit having a size smaller than the TU. Each of the multiple sub-TUs can also be a basic unit for processing associated with a residual block. In other words, a transform unit partition generated from a partitioning operation of a transform unit can also be a transform unit.

[0098] - The transformation may include one or more of the primary transformation and the secondary transformation, and the inverse transformation may include one or more of the primary inverse transformation and the secondary inverse transformation.

[0099] Parameter set: The parameter set can correspond to the header information in the structure of the bit stream.

[0100] - The parameter set may include at least one of the following: video parameter set (VPS), sequence parameter set (SPS), picture parameter set (PPS), adaptive parameter set (APS), or decoding parameter set (DPS), or a combination thereof.

[0101] Information transmitted via signals using a parameter set can be applied to a screen referencing that parameter set. For example, information in a VPS can be applied to a screen referencing a VPS. Information in an SPS can be applied to a screen referencing an SPS. Information in a PPS can be applied to a screen referencing a PPS. A parameter set can reference a higher-level parameter set. For example, a PPS can reference an SPS. An SPS can reference a VPS.

[0102] - In addition, the parameter set may include parallel block group information, stripe header information, and parallel block header information. A parallel block group can refer to a group or stripe that includes multiple parallel blocks.

[0103] Most Probable Mode (MPM): MPM can refer to an intra-prediction mode that has a high probability of intra-prediction for the target block.

[0104] - One or more different MPMs can be determined based on the encoding parameters associated with the target block and the attributes of the entities associated with the target block.

[0105] - One or more MPMs can be determined based on the intra-prediction modes of a reference block. A reference block can include multiple reference blocks. One or more different MPMs can be determined based on which intra-prediction modes have been used for one or more reference blocks. Reference blocks can include spatially neighboring blocks.

[0106] MPM List: An MPM list can be a list that includes one or more MPMs. The number of MPMs in the MPM list can be predefined.

[0107] MPM Index: The MPM index indicates the MPM used for intra-frame prediction of the target block among one or more MPMs in the MPM list.

[0108] MPM Usage Indicator: The MPM usage indicator indicates whether the MPM list has been used for the prediction of the target block.

[0109] Prediction mode: The prediction mode can be information indicating the prediction method used for the target block, such as a mode to be used for intra-frame prediction or a mode to be used for inter-frame prediction. The prediction mode can indicate one of the prediction-related modes, as will be described in the embodiments. Furthermore, the prediction mode can include at least one of intra-frame mode, inter-frame mode, intra-frame block copy mode, or a combination thereof.

[0110] Reference image list: The reference image list can be a list that includes one or more reference images that will be used for prediction of the target block.

[0111] - The list of reference images can include multiple lists. Multiple lists of reference images can include list 0 (List0; L0), list 1 (List1; L1), etc.

[0112] - For inter-frame prediction of a target block, one or more lists of reference images can be used. In the names of information related to inter-frame prediction, parts such as "L0" and "L1" may refer to the list of reference images associated with that information.

[0113] Reference image (reference frame): The reference image (frame) can be an image referenced for predicting the target block. Optionally, the reference frame can be an image that includes the reference block. The reference frame can include images before the target image, the target image, and images after the target image.

[0114] Reference Image Index: A reference image index can be an index that indicates a reference image (picture) among one or more reference images in a list of reference images used for predicting the target block.

[0115] Reference block: A reference block can be a block that is referenced for encoding / decoding the target block (such as prediction and filtering). For example, a reference block can include reference samples that are referenced to derive prediction samples, and can refer to a block that provides information for decoding the target block.

[0116] Reference sample: A reference sample can be a sample used for encoding / decoding a target block (such as prediction and filtering).

[0117] Inter-frame prediction indicator: The inter-frame prediction indicator can indicate the direction of inter-frame prediction for a target block. Inter-frame prediction can be one of unidirectional and bidirectional prediction. Optionally, the inter-frame prediction indicator can represent the number of reference images used to generate prediction blocks for the target block. Optionally, the inter-frame prediction indicator can represent the number of prediction blocks used for inter-frame prediction of the target block. The reference direction can refer to the inter-frame prediction indicator. For example, the inter-frame prediction indicator can represent one of a unidirectional indicator and a bidirectional indicator. Optionally, in an inter-frame mode that uses only reference images in reference image list L0, the inter-frame prediction indicator can have a "0" as a first value; in an inter-frame mode that uses only reference images in reference image list L1, the inter-frame prediction indicator can have a "1" as a second value; and in an inter-frame mode that uses at least two of the reference images in reference image list L0 and reference images in reference image list L1, the inter-frame prediction indicator can have a "2" as a third value.

[0118] Prediction list utilization flag: The prediction list utilization flag for a specific reference image list indicates whether at least one reference image in the specific reference image list is used to generate a prediction block for the target block. For example, a prediction list utilization flag of "0" indicates that reference images from the specific reference image list are not used to generate the prediction block. A prediction list utilization flag of "1" indicates that reference images from the specific reference image list are used to generate the prediction block.

[0119] - Inter-frame prediction indicators can be derived using prediction list utilization flags. Conversely, prediction list utilization flags can be derived using inter-frame prediction indicators. For example, inter-frame prediction indicators can be derived using prediction list utilization flags from multiple reference image lists. When an inter-frame prediction indicator indicates that a specific reference list among the multiple reference image lists is used, the prediction list utilization flag for the specific reference list indicated by the inter-frame prediction indicator can be set to "1", while the prediction list utilization flags for the remaining reference image lists not indicated by the inter-frame prediction indicator can be set to "0".

[0120] Reference direction: The reference direction can indicate a list of reference images used for prediction of the target block. For example, the reference direction can indicate one or more of reference image list L0 and reference image list L1.

[0121] - The reference direction indicates only the list of reference images used for prediction of the target block, and may not indicate that the orientation of the reference images in the list is limited to either the forward or backward direction. In other words, each of the reference image lists L0 and L1 may include both forward and backward images. Here, forward can indicate the direction from the target image to an image preceding the target image. Forward inter-frame prediction can be an inter-frame prediction using an image preceding the target image as a reference image. The backward direction can indicate the direction from the target image to an image following the target image. Backward inter-frame prediction can be an inter-frame prediction using an image following the target image as a reference image.

[0122] - A unidirectional reference direction can mean using one list of reference images. A bidirectional reference direction can mean using two lists of reference images. For example, a reference direction can indicate one of the following: using only reference image list L0, using only reference image list L1, or using both reference image lists. Furthermore, the reference direction can be indicated by an inter-frame prediction indicator.

[0123] Picture Order Count (POC): The POC of an image (picture) can indicate the display order or output order of the images (pictures).

[0124] Motion information: Motion information can be information used to specify a reference block. Motion information may include information used for inter-frame prediction, such as motion vectors (MV), reference image indexes, reference images, inter-frame prediction indicators, and prediction list utilization flags. Furthermore, motion information may include information used in a specific inter-frame prediction mode, such as MV candidates, MV candidate indices, merge candidates, and merge indices. Additionally, motion information may include information related to block vectors as described below. Information related to block vectors may refer to information including at least one of block vectors, block vector candidates, and block vector candidate indices.

[0125] - For inter-frame prediction of a target block, multiple motion information related to multiple lists of reference images can be used separately. Motion information of a specific list of reference images can be used for prediction using that specific list of reference images. Multiple (intermediate) prediction blocks can be derived separately from multiple motion information. The (final) prediction block of the target block can be generated using the statistical values ​​of the multiple (intermediate) prediction blocks.

[0126] MV: Motion vector (MV) can be a two-dimensional (2D) vector used for inter-frame prediction. MV can refer to the offset between the target block and the reference block. Alternatively, MV can indicate the difference between the position of the target block and the position of the reference block.

[0127] - For example, it can be in the form of (mv x , mv y MV is represented in the form of ) xIt can indicate the horizontal component, and mv y It can indicate the vertical component.

[0128] - The zero vector can be (0,0)MV.

[0129] Block Vector (BV): BV can be a two-dimensional (2D) vector used for intra-block replication prediction. BV can refer to the offset between a target block and a reference block in the target image. In other words, BV can indicate the displacement between a target block and a reference block in the target image.

[0130] - For example, similar to MV, it can be in (bv x , bv y BV is represented in the form of ) . bv x It can indicate the horizontal component, and bv y It can indicate the vertical component.

[0131] - The zero vector can be (0,0)MV.

[0132] Motion information candidates: In a specific prediction, motion information of a target block can be selected from motion information candidates determined by a specific scheme. Motion information candidates can refer to the motion information of a reference block, and can refer to the reference block itself containing motion information. Here, the reference block can be a block determined by a specific scheme to select motion information candidates.

[0133] Candidate List: The candidate list can be a list including one or more candidates. For example, the candidate list may include a motion information candidate list, a merge candidate list, an MV candidate list, an MPM list, etc. The candidate list can be generated by both the encoding and decoding devices in the same way. In other words, the candidate list used by the encoding device and the candidate list used by the decoding device can be the same as each other, and the same candidate list can be shared between the encoding and decoding devices. The encoding device can select candidates from the candidates in the candidate list to be used for processing the target block. An indicator indicating the selected candidate can be sent from the encoding device to the decoding device via a signal. The decoding device can specify the candidate to be used for processing the target block from the candidates in the candidate list by utilizing the indicator. Optionally, the encoding and decoding devices can specify the candidate to be used for processing the target block from the candidates in the candidate list based on the same rules.

[0134] Motion information candidate list: The motion information candidate list can refer to a list constructed using one or more motion information candidates.

[0135] Motion Information Candidate Index: A motion information candidate index can be an identifier or indicator that indicates a motion information candidate in the motion information candidate list for prediction of the target block.

[0136] - In a specific inter-frame prediction mode, motion information of the target block can be derived using motion information from additional reconstructed blocks. The additional block can be a neighboring block. In this mode, the motion information of the target block itself may not be transmitted separately; instead, additional information for deriving the motion information of the target block based on the motion information of the additional reconstructed blocks may be transmitted. This additional information may include information indicating blocks within the additional reconstructed blocks (such as motion information candidate indices) whose motion information is used to derive the motion information of the target block.

[0137] For example, inter-frame prediction modes can include AMVP mode, merge mode, skip mode, etc. Motion information candidate indices can be merge indexes or MV candidate indices.

[0138] - In an embodiment, MV can be a part of motion information. In an embodiment, information related to motion information (such as motion information candidates, motion information candidate lists, and motion information candidate indexes) can be replaced with information related to MV (such as MV candidates, MV candidate lists, and MV candidate indexes), and the description of motion information can also be applied to MV.

[0139] Merging: The term "merging" can refer to merging multiple motion information from multiple blocks, or it can refer to applying motion information together with the motion information of an additional block to a target block. In other words, a merging pattern can refer to a pattern for deriving the motion information of a target block from the motion information of neighboring blocks.

[0140] Merge candidate: A merge candidate may refer to a specific (reconstructed) block used to merge the target block, and may refer to the motion information of that specific block. Optionally, a merge candidate may include the motion information of that specific block.

[0141] - Merge candidates for the target block can include spatial merge candidates, temporal merge candidates, history-based candidates, average candidates based on the average of two merge candidates, zero merge candidates, etc.

[0142] Merge candidate list: The merge candidate list can be a list constructed using one or more merge candidates.

[0143] Merge Index: A merge index can be an indicator that identifies a merge candidate in a merge candidate list for predicting the target block. The motion information of the merge candidate indicated by the merge index in the merge candidate list can be used as the motion information of the target block.

[0144] Neighboring blocks: Neighboring blocks can be blocks adjacent to the target block. Neighboring blocks can include spatially adjacent blocks and temporally adjacent blocks. A neighboring block can refer to a reconstructed neighboring block in the reference image. A neighboring block does not necessarily need to be adjacent to the target block.

[0145] Spatial neighbor block: A spatial neighbor block can be a block that is spatially adjacent to the target block.

[0146] - Target blocks and spatially neighboring blocks can be included in the target image.

[0147] - A spatially neighboring block may include a block whose boundary at least a portion is in contact with at least a portion of the boundary of the target block. Optionally, a spatially neighboring block may include a block whose distance from the target block is less than or equal to a specific value.

[0148] - Spatial neighbor blocks can include blocks that are diagonally adjacent to the vertex of the target block.

[0149] - Spatial neighbor blocks can include the top-left block adjacent to the top-left of the target block, the top block adjacent to the top of the target block, the top-right block adjacent to the top-right of the target block, the left block adjacent to the left side of the target block, the right block adjacent to the right side of the target block, the bottom-left block adjacent to the bottom-left of the target block, the bottom block adjacent to the bottom of the target block, and the bottom-right block adjacent to the bottom-right of the target block.

[0150] Temporally adjacent blocks: Temporally adjacent blocks can be blocks that are temporally adjacent to the target block.

[0151] - Temporally adjacent blocks can include co-occurring blocks (COL blocks). A COL block can be a block in a reconstructed image stored in a reference image (picture) buffer. A co-occurring image (co-occurring picture: COL picture) can refer to a picture that includes a COL block. A COL picture can be an image included in a list of reference images.

[0152] - COL blocks can be determined based on the position of target blocks in the target image. The fact that two blocks are "temporally adjacent" can imply that the positions of the two blocks meet specific conditions.

[0153] - The position of the COL block in the COL frame can be the same as the position of the target block in the target image. Optionally, the position of the COL block in the COL frame can correspond to the position of the target block in the target image. Here, the positions of corresponding blocks can mean that the areas of the blocks are the same, that the area of ​​a block is included in the area of ​​the additional block, or that a block occupies a specific position in the additional block.

[0154] For example, the position of a COL block in a COL image can be the same as the position of a target block in the target image. Alternatively, a COL block can be a block in a COL image that includes COL samples. A COL sample can be a sample point with the same coordinates as a specific sample point in the target block.

[0155] - A temporally neighboring block can be a block that is spatially adjacent to the target block in time.

[0156] Neighboring samples: Neighboring samples can refer to samples within neighboring blocks. Neighboring samples can include prediction samples, reconstruction samples, residual samples, and decoding samples.

[0157] Search Range: The search range can refer to the 2D range of the MV search performed during inter-frame prediction. For example, when the best MV is derived to process the target block, the best MV is selected from the MVs within the indicated search range.

[0158] Transform coefficients: Transform coefficients can be coefficients generated by performing a transformation on the residual block. Alternatively, transform coefficients can be coefficient values ​​generated by performing dequantization on the quantization level.

[0159] Quantization level: The quantization level can be an integer number used as the input for dequantization.

[0160] Quantization: Quantization can be a process that generates quantization levels for transform coefficients. Quantization levels can be generated by applying quantization to transform coefficients. Transformation can also be considered as part of quantization.

[0161] Dequantization: Dequantization can be a process of multiplying a factor by a quantization level. (Reconstructed) transform coefficients can be generated by applying dequantization to the quantization level.

[0162] Quantization parameter (QP): QP can refer to a factor used when generating the quantization levels of the transform coefficients. Additionally, QP can refer to a factor used when generating the (reconstructed) transform coefficients for the quantization levels used in inverse quantization. Optionally, QP can be a value mapped to the quantization step size.

[0163] Incremental QP: Incremental QP can be the difference between the QP predicted through specific processing and the QP of the target block. In other words, the QP of the target block can be the sum of the predicted QP and the incremental QP.

[0164] Quantization matrix: A quantization matrix is ​​a matrix used in quantization or dequantization to improve the subjective or objective quality of an image.

[0165] Quantization matrix coefficients: Quantization matrix coefficients can be each element in the quantization matrix.

[0166] Scan: A scan can refer to a method of arranging values ​​in a block or matrix. These values ​​can be coefficients. For example, a scan can refer to the arrangement of values ​​in 2D or 1D form, and the rearrangement of values ​​in 1D or 2D form. A reverse scan can refer to the arrangement (or rearrangement) that is the opposite of the arrangement performed in a scan.

[0167] Non-zero transform coefficients: Non-zero transform coefficients can refer to transform coefficients with non-zero values ​​or quantization levels with non-zero values.

[0168] Bitstream: A bitstream can refer to a stream / sequence of bits that includes encoded information generated by encoding an image. A bitstream may include information dependent on specific syntax elements. For example, the information may include syntax elements. An encoding device can generate a bitstream that includes multiple pieces of information dependent on specific syntax elements. A decoding device can extract multiple pieces of information from the bitstream based on specific syntax elements.

[0169] Signaling: Signaling information can refer to sending information from an encoding device to a decoding device via a bitstream. For example, the information may include syntax elements. Optionally, signaling can refer to the process by which the encoding device includes information in a bitstream. Information signaled by the encoding device can be used by the decoding device. In signaling, the bitstream can be transmitted over a network and can be included in a storage / recording medium. In embodiments, a description indicating that information is signaled may include: 1) the encoding device determines and generates information to signal the information; 2) the encoding device generates encoded information by encoding the information; 3) the (encoded) information is sent from the encoding device to the decoding device via a bitstream; 4) the decoding device obtains the information by decoding the encoded information; and 5) the decoding device determines and generates information by signaling the information.

[0170] Encoding devices generate encoded information by encoding data. Encoded information can be transmitted as a signal via a bitstream. Decoding devices retrieve information by decoding the encoded information.

[0171] Signaling information for a specific target can mean that multiple messages are used for that specific target, and that the processing indicated by the multiple messages is applied to that specific target. For example, signaling information at the unit level can mean using / processing corresponding information for each specific unit.

[0172] - Information sent by signaling may include one or more sub-messages. Sending specific information by signaling can mean sending each of the one or more sub-messages included in the specific information by signaling.

[0173] Selective signaling: - The selective transmission of information via signals. Selective signaling of information can refer to the process by which an encoding device selectively includes information in a bitstream (depending on specific conditions). Selective signaling of information can also refer to the process by which a decoding device selectively retrieves information from a bitstream (depending on specific conditions).

[0174] Skipping signal transmission: - Signal transmission of information can be skipped. Skipping signal transmission of information can refer to the process by which the encoding device does not include information in the bitstream (depending on specific conditions). Skipping signal transmission of information can also refer to the process by which the decoding device does not obtain information from the bitstream (depending on specific conditions). In an embodiment, the decoding device can derive the skipped signal transmission information by utilizing other information.

[0175] Symbols: A symbol can refer to at least one piece of information about a target unit, such as the syntax elements, coding parameters, quantization level, and transform coefficients of the target unit or block. Additionally, a symbol can refer to the target of entropy coding or the result of entropy decoding.

[0176] Entropy coding: Entropy coding is a technique that allocates fewer bits to more frequently occurring symbols and more bits to less frequently occurring symbols. Since symbols are represented in this way, the size of the bitstream indicating symbols can be reduced.

[0177] Entropy coding can be implemented using any method such as Variable Length Coding (VLC) and Context-Adaptive Binary Arithmetic Coding (CABAC). For example, entropy coding can be performed using variable-length tables in Variable Length Coding (VLC). In CABAC, for example, binarization methods for symbols and probabilistic models for symbol / binary numbers can be derived to perform entropy coding, and arithmetic coding using context can be performed.

[0178] Entropy decoding: In entropy decoding, the processes performed in entropy encoding can be reversed. Symbols can be generated by performing entropy decoding on a bitstream.

[0179] Parsing: Parsing can refer to determining the value of a syntax element by performing entropy decoding on the encoded information of a bitstream. Alternatively, parsing can refer to entropy decoding itself.

[0180] Statistical values: Values ​​of multiple pieces of information related to a specific entity (multiple entities) described in the embodiments can be used as inputs for specific operations. Statistical values ​​can be values ​​derived by performing specific operations on values ​​related to such a specific entity (multiple entities). For example, statistical values ​​of multiple pieces of specific information can indicate one or more values ​​generated by the average, weighted average (or weighted mean), weighted sum, minimum, maximum, mode, median, interpolated values, sum of products, and product of sums of the values ​​of multiple pieces of specific information. Furthermore, information (such as constants, variables, and encoding parameters) having specific values ​​determined through operations (computations) in the embodiments can have specific statistical values ​​according to the embodiments.

[0181] Encoding parameters In this embodiment, the encoding parameters may be information required for encoding and decoding. The encoding parameters may include information transmitted by signals from the encoding device to the decoding device, information calculated / derived during the encoding process described in this embodiment, and information for processing the encoding described in this embodiment.

[0182] In embodiments, encoding parameters may include one or more of the following: CTU size, cell size, cell type, cell shape, cell depth, minimum cell size, maximum cell size, maximum cell depth, minimum cell depth, cell partitioning information, quadtree (QT) partitioning information, binary tree (BT) partitioning information, BT partitioning direction, BT partitioning format, ternary tree (TT) partitioning information, TT partitioning direction, TT partitioning format, multi-type tree (MTT) partitioning information, combination of MTT partitions, MTT partitioning direction, MTT partitioning format, prediction mode, intra-frame prediction mode, luma intra-frame prediction mode, chroma intra-frame prediction mode, intra-frame partitioning information. Inter-frame partitioning information, coding block partitioning information, prediction block partitioning information, transform block partitioning information, reference sample line index, reference sample filtering method, reference sample filter tap, reference sample filter coefficient, prediction block filtering method, prediction block filter tap, prediction block filter coefficient, prediction block boundary filtering method, prediction block boundary filter tap, prediction block boundary filter coefficient, inter-frame prediction mode, motion information, motion vector (MV), motion vector difference (MVD), MVD resolution, MV value, MV representation accuracy, reference image list, reference image, reference image index, inter-frame prediction direction, inter-frame prediction indicator, prediction list utilization flag, frame order count (POC), MV candidate, MV candidate index, MV candidate list Information on Advanced Motion Vector Prediction (AMVP) mode usage, merge candidates, merge index, merge candidate list, merge mode usage, motion information correction information, skip mode usage, intra-frame block copy mode usage, block vector (BV), block vector difference (BVD), BVD resolution, BV magnitude, BV representation precision, BV candidates, BV candidate index, BV candidate list, interpolation filter taps, interpolation filter coefficients, transform type, transform size, transform selection information, primary transform usage information, secondary transform usage information, primary transform selection information, secondary transform selection information, residual block presence information, coding block style, coding block flag, quantization parameter (QP), increment QP, quantization matrix. Information on deblocking filters, including: deblocking filter usage information, deblocking filter coefficients, deblocking filter taps, deblocking filter strength, deblocking filter shape / form, adaptive sample offset usage information, adaptive sample offset values, adaptive sample offset categories, and adaptive sample offset types. Information on adaptive loop filters, including: adaptive loop filter coefficients, adaptive loop filter taps, and adaptive loop filter shape / form. Binarization / debinarization methods are also included. Context models are defined, including: context model determination methods, context model update methods, normal mode usage information, bypass mode usage information, effective coefficient flags, last effective coefficient flags, coefficient group cell encoding flags, last effective coefficient position, and a flag indicating whether a coefficient value is greater than 1.Flags indicating whether the indicator coefficient value is greater than 2, flags indicating whether the indicator coefficient value is greater than 3, remaining coefficient value information, sign information, context binary number, bypass binary number, reconstructed sample points, reconstructed luminance sample points, reconstructed chrominance sample points, residual sample points, residual luminance sample points, residual chrominance sample points, transform coefficient, luminance transform coefficient, chrominance transform coefficient, transform coefficient level, luminance transform coefficient level, chrominance transform coefficient level, transform coefficient level scanning method, quantization level, quantized luminance level, quantized chrominance level, size of the MV search range on the decoding device side, and M on the decoding device side. The V search range shape, MV search iteration count on the decoding device side, frame type, stripe identification information, stripe type, stripe partition information, parallel block group identification information, parallel block group type, parallel block group partition information, parallel block identification information, parallel block type, parallel block partition information, bit depth, input sample bit depth, reconstructed sample bit depth, residual sample bit depth, transform coefficient bit depth, quantization level bit depth, mapping availability information, luminance signal information, chrominance signal information, target block color space, residual block color space, and temporal layer information.

[0183] In addition, the encoding parameters may also include 1) the value of information that may be included in the encoding parameters, 2) a combination of multiple pieces of information that may be included in the encoding parameters, 3) statistical values ​​of information that may be included in the encoding parameters, 4) information related to the encoding parameters, 5) information used to calculate / derive the encoding parameters, and 6) information calculated / derived using the encoding parameters.

[0184] In an embodiment, "X usage information" can be information indicating whether "X is used / applied / executed". Optionally, "X usage information" can be information indicating whether "X is available". For example, "specific mode usage information" can be information indicating whether a specific mode is used. Mode information can indicate the mode used for the target block among the modes described in the embodiment. In an embodiment, specific mode usage information can be replaced with mode information, and the description of specific mode usage information can also be applied to mode information. "X usage information" and "X indicator" can be used interchangeably.

[0185] In embodiments, encoding parameters and syntax elements may correspond to each other. For example, a syntax element according to an embodiment may be used as an encoding parameter, and the encoding parameter may be sent as a syntax element by a signal.

[0186] In the embodiments, "X existence information" can be regarded as "information indicating whether X exists" or "information indicating whether information indicating X exists in the bitstream".

[0187] In an embodiment, "X selection information" can be information indicating one of the candidates or methods for X. "X selection information" can be considered as an "X index".

[0188] In an embodiment, the partitioning form of a particular tree can indicate one of symmetric partitioning and asymmetric partitioning, and can indicate one of QT, BT, TT, and no partitioning. The partitioning direction of a particular tree can indicate one of horizontal and vertical directions.

[0189] In an embodiment, when the encoding parameter has one of a plurality of values, "encoding parameter" can be replaced with "information indicating whether the encoding parameter has a specific value among a plurality of values ​​that can be used for the encoding parameter".

[0190] In an embodiment, when the encoding parameter has one of a plurality of values, "encoding parameter" can be replaced with "indicating whether the encoding parameter indicates information of a specific target among a plurality of targets".

[0191] In this embodiment, the encoding parameters may include at least one of the target frame type and the target stripe type. The target frame type may be one of I-frame, B-frame, and P-frame. The target stripe type may be one of I-band, B-band, and P-band.

[0192] - When the target image to be encoded is an I-band, inter-frame coding that references other frames can be avoided; instead, the target image can be encoded using data within the image itself. For example, intra-frame prediction can be used only for I-band coding.

[0193] When the target image is a P-strip, it can be encoded using inter-frame predictions of a reference strip existing in only one direction. Here, one direction can be either the forward or backward direction.

[0194] When the target image is a B-band, it can be encoded by using inter-frame prediction with reference strips present in both directions, or by using inter-frame prediction with reference strips present in one of the forward and backward directions. Here, the two directions can include the forward and backward directions.

[0195] P-strips and B-strips encoded and / or decoded using reference stripes can be considered as images that allow for inter-frame prediction.

[0196] Systems for video encoding and decoding Figure 1 A system for video encoding and decoding according to an embodiment is shown.

[0197] System 100 may include at least one or a combination of encoding device 110 or decoding device 150.

[0198] Each of the encoding device 110 and the decoding device 150 may be a computer or an electronic device.

[0199] Structure of encoding devices The encoding device 110 may include a processor 120, a memory 140, and a communicator 149.

[0200] The processor 120, memory 140 and communicator 149 can be connected to each other via a bus.

[0201] Processor 120 may be a semiconductor device that executes instructions or computer-executable code, such as a central processing unit (CPU). Processor 120 may be at least one hardware processor.

[0202] In an embodiment, the processor 120 may generate and process information input to or output from the encoding device 110 or used in the encoding device 110, and may perform comparisons, determinations, etc., related to the information.

[0203] Processor 120 may include multiple components. These components may include partitioner 122, subtractor 124, transformer 125, quantizer 126, dequantizer 127, inverse transformer 128, adder 129, filter 130, and entropy encoder 139.

[0204] At least some of the aforementioned components may be program modules. Program modules may be included in the encoding device 110 in the form of operating systems, applications, and other program modules. Program modules may be instructions or computer-executable code stored in memory 140 and executed by processor 120.

[0205] The memory 140 may include various types of volatile and non-volatile storage media. For example, the memory 140 may include memory such as read-only memory (ROM) and random access memory (RAM).

[0206] The memory 140 may store instructions and computer-executable code for the operation of the encoding device 110, and may also store the information and bit streams described in the embodiments. The memory 140 may include a reference screen buffer 141.

[0207] The communicator 149 can perform functions related to communication of information in the encoding device 110. For example, the communicator 149 can send a bit stream to the decoding device 150.

[0208] In the names of components of encoding device 110, "-cell" can be used instead of "-device" or "-device". Memory 140 can be designated as a storage cell.

[0209] Operation of encoding devices Encoding device 110 can sequentially encode one or more images (pictures) of a video.

[0210] The memory 140 can store the original image. The original image can be used as the target image in the encoding device 110.

[0211] The processor 120 can generate a bitstream including encoded information by performing encoding on the target image, and can store the generated bitstream in the memory 140. The generated bitstream can be stored in a computer-readable storage medium and can be transmitted by the communicator 149 to the communicator 189 of the decoding device 150 via a wired and / or wireless transmission medium.

[0212] The partitioner 122 can determine the target block by partitioning the target image.

[0213] Predictor 123 can determine the prediction mode of the target block. Predictor 123 can generate a prediction block for the target block by performing a prediction corresponding to the prediction mode.

[0214] The prediction mode for the target block can be one of the available prediction modes. For example, available prediction modes may include intra-frame prediction, inter-frame prediction, and IBC prediction.

[0215] For example, when the prediction mode corresponds to intra-frame prediction, predictor 123 can generate a prediction block for the target block by performing intra-frame prediction on the target block.

[0216] For example, when the prediction mode corresponds to inter-frame prediction, predictor 123 can generate a prediction block for the target block by performing inter-frame prediction on the target block.

[0217] For example, when the prediction mode corresponds to IBC, predictor 123 can generate a prediction block for the target block by performing IBC on the target block.

[0218] Subtractor 124 can generate a residual block of the target block. The residual block can be the difference between the original block and the predicted block. The original block can be a region indicated by the target block in the original image. Alternatively, the residual block can refer to a block generated by applying one or more of a transformation and quantization to the difference between the original block and the predicted block.

[0219] Transformer 125 can generate transformation coefficients by performing a transformation on the residual block.

[0220] Transformer 125 can perform a transformation using one of a variety of transformation methods.

[0221] For example, various transformation methods can include Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Karhunen-Loeve Transform (KLT), and transformations based on each transform.

[0222] Transform skip mode can be a mode that uses both the reconstructed residual block and the prediction block without any transform or inverse transform performed to generate the reconstructed block. When transform skip mode is applied to the target block, the transform and inverse transform of the target block can be skipped, and only the quantization and inverse quantization (dequantization) of the target block can be performed.

[0223] Quantizer 126 can generate quantization levels by applying quantization using quantization parameters to transform coefficients. In this embodiment, the quantization level may also be referred to as a "transform coefficient".

[0224] The entropy encoder 139 generates encoded information by performing entropy encoding based on a probability distribution of information used for image decoding. The bitstream may include the encoded information.

[0225] Information used for image decoding may include quantization levels, syntax elements, etc., calculated by quantizer 126.

[0226] The probability distribution can be determined based on the quantization level and encoding parameters.

[0227] The entropy encoder 139 can use a scan to change quantization levels in 2D block form to 1D vector form in order to perform encoding on the quantization levels. During the scan, it can be determined which of the following scans—upper right diagonal scan, vertical scan, and horizontal scan—to use based on encoding parameters such as the block size and the block's intra-frame prediction mode.

[0228] When encoding a target image / block, predictor 123 uses a reference image / block to perform prediction. The encoded target image / block can be used as a reference image / block for another image / block to be processed subsequently. Therefore, processor 120 can perform reconstruction on the encoded target block and can store the reconstructed image, including the reconstructed target block generated by reconstruction, as a reference image in reference image buffer 141. Inverse quantization and inverse transform can be performed on the encoded target block to perform reconstruction.

[0229] The dequantizer 127 can generate dequantized transformation coefficients by performing dequantization on the quantization level.

[0230] The inverse transformer 128 can generate inverse quantization and inverse transform coefficients by performing an inverse transform on the inverse quantization transform coefficients. In embodiments, the inverse quantization and / or inverse transform coefficients may refer to coefficients on which at least one or a combination of inverse quantization and inverse transform has been applied. The inverse quantization and inverse transform coefficients may be reconstructed residual blocks.

[0231] Adder 129 can generate a reconstruction block by adding the prediction block and the reconstruction residual block to each other.

[0232] The reconstructed block can be processed by filter 130. Filter 130 can be configured to apply one or more of a plurality of filters to the target. Each of the plurality of filters can be a loop filter. The target can be a reconstructed sample, a reconstructed block, or a reconstructed image.

[0233] The reference frame buffer 141 can store reconstructed blocks / images provided from the filter 130. The reconstructed image can be an image that includes reconstructed blocks. Alternatively, the reconstructed image can be an image composed of reconstructed blocks.

[0234] The reference frame buffer 141 can provide the stored reconstructed image as a reference image to the predictor 123. From the viewpoint of storing the decoded (i.e., reconstructed) image (frame), the reference frame buffer 141 can also be referred to as the "decoded frame buffer (DPB)".

[0235] Structure of decoding device The decoding device 150 may include a processor 160, a memory 180, and a communicator 189.

[0236] The description of the processor 120, memory 140, and communicator 149 associated with the encoding device 110 can also be applied to the processor 160, memory 180, and communicator 189 associated with the decoding device 150. Repeated descriptions will be omitted here.

[0237] Processor 160 may include multiple components. These components may include entropy decoder 161, partitioner 162, predictor 163, dequantizer 167, inverse transformer 168, adder 169, and filter 170.

[0238] The memory 180 may include a reference screen buffer 181.

[0239] The communicator 189 can perform functions related to communication with information in the decoding device 150. For example, the communicator 189 can receive a bit stream from the encoding device 110.

[0240] In the names of components of the decoding device 150, "-unit" can be used instead of "-device" or "-device". The memory 180 can be designated as a storage unit.

[0241] Operation of decoding devices The communicator 149 of the encoding device 110 can send the bit stream generated by the encoding device 110 to the decoding device 150. Optionally, a computer-readable storage medium storing the bit stream can send the bit stream generated by the encoding device 110 to the decoding device 150.

[0242] The communicator 189 can receive a bit stream from the encoding device 110 via a wired / wireless transmission medium. The received bit stream can be stored in the memory 180.

[0243] Processor 160 may obtain bit streams from memory 180 or computer-readable storage media.

[0244] Bitstreams may include encoded information.

[0245] The entropy decoder 161 can generate information for image decoding by performing entropy decoding on the encoded information of the bitstream based on a probability distribution.

[0246] Information used for image decoding can include quantization levels, syntax elements, etc.

[0247] The entropy decoder 161 can use scanning to transform quantization levels in 1D vector form into 2D block form in order to perform decoding on the quantization levels. During scanning, it can be determined which of the following scans—upper right diagonal scan, vertical scan, and horizontal scan—will be used based on coding parameters such as the block size and the block's intra-frame prediction mode.

[0248] The entropy decoder 161 can provide syntax elements to other components of the processor 160, such as the partitioner 162.

[0249] A common description of the relationship between components of the encoding device and components of the decoding device. Decoding device 150 uses the bitstream generated by encoding device 110 to perform decoding. Encoding device 110 can use the reconstructed image derived in decoding device 150 instead of the original image not provided to decoding device 150 to encode the target block. Therefore, encoding device 110 and decoding device 150 need to be able to generate reconstructed blocks / images in the same way. In this regard, the description of partitioner 122, predictor 123, dequantizer 127, inverse transformer 128, adder 129, filter 130, and reference frame buffer 141 of encoding device 110 disclosed in the embodiments can be applied to partitioner 162, predictor 163, dequantizer 167, inverse transformer 168, adder 169, filter 170, and reference frame buffer 181 of decoding device 150, respectively. Repeated descriptions will be omitted here.

[0250] Furthermore, each of the partitioner 122, predictor 123, dequantizer 127, inverse transformer 128, adder 129, and filter 130 of the encoding device 110 can generate information about the syntax elements for processing a specified target. Furthermore, each of the partitioner 162, predictor 163, dequantizer 167, inverse transformer 168, adder 169, and filter 170 of the decoding device 150 can use the information about the syntax elements to perform processing on the target (the same processing performed by the encoding device 110).

[0251] As described above, corresponding components of the encoding device 110 and the decoding device 150 can perform the same or corresponding functions. In embodiments, the processor may refer to the processor 120 of the encoding device 110 and / or the processor 160 of the decoding device 150. For example, in functions related to prediction, the processor may refer to the predictor 123, the subtractor 124, and the adder 129, and may also refer to the predictor 163 and the adder 169. In functions related to transformation, the processor may refer to the transformer 125 and the inverse transformer 128, and may also refer to the inverse transformer 168. In functions related to quantization, the processor may refer to the quantizer 126 and the dequantizer 127, and may also refer to the dequantizer 167. In functions related to entropy encoding / decoding, the processor may refer to the entropy encoder 139 and / or the entropy decoder 161. In functions related to filtering, the processor may refer to the filter 130 and / or the filter 170. The memory may refer to the memory 140 of the encoding device 110 and / or the memory 180 of the decoding device 150. The reference frame buffer may refer to the reference frame buffer 141 of the encoding device 110 and / or the reference frame buffer 181 of the decoding device 150. The communicator may refer to the communicator 149 of the encoding device 110 and / or the communicator 189 of the decoding device 150.

[0252] Partitions of the units that make up an image Figure 2 An image partitioning structure according to an embodiment is shown.

[0253] Figure 2 An example of a cell being divided into multiple sub-cells can be illustrated.

[0254] A CU can be used as the basic unit for image encoding and decoding. Furthermore, a CU can be used as the basic unit for prediction, transform, quantization, inverse quantization, inverse transform, entropy coding, and entropy decoding.

[0255] A CU can be used as a unit for applying prediction modes. In other words, it can be determined which of the available prediction modes available for each CU will be applied to the encoding. For example, available prediction modes may include intra-frame prediction, inter-frame prediction, and intra-block copy (IBC) prediction.

[0256] The target image 200 can be sequentially partitioned into units of CTUs. For each CTU, a partitioning structure can be determined. A CTU may be partitioned into CUs depending on the partitioning structure. Optionally, a CTU may be used as a CU. The size of a CTU can be the largest CU size.

[0257] Each CU can have depth information. Depth information can refer to the depth of the CU and can also represent the size of the CU. The depth of a CTU can be 0. The depth of a CU generated by partitioning a CTU can be 1. When a parent CU is partitioned into child CUs, the depth of the child CU can be increased by 1 from the depth of the parent CU. The number of CUs generated by partitioning can be a positive integer equal to or greater than 2, including 2, 4, 8, 16, etc. Depending on the number of child CUs, at least one of the horizontal and vertical dimensions of the child CUs generated by partitioning the parent CU can be less than at least one of the horizontal and vertical dimensions of the parent CU.

[0258] The CUs can be recursively partitioned in the same way until a predefined maximum depth or a predefined minimum size is reached. The depth of the minimum coding unit (SCU) can be the predefined maximum depth, and the size of the SCU can be the predefined minimum size. The size of the SCU can be the minimum CU size.

[0259] For example, the range of CU depths can correspond to values ​​from 0 to 3. Depending on the CU depth, the CU can have dimensions ranging from 64×64 to 8×8. A CTU with a depth of 0 can be a 64×64 block. 0 can be the minimum depth. An SCU with a depth of 3 can be an 8×8 block. 3 can be the maximum depth. Depth 0 can represent a CTU as a 64×64 block. Depth 1 can represent a CU as a 32×32 block. Depth 2 can represent a CU as a 16×16 block. Depth 3 can represent an SCU as an 8×8 block.

[0260] The partition information of a CU indicates whether the CU has been partitioned. The partition information can be a 1-bit flag. All CUs except SCUs can include partition information. For example, the partition information of a CU that is not further partitioned can be "0" as the first value, while the partition information of a partitioned CU can be "1" as the second value.

[0261] Quadtree (QT) partitioning refers to dividing a parent CU into four child CUs. When a parent CU is partitioned into four child CUs, the horizontal and vertical dimensions of each child CU can be half the horizontal and vertical dimensions of the parent CU, respectively.

[0262] Binary tree (BT) partitioning refers to dividing a CU into two CUs. For example, when a parent CU is partitioned into two child CUs, the horizontal or vertical dimension of each child CU can be half the horizontal or vertical dimension of the parent CU.

[0263] Ternary tree (TT) partitioning refers to dividing a single CU (Computer Unit) into three CUs. For example, when a parent CU is partitioned into three child CUs, the three child CUs can be generated by partitioning the parent CU's horizontal or vertical dimensions in a 1:2:1 ratio. The horizontal or vertical dimensions of the child CUs can be 1 / 4, 1 / 2, and 1 / 4 of the parent CU's horizontal or vertical dimensions, respectively.

[0264] exist Figure 2 In this configuration, the QT partition is applied to the first CTU. The QT partition, BT partition, and TT partition are applied to the second CTU.

[0265] To partition a CTU, at least one of the following different partition types can be applied to the CTU: QT partition, BT partition, and TT partition. Different partition types can be applied based on specific priorities.

[0266] For example, QT partitions can be preferentially applied to CTUs. CUs that cannot be further partitioned with QT partitions can correspond to leaf nodes of QT. A CU that is a leaf node of QT can be the root node of a BT and / or TT. A CU that is a leaf node of QT can be partitioned in BT or TT form, or it can be left unpartitioned. In this case, the QT partitions may not be applied again to CUs generated by applying BT or TT partitions to CUs that are leaf nodes of QT.

[0267] QT partitioning information can be used to signal the partitioning of the CU corresponding to each node of the QT. QT partitioning information can be a flag. The QT partitioning information of a cell can indicate whether the cell is partitioned in QT format. A "0" as the first value of the QT partitioning information indicates that the CU is not partitioned in QT format. QT partitioning information with a first value can refer to Multi-Type Tree (MTT) partitioning. MTT partitioning can include BT partitioning and TT partitioning. A "1" as the second value of the QT partitioning information indicates that the CU is partitioned in QT format.

[0268] There may be no priority between BT and TT partitions. That is, CUs corresponding to leaf nodes of QT can be partitioned in either BT or TT format. Furthermore, CUs generated by BT or TT partitions can be further partitioned in either BT or TT format, or they may not be further partitioned.

[0269] The CU corresponding to a leaf node of QT can be the root node of MTT. For each CU corresponding to a node of MTT, the CU may also include partition direction information and partition format information in MTT format.

[0270] Partition direction information indicates the partitioning direction of an MTT partition. A "0" as the first value of the partition direction information indicates that the CU is partitioned horizontally. A "1" as the second value of the partition direction information indicates that the CU is partitioned vertically.

[0271] Partition format information indicates the partition format used for Multi-Type Tree (MTT) partitioning. A "0" as the first value of the partition format information indicates that the CU is partitioned in TT format. A "1" as the second value of the partition format information indicates that the CU is partitioned in BT format.

[0272] Here, each of the partition direction information and partition format information mentioned above can be a flag with a specific length (e.g., 1 bit).

[0273] The partition information of a CU may also include QT partition information, partition direction information, and partition format information.

[0274] A CU that has not been further partitioned through QT, BT, or TT partitioning can be used as a unit for specific processes such as prediction, transform, quantization, inverse quantization, inverse transform, entropy coding, and entropy decoding. In other words, a CU may not be further partitioned for a specific process. Therefore, the partitioning information used to partition a CU into PUs and / or TUs may not exist in the bitstream.

[0275] Conversely, when the size of the CU is larger than the maximum TU size, the CU can be recursively partitioned until the size of the CU becomes smaller than or equal to the maximum TU size. For example, when the size of the CU is 64×64 and the maximum TU size is 32×32, the CU can be partitioned into four 32×32 TUs to perform the transformation. Similarly, when the size of the CU is 32×64 and the maximum TU size is 32×32, the CU can be partitioned into two 32×32 TUs to perform the transformation.

[0276] In this scenario, it's not necessary to separately send a signal indicating whether the CU has been partitioned for transformation. Instead, partitioning can be determined by comparing the CU's dimensions (horizontal / vertical) with the maximum TU's dimensions (horizontal / vertical). For example, if the CU's horizontal dimension is greater than the maximum TU's horizontal dimension, the CU can be vertically bisected. Similarly, if the CU's vertical dimension is greater than the maximum TU's vertical dimension, the CU can be horizontally bisected.

[0277] For example, the minimum size of a CU can be 4×4. For example, the maximum size of a transform block can be 64×64. For example, the minimum size of a transform block can be 4×4. The minimum size of a QT can be the minimum size of the CU corresponding to a leaf node of the QT. The maximum depth of an MTT can be the maximum depth of the path from the root node to a leaf node of the MTT.

[0278] The maximum size of BT can represent the maximum size of the CU corresponding to each node of BT, and the maximum size of TT can represent the maximum size of the CU corresponding to each node of TT. The minimum size of BT and / or the minimum size of TT can be set as the minimum size of the CU.

[0279] When the depth of the CU corresponding to a node in the MTT is equal to the maximum depth of the MTT, the CU may not be partitioned in BT form and / or TT form.

[0280] Based on the various sizes and depths of the CU described above, each piece of information described in the embodiments may or may not exist in the bitstream.

[0281] Information about the maximum or minimum size described in the embodiments can be transmitted via signals at a level higher than the CU. In the embodiments, levels higher than the CU may include video level, sequence level, picture level, sub-picture level, parallel block group level, parallel block level, stripe level, etc.

[0282] The information described in the embodiments can be transmitted separately using signals for different types of stripes. Different types of stripes may include intra-frame stripes and inter-frame stripes.

[0283] Block processing depends on block attributes Whether to apply / execute a specific process described in the embodiments can be determined based on the attributes of a block associated with a specific process. Whether to apply / execute a specific process described in the embodiments can be determined based on whether the attributes of a block associated with a specific process satisfy specific conditions. For example, a block may include a target block, neighboring blocks, and reference blocks. The block may include other blocks described in the embodiments. A block may be one of the blocks and units described in the embodiments.

[0284] The blocks for a particular process described in the application examples may have a square shape or a non-square shape.

[0285] In this embodiment, the attributes of a block may include its size. When specific conditions related to the block size are met, specific processes described in this embodiment may be applied / performed.

[0286] In an embodiment, specific conditions may include a minimum block size condition and a maximum block size condition. The block applying the minimum block size condition and the block applying the maximum block size condition may be different from each other.

[0287] In the embodiments, the minimum block size and / or maximum block size for a specific process can be predefined.

[0288] In an embodiment, the processing according to the embodiment may be applied / executed when the block size is equal to or greater than the minimum block size and / or when the block size is less than or equal to the maximum block size. Optionally, in an embodiment, the processing according to the embodiment may be applied / executed when the block size is greater than the minimum block size and when the block size is less than the maximum block size.

[0289] In this embodiment, the processing according to the embodiment may be applied / executed only when the block size is equal to or greater than the minimum block size and less than or equal to the maximum block size. Alternatively, the processing according to the embodiment may be applied / executed only when the block size is greater than the minimum block size and less than or equal to the maximum block size. Alternatively, the processing according to the embodiment may be applied / executed only when the block size is equal to or greater than the minimum block size and less than the maximum block size. The processing according to the embodiment may be applied / executed only when the block size is greater than the minimum block size and less than the maximum block size.

[0290] In this embodiment, the processing according to the embodiment may be applied / executed only if the block size is a predefined block size.

[0291] In embodiments, block size can be determined based on various methods. For example, block size can represent the horizontal or vertical dimension of the block. Block size can represent both the horizontal and vertical dimensions of the block. Block size can represent the area of ​​the block. Block size can represent 1) the result of a known equation, 2) the result of an equation according to the embodiment, or 3) a statistical value using the horizontal and vertical dimensions of the block.

[0292] Furthermore, for the first size, the processing according to the first embodiment can be applied / executed, and for the second size, the processing according to the second embodiment can be applied / executed.

[0293] In embodiments, the block size can be 2×2, 4×4, 8×8, 16×16, 32×32, 64×64, or 128×128. Optionally, in embodiments, the block size can be (2 SIZE X )×(2 SIZE Y (or similar.) SIZE X It can be one of 1 or a larger integer. SIZE Y It can be one of 1 or a larger integer.

[0294] Predictive information used for prediction Prediction information can be used to generate prediction blocks for the target block.

[0295] Encoding device 110 can generate prediction information required for prediction and can generate a bit stream including the prediction information. The prediction information can be transmitted from encoding device 110 to decoding device 150 via the bit stream. Decoding device 150 can obtain the prediction information from the bit stream and can generate a prediction block by performing prediction on the target block using the prediction information.

[0296] The prediction information may include intra-frame prediction information, inter-frame prediction information, and IBC prediction information. In embodiments, the prediction information may be replaced by intra-frame prediction information, inter-frame prediction information, and / or IBC prediction information. Intra-frame prediction may include the information for intra-frame prediction described in the embodiments. Inter-frame prediction may include the information for inter-frame prediction described in the embodiments. IBC information may include the information for IBC prediction described in the embodiments.

[0297] Intra-frame prediction Figure 3 The structure of intra-frame prediction according to an embodiment is shown.

[0298] Intra-frame prediction can be performed using reference samples and coding parameters for the target block. The reference samples can be (reconstructed) samples within a (reconstructed) reference block. Optionally, intermediate prediction samples can be generated using samples such as reconstructed samples described in the embodiments, and then the intermediate prediction samples can be used to generate reference samples. Processes such as filtering described in the embodiments can be applied when generating reference samples.

[0299] The reference block can be a (spatial) neighboring block of the target block. The coding parameters can be the coding parameters used for the target block and / or the coding parameters used for the reference block. In intra-frame prediction, the reference sample can be a neighboring sample.

[0300] Predicted blocks can be generated by performing intra-frame prediction on the target block based on reference points of the target image and information related to those reference points, according to an intra-frame prediction mode. The size of the target block and the size of the predicted block can be equal.

[0301] In this embodiment, the prediction block may be a PU. Alternatively, the prediction block may correspond to the CU or TU described in the embodiment. The prediction block may have a square or rectangular shape.

[0302] Intra-frame prediction modes can be represented by at least one of mode number, mode value, mode angle, and mode direction. Figure 3The lower right portion shows the prediction directions for multiple intra prediction modes used for the target block. Among these intra prediction modes, all except DC mode and planar mode can be directional modes. A directional mode can be an intra prediction mode with a specific direction or angle. The intra prediction mode for the target block can be selected from directional modes and non-directional modes.

[0303] In the rectangle indicating the target block in the lower right part of the figure, the number "0" can indicate a planar mode as a non-directional intra-prediction mode. The number "1" can indicate a DC mode as a non-directional intra-prediction mode. In the rectangle indicating the target block in the lower right part of the figure, the arrow from the center of the rectangle to the outside indicates the prediction direction of the directional intra-prediction mode. Furthermore, the numbers indicated near the arrows can represent examples of mode values ​​assigned to the intra-prediction mode or the prediction direction of the intra-prediction mode.

[0304] Intra-prediction can be performed based on the intra-prediction mode used for the target block. One of the intra-prediction modes available for the target block can be used as the intra-prediction mode for the target block.

[0305] The number of intra-prediction modes available for the target block can be a predefined value. Alternatively, the number of intra-prediction modes available for the target block can be determined based on the attributes of the prediction block. For example, the attributes of the prediction block can include coding parameters such as shape, size, and color components.

[0306] For example, in Figure 3 In the image, the directional patterns indicated by the dashed lines (i.e., directional patterns with numbers corresponding to any one of -14 to -1 or any one of 67 to 80) can be applied only for predictions of non-square blocks. Therefore, the number of intra-frame prediction patterns available for predictions of square blocks can be 67 (planar mode, DC mode, 65 directional patterns).

[0307] For example, the number of available intra-prediction modes can vary depending on which of the luma or chroma signals corresponds to the color component of the block. The number of available intra-prediction modes for the luma component block can be greater than the number of available intra-prediction modes for the chroma component block.

[0308] Intra-frame prediction modes can include horizontally-below mode, horizontal mode, vertical mode, and vertically-right mode. Horizontally-below mode can be an intra-frame prediction mode located below the horizontal mode. Vertically-right mode can be a mode located to the right of the vertical mode. For example, in... Figure 3In the above, the mode value for the horizontal mode can be 18. The mode value for the vertical mode can be 50. Intra-prediction modes (each of which has a mode value of 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, and 66) can be vertical right-side modes. Intra-prediction modes with a mode value of 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, and 17 can be horizontal bottom-side modes.

[0309] The number of intra-prediction modes and the mode number of each intra-prediction mode described above are merely examples. The number of intra-prediction modes and the mode number of each intra-prediction mode may be defined differently depending on the embodiments, their implementation methods, and / or necessity.

[0310] When generating a prediction block for a target block in the intra-frame prediction mode of planar mode, the sample value of the prediction sample can be generated by using the weighted sum of the upper reference sample of the target sample, the left reference sample of the target sample, the upper right reference sample of the target block, and the lower left reference sample of the target block, based on the position of the prediction sample in the prediction block.

[0311] When the intra-frame prediction mode is DC mode, a prediction block can be generated based on the average of the sample values ​​of multiple reference samples. These reference samples can include the top reference sample and the left reference sample of the target block. The values ​​of the prediction samples in the prediction block can be determined based on the average of the sample values ​​of the multiple reference samples. Furthermore, filtering using the values ​​of the reference samples can be performed on specific rows and / or specific columns in the target block. A specific row can be one or more upper rows adjacent to the upper reference sample. A specific column can be one or more left columns adjacent to the left reference sample.

[0312] When the intra-frame prediction mode is directional mode, the prediction block can be generated using the top reference sample, left reference sample, top right reference sample, and / or bottom left reference sample of the target block.

[0313] The intra-prediction mode of the target block can be determined using the intra-prediction modes of its neighboring blocks. Information for determining the intra-prediction mode of the target block can be transmitted using signals.

[0314] For example, when the intra-prediction mode of the target block and the intra-prediction mode of the neighboring blocks are the same, an indicator indicating that the intra-prediction mode of the target block and the intra-prediction mode of the neighboring blocks are the same can be sent by signaling.

[0315] For example, an indicator can be sent to indicate an intra-prediction mode that is the same as the intra-prediction mode of the target block among the intra-prediction modes of multiple neighboring blocks.

[0316] For example, when the intra prediction modes of the target block and neighboring blocks are different from each other, an indicator indicating the intra prediction mode of the target block can be sent using signals. Alternatively, information for deriving the intra prediction mode of the target block based on the intra prediction modes of neighboring blocks can be sent using signals.

[0317] Reference samples used for intra-frame prediction of the target block may include lower left reference samples, left side reference samples, upper left reference samples, upper top reference samples, upper right reference samples, etc.

[0318] For example, a left reference sample can be a reconstructed reference sample adjacent to the left side of the target block. A top reference sample can be a reconstructed reference sample adjacent to the top of the target block. A top-left reference sample can be a reconstructed reference sample adjacent to the top-left diagonal of the target block. A bottom-left reference sample can be a sample located below a left reference sample on the same line as the left-side reference sample line formed by the left reference samples. A top-right reference sample can be a sample located to the right of an upper reference sample on the same line as the upper-side reference sample line formed by the upper reference samples.

[0319] Reference samples for intra-prediction of the target block can be determined based on the intra-prediction mode of the target block. One or more reference samples can be used to determine the sample values ​​for the prediction samples of the prediction block. Figure 3 In the diagram, the direction of the intra-prediction mode, indicated by the arrow, can indicate the direction from the predicted sample to each reference sample. The direction of the intra-prediction mode can represent the dependency between the reference sample and the predicted sample. For example, depending on the intra-prediction mode, the sample value of a particular reference sample can be used as the sample value of at least one sample in the prediction block. Here, the particular reference sample and at least one sample in the prediction block can be samples that are specified as a straight line in the direction of the corresponding intra-prediction mode. In other words, the sample value of the particular reference sample can be copied to the sample value of the predicted sample located in the direction opposite to the direction of the intra-prediction mode. Optionally, the sample value of the predicted sample in the prediction block can be the sample value of the reference sample located in the direction of the intra-prediction mode based on the position of the predicted sample.

[0320] Reference points used for intra-frame prediction are not limited to those points that are exactly adjacent to the target block. For example... Figure 3 As shown, for intra-frame prediction of the target block, at least one or a combination of reference sample lines 0 to 3 can be used.

[0321] Figure 3Each reference sample line can include one or more reference samples. When the reference sample line number is smaller, it indicates a line closer to the target block. Reference sample line 0 can be a line with reference samples that are exactly adjacent to the target block. When the top-left coordinate of the target block is (X, Y), the horizontal length of the target block is W, and the vertical length of the target block is H, the reference sample on reference sample line 0 can be a sample with x-coordinate X-1 or y-coordinate Y-1. Here, the y-coordinate of a reference sample with x-coordinate X-1 can be Y-1 to Y+2H. The x-coordinate of a reference sample with y-coordinate Y-1 can be X-1 to X+2W. Reference samples in reference sample line A can be samples with x-coordinate XA-1 or y-coordinate YA-1. Here, the y-coordinate of a reference sample with x-coordinate XA-1 can be YA-1 to Y+2H+A. The x-coordinate of a reference sample with y-coordinate YA-1 can be XA-1 to X+2W+A. “A” can be 1, 2 or 3.

[0322] Samples in fragments A and F can be derived by filling with samples that are closest to fragments B and E, respectively, instead of being obtained from reconstructed neighboring blocks.

[0323] A reference sample line index can indicate the reference sample line among multiple reference sample lines used for intra-frame prediction of the target block. For example, a reference sample line index can have a value corresponding to one of 0 to 3. The reference sample line index can be transmitted by signal.

[0324] When inter-color component intra-frame prediction is used for a target block, a predicted block for the second color component can be generated based on the reconstructed block of the first color component of the target block. For example, the first color component can be the luma component, and the second color component can be the chroma component.

[0325] For intra-frame prediction between color components, parameters between the first and second color components can be derived based on a template. For example, the parameters can be those of a linear model.

[0326] For example, the template may include an upper reference sample and / or a left reference sample of the target block, and may include an upper reference sample and / or a left reference sample of the reconstructed block of the first color component corresponding to the reference sample.

[0327] When parameters are exported, a predicted block for the second color component can be generated for the target block by applying the reconstructed block of the first color component to the linear model. Depending on the image format or the type of inter-color component intra-frame prediction, subsampling / downsampling can be performed on the neighboring samples of the reconstructed block of the first color component and the reconstructed block of the first color component. When subsampling is performed, the corresponding samples exported through subsampling can be used to perform parameter export and inter-color component intra-frame prediction.

[0328] Intra-fragmented sub-partitioning (ISP) prediction refers to the sequential execution of intra-fragmentation over multiple sub-blocks generated by partitioning a target block. In ISP prediction, the target block can be partitioned into two or four sub-blocks in the horizontal and / or vertical directions. The sub-blocks generated from the partitions can be reconstructed sequentially. When intra-fragmentation is performed on each sub-block, a sub-prediction block can be generated for each sub-block. Furthermore, when inverse quantization and / or inverse transform is performed on each sub-block, a sub-residual block can be generated for the corresponding sub-block. The corresponding reconstructed sub-block can be generated by adding the sub-prediction block to the sub-residual block. The reconstructed sub-block can be used as reference samples for intra-fragmentation of other sub-blocks that will be processed subsequently.

[0329] When performing prediction for a target block, it can be determined whether samples included in the reconstructed neighboring blocks can be used as reference samples for the target block. When there are unusable samples among the samples in the neighboring blocks that cannot be used as reference samples for the target block, the sample value of each unusable sample can be replaced by a value generated based on the copying and / or interpolation of the sample value using at least one sample included in the reconstructed neighboring blocks. When the sample value of a sample is replaced by a value generated by copying and / or interpolation, the corresponding sample can be used as a reference sample for the target block.

[0330] When performing intra-frame prediction, the sample values ​​of prediction samples in the prediction block can be determined based on the sample values ​​of reference samples. The position of the reference sample can be specified by the position of the prediction sample and the direction of the intra-frame prediction mode. When the position specified by the position of the prediction sample and the direction of the intra-frame prediction mode is an integer position, the sample value of a reference sample indicated by the integer position can be used to determine the sample value of the prediction sample in the prediction block. When the position specified by the position of the prediction sample and the direction of the intra-frame prediction mode is not an integer position, an interpolated reference sample can be generated based on the two reference samples closest to the specified position. The sample value of the interpolated reference sample can be used to determine the sample value of the prediction sample. In other words, when the position specified by the position of the prediction sample and the direction of the intra-frame prediction mode indicates the position between two reference samples, an interpolated sample value can be generated based on the sample values ​​of the two samples.

[0331] Inter-frame prediction Figure 4 The structure of inter-frame prediction for explaining inter-frame prediction processing according to an embodiment is shown.

[0332] Figure 4 The rectangle shown can represent an image (screen). Furthermore, in Figure 4 In the image, the arrow can specify the prediction direction.

[0333] Based on the encoding type, the individual images that make up a video can be classified into I-frames (i.e., intra-frame frames), P-frames (i.e., one-way predictive frames), and B-frames (i.e., two-way predictive frames). Encoding and decoding of each frame can be performed according to its encoding and decoding type.

[0334] When the target frame is an I-frame, information from the target frame can be used to perform encoding and decoding of the target frame without referencing inter-frame prediction of other images. For example, intra-frame prediction and / or IBC prediction can be used to perform encoding and decoding of the I-frame.

[0335] Encoding and decoding of P-frames and B-frames can be performed through at least one of intra-frame prediction, IBC prediction, and inter-frame prediction using a reference image.

[0336] When the target frame is a P-frame, one-way inter-frame prediction using a list of reference images can be used to perform encoding and decoding of the target frame.

[0337] When the target frame is a B-frame, the encoding and decoding of the target frame can be performed using one-way inter-frame prediction or two-way inter-frame prediction with two reference image lists.

[0338] In the following, inter-frame prediction for a target block in inter-frame mode according to an embodiment will be described in detail.

[0339] When the prediction mode of the target block is inter-frame mode, inter-frame prediction can be performed on the target block. The target block can be a prediction block or a partition prediction block.

[0340] Inter-frame prediction can be performed using reference images and motion information. In inter-frame prediction, a reference image index can be used to select a reference image (reference frame), and motion information can be used to determine a reference block within the reference image corresponding to the target block. The determined reference block can then be used to generate a predicted block for the target block.

[0341] Motion information can be derived using encoding parameters, etc. For example, motion information can be derived using motion information from reconstructed neighboring blocks, motion information from COL blocks, and / or motion information from blocks adjacent to COL blocks.

[0342] In this embodiment, the candidate list can be used for inter-frame prediction. The candidate list may include multiple candidates. An index indicating the candidate in the candidate list used for inter-frame prediction of the target block can be transmitted by signaling. The candidate list can be derived by the encoding device 110 and the decoding device 150 in the same manner based on the same information. Here, the same information may include the reconstructed image and the reconstructed block. Furthermore, in order to execute the candidate selection using the index, the order of the candidates in the candidate list needs to be consistent.

[0343] In this embodiment, prediction for a target block can be performed by utilizing the motion information of spatial candidates or temporal candidates as the motion information of the target block. The motion information of spatial candidates can be referred to as spatial motion information. The motion information of temporal candidates can be referred to as temporal motion information.

[0344] Spatial candidates can be reconstructed spatial neighbor blocks that are spatially adjacent to the target block.

[0345] Spatial candidates can be 1) blocks that exist in the target image, 2) blocks that have been reconstructed by decoding, and 3) blocks that are adjacent to the target block.

[0346] Spatial candidates can include the target block's left block, top block, bottom left block, top right block, and top left block.

[0347] Temporal candidates can be reconstructed time neighbors that exist in the reconstructed COL image (picture) and correspond to the target block.

[0348] In this embodiment, the motion information of spatial candidates can be the motion information of blocks including spatial candidates. The motion information of temporal candidates can be the motion information of blocks including temporal candidates.

[0349] In inter-frame prediction, COL blocks can be identified for the target block. The region of the target block in the target image and the region of the COL block in the COL frame can be the same. In other words, a COL block can be a block that occupies a specific area in the COL frame. The specific area can be the region in the COL frame that corresponds to the region of the target block.

[0350] The time candidate can indicate the position inside the COL block within the COL screen and / or the position outside the COL block.

[0351] For example, a COL block can include a first COL block and a second COL block. The top-left coordinate of the COL block is (xP, yP) and the size of the COL block is (nPSW, nPSH). The first COL block can be the block occupying coordinates (xP+nPSW, yP+nPSH). The second COL block can be the block occupying coordinates (xP+(nPSW>>1), yP+(nPSH>>1)). When the first COL block is unavailable, the second COL block can be selectively used as the COL block.

[0352] The MV of the target block can be determined based on the MV of the COL block. Scaling can be performed on the MV of the COL block. The scaled MV of the COL block can be used as the MV of the target block or the predicted MV. Optionally, the MVs of temporal candidates listed in the candidate list related to inter-frame prediction can be scaled MVs.

[0353] The ratio of the scaled MV to the COL block's MV can be the same as the ratio of the first time distance to the second time distance. The first time distance can be the distance between the reference image of the target block and the target image. The second time distance can be the distance between the reference image of the COL block and the COL image.

[0354] The scheme for deriving motion information can be determined based on the inter-frame prediction mode of the target block. For example, AMVP mode, merge mode, skip mode, merge mode with MVD, sub-block merge mode, GPM, combined inter-intra-frame prediction (CIIP) mode, or affine inter-frame prediction mode can be used as the inter-frame prediction mode. The various inter-frame prediction modes will be described in the following embodiments.

[0355] AMVP mode When the AMVP pattern is used as the prediction pattern, a list of MV candidates, including one or more MV candidates, can be created using spatial MV candidates, temporal MV candidates, history-based MV candidates, zero vectors, etc. At least one of spatial MV candidates, temporal MV candidates, zero vectors, etc., can be determined and used as an MV candidate.

[0356] Spatial candidates can include reconstructed spatial neighbor blocks. The motion vectors (MVs) of the reconstructed spatial neighbor blocks can be referred to as spatial motion vector (MV) candidates. Temporal candidates can include COL blocks and blocks adjacent to COL blocks. The MV of a COL block or the MV of a block adjacent to a COL block can be referred to as temporal motion vector (MV) candidates. History-based MV candidates can be MVs included in a list of MVs of other blocks that were encoded / decoded before the target block was encoded / decoded.

[0357] Encoding device 110 can use an MV candidate list to determine the MV to be used for encoding the target block within a search range. The maximum number of MV candidates in the MV candidate list can be predefined. N can represent the predefined maximum number. For example, N can be 2. Optionally, the maximum number of candidates can be signaled from the encoding device to the decoding device, or it can be derived by the decoding device. Encoding device 110 can determine MV candidates from the MV candidates in the MV candidate list that will be used as the predicted MV for the target block. The MV to be used for encoding the target block can be the MV that can be encoded with minimal cost. Encoding device 110 can determine whether an AMVP pattern will be used to encode the target block and can generate AMVP pattern usage information indicating whether an AMVP pattern is used.

[0358] Inter-frame prediction information may include 1) AMVP mode usage information, 2) MV candidate index, 3) MVD, 4) MVD resolution information, 5) reference orientation, and 6) reference image index, and may also include residual blocks. The inter-frame prediction information can be transmitted from the encoding device 110 to the decoding device 150 via a bitstream.

[0359] Decoding device 150 can obtain AMVP mode usage information from the bitstream. When the AMVP mode usage information indicates that the AMVP mode is used, decoding device 150 can obtain MV candidate index, MVD, MVD resolution information, reference orientation, and reference image index from the bitstream. Among the MV candidates included in the MV candidate list, the MV candidate indicated by the MV candidate index can be selected as the predicted MV of the target block.

[0360] MVD can represent the difference between the MV actually used for inter-frame prediction of the target block and the predicted MV. Encoding device 110 can derive a predicted MV that is closer to the MV actually used for inter-frame prediction of the target block in order to use an MVD with the smallest possible value. Decoding device 150 can derive the MV of the target block by adding the MVD to the predicted MV. In other words, the MV of the target block derived by decoding device 150 can be the sum of the MVD and the predicted MV candidate.

[0361] Furthermore, the encoding device 110 can generate MVD resolution information. The MVD resolution information can be used to adjust the MVD resolution. The decoding device 150 can use the MVD resolution information to adjust the MVD resolution.

[0362] Furthermore, the encoding device 110 can calculate the MVD based on an affine model. The affine control points MV of the target block can be derived based on the sum of the affine control point MV candidates and the MVD. The MV of each sub-block in the target block can be derived using the affine control points MV.

[0363] Merge mode When using the merge mode, motion information from spatial candidates, motion information from temporal candidates, etc., can be used to create a merge candidate list that includes multiple merge candidates. Motion information can include 1) motion vector (MV), 2) reference image index, 3) reference orientation, etc. Each merge candidate can be a piece of motion information.

[0364] Merge candidates may include 1) spatial merge candidates generated based on spatial candidates, 2) temporal merge candidates generated based on temporal candidates, 3) historical merge candidates, 4) average merge candidates, and 5) zero merge candidates.

[0365] Historically based merge candidates can include information from multiple motion information of other blocks that were encoded / decoded before the target block was encoded / decoded.

[0366] The average merge candidate can be a merge candidate generated based on the average of two merge candidates in the merge candidate list.

[0367] Zero-vector motion information can be zero-vector motion information. Zero-vector motion information can be motion information where MV is a zero-vector.

[0368] Merging candidates can be added to the merging candidate list according to a predefined method and a predefined order, so that the merging candidate list has a preset number of merging candidates. The encoding device 110 and the decoding device 150 can construct the same merging candidate list by using the predefined method and predefined order.

[0369] Encoding device 110 can select a merge candidate from a merge candidate list to be used for encoding the target block. Encoding device 110 can determine whether a merge mode will be used to encode the target block and can generate merge mode usage information indicating whether a merge mode is used.

[0370] Inter-frame prediction information may include 1) merging mode usage information, 2) merging index, and 3) correction information, and may include residual blocks. The inter-frame prediction information can be transmitted from the encoding device 110 to the decoding device 150 as a bit stream.

[0371] Decoding device 150 can obtain merge mode usage information from the bitstream. When the merge mode usage information indicates that a merge mode is being used, decoding device 150 can obtain information related to the merge mode from the bitstream, such as the merge index.

[0372] Encoding device 110 can select the best merge candidate from merge candidates included in the merge candidate list, and can set the value of the merge index to indicate the selected merge candidate.

[0373] The correction information can be used to correct the MV. Encoding device 110 can generate the correction information. Decoding device 150 can derive the corrected MV by correcting the MV of the merge candidate selected by the merge index based on the correction information. The corrected MV can be used as the MV of the target block.

[0374] In an embodiment, the correction information may include MVD (Modular Value Determination). The correction information may include one or more of correction usage information, correction direction information, and correction size information. The correction usage information may indicate whether correction of the MV is used. A merging mode that performs MV correction based on the correction information may be referred to as a merging mode with MVD.

[0375] In merge mode, predictions for the target block can be performed using merge candidates indicated by the merge index from among the merge candidates included in the merge candidate list.

[0376] The motion information of the target block can be derived using 1) MV, 2) reference image index, 3) reference direction, etc., of the merge candidates indicated by the merge index.

[0377] In an embodiment, the merge candidates in the merge candidate list can be specific modes for deriving inter-frame prediction information. Each merge candidate can be information indicating a specific mode for deriving inter-frame prediction information. Inter-frame prediction information for a target block can be derived based on the specific mode indicated by the merge candidate. In this regard, a specific mode can be considered a specific inter-frame prediction information deriving mode or a specific motion information deriving mode. A specific mode can include a series of processes for deriving inter-frame prediction information.

[0378] Inter-frame prediction information for the target block can be derived based on a specific mode indicated by the merge candidate selected by the merge index from the merge candidate list. For example, the specific mode may include a sub-block motion information derivation mode and an affine motion information derivation mode, and may also include other modes for deriving motion information as described in the embodiments.

[0379] Skip mode can be a mode that does not use residual blocks. In other words, when using skip mode, the reconstructed block can be the same as the predicted block. The description of merge mode in the embodiments can also be applied to skip mode. The difference between merge mode and skip mode lies in whether residual blocks are sent and used. In other words, except for not sending / using residual blocks, skip mode can be similar to merge mode, and the description of merge mode can also be applied to skip mode.

[0380] Sub-block merging mode can be a mode for deriving motion information of target sub-blocks within a target block. When applying sub-block merging mode, a list of sub-block merging candidates can be created using affine control point motion vector merging candidates and / or sub-block-based temporal merging candidates. Sub-block-based temporal merging candidates can be the motion information of the COL sub-blocks of the target sub-block.

[0381] In GPM, two motion information points of the target block can be used to generate a first prediction block and a second prediction block. For each coordinate of the target block, the weighted sum of the first prediction sample points of the first prediction block and the second prediction sample points of the second prediction block can be used to generate the final prediction block of the final prediction block.

[0382] Here, the first weight of the first predicted sample and the second weight of the second predicted sample can be determined based on the boundaries of the GPM. The boundaries can indicate the partition lines used to partition the target block. According to the boundaries, the target block can be partitioned into a first partition region and a second partition region.

[0383] When the distance between the final predicted sample point and the boundary is less than or equal to the reference value, the value of the final predicted sample point in the final prediction block can be determined by the weighted sum of the first predicted sample point in the first prediction block and the second predicted sample point in the second prediction block. When the distance between the final predicted sample point and the boundary is greater than the reference value, one of the first weight and the second weight can be 1, and the other can be 0.

[0384] The Combined Inter-Frame Intra-Frame Prediction (CIIP) mode can be a mode that uses a weighted sum of prediction samples generated by inter-frame prediction and prediction samples generated by intra-frame prediction to derive prediction samples for the target block.

[0385] In the above mode, improvements can be made to the derived motion information itself, and the improved motion information can be used as the motion information for the target block. For example, blocks within a specific region determined based on the derived motion information can be searched, and the motion information of the block with the minimum sum of absolute differences (SAD) among the found blocks can be used as the improved motion information for the target block. The specific region can be a square region within a reference image specified by the motion information. The point indicated by the motion information can be the center of the specific region.

[0386] In the aforementioned mode, optical flow can be used to perform compensation for the predicted samples derived through inter-frame prediction.

[0387] Figure 5 The order in which spatial candidates are added to the candidate list is shown according to an embodiment.

[0388] exist Figure 5 The text describes the locations of spatial candidates.

[0389] The large block at the center of the attached diagram indicates the target block. The five smaller blocks adjacent to the target block represent spatial candidates.

[0390] The target block's coordinates can be (xP, yP), and the target block's size can be (nPSW, nPSH).

[0391] Spatial candidate A0 can be the block that is adjacent to the lower left of the target block. A0 can be the block that occupies the sample point corresponding to the coordinates (xP-1, yP+nPSH).

[0392] Spatial candidate A1 can be the block to the left of the target block. A1 can be the bottommost block among the blocks to the left of the target block. Optionally, A1 can be the block to the top of A0. A1 can be the block that occupies the sample point corresponding to the coordinates (xP-1, yP+nPSH).

[0393] Spatial candidate B0 can be the block that is adjacent to the top right of the target block. B0 can be the block that occupies the sample point corresponding to the coordinates (xP+nPSW, yP-1).

[0394] Spatial candidate B1 can be the block adjacent to the top of the target block. B1 can be the rightmost block among the blocks adjacent to the top of the target block. Optionally, B1 can be the block adjacent to the left of B0. B1 can be the block that occupies the sample point corresponding to the coordinates (xP+nPSW-1,yP-1).

[0395] Spatial candidate B2 can be the block that is the top-left neighbor of the target block. B2 can be the block that occupies the sample point corresponding to the coordinates (xP-1, yP-1).

[0396] like Figure 5 As shown, when adding spatial candidates to the candidate list, the order B1, A1, B0, A0, and B2 can be used. That is, available spatial candidates can be added to the candidate list in the order of B1, A1, B0, A0, and B2. Figure 5 middle, Figure 5 The order in which spatial candidates are added to the merge candidate list in the diagram is merely an example.

[0397] The candidate list can include motion information candidate list, merged candidate list, MV candidate list, BV candidate list, MPM list, etc.

[0398] To add a spatial or temporal candidate to the candidate list, it can be determined whether the spatial or temporal candidate is available. When a candidate block exists outside the boundaries of an image, strip, parallel block, etc., the availability of the corresponding candidate block can be set to "false". The description "the availability of the corresponding candidate block can be set to false" can mean that the corresponding candidate block is designated as "unavailable".

[0399] The maximum number of candidates in the candidate list can be set. N can represent the preset maximum number of candidates. The preset maximum number of candidates can be signaled via parameter sets, headers, etc. For example, the maximum number of candidates in the candidate list for a target block in a strip can be set via the strip header. For example, N can be essentially 5.

[0400] IBC mode IBC mode can be an intra-frame block copy prediction mode, in which predicted blocks for the target block are generated by referencing previously reconstructed regions in the target image. In this sense, IBC mode can be referred to as the current image reference mode. A block vector (BV) can be used to specify the previously reconstructed region.

[0401] IBC mode usage information can be used to determine whether the target block is encoded / decoded in IBC mode. Encoding device 110 can determine whether IBC mode will be used to encode the target block and can generate IBC mode usage information indicating whether IBC mode is used. Decoding device 150 can obtain IBC mode usage information from the bitstream.

[0402] In IBC mode, predicted blocks for a target block can be generated based on BV. BV can specify a reference block. BV can refer to the displacement between the target block and the reference block. The reference block can be a block in the target image (picture). The description of MV in the embodiments can also be applied to BV.

[0403] IBC modes can include skip mode, merge mode, AMVP mode, etc. The descriptions of AMVP mode, merge mode, and skip mode in the embodiments can be similarly applied to the AMVP mode, merge mode, and skip mode of IBC modes, respectively.

[0404] In skip or merge mode, a merge candidate list can be constructed, and the merge index can specify one of the merge candidates in the merge candidate list. The BV of the specified merge candidate can be used as the BV of the target block.

[0405] In the AMVP model, BVD can be used. The description of MVD in the examples can also be applied to BVD.

[0406] In IBC mode, the reference block can be limited to a block in a previously reconstructed region of the target image. Optionally, the reference block can be included in the target CTU or in at least one of the left-hand CTUs. For example, the BV value can be restricted such that the reference block is located in a specific region. The specific region can be a region corresponding to three blocks of a specific size that are encoded / decoded before the block of a specific size including the target block is encoded / decoded. The specific size can be 64×64.

[0407] Transformation and Quantization Quantization levels can be generated by performing a transform / quantization on the residual block. A residual block can refer to the difference between the original block and the predicted block. Reconstructed residual blocks can be generated by performing inverse quantization and / or inverse transform on the quantization levels. Reconstructed residual blocks can refer to the difference between the reconstructed block and the predicted block.

[0408] When performing a transformation or inverse transformation, a separable transformation or a two-dimensional (2D) non-separable transformation can be performed on the residual block. A separable transformation can be a transformation that performs a 1D transformation on the residual block in each of the horizontal and vertical directions.

[0409] Transform kernels used for transformations may include: 1) various DCT kernels, such as DCT type 2 (DCT-II); 2) DST kernels; and 3) kernels derived through training. For 1D transforms, in addition to DCT-II, DCT type and DST type may also include DCT-V, DCT-VIII, DST-I, and DST-VII.

[0410] Transform sets can be used when determining the DCT type, DST type, or kernel derived through training for the transform. Each transform set can include multiple transform candidates. Each transform candidate can be a kernel derived through a DCT type, DST type, or training.

[0411] Encoding device 110 can perform transform and inverse transform using transform candidates included in each transform set. Decoding device 150 can perform inverse transform using transform candidates included in each transform set. Transform selection information can be transmitted by signaling, indicating which of the multiple transform candidates included in the transform set applied to the residual block will be used. Transform selection information may include vertical transform selection information and horizontal transform selection information. Vertical transform selection information can indicate which of the transforms belonging to the transform set will be used for vertical transform. Horizontal transform selection information can indicate which of the transforms belonging to the transform set will be used for horizontal transform.

[0412] The transformation may include at least one of a primary transformation and a secondary transformation. Primary transformation coefficients can be generated by performing a primary transformation on the residual block, and secondary transformation coefficients can be generated by performing a secondary transformation on the transformation coefficients. Here, the transformation coefficients may include both primary and secondary transformation coefficients.

[0413] Primary transformation can refer to multiple transformation selection (MTS) that applies different transformations to various 1D directions (i.e., vertical and horizontal directions).

[0414] Secondary transformations can be transformations used to increase the energy concentration of the transformation coefficients generated by the primary transformation. Secondary transformations can be 1) separable transformations such as the primary transformation, and 2) 2D non-separable transformations. 2D non-separable transformations can refer to low-frequency non-separable transformations (LFNST) or non-separable primary transformations (NSPT).

[0415] NSPT can be applied to specific block sizes (such as 4×4, 4×8, 8×4, 4×16, 16×4, 8×8, 8×16 and 16×8) for intra-frame coding.

[0416] The primary transformation can be performed using at least one of several predefined transformation methods. In the example, the predefined transformation methods may include DCT, DST, KLT, etc. Furthermore, depending on the transformation kernel functions that define the DCT and DST, the primary transformation can be a transformation with various transformation types. For example, depending on the multiple transformation kernels, the primary transformation may include multiple transformations such as DCT-2, DCT-4, DCT-5, DCT-7, DCT-8, DST-1, DST-2, DST-4, DST-7, and DST-8.

[0417] In an embodiment, the transform type can be determined based on coding parameters associated with the target block. For example, the transform type can be determined based on one or more of the following: 1) the prediction mode of the target block (e.g., one of intra-frame prediction and inter-frame prediction), 2) the size of the target block, 3) the shape of the target block, 4) the intra-frame prediction mode of the target block, 5) the components of the target block (e.g., one of luma component and chroma component), and 6) the partition type applied to the target block (e.g., one of QT, BT, TT, and non-separable type).

[0418] For example, in the case of a primary transformation, a transformation set can be defined even in a secondary transformation. The method for deriving and / or determining the transformation set according to the embodiments can also be applied to both secondary and primary transformations.

[0419] In this embodiment, primary and / or secondary transformations can be determined for a specific target. Transformation selection information may include transformation target information. The transformation target information may refer to a target to which primary and / or secondary transformations have been applied.

[0420] For example, primary and / or secondary transformations can be applied to one or more signal components in the luminance and chrominance components.

[0421] In this embodiment, the transform selection information may include primary transform usage information and secondary transform usage information. The primary transform usage information may indicate whether the primary transform is applied to the residual block of the target block. The secondary transform usage information may indicate whether the secondary transform is applied to the residual block of the target block.

[0422] In an embodiment, it can be determined whether to apply primary and / or secondary transformations based on the encoding parameters of the target block / neighboring block (such as the size and shape of the target block / neighboring block).

[0423] In an embodiment, the transform selection information may include primary transform selection information and secondary transform selection information. The primary transform selection information indicates the transform method to be applied to the residual block from among a variety of transform methods available in the primary transform. The primary transform selection information may be a primary transform index. The secondary transform selection information indicates the transform method to be applied to the transform coefficients from among a variety of transform methods available in the secondary transform. The secondary transform selection information may be a secondary transform index.

[0424] In embodiments, transformation methods for primary and secondary transforms can be derived separately based on specific information such as encoding parameters. For example, encoding parameters may include encoding parameters for the target block / neighboring block.

[0425] In this embodiment, information related to the transformation, such as transformation selection information and sub-information of the transformation selection information, can be transmitted via signals to a specific target. For example, the specific target may be a CU.

[0426] Information related to the transformation (such as transformation selection information) and sub-information of the transformation selection information can be exported for a specific target. For example, the specific target can be a CU.

[0427] Quantization levels can be generated by performing quantization on the results generated by performing primary and / or secondary transformations or on the residual block.

[0428] The above description of the transformation can also be applied to the inverse transformation. In this application, the inverse processing of the transformation description can be performed in the inverse transformation. The term "transformation" in the names of components related to the transformation can be changed to the term "inverse transformation." Furthermore, the input of the transformation can be regarded as the output of the inverse transformation. The output of the transformation can be regarded as the input of the inverse transformation. The decoding device 150 can acquire transformation-related information such as transformation selection information, and can perform the inverse processing of the transformation-related processing indicated by the transformation-related information by using the transformation-related information.

[0429] A target block can include multiple sub-blocks. Each sub-block can be defined based on the minimum block size or minimum block shape. A target block can be partitioned into multiple sub-blocks, each of which can include coefficients such as 4×4, 2×8, and 8×2. A target block can be a transform block. Transform coefficients or quantization levels can be represented by the shape of the block. Transform coefficients can be quantized transform coefficients.

[0430] Transform coefficients or quantization levels can be scanned using at least one scan type selected from diagonal scan, vertical scan, and horizontal scan. A diagonal scan can be either an upper right diagonal scan or a lower left diagonal scan.

[0431] For example, the coefficients of a block can be changed or arranged in 1D vector form by using a diagonal scan. A vertical scan can be an operation that scans coefficients with a 2D block shape in the column direction. A horizontal scan can be an operation that scans coefficients with a 2D block shape in the row direction.

[0432] The scan type of coefficients can be determined based on coding parameters such as intra-prediction mode, block size, and block shape. For example, it can be determined which scan—diagonal, vertical, or horizontal—will be used based on coding parameters such as intra-prediction mode, block size, and block shape. A block can be a transform unit.

[0433] Depending on the type of scan, a scan can begin at a specific starting point and can end at a specific ending point.

[0434] During a scan, the scan order corresponding to the scan type can be applied to the sub-blocks first. Next, the scan order corresponding to the scan type can be applied to the transform coefficients or quantization levels within each sub-block.

[0435] Encoding device 110 can generate a bit stream that includes entropy-coded transform coefficients or entropy-coded quantization levels by performing entropy coding on transform coefficients or quantization levels.

[0436] Decoding device 150 can obtain entropy-coded transform coefficients or entropy-coded quantization levels from the bitstream, and can generate transform coefficients or quantization levels by performing entropy decoding on the entropy-coded transform coefficients or entropy-coded quantization levels. The coefficients can be arranged in 2D blocks via inverse scanning. The inverse scanning arrangement can be a rearrangement that is the reverse of the scanning arrangement.

[0437] The inverse scan transform coefficients or the quantization level of the inverse scan can be generated by inverse scanning the coefficients. Here, the inverse scan type can include diagonal scan, vertical scan, and horizontal scan, and the inverse scan type of the inverse transform corresponding to the scan type of the transform can be selected.

[0438] Decoding device 150 can perform dequantization on the (inverse scan) coefficients. Depending on whether a secondary inverse transform is to be performed, a secondary inverse transform can be performed on the result of the dequantization. Furthermore, depending on whether a primary inverse transform is to be performed, a primary inverse transform can be performed on the result of the secondary inverse transform. Reconstructed residual blocks can be generated by selectively performing both secondary and primary inverse transforms on the coefficients.

[0439] Filtering To improve image quality, filtering can be performed on blocks. Filtering can be used to determine or update the values ​​of target samples.

[0440] The target sample point can be one of the sample points described in the embodiments. For example, the target sample point can be one or more of the sample points described in the embodiments, such as prediction sample points, reference sample points, residual sample points, reconstruction sample points, and filtered reconstruction sample points.

[0441] The target sample point can be a sample point within one or more of the following: target image, target strip, target CTB, target block, reference sample point line, and template. The target block can be one of the above-mentioned blocks. For example, the target block can be one or more of the blocks described in the embodiments, such as transform block, prediction block, reference block, residual block, and reconstruction block.

[0442] In the embodiments, the filtering process described as being applied to one object can also be applied to another object. For example, the filtering process described in a specific loop filter can also be applied to transform blocks, prediction blocks, reference blocks, residual blocks, etc.

[0443] Specific types of filtering can be used for filtering processes according to embodiments. Filter types may include filter taps (or filter tap lengths), filter shapes, filter strengths, filter coefficients (or weights), and offsets for each filter.

[0444] A filter tap can represent the number of input samples used for the filter. Input samples may include target samples. Optionally, input samples may include specific values ​​determined for the target samples. Input samples may include one or more reference samples. One or more reference samples may be determined based on the properties of the target block described in the embodiments. Properties may include encoding parameters. For example, the properties of the target sample may include the position of the target sample. One or more reference samples may be specified based on their position relative to the position of the target sample.

[0445] The filter shape can be represented by the shape formed by the input samples. A specific value determined for a target sample can be considered as the target sample. In other words, when a specific value determined for a target sample is used as an input sample for the filter, the target sample can also be considered to form the filter shape.

[0446] The samples whose values ​​are determined by filtering can include multiple samples. The filter strength can represent the range of samples whose values ​​are determined by filtering. The filter strength can be either a strong filter strength or a weak filter strength. The number of samples whose values ​​are determined by a strong filter strength can be greater than the number of samples whose values ​​are determined by a weak filter strength. Optionally, the filter strength can represent the range of values ​​changed by filtering. The range of sample values ​​changed by a strong filter strength can be wider than the range of sample values ​​changed by a weak filter strength.

[0447] Filter coefficients can be the coefficients or weights of the input samples.

[0448] An offset can be a specific value that will be added to the result of a computation (such as a weighted sum) performed using the values ​​of the input samples and coefficients.

[0449] Filtering, interpolation, and sampling all share the characteristic that the values ​​of the sample points are updated. Therefore, the description of any one of filtering, interpolation, and sampling in the embodiments can also be applied to another of filtering, interpolation, and sampling. Here, sampling can include at least one of upsampling, downsampling, and subsampling.

[0450] Filtering may include filtering performed by predictor 123, predictor 163, etc.

[0451] When encoding a target block, prediction errors may exist between the original samples in the original block and the predicted samples in the prediction block. To reduce prediction errors, filtering can be performed on at least one of the predicted samples in the prediction block and the reference samples used for prediction.

[0452] For example, in intra-frame prediction, reference samples may include one or more of the following: top-left reference sample, top-right reference sample, left-side reference sample, and bottom-left reference sample. Filtering of prediction samples can be performed by applying specific weights to the prediction sample, left-side reference sample, top-side reference sample, and / or top-left reference sample, respectively.

[0453] Filtering can be performed on at least one of the predicted sample and the reference sample based on the attributes of the target block and the attributes of the predicted sample. For example, whether to perform filtering, the type of filter, the region to which the filter is applied, the filter weights, the reference sample, the range of the reference sample, and the location of the reference sample can be determined based on the attributes of the target block and the attributes of the predicted sample.

[0454] For example, the attributes of the target block may include information related to the target block described in the embodiments, such as the target block's 1) size, 2) prediction mode, 3) intra-frame prediction mode, 4) reference sample line, 5) sample value, and 6) coding parameters.

[0455] For example, the attributes of the predicted sample may include information related to the predicted sample described in the embodiments, such as 1) the sample value of the predicted sample and 2) the position of the predicted sample in the target block, and may include encoding parameters related to the predicted sample.

[0456] Filtering may include loop filtering performed by filter 130, filter 170, etc.

[0457] Figure 6 Several loop filters are shown based on the example.

[0458] Multiple loop filters used for loop filtering may include one or more of the following: Luminance Map with Chroma Scaling (LMCS), Deblocking Filter, Sample Adaptive Offset (SAO), and Adaptive Loop Filter (ALF).

[0459] Multiple loop filters can be connected sequentially to each other. For example, multiple loop filters can be connected in the order of LMCS, deblocking filter, SAO, and ALF. Furthermore, multiple loop filters can be connected in a sequence of all available permutations of the multiple loop filters. The output of one of the multiple loop filters can be used as the input of the next filter.

[0460] like Figure 6 As shown, an input image can be fed into a first filter. The input image can be a block as described in the embodiments. For example, the input image can be a reconstructed block generated by adder 129 or adder 169. The output of one filter can be the input of the next filter. The output image can be generated by the last filter. The output image can be a filtered block as described in the embodiments. For example, the output image can be a filtered reconstructed image generated by filter 130 or filter 170.

[0461] The target block can represent the image input to the filter. The target block for filtering can represent the image output from the filter.

[0462] LMCS can include luminance signal mapping of the luminance signal of the target block and chrominance signal scaling of the chrominance signal of the target block.

[0463] Luminance signal mapping can perform codeword reallocation for luminance signals.

[0464] Luminance signal mapping can include forward mapping and backward mapping. In forward mapping, the existing dynamic range can be partitioned into multiple intervals. The mapped dynamic range can be determined by performing codeword reallocation on the input image using a linear model for the corresponding intervals. In backward mapping, a reverse mapping is performed from the mapped dynamic range back to the existing dynamic range.

[0465] Chroma scaling can correct chroma signals based on the relationship between luminance signals and their corresponding chroma signals.

[0466] Forward mapping can be performed between inter-frame prediction and reconstruction of the luma signal, and between inter-frame prediction and chroma scaling of the luma signal. Inverse mapping can be performed between reconstruction and loop filtering of the luma signal. Chroma scaling can be performed between inverse transform and reconstruction of the chroma signal.

[0467] This structure allows for the inverse quantization, inverse transform, prediction, and reconstruction of luminance and chrominance signals within the mapped dynamic range. It also enables loop filtering, inter-frame prediction, intra-frame prediction, and reconstruction of chrominance signals within the existing dynamic range.

[0468] Deblocking filters remove block distortion that occurs at the boundaries between blocks in a reconstructed image. For example, a block can be a transform block. Furthermore, a block can be a sub-block of a specific block as described in the embodiments. Here, the boundary between blocks can refer to the sample points adjacent to the boundary between blocks.

[0469] Deblocking filters can be applied to both vertical and horizontal boundaries between blocks. After filtering the vertical boundaries between blocks, filtering can be applied again to the horizontal boundaries between the filtered blocks.

[0470] Deblocking filters can be applied selectively. The decision to apply a deblocking filter to a target block can be determined based on at least one of the following: a specific number of samples included in a specific number of columns or rows in the target block, and a specific number of samples included in a specific number of columns or rows in neighboring blocks adjacent to a specific boundary.

[0471] When a deblocking filter is to be applied to a target block, the filter to be applied can be determined based on the required intensity of the deblocking filter. In other words, a filter selected from several other filters based on the intensity of the deblocking filter can be applied to the target block. These multiple filters can include one of the following: a long-tap filter, a strong filter, a weak filter, or a Gaussian filter.

[0472] The maximum length of the deblocking filter can be determined based on the properties of the target block, such as the size of the target block, the components of the target block, and the coding parameters of the target block.

[0473] SAO can perform compensation for distortion between the original and reconstructed images based on samples. To compensate, SAO can apply an appropriate offset to the sample values ​​of the samples. That is, an offset can be added to the sample values.

[0474] Offsets can be determined for a target block. For example, offsets can be determined for each component of the CTB. The determined offsets can be applied to samples within a specific component of the CTB.

[0475] SAO can include SAOs that include edge offset (EO) and SAOs that use band offset (BO). Depending on the characteristics of the samples in a particular block, such as a CTU, it can be determined whether to perform an SAO using EO or an SAO using BO.

[0476] In SAO using EO, distortion correction in samples can be performed based on the orientation of edges within the target block. EO style categories can include horizontal, vertical, 135-degree diagonal, and 45-degree diagonal styles. For the target block, information indicating the style category to be applied and multiple offsets for that style category can be signaled. The number of offsets can be up to four. For target samples within the target block, adjacent samples can be determined based on the orientation of the style category. The offset to be applied to the target sample can be determined using the styles of the adjacent samples.

[0477] In the offset using BO, the luminance values ​​in the target block can be classified into specific bands, thus correcting distortion in the samples. The bit depth of the input image can be divided into m intervals. For example, m could be 32. The specific band can be n consecutive intervals among the m intervals. For example, n could be 4. N offsets for the n intervals can be transmitted using signals. Furthermore, information indicating the first interval selected from the n intervals among the m intervals can be transmitted using signals. The offset of the interval corresponding to the target sample can be added to the sample value of the target sample in the target cell.

[0478] ALF can perform compensation for distortion between the reconstructed image and the original image.

[0479] The filter coefficients of ALF can be transmitted as a signal via a bitstream.

[0480] The filter shape of an ALF can be determined by the components of the target block. For example, a 7×7 diamond filter can be used for the luminance component. A 5×5 diamond filter can be used for the chrominance component.

[0481] In ALF, the characteristics of a specific block can be determined, and the category of that block can be determined based on those characteristics. In other words, ALF determination and category determination can be performed on a 4×4 block basis. Filter coefficients can be calculated based on the category. A specific block can be a block with a 4×4 size.

[0482] Based on the orientation and activity determined using the gradient of a specific block, one of 25 categories can be chosen as the category of a particular block. Depending on the gradient of the particular block, rotational transformations, vertical symmetric transformations, and / or diagonal symmetric transformations of the filter can be applied.

[0483] Information related to whether or not ALF is applied can be sent via signals to specific units such as CTB.

[0484] The available signal transmission indicates the index of the filter in the available filters that will be applied to a specific unit. Here, the available filters may include fixed filters and filters configured using a parameter set. For example, the parameter set may be an adaptive parameter set (APS). Fixed filters may be defined equivalently by the encoding device 110 and the decoding device 150. The filter coefficients of filters configured using a parameter set can be determined based on the encoding parameters.

[0485] Entropy encoding and entropy decoding Figure 7 The entropy encoding and entropy decoding are shown based on the example.

[0486] exist Figure 7 The upper part shows the entropy encoding process performed by the entropy encoder 139.

[0487] The entropy encoder 139 may include a context modeler, a binarization unit, and an entropy coding unit. The context modeler may include a context selection unit and a context memory.

[0488] A binarization unit generates a binary number for each syntax element by performing binarization on the syntax elements of the target block. Binarization can be a process that converts syntax elements into binary form.

[0489] Information about syntax elements and binary numbers can be provided from the binarization unit to the context selection unit.

[0490] The context modeler can perform context updates.

[0491] Context can represent the probability of occurrence of each binary number of the encoded syntax element.

[0492] The context modeler can perform context updates to apply the current probability information to the encoding of the binary numbers of the syntax elements of the target block. The updated context can be stored in the context memory. Here, the updated context corresponding to the syntax elements of the target block (the binary numbers in the syntax elements of the target block) can be derived by the context modeler.

[0493] The context selection unit can select a context corresponding to the binary number of the syntax elements of the target block. The selected context can be loaded from the context memory and can be used as an updated context for entropy encoding of the binary number of the syntax elements of the target block.

[0494] The updated context can be used for entropy encoding of the syntactic elements of the target block.

[0495] An entropy encoder can generate encoded information of the syntax elements of a target block by performing entropy encoding using the generated binary number and the updated context, and can generate a bit stream including the encoded information. The entropy encoder can use at least one of arithmetic encoding methods and bypass encoding methods.

[0496] exist Figure 7 The lower part shows the entropy decoding process performed by the entropy decoder 161.

[0497] The entropy decoder 161 may include a context modeler, an entropy decoding unit, and a debinarization unit. The context modeler may include a context selection unit and a context memory.

[0498] The context modeler can perform context updates.

[0499] The context can represent the probability of occurrence of each binary number of a decoded syntax element.

[0500] The context modeler can perform context updates to apply the currently decoded probability information to the entropy decoding of the binary numbers of the syntax elements of the target block. The updated context can be stored in the context memory. Here, the updated context corresponding to the syntax elements of the target block (the binary numbers in the syntax elements of the target block) can be derived by the context modeler.

[0501] The context selection unit can select a context corresponding to the binary number of the syntax elements of the target block. The selected context can be loaded from the context memory and can be used as an updated context for entropy decoding of the binary number of the syntax elements of the target block.

[0502] The updated context can be used for entropy decoding of the syntactic elements of the target block.

[0503] An entropy decoder can generate the binary representation of the syntax elements of a target block by entropy decoding the encoded information of the bitstream based on an updated context. The entropy decoder can use at least one of arithmetic decoding methods and bypass decoding methods.

[0504] A debinarization unit can obtain the syntax elements of a target block by performing debinarization on at least one of the generated binary numbers. Debinarization can be a process that converts at least one binary number in the binary number into the form of a syntax element.

[0505] Information about syntax elements and binary numbers can be provided to the context selection unit from the debinarization unit.

[0506] Each syntax element may be one of the encoding parameters described in the embodiment.

[0507] Methods for binarization, debinarization, entropy coding, and entropy decoding In an embodiment, in order to perform signal transmission for specific information, one or more of the binarization methods, debinarization methods, entropy coding methods, and entropy decoding methods listed below can be used.

[0508] - Signed zero-order exponent Golomb binarization / inverse binarization method (abbreviated as se(v)) - Signed k-order Exp_Golomb binarization / inverse binarization method (abbreviated as sek(v)) - A 0th-order Exp_Golomb binarization / debinarization method for unsigned positive integers (abbreviated as ue(v)). - A k-order Exp_Golomb binarization / inverse binarization method for unsigned positive integers (abbreviated as uek(v)). - Fixed-length binarization / inverse binarization method (abbreviated as f(n)) - Truncation Rice binarization / inverse binarization method or truncation univariate binarization / inverse binarization method (abbreviated as tu(v)) - Truncation of binary binarization / inverse binarization method (abbreviated as tb(v)) -Context-adaptive arithmetic coding / decoding method (abbreviated as ae(v)) - A string of bits in bytes (abbreviated as b(8)) - Signed integer binarization / debinarization method (abbreviated as i(n)) - Unsigned positive integer binarization / debinarization method (abbreviated as u(n)) (where "u(n)" can represent a fixed-length binarization / debinarization method).

[0509] - Univariate binarization / inverse binarization methods Adaptive execution of processing in the embodiments The above-described processing can be performed using the same and / or corresponding methods in both the encoding device 110 and the decoding device 150. Furthermore, in image encoding and / or decoding, combinations of one or more of the methods described in the foregoing embodiments can be used.

[0510] The application order in the encoding device 110 and the decoding device 150 may be different. Alternatively, the application order in the encoding device 110 and the decoding device 150 may be (at least partially) the same.

[0511] The processing according to the embodiments can be performed on each processing target. The processing according to the embodiments can be performed equivalently on a particular target. For example, such a particular target could be a luminance signal and / or a chrominance signal.

[0512] Processing according to the embodiments can be selectively applied / executed based on specific conditions or specific objectives.

[0513] In embodiments, processes according to embodiments can be selectively applied / executed based on time layers. Time layer information for a specific process can be information indicating the time layer to which the process can be applied / executed. Time layer information can be sent via signals for a specific process. The time layer information can represent the lowest and / or highest layer to which a specific process can be applied, and can indicate the specific layer to which a specific process can be applied / executed. Optionally, a fixed time layer can be defined to which the process according to embodiments is applied / executed.

[0514] In an embodiment, the type of processing applied / executed according to the embodiment can be defined, and it can be determined whether to apply / execute processing according to the embodiment based on the defined type. Types may include frame type, strip type, parallel block type, etc.

[0515] As described in the embodiments, when applying / performing specific processing to a specific target, specific conditions may be required, and the specific processing may be performed under specific determinations. Determining whether a specific condition is met based on a specific encoded parameter, or when making a specific determination based on a specific encoded parameter, the specific encoded parameter can be interpreted as being replaced by another encoded parameter. In other words, the specific conditions or encoded parameters affecting specific decisions described in the embodiments can be considered exemplary, and it can be understood that, in addition to the specified encoded parameter, one or more other encoded parameters or a combination of one or more other encoded parameters perform the function of the specified encoded parameter.

[0516] The processing in the embodiments can be applied / executed based on the size of at least one of the blocks described in the embodiments. For example, a block may include an encoding block, a prediction block, a transform block, a reference block, a current block, and a target block. Optionally, in the embodiments, a block may include adjacent blocks. Here, the size can be defined as the minimum and / or maximum size for which the embodiments are applied, and can be defined as a fixed size for the processing in the embodiments. Furthermore, for the processing in the case of the embodiments, a first embodiment can be applied to a first size, and a second embodiment can be applied to a second size. That is, the processing of the embodiments can be applied in combination according to the size. In addition, the processing in the embodiments can be applied only when the size of the target is equal to or greater than the minimum size and less than or equal to the maximum size. That is, the processing in the embodiments can be applied only when the block size falls within a specific range.

[0517] The following sections will disclose improvements related to image encoding and decoding implemented in image encoding and image decoding devices.

[0518] In this disclosure, "the case where the indicator indicating whether a specific method is performed is true" can refer to the case where the indicator indicates whether a specific method is performed based on the prediction pattern, motion information, encoding parameters, and / or position. Similarly, "the case where the indicator indicating whether a specific method is performed is false" can refer to the case where the indicator indicating whether a specific method is performed is not true.

[0519] In this disclosure, "the case of executing a specific method / pattern" can refer to the case where the indicator indicating whether to execute a specific method / pattern has a first value. The first value can be true or 1.

[0520] In this disclosure, when a particular mode (or particular method) is not enabled in a particular unit, at least one of the syntax elements for the particular mode (or particular method) may be omitted in the particular unit and its sub-units for signal transmission / encoding / decoding.

[0521] A specific unit can be at least one of a video sequence, a frame, a subframe, a strip, a parallel block, a block, a coding tree unit (CTU), a coding unit (CU), a prediction unit (PU), a transform unit (TU), a coding block (CB), a prediction block (PB), or a transform block (TB); and the syntax element describing whether a specific mode is enabled in a specific unit can refer to at least one of a video parameter set, a decoding parameter set, a sequence parameter set, an adaptive parameter set, a frame parameter set, a frame header, a subframe header, a strip header, a parallel block group header, a parallel block header, a block, a CTU, a CU, a PU, a TU, a CB, a PB, or a TB.

[0522] A subunit can be a unit at a lower level than a specific unit. In other words, a subunit can refer to a region belonging to a specific unit or a unit that references information about a specific unit. For example, if the specific unit is a stripe, its subunits can be CTUs, CUs, PUs, or TUs that belong to and reference syntax elements of the corresponding stripe.

[0523] When a statement says that a particular mode (or method) is executed only under certain conditions, it can mean that the particular mode (or method) is enabled only when those conditions are met.

[0524] In this disclosure, "motion information improvement" and "motion information refinement" can be used interchangeably and can be substituted for each other.

[0525] In this disclosure, "motion information improvement value" and "motion information refinement vector" can be used interchangeably and can be substituted for each other.

[0526] In this disclosure, deriving second motion information from first motion information can mean obtaining second motion information by refining the first motion information.

[0527] In addition, inter-frame coding mode, IBC mode and intra-frame template matching mode share the following common feature: they use vectors to indicate the reference block that will be used to predict the target block within the reconstruction region used for target block prediction.

[0528] Therefore, in this embodiment, the inter-frame coding mode can be replaced with IBC mode or intra-frame template matching mode. The description of the inter-frame coding mode can also be applied to IBC mode or intra-frame template matching mode. Furthermore, information related to the inter-frame coding mode can be used as information related to IBC mode or intra-frame template matching mode.

[0529] Additionally, in the embodiments, the motion vector (MV) in the inter-frame coding mode can be replaced with the block vector (BV) in the IBC mode. The description for using MV for the target block can also be applied to the case where BV is used for the target block, and MV can be replaced with BV. Furthermore, information related to MV can be considered as information related to BV, and the description of information related to MV can be applied to information related to BV.

[0530] 1. Adaptive Motion Vector Resolution (AMVR) and Adaptive Block Vector Resolution (ABVR) In embodiments, the term "resolution" may refer to "motion vector resolution".

[0531] In adaptive motion vector resolution, the resolution of the motion vector difference can be adjusted in block units.

[0532] Adaptive motion vector resolution information indicates the resolution of the motion vector difference. The resolution of the motion vector difference of the target block can be determined by transmitting / encoding / decoding the adaptive motion vector resolution information.

[0533] Information related to adaptive motion vector resolution may include at least one of the following: an indicator indicating whether adaptive motion vector resolution is used, an index indicating the motion vector resolution, the number of resolution candidates in the adaptive motion vector resolution, and the type of resolution candidates in the adaptive motion vector resolution. For example, if the indicator indicating whether adaptive motion vector resolution is used in the target block is true, the index of the motion vector resolution in the target block can be signaled / encoded / decoded.

[0534] The resolution of motion vectors applicable to each block can be the same or different from each other. For example, the resolution of motion vectors applicable to a target block can be determined based on at least one of the target block's encoding parameters, motion information, and mode information.

[0535] Adaptive motion vector resolution can improve coding efficiency by adjusting the resolution of the motion vector difference.

[0536] For example, the adjusted resolution can be one of 16 pixels, 8 pixels, 4 pixels, full pixels, half pixels, and quarter pixels. However, this disclosure is not limited to a specific resolution, but can employ various motion vector resolutions.

[0537] When the adjusted resolution is n pixels, if the component value of the motion vector difference changes by 1, the position indicated by the motion vector difference can change by n pixels. In other words, when the adjusted resolution of the target block is n pixels, each component of the motion vector difference can indicate the reference block in units of n pixels.

[0538] For example, if the motion vector difference to be applied to the target block is (a, b) and the adjusted resolution is p pixels, (a / p, b / p) can be encoded instead of (a, b). In other words, the motion vector difference transmitted / encoded in the encoding device can be (a / p, b / p). The decoding device can derive the original motion vector difference (a, b) again by multiplying the transmitted motion vector difference (a / p, b / p) by p.

[0539] In the following text, the description of adaptive motion vector resolution can be applied equivalently to adaptive block vector resolution.

[0540] 2. Collective Partitioning (GPM) Geometric partitioning refers to a method that determines one or more partition boundaries in the form of straight lines or curves to divide a target block into multiple sub-blocks, and performs a weighted summation on the prediction block (or reference block) of each of the multiple sub-blocks generated for partitioning using a weighted graph determined based on the partition boundaries. For example, a target block can be partitioned into two sub-blocks by straight lines with arbitrary directions or angles.

[0541] In geometric partitioning mode, individual sub-blocks can be predicted based on various patterns.

[0542] In some embodiments, a prediction block may refer to a prediction block obtained from unidirectional inter-frame prediction and / or bidirectional inter-frame prediction, or a reference block in at least one direction of unidirectional inter-frame prediction and / or bidirectional inter-frame prediction. In another embodiment, a prediction block may refer to a prediction block obtained from intra-frame prediction. In yet another embodiment, one prediction block may be an inter-frame predicted block, and the other prediction block may be an intra-frame predicted block. Furthermore, IBC mode, intra-template matching mode, or affine mode can be used to generate prediction blocks for each sub-block generated in geometric partitioning mode.

[0543] Figure 8 This shows the partition boundaries in the geometric partitioning pattern; Figure 9 This shows the partition boundaries, partition offsets, and partition angles in the geometric partitioning mode; Figure 10The weight map for each prediction block along a specific boundary in the geometric partitioning pattern is shown.

[0544] In geometric partitioning mode, partition boundaries can be determined based on at least one of partition offset and partition angle. Information regarding partition boundaries in geometric partitioning mode can be information used to specify at least one of partition offset and partition angle.

[0545] Information regarding the geometric partitioning pattern of the target block may include at least one of an indicator indicating whether the geometric partitioning pattern is executed, information about the partition boundaries, and information for generating prediction blocks corresponding to the sub-blocks of the partition. Furthermore, information regarding the geometric partitioning pattern may include information for combining the prediction blocks corresponding to the individual sub-blocks, such as blending information.

[0546] Indicators can be used to send / encode / decode signals to indicate whether a geometric partitioning mode is to be performed on the target block.

[0547] When performing a geometric partitioning mode on a target block, information about the partition boundaries in the target block can be sent / encoded / decoded using signals.

[0548] In some embodiments, the image encoding device can be configured to indicate multiple partition types (e.g., Figure 8 The index information of one partition type of the target block is encoded in the partition type shown, and the encoded index information is sent to the image decoding device; the image decoding device can determine the partition type of the target block among multiple partition types based on the index information.

[0549] In another embodiment, information indicating at least one of partition offset and partition angle (or direction) can be encoded. The image decoding device can specify the partition boundaries of the target block based on this information received from the image encoding device.

[0550] Optionally or additionally, information regarding the partition boundaries of the target block can be determined based on the motion information of the target block, the intra-coding method information of the target block, the coding parameters of the target block, the intra-coding method information of neighboring blocks, the coding parameters of neighboring blocks, neighboring samples of the target block, and at least one neighboring sample of a reference block of the target block. Furthermore, when the target block is a chroma component block, the luma component samples of the target block (at least one sample of the luma component block or region corresponding to the target block) can be additionally considered. The motion information or intra-coding method information of the target block can be information obtained by template matching for the target block. Optionally, the motion information or intra-coding method information of the target block can be motion information or intra-coding method information used for prediction of each sub-block.

[0551] When performing a geometric partitioning mode on a target block, signal transmission / encoding / decoding can be used to generate prediction information for at least one sub-block in the partitioned sub-blocks.

[0552] Optionally or additionally, prediction information about at least one sub-block can be determined based on at least one of the following: motion information of the target block, intra-frame coding method information of the target block, coding parameters of the target block, intra-frame coding method information of neighboring blocks, coding parameters of neighboring blocks, neighboring samples of the target block, and neighboring samples of a reference block of the target block. Furthermore, when the target block is a chroma component block, luma component samples of the target block (at least one sample from the luma component block or region corresponding to the target block) can be additionally considered.

[0553] The motion information or intra-coding method information of the target block can be a motion vector or block vector obtained by template matching over the target block. Optionally, the intra-coding method information of the target block can be information indicating whether IBC mode or intra-template matching mode is applied to the target block. For example, the prediction information available for each sub-block can be determined based on the type of intra-coding method applied to or allowed for the target block.

[0554] The intra-frame coding method information for neighboring blocks may include information about the geometric partitioning pattern. For example, when the geometric partitioning pattern is applied to neighboring blocks, at least one of the following can be derived from the correspondence information of neighboring blocks: information about the partition boundaries of the target block, prediction information about at least one sub-block, and information for combining the prediction blocks corresponding to each sub-block.

[0555] If inter-frame prediction is used to predict at least one sub-block from the sub-block partitioned from the target block, then at least one of the motion information of the target block can be refined.

[0556] At least one of the following information can be transmitted / encoded / decoded using signals: information indicating whether to refine the motion information of a target block for which a geometric partitioning pattern has been applied, and information specifying the motion information or orientation of the target to be refined. Optionally, the above information can be determined based on the motion information of the target block, the intra-frame coding method information of the target block, the coding parameters of the target block, the intra-frame coding method information of neighboring blocks, the coding parameters of neighboring blocks, the neighboring samples of the target block, and at least one neighboring sample of a reference block of the target block. Furthermore, when the target block is a chroma component block, the luma component samples of the target block (at least one sample of the luma component block or region corresponding to the target block) can be additionally considered.

[0557] Motion information refinement can be performed through template matching, bilateral matching, application (addition) of motion offset, and at least one of specific operations.

[0558] A specific operation may include at least one of mirroring, scaling, and copying, but the type of operation is not limited to this specific example.

[0559] The result of mirroring a specific motion vector MV can be -MV.

[0560] The result of scaling a specific motion vector MV can be a motion vector whose magnitude is refined by considering the following POC intervals: the POC interval between the image to which the block with the motion vector MV belongs and the reference image indicated by the MV; and the POC interval between the image to which the target block belongs and the current reference image. In this case, the direction of the specific motion vector MV and the direction of the motion vector obtained by applying scaling to the MV can be the same as or opposite to each other.

[0561] For example, the result of copying a specific motion vector MV can be MV.

[0562] If the target block is based on bidirectional prediction, the motion information along the LX direction (one of the L0 and L1 directions) can be refined first, and then the corresponding information can be used to refine the motion information along the L(1-X) direction.

[0563] For example, refining motion information only along the LX direction can mean, but is not limited to, searching for a motion information refinement vector that minimizes the matching cost in the LX direction.

[0564] At this point, X can be 0, 1, or a positive integer.

[0565] For example, information about X can be determined through encoding / decoding.

[0566] For example, the LX direction can refer to the direction along which the L0 and L1 directions result in a lower / higher matching cost.

[0567] Matching cost may refer to, but is not limited to, at least one of template matching cost and bilateral matching cost.

[0568] For example, after determining the motion information refinement vector MV along the LX direction diff Subsequently, the motion information refinement vector along the L(1-X) direction can be determined as MV. diff For example, after determining the motion information refinement vector MV along the LX direction... diff Subsequently, the motion information compensation vector in the L(1-X) direction can be determined by adjusting the MV relative to the motion information along the L(1-X) direction. diff The motion vector obtained by scaling.

[0569] If the target block is based on bidirectional prediction, then only the motion information along the LX direction, which is one of the L0 and L1 directions, can be refined.

[0570] If the target block is based on bidirectional prediction, indicators that specify whether motion information refinement is performed along one or both directions can be encoded / decoded. For example, indicators that specify whether refinement is performed along each direction can be encoded / decoded.

[0571] The matching cost for specific motion information can be replaced with the matching cost for motion information after replacing the motion vector in the corresponding motion information with ROUNDED_MV.

[0572] For example, the matching cost for a specific motion vector can be replaced with the matching cost for ROUNDED_MV.

[0573] ROUNDED_MV can be obtained through ROUND(MV, TARGET_PRECISION).

[0574] Here, ROUND(MV,TARGET_PRECISION) can be a function that rounds MV to TARGET_PRECISION units, and MV can refer to a motion vector of specific motion information.

[0575] For example, when the original precision of MV is ORIG_PRECISION, ROUND(MV, TARGET_PRECISION) can produce the same result as by changing the precision of MV to TARGET_PRECISION and then changing the precision back to ORIG_PRECISION.

[0576] ORIG_PRECISION and TARGET_PRECISION can respectively refer to (but are not limited to) one of 16 pixels, 8 pixels, 4 pixels, full pixels, half pixels and quarter pixels.

[0577] 3. Template matching Figure 11 An example of template matching is shown.

[0578] In template matching, the motion information of the target block can be determined and / or modified based on the result of calculating the cost function between the target template of the target block and the reference template of the reference block.

[0579] In some embodiments, template matching can be used to determine the motion information of a target block. Based on the cost calculation between the target template and each reference template, the motion information corresponding to the displacement from the position of the target block to the position of the reference block with the lowest cost reference template can be used as the motion information of the target block.

[0580] In another embodiment, template matching can be used to modify or improve the motion information of the target block. For example, a reference template for a reference block located at the position indicated by the initial motion information for the target block and at a position separated from that position by a predetermined direction and interval can be determined. Then, based on the cost calculation between the target template and each reference template, the motion information corresponding to the displacement from the target block to the reference block with the lowest cost reference template is determined as the final motion vector of the target block.

[0581] In another embodiment, template matching can be used to reorder motion information candidates included in a list of motion information candidates for a target block. For example, a reference template for a reference block is configured at the location indicated by each of the motion information candidates. The motion information candidates are then reordered in ascending order of cost based on a cost calculation between the target template and the reference template. This reordering allows candidates with lower indices to be assigned higher probabilities of being selected as motion information for the target block.

[0582] In template matching, a reference block may include at least one of the following: a reference block pointed to by the initial motion information, a reference block pointed to by the motion information derived during the template matching search process, and a reference block pointed to by the motion information finally improved by template matching.

[0583] Here, the initial motion information can be the motion information of the target block sent from the encoding device to the decoding device via a signal. The motion information improved by template matching can be the motion information with the lowest matching cost derived during the template matching search process. However, the method for deriving motion information is not limited to the above criteria.

[0584] The template matching method may include at least one of inter-frame template matching mode and intra-frame template matching mode.

[0585] Inter-frame template matching mode can refer to a template matching method that configures a reference block based on at least one of the predicted samples, reconstructed samples, and residual samples of a reference frame that has been reconstructed before the target frame.

[0586] Intra-frame template matching mode can refer to a template matching method that configures a reference block based on at least one of the predicted samples, reconstructed samples, and residual samples of the target image.

[0587] Template matching cost can be obtained by calculating the cost using a cost function between the templates of the target block and the reference block used in template matching. Template matching cost can refer to the template matching cost of the displacement between the target block and the reference block used in template matching; that is, the template matching cost of motion information.

[0588] When the target block is a chroma component block, the template matching cost can be determined based on the sample values ​​of the reconstructed luminance component region corresponding to the chroma component block.

[0589] As an exemplary embodiment, when the target block is a chroma component block, the template matching cost may refer to the cost between the sample points within the luminance component region corresponding to the reference chroma component block determined based on the initial motion information of the target chroma component block and the sample points within the luminance component block corresponding to the target chroma component block. Here, the initial motion information may be a zero vector or a motion vector obtained by scaling a predetermined motion vector of the luminance component region corresponding to the target chroma component block based on the chroma component format. Here, the initial motion information is a zero vector indicating that the position of the target chroma component block in the target image is the same as the position of the reference chroma component block in the reference image.

[0590] The luma component region corresponding to the target chroma component block can be considered as a luma component block, and template matching can be performed on the luma component block. Then, the motion information of the target chroma component block can be determined based on the motion information exhibiting the minimum cost.

[0591] When there are multiple luminance component blocks in the luminance component region corresponding to the target chrominance component block, the motion information of the target chrominance component block can be determined by performing template matching on at least one of the multiple luminance component blocks.

[0592] In another embodiment, the template matching cost of the target chromaticity component block can be obtained by calculating the cost between the luminance component region corresponding to the reference chromaticity component block or the reference template of the reference chromaticity component block and the luminance component region corresponding to the target template of the target chromaticity component block or the target template of the target chromaticity component block.

[0593] 3.1 Template Configuration The templates used for template matching can include a target template and a reference template.

[0594] The target template can be configured using a reference region that includes sample points around the target block. The reference region of the target block can include at least one of the sample points located in the lower left, left, upper left, upper, and upper right regions of the target block.

[0595] In some embodiments, the target template may be the same as the reference area of ​​the target block.

[0596] In some other embodiments, when a target template is configured for template matching, a subset of sample points within a reference region of the target block can be selected. The selected sample points can then be used to configure the target template.

[0597] A reference template can be configured using a reference region that includes neighboring samples of the reference block. The reference region of the reference block can be a region corresponding to the reference region of the target block. For example, the reference region of the reference block can include at least one of the samples located in the lower left, left, upper left, upper, and upper right regions of the reference block.

[0598] In some embodiments, the reference template used for template matching may be the same as the reference region of the reference block. For example, the sample points of the reference template based on the reference block may be sample points corresponding to the sample points of the target template based on the target block.

[0599] In some other embodiments, when a reference template is configured for template matching, a subset of sample points within a reference region of a reference block can be selected. These selected sample points can then be used to configure the reference template. For example, the sample points selected for configuring the reference template based on the reference block may correspond to sample points selected for configuring the template based on the target block.

[0600] In the inter-frame template matching mode, each of the reference block, reference template, and reference region consists of at least one of the predicted sample, reconstructed sample, and residual sample of the reference frame reconstructed before the target frame.

[0601] In intra-frame template matching mode, each of the reference block, reference template, and reference region consists of at least one of the predicted sample, reconstructed sample, and residual sample of the target image.

[0602] The target / reference template used for template matching may include at least one of the following: 1) at least one sample in the TMSIZE_LEFT bar adjacent to the left side of the target / reference block; and 2) at least one sample in the TMSIZE_ABOVE bar adjacent to the top of the target / reference block.

[0603] However, the positional relationship between each sample point in the template and the target / reference block, and / or the method of configuring the template, are not limited to the relationships or methods described above.

[0604] Each of TMSIZE_LEFT and TMSIZE_ABOVE can be a positive integer greater than or equal to 0, 1, 2, 3, 4 or 4.

[0605] TMSIZE_LEFT and TMSIZE_ABOVE can be equal to each other. Alternatively, TMSIZE_LEFT and TMSIZE_ABOVE can be different from each other.

[0606] Each of TMSIZE_LEFT and TMSIZE_ABOVE can be a predefined value or a value determined based on information sent / encoded / decoded by signals.

[0607] Each of TMSIZE_LEFT and TMSIZE_ABOVE can be determined based on at least one of the target block's motion information, encoding parameters, size, and prediction mode.

[0608] For example, when the width of the target block is W and its height is H, if the smaller (or larger) of W and H is less than (or greater than) TMSIZE_THRES, then TMSIZE1 can be used as TMSIZE_LEFT / TMSIZE_ABOVE; otherwise, TMSIZE2 can be used as TMSIZE_LEFT / TMSIZE_ABOVE.

[0609] TMSIZE1 and TMSIZE2 can be predefined values.

[0610] TMSIZE1 and TMSIZE2 can be, but are not limited to, 0, 1, 2, 4 or positive integers.

[0611] To configure the template, interpolation filtering can be performed.

[0612] The interpolation filter used to configure the template may be the same as or different from the interpolation filter used for motion compensation of the predicted block.

[0613] Differences between interpolation filters can arise from variations in at least one of the interpolation filter type, interpolation filter taps, and interpolation filter coefficients; however, the criteria used to determine differences between interpolation filters are not limited to specific factors.

[0614] For example, to reduce computational complexity, interpolation filters with fewer taps than those used for motion compensation of predicted blocks can be used to configure templates.

[0615] In another embodiment, reference sample filtering can be performed to configure the template.

[0616] The reference sample filter used to configure the template may be the same as or different from the reference sample filter used for intra-frame prediction of the prediction block.

[0617] For example, differences between reference sample filters may arise from variations in at least one of the reference sample filter type, reference sample filter taps, and reference sample coefficients; however, the criteria used to determine differences in reference sample filters are not limited to specific factors.

[0618] For example, to reduce computational complexity, a reference sample filter with fewer taps than the reference sample filter used for intra-frame prediction of the prediction block can be used to configure the template.

[0619] Alternatively, multiple target templates can be used to perform template matching on a target block. Template matching using multiple target templates can be performed as follows.

[0620] By performing template matching using each individual target template, the matching cost for each target template can be calculated. Regarding the matching cost for a target template, the lowest cost obtained from performing template matching using the target template can be referred to as the cost for the reference block or the cost indicating the motion information of the reference block.

[0621] Based on the matching cost for each target template, a predicted block for the target block can be generated using one of the reference blocks determined by motion information derived from each target template. For example, the reference block with the lowest matching cost can be determined as the predicted block for the target block.

[0622] Optionally, based on the matching cost for each target template, a predicted block for the target block can be generated using at least two reference blocks determined by motion information derived from each target template.

[0623] For example, a predicted block for a target block can be generated by a weighted sum of multiple reference blocks. The weights to be applied to each reference block can be determined based on the matching cost for each reference block. For example, higher weights can be applied to reference blocks with lower matching costs.

[0624] Whether to use a reference block determined by motion information derived from a specific target template to generate a prediction block for the target block can be determined based on at least one of the following: the size of the target block, the motion information of the target block, the encoding mode of the target block, the encoding parameters of the target block, the matching cost for the specific target template, and the matching cost for at least one other target template for the target block.

[0625] In an exemplary embodiment, it can be determined whether to use a reference block determined by motion information derived from a first target template to generate a predicted block of a target block based on the matching cost of a second target template that has the lowest matching cost among a plurality of target templates. For example, if the cost difference between the first target template and the second target template is less than or equal to a predetermined threshold, a reference block derived from the first target template can be used.

[0626] When performing template matching using multiple target templates, the size of the template regions used can always be the same.

[0627] Optionally, when template matching is performed using multiple target templates, the sizes of the template regions used can be different. In this case, the matching cost for each target template can be replaced with the result obtained by multiplying or dividing the matching cost for the corresponding target template by a predetermined value.

[0628] The predetermined value can be determined based on at least one of a plurality of target templates for the target block (e.g., the size of a region of at least one of a plurality of target templates for the target block), the size of the target block, the motion information of the target block, the encoding mode of the target block, and the encoding parameters of the target block.

[0629] If the cost function used to calculate the matching cost is based on the sum of errors among the samples within the target template rather than the average of the errors, the matching cost can increase proportionally with the size of the template region. In this case, directly comparing the matching costs for a first template and a second template occupying regions of different sizes may be inappropriate. Therefore, by taking into account the size differences or ratios between template regions, the matching costs calculated for template regions of different sizes can be normalized to the scale of regions of the same size. For example, the matching cost can be scaled to a predetermined value determined by taking into account the size of the template region. Scaling can be performed by multiplication or division between the matching cost and the predetermined value.

[0630] Figure 12 The target template that can be used for the target block is shown.

[0631] The target template can be configured using one or more lines consisting of reconstructed sample points around the target block.

[0632] For example, such as Figure 12 As shown in (a), at least one of the templates labeled 0 to 7 can be used as the target template for the target block.

[0633] Optionally, such as Figure 12 As shown in (b), at least one of the templates labeled 0 to 3 can be used as the target template for the target block.

[0634] 3.2 Template matching of sub-block units Template matching can be performed on sub-block units. The target block can be divided into N×M sub-block units, and template matching can be performed on the sub-block units to determine motion information.

[0635] N and M can be 2, 4, 8, 16, 32, 64, or positive integers. N and M can be predefined values. Optionally, information about at least one of N and M can be transmitted / encoded / decoded using signals, and at least one of N and M can be determined based on this information. Optionally, at least one of N and M can be determined based on at least one of the following: the size of the target block, the motion information of the target block, the intra-frame coding method information of the target block, and the coding parameters of the target block.

[0636] If the target block is predicted in template matching mode using sub-block units, each sub-block can be treated as a block, and then template matching can be performed.

[0637] In some embodiments, the target template for each sub-block can be constructed using neighboring samples of the target block and may exclude samples within the target block. In this case, template matching and prediction for each sub-block can be performed in parallel.

[0638] Figure 13 This is an exemplary embodiment illustrating the target template for each sub-block in a template matching pattern based on sub-block units.

[0639] like Figure 13 As shown, the reconstructed sample points adjacent to the left and / or top of each sub-block can be used as target templates.

[0640] If the reconstructed sample points immediately adjacent to the sub-block do not exist within the target block, then reconstructed sample points adjacent to the target block, other than those in the sub-block, can be used instead as sample points for forming the target template of the sub-block. Alternatively, reconstructed reference sample points around previously predicted or reconstructed sub-blocks can be used.

[0641] In some other embodiments, a target template for each sub-block can be constructed by including sample points within the target block; in this case, template matching and prediction for each sub-block can be performed sequentially.

[0642] Figure 14 This is another exemplary embodiment illustrating the target template for each sub-block in a template matching pattern based on sub-block units.

[0643] like Figure 14 As shown, the target template for each sub-block can be composed of samples adjacent to the corresponding sub-block. In this embodiment, each sub-block can be predicted or reconstructed at the sub-block unit level. Therefore, when template matching is performed for a specific sub-block, predicted or reconstructed samples may exist in sub-blocks where template matching has already been performed. Thus, the target template for a specific sub-block can be composed of predicted or reconstructed samples from previously predicted or reconstructed sub-blocks.

[0644] In sub-block-based template matching, template matching can be performed using sub-block units from each of two or more (N, M) combinations.

[0645] The final (N, M) combination is determined based on the matching cost calculated for each (N, M) combination, and the target block can be predicted based on the result of template matching performed using the determined (N, M) combination.

[0646] Information about template matching patterns at the sub-block level can be sent / encoded / decoded using signals. For example, an indicator can be sent / encoded / decoded to show whether a template matching pattern is performed on the target block at the sub-block level.

[0647] Optionally, whether to perform template matching mode on the target block at the sub-block unit level can be determined based on at least one of the following: motion information of the target block, intra-frame coding method information of the target block, coding parameters of the target block, intra-frame coding method information of neighboring blocks, coding parameters of neighboring blocks, neighboring samples of the target block, and at least one neighboring sample in the reference block of the target block, without involving the signal transmission / encoding / decoding of information. If the target block is a chroma component block, the luma component samples of the target block (at least one sample in the luma component block or region corresponding to the target block) can also be considered.

[0648] For example, when the prediction mode of the target block is inter-frame prediction mode, template matching at the sub-block unit level can be performed, and affine mode and / or template matching mode are applied to the target block. Optionally, template matching at the sub-block unit level can be performed based on the motion information of the target block generated by template matching at the full block unit level or by signal transmission. For example, when the merge index is greater than or equal to a certain value, template matching at the sub-block unit level can be performed to refine the motion information of the target block.

[0649] In another example, when the intra-coding method information of the target block indicates IBC mode or intra-template matching mode and indicates that prediction is performed at the sub-block unit level, template matching can be performed at the sub-block unit level. Optionally, when the merge index of the target block to which IBC merge mode is applied is greater than or equal to a specific value, template matching at the sub-block unit level can be performed to refine the block vector of the target block.

[0650] When predicting a target block using a template matching pattern, information about the template matching pattern of sub-block units can be sent / encoded / decoded using signals.

[0651] Alternatively, information about the template matching pattern of the sub-block unit can be sent / encoded / decoded using signals, regardless of whether the target block is predicted in a template matching pattern.

[0652] Whether to perform template matching mode for sub-block units on the target block and / or whether to send / encode / decode information about template matching mode for sub-block units can be determined based on at least one of the following: the target block's coding mode, motion information of the target block, intra-frame coding method information of the target block, coding parameters of the target block, intra-frame coding method information of neighboring blocks, coding parameters of neighboring blocks, neighboring samples of the target block, and at least one neighboring sample of the target block's reference block. If the target block is a chroma component block, the luma component samples of the target block (at least one sample in the luma component block or region corresponding to the target block) can also be considered.

[0653] For example, template matching mode by sub-block units can only be performed on a target block if the size of the target block is greater than or equal to a threshold, and information about the template matching mode by sub-block units can be sent / encoded / decoded using signals. The threshold can be 4, 8, 16, 32, 64, 128, or a positive integer.

[0654] 3.3 Template Matching in Geometric Partitioning Patterns If the prediction mode of the target block is a geometric partitioning mode, the template to be used to generate the prediction block for each partition region can be determined based on the partitioning information of the geometric partitioning mode.

[0655] The partition information of a geometric partitioning pattern refers to the information used to distinguish partition regions in the geometric partitioning pattern, and may be, but is not limited to, at least one of partition angle, partition offset, and index used to specify partition information from the partition information candidate list.

[0656] The partition angle can refer to, but is not limited to, the angle between the partition line and the X-axis or Y-axis.

[0657] Partition offset can refer to the distance from a specific sample point location relative to the target block to a straight line. For example, a specific sample point location relative to the target block can be, but is not limited to, the center of the target block, the upper left of the target block, the lower left of the target block, the upper right of the target block, and the lower right of the target block.

[0658] Figure 15 This is an exemplary diagram illustrating a method for configuring a target template for each sub-block of a target block to which a geometric partitioning pattern has been applied.

[0659] For example, if the prediction mode of the target block is a geometric partitioning mode and the partitioning angle is less than or equal to a specific angle, then the prediction block for the X_PARTITION partition region can be configured with a template using only one of the left or upper sample set based on the target block.

[0660] For example, if the prediction mode of the target block is a geometric partitioning mode and the partitioning angle is greater than or equal to a specific angle, then the prediction block for the X_PARTITION partition region can be configured with a template using only one of the left or upper sample sets based on the target block.

[0661] X_PARTITION can be 1, 2, or a positive integer.

[0662] like Figure 15 As shown, the area above the partition line can be referred to as the first partition area, and the area below it can be referred to as the second partition area. The partition line is a straight line specified by the partition information, which means the straight line that serves as the boundary between partition areas in the geometric partitioning pattern.

[0663] The template for the prediction block in the first partition region may consist only of samples belonging to the upper template region relative to the target block, and the template for the prediction block in the second partition region may consist only of samples belonging to the left template region relative to the target block.

[0664] Figure 16 This is another exemplary diagram illustrating a method for configuring a target template for each sub-block of a target block to which a geometric partitioning pattern has been applied.

[0665] For example, if the target block is based on a geometric partitioning pattern and the partitioning angle is less than / greater than or equal to a specific angle, the template for the predicted block of the X_PARTITION partition region can be configured using either the left sample set or the upper sample set based on the target block.

[0666] When Figure 16 When partitioning the target block as shown, the template for the prediction block of the first partition region can be composed of sample points belonging to the left template region and the upper template region relative to the target block, and the template for the prediction block of the second partition region can be composed only of sample points belonging to the left template region relative to the target block.

[0667] Figure 17 This is yet another exemplary diagram illustrating a method for configuring a target template for each sub-block of a target block to which a geometric partitioning pattern has been applied.

[0668] For example, if the prediction mode of the target block is a geometric partitioning mode and partitioning information is specified, the partitioning information can be extended to the template region, and the template for the prediction block for the corresponding region can be configured using only the samples of the region corresponding to each extended partition region.

[0669] When Figure 17 When partitioning the target block as shown, the template for the prediction block of the first partition region can be composed of samples belonging to the first extended partition region, and the template for the prediction block of the second partition region can be composed only of samples belonging to the second extended partition region.

[0670] 3.4 Subsampling in Template Matching In template matching, subsampling can be performed to configure the template, calculate the cost, etc. Furthermore, subsampling can be performed on the search region used for template matching.

[0671] ■ Subsampling for configuring templates Subsampling can be used to configure a template for template matching. In other words, when configuring a template, all samples in the reference region can be used, or only a portion of the samples in the reference region can be used. Here, the reference region can refer to at least one of the region referenced by the target block and the region referenced by the reference block used to configure the template.

[0672] When configuring a template for template matching using only a portion of the sample points located in the reference region, subsampling can be performed on the entire reference region or a portion of the reference region.

[0673] When configuring a template for template matching using only a portion of the sample points located in the reference region, the reference region can be partitioned into two or more regions.

[0674] In some embodiments, the partitioned region can be one of the following: 1) a region where subsampling is performed, 2) a region used for configuring a template without involving subsampling, and 3) a region not used for configuring a template. A template for template matching can be configured using samples selected from region 1) through subsampling and samples within region 2).

[0675] In one example, the region corresponding to case 1) could be the region within the reference region that belongs to the left and / or upper left of the block.

[0676] In another example, the region corresponding to case 1) could be the region above and / or to the upper left of the block within the region of the reference area.

[0677] In some other embodiments, the partitioned region can be one of the following: 1) a region where subsampling is performed and 2) a region not used for template configuration. The template for template matching can be configured using the sample points selected by subsampling within region 1).

[0678] ■ Subsampling for cost calculation When calculating the cost between a target template and a reference template in template matching, all samples from each template can be used, or only a subset of samples from each template can be used. In other words, the cost between templates can be calculated using only a subset of samples from each template.

[0679] To calculate the cost between templates using only a subset of samples from the template, subsampling can be performed on the entire or a portion of the template region.

[0680] When performing cost calculations between templates using only a subset of samples from the template, the template region used for template matching can be partitioned into two or more regions.

[0681] In some embodiments, the partitioned region can be one of the following: 1) a region where subsampling is performed, 2) a region used to compute the cost function without subsampling, and 3) a region not used to compute the cost function. The cost function between templates for template matching can be computed using samples selected by subsampling from region 1) and samples within region 2).

[0682] In another embodiment, the partitioned region can be one of 1) a region where subsampling is performed and 2) a region not used for calculating the cost function. The cost function between templates for template matching can be calculated using samples selected by subsampling within region 1).

[0683] ■ Subsampling of the search area Template matching search processing can be performed using a reference template included in a search region within a predetermined range from the sample point position indicated by the initial motion vector of the target block. In this case, the search processing can use all samples / positions within the search region, or it can select only a portion of the samples / positions within the search region. The search and / or matching cost can be calculated only for the selected samples / positions or for the motion information indicating the selected samples / positions.

[0684] When performing template matching search processing using only a portion of the sample points / locations within the search area, subsampling can be performed on the entire or a portion of the search area.

[0685] When performing template matching search processing using only a subset of samples / locations within the search area, the search area can be divided into two or more regions.

[0686] In some embodiments, the partitioned region can be one of the following: 1) a region where subsampling is performed, 2) a region where search processing is performed without subsampling, and 3) a region where search processing is not performed. Template matching search processing can be performed on pixels and / or selected locations within region 1) selected by subsampling and on sample points / locations within region 2). Optionally, template matching search processing can be performed on sample points / locations within region 1) selected by subsampling and on motion information indicating sample points / locations within region 2).

[0687] In some other embodiments, the partitioned region can be one of 1) a region where subsampling is performed and 2) a region where no search processing is performed. Template matching search processing can be performed using sample points / locations selected by subsampling within region 1). Alternatively, template matching search processing can be performed on motion information indicating sample points / locations selected by subsampling within region 1.

[0688] Figure 18 and Figure 19Various embodiments of subsampling methods for template matching are shown.

[0689] Figure 18 and Figure 19 The shaded sample points (or positions) in the text represent sample points (or positions) selected through subsampling.

[0690] By means of, such as Figure 19 Subsampling is performed on all or part of the reference area shown, and the template is configured using only the sample points (or locations) selected through subsampling.

[0691] By means of, such Figure 19 The cost calculation is performed by subsampling all or part of the template region matched by the template shown, and only by subsampling the selected sample points (or locations).

[0692] In addition, it can be used for such Figure 18 The template matching search region shown is subsampled in all or part, and the search and / or matching cost can be calculated only for the selected pixels and / or locations or only for the motion information indicating the selected pixels and / or locations.

[0693] 3.5 Template Matching Search Method By using first motion information encoded in or decoded from a bitstream as initial motion information, a first search step can be performed to derive second motion information, which is a refinement of the first motion information. A second search step, performed after the first search step, can use the second motion information as initial motion information.

[0694] If the initial motion information (e.g., the initial motion vector or the initial block vector) is not motion information in integer pixel units (i.e., in fractional pixel units), the search can be performed using the result obtained by rounding (or truncating or rounding up) the initial motion information.

[0695] To generate a reference template at the location indicated by motion information obtained by adding a specific offset to initial motion information in fractional pixel units during the search process, an interpolation filter must be applied to samples at integer pixel locations to generate samples at fractional pixel locations. However, if the initial motion information is limited to integer pixel units, the computational complexity can be reduced because interpolation at fractional pixel locations is not required during the search process.

[0696] ■ Definition of Search The search can be performed using the calculation of the cost function to determine the similarity between NUM_TEMPLATE_COMPARE templates.

[0697] The search may include the process of determining at least one piece of motion information that meets specific conditions within a specific search range. The motion information of the target block may be determined and / or modified based on the at least one piece of motion information determined through the search.

[0698] Motion information that meets specific conditions can represent, but is not limited to, motion information with the lowest matching cost among motion information within the search range.

[0699] ■ Cost Function The cost function used to calculate the cost can refer to a function that determines the similarity between at least one sample in the target template and at least one sample in the reference template.

[0700] The similarity between the first and second values ​​can be determined using at least one of the following operations: 1) the difference between the two values, 2) the ratio between the two values, and 3) comparing the difference between the two values ​​with a specific value.

[0701] In one embodiment of the operation of comparing the difference between two values ​​with a specific value, a method can be used that assigns multiple similarity values ​​based on which segment the difference between the two values ​​belongs to, where the segment is defined by multiple thresholds. For example, when using a single threshold, the method can be used such that if the difference between the two values ​​is less than or equal to the threshold, the method assigns a first similarity value (e.g., 1), otherwise assigns a second similarity value (e.g., 0).

[0702] The cost function can be one or more of the following: Sum of Absolute Differences (SAD), Sum of Absolute Transformed Differences (SATD), Sum of Absolute Differences with Mean Removed (MR-SAD), Mean Squared Error (MSE), and Sum of Squared Errors (SSE). However, the cost function is not limited to the terms listed above.

[0703] The cost function used in template matching can be predefined or determined based on information sent / encoded / decoded by signals.

[0704] The cost function used in template matching can be determined based on at least one of the following: whether bilateral matching is performed, the conditions associated with bilateral matching, whether bidirectional prediction with CU weights is performed, and the size of the target block.

[0705] In some embodiments, MR-SAD can be used as a cost function for template matching if the target block satisfies one or more of the enable conditions for bilateral matching described below, or if bilateral matching is performed on the target block.

[0706] In other embodiments, SAD can be used as a cost function for template matching if the target block does not meet one or more of the enable conditions for bilateral matching, or if bilateral matching is not performed on the target block.

[0707] In another embodiment, the type of cost function in template matching can be determined based on specific conditions used to determine the type of cost function in bilateral matching. The type of cost function in bilateral matching can be determined based on whether specific conditions are met in bilateral matching. In this case, the type of cost function in template matching can be determined based on whether the enabling conditions of bilateral matching and the specific conditions used to determine the type of cost function in bilateral matching are met.

[0708] For example, if the target block meets the enable conditions for bilateral matching and the specific conditions for determining the cost function for bilateral matching, then MR-SAD can be used as the cost function for template matching; otherwise, SAD can be used as the cost function for template matching.

[0709] In another embodiment, MR-SAD can be used as the cost function in template matching if the target block meets the enable conditions for bilateral matching and bidirectional prediction with CU weights is performed, or if the number of samples in the target block is greater than a certain value; otherwise, SAD can be used as the cost function in template matching.

[0710] ■ Search Scope (Search Area) The search area can be a specific range centered on the location indicated by the initial motion information. In other words, the center of the search area can be the location indicated by the initial motion information.

[0711] Optionally, the search range can be a specific range where the position indicated by the initial motion information is the upper left point of the search range. In other words, the upper left point of the search range can be the position indicated by the initial motion information.

[0712] Optionally, the search range may include the region of neighboring sample point locations in at least one of the following directions relative to the target block: lower left, left, upper left, upper, and upper right.

[0713] The search area can have a rectangular shape with a horizontal length of SR_X and a vertical length of SR_Y. Alternatively, the search area can have a rhombus shape with a horizontal length of SR_X and a vertical length of SR_Y. Alternatively, the search area can have a hexagonal shape obtained by excluding the rectangular area at the lower right of the rectangle. However, the shape and size of the search area are not limited to a particular embodiment.

[0714] Each of SR_X and SR_Y can be a positive integer. Each of SR_X and SR_Y can have a predefined value or a value determined based on information sent / encoded / decoded by signals.

[0715] Initial motion information can be determined based on at least one of the following: motion information of the target block, encoding parameters of the target block, motion vector of the target block, reference image of the target block, block vector of the target block, predicted motion vector of the target block, predicted block vector of the target block, motion information of at least one neighboring block of the target block, merging candidate of the target block, motion vector difference of the target block, and block vector difference of the target block.

[0716] ■Search Method The search method can be defined based on at least one of the units of search style, search resolution, search range, initial motion information, and derived motion information.

[0717] The search pattern can be one of the following: diamond, cross, or full search. However, the search pattern is not limited to the patterns listed above.

[0718] When (0, 0) indicates the position indicated by the initial motion information, a search using a diamond pattern can correspond to searching for one or more of the positions (0, 2×RR), (RR, RR), (2×RR, 0), (RR, -RR), (0, -RR), (-RR, -RR), (-RR, 0), (-RR, RR), and (0, 0).

[0719] When (0,0) represents the position indicated by the initial motion information, a cross-shaped search can correspond to searching for one or more of the positions (0,RR), (RR,0), (0,-RR), (-RR,0), and (0,0).

[0720] RR can be the search resolution or a value determined based on the search resolution. RR can be a predefined positive integer.

[0721] A full search can be used to search all locations within a predefined search range.

[0722] For example, when FS_i has values ​​from -FS_X to FS_X and FS_j has values ​​from -FS_Y to FS_Y, a full-search style search corresponds to searching for the (FS_i×RR, FS_j×RR) position. Here, (0, 0) can be a position indicated by the initial motion information. However, the search range is not limited to a specific position. Each of FS_X and FS_Y can be a predefined positive integer.

[0723] The search resolution can be one of 4 pixels, full pixels, half pixels, or a quarter pixels. However, the search resolution is not limited to a specific pixel.

[0724] The search resolution can be predefined, determined based on at least one of the information about the adaptive motion vector resolution, or determined based on the value sent / encoded / decoded by the signal.

[0725] The cells on which motion information is derived can include whole blocks and sub-blocks.

[0726] To determine the search method defined by search style, search resolution, etc., the following can be considered: motion information of the target block, encoding parameters of the target block, size of the target block, prediction mode of the target block, reference image of the target block, at least one sample value in the target block, target template, at least one sample value in the target template, and at least one region of the target template.

[0727] Figures 20 to 25 Each illustrates a search method for template matching according to one embodiment.

[0728] Based on the motion information of the target block, encoding parameters, prediction mode, and adaptive motion vector resolution, it is possible to determine Figures 20 to 25 The table shown contains specific columns. You can use the search style and search resolution corresponding to the rows marked "v" in descending order from the top to the bottom of the selected columns.

[0729] For example, in Figure 20 In the process, if the target block is in AMVP mode and the resolution determined by the adaptive motion vector resolution is 4 pixels, a search using a diamond pattern with a 4-pixel search resolution can be performed, followed by a search using a cross pattern with a 4-pixel search resolution.

[0730] ALT_IF can represent the index of an adaptive interpolation filter. An interpolation filter can be applied to compute the pixel value at a sample location with a specific resolution. The adaptive interpolation filter can be an interpolation filter selected from multiple interpolation filters by its index. In other words, when applying an adaptive interpolation filter, different interpolation filters can be used depending on the index to compute the pixel value at a sample location with a specific resolution.

[0731] For example, a specific resolution can be half a pixel. However, a specific resolution is not limited to half a pixel.

[0732] For example, the interpolation filter determined by the index can be one of a 6-tap interpolation filter and an 8-tap interpolation filter. However, the method for determining the interpolation filter is not limited to the method described above.

[0733] Figure 26This illustrates a method for configuring a first template (target template) in affine mode. Figure 27a and Figure 27b This illustrates a method for configuring a second template (reference template) in affine mode.

[0734] CPMV can refer to the Affine Control Point Motion Vector (CPMV) used in affine mode. CPMV can be used to derive the MV for each sub-block within a target block.

[0735] Reference Figure 26 The target template (A0 to A3 and L0 to L3) can be composed of sub-blocks that are adjacent to the top and left side of the target block (target CU).

[0736] Reference Figure 27a In one embodiment, the reference templates (A0 to A3 and L0 to L3) corresponding to the target template can be determined using the motion vector (or block vector) determined for each reference template position based on the target block CPMV. For example, a sub-block at a position indicated by the motion vector of the reference template from the target template position can be configured as a reference template.

[0737] Reference Figure 27b In another embodiment, the motion vector (or block vector) of the sub-block closest to each target template within the target block can be used to determine the reference template (A0 to A3 and L0 to L3) corresponding to the target template. For example, the CPMV of the target block can be used to determine the motion vector (or block vector) of the sub-block closest to the target template. The sub-block at the position indicated by the determined motion vector from the target template position can be configured as the reference template.

[0738] If the target block is in affine mode, it can be partitioned into sub-blocks of width N and height M. Motion information for each sub-block can be determined based on at least one of the target block's motion information, encoding parameters, and dimensions. Template matching cost for the target block can be determined based on at least one of the template matching costs for the sub-blocks of each partition. For example, the template matching cost for the target block can be the sum of the template matching costs for the sub-blocks of the partition or the average of the template matching costs for each sub-block.

[0739] N and M can be 2, 4, 8 or positive integers.

[0740] N and M can have predefined values, or they can have values ​​determined based on information sent / encoded / decoded by signals.

[0741] 3.6 Template Matching in Bilateral Prediction Blocks The motion information of a target block can be determined based on the motion information of neighboring blocks. This attribute can be expressed as "the target block inherits motion information from its neighboring blocks".

[0742] For example, if the target block is in merge mode, a merge candidate can be specified from the merge candidate list based on the merge index, and the motion information of the specified merge candidate can be used as the motion information of the target block.

[0743] If the target block is in AMVP mode, an MV candidate can be specified from the MV candidate list based on the MV candidate index, and the motion information of the specified MV candidate can be used as the motion information of the target block.

[0744] If bidirectional prediction is performed on a target block, for example, if motion information inherited from neighboring blocks indicates bidirectional prediction, then an embodiment of performing template matching in the target block may include the following steps.

[0745] ■Step 1: Perform template matching for each of the L0 and L1 directions, and calculate the template matching costs C0 and C1 for the motion information determined for the L0 and L1 directions.

[0746] At this point, when performing template matching for each direction, motion information in other directions is not considered, and template matching can be performed in the same way as template matching in unidirectional prediction along the corresponding direction.

[0747] If the target block satisfies the predefined conditions, then MR-SAD can be used as the cost function; otherwise, SAD can be used as the cost function.

[0748] Here, the predefined conditions can be based on at least one of the following: whether a local illumination compensation mode is performed in the target block; whether bidirectional prediction with CU weights is performed in the target block; an indicator indicating whether a model-based prediction method is performed in the target block; whether bilateral matching is performed in the target block; an indicator indicating whether bilateral matching is performed in the target block; motion information of the target block; size of the target block; encoding parameters of the target block; motion information of the target block's neighboring blocks; encoding parameters of the target block's neighboring blocks; and the type of cost function in template matching in the target block's neighboring blocks.

[0749] In some embodiments, MR-SAD can be used as the cost function if the condition for whether to perform a model-based prediction method in the target block is true or the indicator indicating whether to perform a model-based prediction method in the target block is true; otherwise, SAD can be used as the cost function.

[0750] In another embodiment, if the condition for whether to perform bilateral matching in the target block is true or the indicator indicating whether to perform bilateral matching in the target block is true; and if the number of samples in the target block is greater than or equal to a specific value, then MR-SAD can be used as the cost function; otherwise, SAD can be used as the cost function.

[0751] The cost function may refer to the cost function used in the search template matching; and / or the cost function used to calculate at least one of C0, C1, and C'. C' will be explained in step 3 later.

[0752] The cost function used in template matching search and the cost function used to compute at least one of C0, C1, and C' can be the same or different. For example, MR-SAD can be used as the cost function in template matching search, and SAD can be used as the cost function for C0, C1, and C'.

[0753] In another embodiment, if the size of the target block is smaller than a specific value, SAD can be used as the cost function in template matching; otherwise, MR-SAD can be used as the cost function in template matching. The size of the block may include at least one of the following: the width of the block, the height of the block, (the sum of the width and height of the block), and (the product of the width and height of the block).

[0754] In another embodiment, if BCW is not performed on the target block or the same weights are applied to the reference blocks in the L0 and L1 directions in BCW, SAD can be used as the cost function in template matching; otherwise, MR-SAD can be used as the cost function.

[0755] In another embodiment, if Local Illuminance Compensation (LIC) mode is not performed on the target block, SAD can be used as the cost function in template matching; otherwise, MR-SAD can be used as the cost function. The Local Illuminance Compensation mode can be a mode in which at least one of weights and offsets is derived by calculating the correlation between the template of the target block and the template of the reference block, and the derived weights or offsets are applied to all or part of the target block (or the reference block of the target block). Here, weights refer to parameters used in multiplication operations with the reference block of the target block, and offsets refer to parameters used in addition operations with the reference block of the target block. In the Local Illuminance Compensation mode, the sample value P can be modified as follows.

[0756] P' = w P + o (where w represents the weight and o represents the offset).

[0757] ■Step 2 : If C0 < C1, then a new target template T' is generated using the target template and the template in the L0 direction.

[0758] For example, when the target template is T, the reference template in the L0 direction is T0, and the reference template in the L1 direction is T1, the new target template T’ can be determined as follows: T’ = w T T + w T0 T0 w T and w T0 can be predefined values respectively.

[0759] w T can be a positive number, and w T0 can be a negative number. w T and w T0 can be values determined based on whether bidirectional prediction with CU weights is performed in the target block; and / or the weights in the bidirectional prediction with CU weights.

[0760] For example, w T can be 2, and w T0 can be -1. If bidirectional prediction with CU weights is not performed in the target block, or if the weights in the bidirectional prediction with CU weights for the L0 and L1 directions are the same, then w T can be 2, and w T0 can be -1.

[0761] For example, if bidirectional prediction with CU weights is performed in the target block, then w T can be a value determined based on the weight for the L1 direction. w T0 can be a value determined by dividing the weight for the L0 direction by the weight for the L1 direction.

[0762] If C0 > C1, the target template and the template in the L1 direction can be used to generate the new target template T’.

[0763] If the values of C0 and C1 are the same, it can be considered as either the case of C0 < C1 or C1 > C0, and the process of step 2 can be executed.

[0764] ■Step 3: If C0 < C1, template matching is performed for the L1 direction with the target template as T’, and the template matching cost C’ of the motion information determined for the L1 direction is calculated.

[0765] If C0 > C1, template matching is performed for the L0 direction with the target template as T', and the template matching cost C’ of the motion information determined for the L0 direction is calculated.

[0766] If the value of C0 is the same as the value of C1, it is considered a case of C0 < C1 or C1 > C0, and the process of step 3 can be performed.

[0767] ■Step 4: The motion information of the target block can be changed to unidirectional motion information based on at least one of the values of C’, C0, or C1.

[0768] For example, if the value obtained by multiplying C’ by a predetermined value w C' is greater than the smaller of C0 or C1 (or the value obtained by multiplying the smaller by the predetermined value w CX ), the motion information of the target block is changed to motion information indicating unidirectional prediction in the L0 direction or the L1 direction. Here, the result of multiplying the two values can represent one of A × B or the rounded, ceiling, or truncated value of A × B.

[0769] For example, if C0 < C1, the motion information of the target block can be changed to motion information indicating unidirectional prediction in the L0 direction. Alternatively, the motion information for the L1 direction can be considered unavailable in the target block.

[0770] If C0 > C1, the motion information of the target block can be changed to motion information indicating unidirectional prediction in the L1 direction. Alternatively, the motion information for the L0 direction can be considered unavailable in the target block.

[0771] If the value of C0 is the same as the value of C1, it is considered a case of C0 < C1 or C1 > C0, and the process of step 4 can be performed.

[0772] w c' and w CX can be predefined values respectively.

[0773] w c' and w CX can be values determined based on whether bidirectional prediction with CU weights is performed in the target block; and / or the weights in bidirectional prediction with CU weights.

[0774] For example, w c' can be 1 / 2, and w CX can be 9 / 8.

[0775] For example, w c' can be a value determined based on the weight in the L0 direction or the L1 direction in bidirectional prediction with CU weights. If C0 < C1, w c' can be the weight in the L1 direction in bidirectional prediction with CU weights or a value determined based on that weight. If C1 < C0, w c'It can be the weight in the L0 direction of a bidirectional prediction with CU weights or a value determined based on that weight.

[0776] Steps 2 through 4 can only be executed if the target block meets the predefined conditions.

[0777] For example, steps 2 through 4 can be performed only if the target block is based on bilateral prediction; and if bilateral matching is not performed in the target block, or if the target block does not meet the conditions for enabling bilateral matching.

[0778] Furthermore, when refining the first motion information using template matching with the first frame as a reference frame, second motion information using the second frame as a reference frame can be derived from the first motion information. Then, template matching can be used to perform further refinement on the second motion information.

[0779] A first reference block indicated by first refined motion information can be used to generate a prediction block for a target block, wherein the first refined motion information is the result of refining the first motion information using template matching.

[0780] Furthermore, a second reference block indicated by the second refined motion information can be used to generate a prediction block for the target block, wherein the second refined motion information is the result of refining the second motion information using template matching.

[0781] When generating the prediction block for the target block, a weighted summation of the first and second reference blocks can be performed.

[0782] Here, the first screen and the second screen can be screens from the reference screen list of the target block in the L0 direction and / or screens from the reference screen list of the target block in the L1 direction, respectively.

[0783] When deriving second motion information from first motion information, the result obtained by scaling the first motion vector based on the POC of the first reference frame, the POC of the second reference frame, and the POC of the target frame can be used as the motion vector of the second motion information.

[0784] For example, the motion vector of the second motion information may have the same magnitude as the motion vector of the first motion information, but in the opposite direction.

[0785] The first motion information and the second motion information may differ in at least one aspect of the reference frame, the reference frame index, and the motion information.

[0786] Optionally, the first motion information and the second motion information can be the same, except for the reference frame, the reference frame index, and the motion information.

[0787] So far, a method has been described that derives second motion information from first motion information and generates a prediction block for a target block based on the first and second motion information. However, in the same manner as the method for deriving the second motion information and the method for generating the prediction block for the target block, Nth motion information can be derived, and when generating the prediction block for the target block, a reference block indicated by the Nth motion information can be used. N can be 2, 3, or a positive integer.

[0788] 4. Two-sided matching In bilateral matching, reference blocks in the L0 direction and reference blocks in the L1 direction are used as templates, and the motion information of the target block can be determined and / or changed based on the calculation of the cost function between the two templates.

[0789] The reference block may include at least one of the following: 1) the reference block pointed to by the initial motion information; 2) the reference block pointed to by the motion information derived during the bilateral matching search process; and 3) the reference block pointed to by the motion information finally improved by bilateral matching.

[0790] For example, when configuring a template for bilateral matching, a reference block in the L0 direction and a reference block in the L1 direction can be used as templates.

[0791] The cost of bilateral matching can be correlated with the result obtained by calculating the cost function of the templates of the reference blocks in the L0 direction and the reference blocks in the L1 direction used in bilateral matching.

[0792] If the target block is in IBC mode and two or more reference blocks are used to predict the target block, then two different reference blocks from the target block's reference blocks can be used as templates to perform bilateral matching.

[0793] The following describes bilateral prediction in inter-frame prediction, rather than in IBC mode; however, the technical features of bilateral prediction in inter-frame prediction can be equivalently applied to IBC mode. In this case, the reference blocks in the L0 and L1 directions can be replaced by two reference blocks generated in IBC mode.

[0794] 4.1 Subsampling in bilateral matching ■ Subsampling for configuring templates When configuring a template for bilateral matching, you can select only a subset of pixels and / or positions within the reference blocks in the L0 and L1 directions. The template can be configured using only the selected pixels and / or positions.

[0795] The template used for bilateral matching may correspond to at least one of the templates in the L0 direction and the template in the L1 direction.

[0796] In some embodiments, when configuring a template for bilateral matching, subsampling can be used for reference blocks in the L0 direction and reference blocks in the L1 direction.

[0797] In another embodiment, when configuring a template for bilateral matching, subsampling can be used on a portion of the reference block in the L0 direction and a portion of the reference block in the L1 direction.

[0798] For example, when configuring a template for bilateral matching, the reference block in the L0 direction and the reference block in the L1 direction can each be partitioned into two or more regions. The partitioned regions can be one of the following: 1) a region where subsampling is performed, 2) a region used for template configuration without involving subsampling, and 3) a region not used for template configuration. The template for bilateral matching can be configured using pixels and / or positions selected by subsampling from region 1) and pixels and / or positions within region 2).

[0799] Optionally, the partitioned region can be one of 1) the region where subsampling is performed and 2) the region not used for template configuration. The template for bilateral matching can be configured using pixels and / or positions selected by subsampling in region 1).

[0800] In another embodiment, when configuring a template for bilateral matching, the area used to configure the template may be a portion of a reference block in the L0 direction and a portion of a reference block in the L1 direction.

[0801] In another embodiment, when configuring a template for bilateral matching, pixels (or positions) for configuring the template can be selected from only a portion of the reference block in the L0 direction and a portion of the reference block in the L1 direction.

[0802] The size of a portion of the reference block in the L0 direction can be smaller than the size of a portion of the reference block in the L1 direction.

[0803] For example, the height (vertical dimension) of a portion of the reference block in the L0 direction can be smaller than the height (vertical dimension) of the reference block in the L1 direction.

[0804] For example, the width (horizontal dimension) of a portion of the reference block in the L0 direction can be smaller than the width (horizontal dimension) of the reference block in the L1 direction.

[0805] The size of a portion of the reference block in the L1 direction can be smaller than the size of the reference block in the L0 direction.

[0806] For example, the height (vertical dimension) of a portion of the reference block in the L1 direction can be smaller than the height (vertical dimension) of the reference block in the L0 direction.

[0807] For example, the width (horizontal dimension) of a portion of the reference block in the L1 direction can be smaller than the width (horizontal dimension) of the reference block in the L0 direction.

[0808] ■ Subsampling for cost calculation When performing cost calculations between templates in bilateral matching, you can select only a subset of pixels and / or locations within the template region. Cost calculations can be performed only for the selected pixels and / or locations.

[0809] In some embodiments, when calculating the cost between templates in a bilateral matching process, subsampling can be performed on the template regions in the L0 direction and the template regions in the L1 direction.

[0810] In another embodiment, when calculating the cost between templates in bilateral matching, subsampling can be performed on a portion of the template region in the L0 direction and a portion of the template region in the L1 direction.

[0811] Optionally, when performing cost calculation between templates in bilateral matching, the template regions in the L0 direction and the template regions in the L1 direction can be partitioned into two or more regions, respectively. The partitioned regions can be one of the following: 1) regions where subsampling is performed, 2) regions used for cost calculation without involving subsampling, and 3) regions not used for cost calculation. Cost calculation between templates in bilateral matching can be performed using pixels and / or positions selected by subsampling in region 1) and pixels and / or positions within region 2).

[0812] Optionally, the partitioned region can be one of 1) the region where subsampling is performed and 2) the region not used for cost calculation. The cost function between templates in bilateral matching can be calculated using pixels and / or locations selected from region 1) through subsampling.

[0813] In another embodiment, when performing cost calculation between templates in bilateral matching, the region used for cost calculation may be a portion of the template in the L0 direction and a portion of the template in the L1 direction.

[0814] The size of a portion of the template in the L0 direction can be different from the size of a portion of the template in the L1 direction.

[0815] In one example, the size of a portion of the template in the L0 direction can be smaller than the size of a portion of the template in the L1 direction.

[0816] For example, the height (vertical dimension) of a portion of the template in the L0 direction can be smaller than the height (vertical dimension) of the template in the L1 direction.

[0817] For example, the width (horizontal dimension) of a portion of the template in the L0 direction can be smaller than the width (horizontal dimension) of the template in the L1 direction.

[0818] In another example, the size of a portion of the template in the L1 direction may be smaller than the size of the template in the L0 direction.

[0819] For example, the height (vertical dimension) of a portion of the template in the L1 direction can be smaller than the height (vertical dimension) of the template in the L0 direction.

[0820] For example, the width (horizontal dimension) of a portion of the template in the L1 direction can be smaller than the width (horizontal dimension) of the template in the L0 direction.

[0821] ■ Subsampling of the search area For example, when performing a bilateral matching search process, only a subset of pixels and / or locations within the search area may be selected. Search and / or matching cost calculations may be performed only for the selected pixels and / or locations. Alternatively, search and / or matching cost calculations may be performed only for motion information indicating the selected pixels and / or locations.

[0822] For example, when performing a bilateral matching search, subsampling can be performed on all or part of the search region.

[0823] Optionally, for example, when performing bilateral matching search processing, the search area can be divided into two or more regions. The partitioned regions can be one of the following: 1) a region where subsampling is performed, 2) a region where search processing is performed without subsampling, and 3) a region where search processing is not performed. Bilateral matching search processing can be performed on pixels and / or locations selected by subsampling within region 1) and pixels and / or locations within region 2). Optionally, bilateral matching search processing can be performed on pixels and / or locations selected by subsampling within region 1) and motion information indicating pixels and / or locations within region 2).

[0824] Optionally, for example, when performing bilateral matching search processing, the search area can be divided into two or more regions. The partitioned regions can be one of 1) a region where subsampling is performed and 2) a region where search processing is not performed. Bilateral matching search processing can be performed on pixels and / or locations selected by subsampling within region 1). Optionally, bilateral matching search processing can be performed on motion information indicating pixels and / or locations selected by subsampling within region 1.

[0825] Figure 18 Various examples of subsampling methods in bilateral matching are shown.

[0826] Figure 18The shaded sample points (or positions) in the text represent sample points (or positions) selected through subsampling.

[0827] like Figure 18 As shown, subsampling can be performed on the reference block region in the L0 direction and the reference block region in the L1 direction, and the template can be configured using only selected sample points (or locations).

[0828] like Figure 19 As shown, subsampling can be performed on a portion of the reference block region in the L0 direction and a portion of the reference block region in the L1 direction, and the template can be configured using only selected sample points (or locations).

[0829] For all or part of the template region of a bilateral match, it can be as follows: Figure 18 The diagram shows the performance of subsampling, and cost calculation can be performed only for selected sample points (or locations).

[0830] For a bilateral matching search region, whether it is all or part, it can be like this: Figure 18 The subsampling shown is performed, and cost calculation can be performed only for selected pixels and / or locations.

[0831] For a bilateral matching search region, whether it is all or part, it can be like this: Figure 18 The subsampling shown is performed, and cost calculation can be performed only for motion information indicating the selected pixels and / or locations.

[0832] 4.2 Conditions for performing bilateral matching Bilateral matching can always be performed. Alternatively, bilateral matching can be performed only when predefined enabling conditions are met.

[0833] For example, when the inter-frame prediction mode is used for the target block and two or more reference blocks are employed, bilateral matching can be performed.

[0834] For example, bilateral matching can be performed when the first direction is different from the second direction, and the first POC interval is the same as the second POC interval. The first direction can be the direction from the target image to the reference image in the L0 direction. The second direction can be the direction from the target image to the reference image in the L1 direction. The first POC interval can be the difference between the POC of the target image and the POC of the reference image in the L0 direction. The second POC interval can be the difference between the POC of the target image and the POC of the reference image in the L1 direction.

[0835] For example, bilateral matching can be performed when the first direction is different from the second direction. For example, bilateral matching can be performed even if the first POC interval is different from the second POC interval, when the first direction is different from the second direction.

[0836] Here, the fact that the first direction and the second direction are different from each other indicates that Equation 1 below is satisfied.

[0837] [Equation 1] (POCt – PCO0)×(POCt – POC1)<0 Here, the fact that the first direction is the same as the second direction indicates that Equation 2 below is satisfied.

[0838] [Equation 2] (POCt-POC0)×(POCt-POC1)>0 In Equations 1 and 2, POCt represents the POC of the target image, POC0 represents the POC of the reference image in the L0 direction, and POC1 represents the POC of the reference image in the L1 direction.

[0839] 4.3 Two-sided matching search method ■ Definition of Search A search can be performed using the calculation of a cost function to determine the similarity between two templates.

[0840] The search may include the process of determining at least one piece of motion information that meets specific conditions within a specific search range. The motion information of the target block may be determined and / or modified based on the at least one piece of motion information determined through the search.

[0841] In other words, motion information that meets specific conditions can refer to, but is not limited to, motion information with the lowest matching cost among the motion information within the search range.

[0842] A search may include the process of determining at least one piece of motion information that satisfies specific conditions within a particular search range. The motion information of the block determined through the search can be used as the motion information of the target block.

[0843] In other words, a block that meets certain conditions can refer to, but is not limited to, the motion information with the lowest matching cost among reference blocks within the search range.

[0844] ■ Cost Function Given two templates for bilateral matching, the cost function can refer to a function that determines the similarity between at least one sample in the first template and at least one sample in the second template.

[0845] The similarity between a first value and a second value can be determined using at least one of the following operations: 1) the difference between two values, 2) the ratio between two values, and 3) comparing the difference between two values ​​with a specific value.

[0846] The cost function can be one or more of the following: Sum of Absolute Differences (SAD), Sum of Absolute Transformed Differences (SATD), Sum of Absolute Differences with Mean Removed (MR-SAD), Mean Squared Error (MSE), and Sum of Squared Errors (SSE). However, the cost function is not limited to the terms listed above.

[0847] The cost function used in bilateral matching can be predefined or determined based on information sent / encoded / decoded by signals.

[0848] In one embodiment, if the size of the target block is smaller than a certain value, SAD can be used as the cost function in bilateral matching; otherwise, MR-SAD can be used as the cost function in template matching. The size of the block may include at least one of the following: the width of the block, the height of the block, (the sum of the width and height of the block), and (the product of the width and height of the block).

[0849] In another embodiment, if BCW is not performed on the target block or the same weights are used for the reference block in the L0 and L1 directions of BCW, SAD or SATD can be used as the cost function in bilateral matching; otherwise, MR-SAD or MR SATD can be used as the cost function.

[0850] In another embodiment, if the local illumination compensation mode is not performed on the target block, the SAD can be used as the cost function in the two-sided matching; otherwise, the MR-SAD can be used as the cost function.

[0851] 4.4 Search Steps for Bilateral Matching If the initial motion information (e.g., the initial motion vector or the initial block vector) is not in integer pixel units (i.e., in fractional pixel units), the initial motion information can be modified by rounding (or truncating or rounding up) the initial motion information and the search can be performed using the modified initial motion vector.

[0852] For example, in order to generate a reference template at the location indicated by motion information obtained by adding a specific offset to the initial motion information in fractional pixel units during the search process, an interpolation filter must be applied to the samples at integer pixel locations to generate samples at fractional pixel locations. However, if the initial motion information is limited to integer pixel units, the computational complexity can be reduced because interpolation at fractional pixel locations is not required during the search process.

[0853] Two-sided matching can include one or more search steps.

[0854] For example, bilateral matching can be configured to sequentially include steps 1) deriving motion information for the entire block and 2) deriving motion information for the sub-blocks of the block. However, the method for deriving motion information in each step and the order of execution of the steps are not limited to the above configuration.

[0855] In each search step of the bilateral matching, the motion information regarding the direction of BM_NUM can be improved within the motion information regarding the L0 direction and the motion information regarding the L1 direction. BM_NUM can be 0, 1, 2, or a positive integer. The BM_NUM directions used in the steps of the bilateral matching can be the same or different from each other.

[0856] For example, when BM_NUM is 1 in a specific search step, LX can be applied only to that specific search step. BM Motion information in the direction is used to improve motion information. Here, X BM It can have a value of 0 or 1. In other words, if X BM If the value is 0, then the L0 direction is set to LX. BM Direction, if X BM If the value is 1, then the L1 direction is set to LX. BM direction.

[0857] For example, if BM_NUM is 1 and X is 1 in a particular search step BM If the value is 0, then with the template and motion information in the L1 direction fixed, a search can be performed only in the L0 direction for a specific search step.

[0858] Signals can be used to send / encode / decode information about X. BM Information.

[0859] Optionally, X BM It can be predefined.

[0860] For example, X BM It can be set to the value in the direction with the larger POC difference from the target image, either the L0 or L1 direction. The POC difference can be the difference between the POC of the reference image and the POC of the target image (in a specific direction).

[0861] Optionally or otherwise, X BM This can be set to the direction from which the template matching cost is higher, either the L0 or L1 direction. Here, the template matching cost in a specific direction can be the template matching cost of motion information in that specific direction.

[0862] The following section will describe the methods used to determine X. BM Various embodiments.

[0863] ■ Used to determine X BM Method In one embodiment, if the first POC difference is greater than the second POC difference, then X BM It can be 0; otherwise, X BM It can be 1. Optionally, if the first POC difference is greater than the second POC difference, then X BM It can be 1; otherwise, X BM It can be 0. The first POC difference can be the difference between the POC of the target image and the POC of the reference image in the L0 direction. The second POC difference can be the difference between the POC of the target image and the POC of the reference image in the L1 direction.

[0864] In another embodiment, X can be determined based on a context model and / or a probabilistic model used for entropy encoding and entropy decoding of motion information and encoding parameters of the target block. BM .

[0865] For example, X can be determined based on at least one of a context model and / or a probabilistic model used for entropy encoding and entropy decoding of inter-prediction indicators that indicate the inter-prediction direction of the target block. BM .

[0866] Based on the context model and / or probabilistic model used in the entropy coding and entropy decoding of inter-frame prediction indicators, the more probable direction in the target block can be selected as LX from the unidirectional prediction in the L0 direction versus the unidirectional prediction in the L1 direction. BM Direction, that is, the direction in which motion information enhancement is performed.

[0867] Optionally, by using the context model and / or probability model used in the entropy coding and entropy decoding of the inter-frame prediction indicator, the more probable direction in the target block between the unilateral prediction in the L0 direction and the unidirectional prediction in the L1 direction can be selected as L(1-X). BM The direction in which motion information enhancement is not performed is the direction in which motion information enhancement is not performed.

[0868] A more likely direction may refer to the direction that uses fewer bits when performing entropy encoding using a contextual model and / or a probabilistic model. Alternatively, a more likely direction may refer to the direction that has a higher probability of yielding a contextual model and / or a probabilistic model relative to that direction.

[0869] In another embodiment, LX can be determined based on the weights in the bidirectional prediction of the target block with CU weights. BM For example, LX BM It can be the direction with higher weight between the L0 and L1 directions. Optionally, LX BM It can be the direction with the lower weight between the L0 direction and the L1 direction.

[0870] In another embodiment, X can be determined based on at least one of the motion information of neighboring blocks and encoding parameters. BM For example, it can be based on... Figure 5 The motion information of at least one of the corresponding neighboring blocks of A0, A1, B0, B1 and B2 and at least one of the encoding parameters are used to determine X in the target block.

[0871] For example, the X of the target block can be determined based on at least one of the inter-prediction indicator in neighboring blocks and the inter-bidirectional prediction weights. BM .

[0872] For example, one or more context models and / or probability models can be used for X BM Entropy encoding and entropy decoding. In multiple context models and / or probabilistic models, the entropy encoding information for the target block can be determined based on at least one of the motion information and encoding information of neighboring blocks. BM Contextual and / or probabilistic models for entropy encoding and entropy decoding.

[0873] X used in the block BM The context model and / or probability model for entropy encoding and entropy decoding can be the same. Optionally, the X used in the block... BM The context model and / or probability model for entropy encoding and entropy decoding can vary based on at least one of the inter-frame prediction direction and inter-frame bidirectional prediction weights of neighboring blocks.

[0874] The same X can be used in the search step of bilateral matching. BM Optionally, different X values ​​can be used in each search step of the bilateral matching process. BM .

[0875] The following text will describe BM_NUM, which represents the number of directions that improve motion information between the L0 and L1 directions.

[0876] ■BM_NUM BM_NUM can be 0, 1, 2 or a positive integer.

[0877] BM_NUM can be predefined.

[0878] BM_NUM in each search step of bilateral matching can be determined based on encoding parameters. Alternatively, BM_NUM can be determined based on at least one of motion information, the search step of bilateral matching, the matching cost in the previous search step, the matching cost for the initial motion information of the current search step, and BM_NUM in the previous search step.

[0879] In one embodiment, BM_NUM in the first search step of bilateral matching can be 1 or 2.

[0880] In one embodiment, the BM_NUM of the current search step can be determined based on the matching cost in the previous search steps.

[0881] For example, if the difference between the matching cost of the initial motion information for the previous search step and the matching cost of the improved motion information for the previous search step is less than COSTDIFF_FORBMNUM, then BM_NUM for the current search step can be 0.

[0882] COSTDIFF_FORBMNUM can be 0, 1, 2, 4, 8, 16 or a positive integer.

[0883] COSTDIFF_FORBMNUM can be determined based on the size of the target block. COSTDIFF_FORBMNUM can be the product of the number of pixels in the target block and a specific value. The specific value can be 0, 1, 2, 4, 8, or a positive integer.

[0884] In one embodiment, if BM_NUM was 0 in a previous search step, then BM_NUM in the current search step can be 0.

[0885] In one embodiment, if the matching cost of the initial motion information of the current search step is less than COSTDIFF_FORBMNUM_INIT, then the BM_NUM of the target block can be 0.

[0886] COSTDIFF_FORBMNUM_INIT can be 0, 1, 2, 4, 8 or a positive integer.

[0887] For example, COSTDIFF_FORBMNUM_INIT can be determined based on the size of the target block. COSTDIFF_FORBMNUM_INIT can be the product of the number of pixels in the target block and a specific value. The specific value can be 0, 1, 2, 4, 8, or a positive integer.

[0888] In one embodiment, when motion improvement is performed during the search phase of bilateral matching, motion improvement in the L0 direction can only be performed if the matching cost of the motion information in the L0 direction of the initial motion information in the current search phase is greater than COSTDIFF_FORBMNUM_INIT.

[0889] In some embodiments, when motion improvement is performed during the search phase of bilateral matching, motion improvement in the L1 direction can only be performed if the matching cost of the motion information in the L1 direction of the initial motion information in the current search phase is greater than COSTDIFF_FORBMNUM_INIT.

[0890] A BM_NUM value of 0 for a specific search step in a bilateral match indicates that motion information improvement is not performed in that specific search step. Alternatively, a BM_NUM value of 0 for a specific search step in a bilateral match can indicate that the specific search step is not performed.

[0891] For example, when performing bilateral matching, BM_NUM in the motion information export step for the entire block can be 1, and BM_NUM in the motion information export step for sub-blocks can be 2. In this case, only LX can be improved in the motion information export step for the entire block. BM Motion information in the direction can be improved in both the L0 and L1 directions during the motion information derivation step for sub-blocks.

[0892] Figure 28 This illustrates a bilateral match according to one embodiment.

[0893] Figure 28 This illustrates the case where BM_NUM is 2 during the step of deriving motion information from the entire block in a bilateral matching process.

[0894] MV0 represents the initial motion information in the L0 direction, and MV1 represents the initial motion information in the L1 direction.

[0895] MV diff This can refer to the improved value of motion information derived through bilateral matching.

[0896] MV 0' and MV 1' It can be motion information derived through bilateral matching.

[0897] In bilateral matching, the magnitude of the motion information improvement value in the L0 direction can be the same as the magnitude of the motion information improvement value in the L1 direction. The direction of the motion information improvement value in the L0 direction can be opposite to the direction of the motion information improvement value in the L1 direction. In other words, equations 3 and 4 below can be established.

[0898] [Equation 3]MV 0' =MV0+MV diff [Equation 4]MV 1' =MV1-MV diff For example, given a target block in IBC mode and using two or more reference blocks, when performing bilateral matching on the target block, the values ​​and directions of the motion information improvements for the first reference block and the motion information improvements for the L1 direction can be the same.

[0899] Subsampling in the above template matching or bilateral prediction can be performed based on at least one of the following: whether template matching or bilateral prediction is performed; an indicator indicating whether template matching or bilateral prediction is performed; motion information of the target block; encoding parameters of the target block; and search steps in template matching.

[0900] For example, the subsampling method can be determined based on at least one of the following: whether template matching is performed; an indicator indicating whether template matching is performed; motion information of the target block; encoding parameters of the target block; and search steps in template matching.

[0901] The subsampling method used in template matching or bilateral matching can be the same for each search step. Alternatively, the subsampling method used when performing template matching and / or bilateral matching can be different for each search step.

[0902] In another example, whether to perform subsampling can be determined based on at least one of the following: whether template matching is performed; an indicator indicating whether template matching is performed; motion information of the target block; encoding parameters of the target block; and search steps in template matching.

[0903] Whether or not subsampling is performed during template matching can be the same for each search step. Alternatively, whether or not subsampling is performed can be different for each search step.

[0904] Whether to perform subsampling in the horizontal direction and whether to perform subsampling in the vertical direction can be the same. Alternatively, whether to perform subsampling in the horizontal direction and whether to perform subsampling in the vertical direction can be different from each other.

[0905] 4.5 Sub-block-based bilateral matching When performing bilateral matching, motion information can always be exported only at the sub-block level. In other words, motion information export can be performed without exporting motion information for the entire block.

[0906] During bilateral matching, signal transmission / encoding / decoding can be used to send / encode / decode derived indicator information about whether motion information is performed at the sub-block unit level.

[0907] When performing bilateral matching, it can be determined whether to refine the motion information into sub-block units based on at least one of the target block size, target block encoding parameters, target block motion information, neighboring block encoding parameters, and neighboring block motion information.

[0908] In some embodiments, when performing bilateral matching, when the width of the target block is W and the height of the target block is H, motion information can only be derived from sub-block units if the larger of W and H is greater than MIN_SIZE_THRES_FOR_SUB. For example, an indicator indicating whether to derive motion information from sub-block units can only be encoded / decoded if the larger of W and H is greater than a first threshold (MIN_SIZE_THRES_FOR_SUB).

[0909] MIN_SIZE_THRES_FOR_SUB can be a predefined value. For example, MIN_SIZE_THRES_FOR_SUB can be 8, 16, 32, 64, 128, or a positive integer.

[0910] As an alternative or additional embodiment, motion information can only be derived from sub-block units if the larger of W and H is less than a second threshold (MAX_SIZE_THRES_FOR_SUB). For example, an indicator indicating whether to derive motion information from sub-block units can only be encoded / decoded if the larger of W and H is less than MAX_SIZE_THRES_FOR_SUB.

[0911] MAX_SIZE_THRES_FOR_SUB can be a predefined value, such as 8, 16, 32, 64, 128, or a positive integer.

[0912] In some embodiments, motion information can be exported in sub-block units only if the smaller of W and H is greater than MIN_SIZE_THRES_FOR_SUB. For example, an indicator indicating whether to export motion information in sub-block units can only be encoded / decoded if the smaller of W and H is greater than MIN_SIZE_THRES_FOR_SUB.

[0913] As an alternative or additional embodiment, the export of motion information in sub-block units may be performed only if the smaller of W and H is less than MAX_SIZE_THRES_FOR_SUB. For example, an indicator indicating whether to export motion information in sub-block units may be encoded / decoded only if the smaller of W and H is less than MAX_SIZE_THRES_FOR_SUB.

[0914] In some embodiments, when the width of the target block is W and the height of the target block is H, the extraction of motion information in sub-block units can only be performed if the result value of W×H (i.e., the area of ​​the target block or the total number of samples in the target block) is greater than a third threshold (MIN_SIZE_AREA_THRES_FOR_SUB). For example, the indicator indicating whether to perform the extraction of motion information in sub-block units can only be encoded / decoded if the result value of W×H is greater than MIN_SIZE_AREA_THRES_FOR_SUB.

[0915] MIN_SIZE_AREA_THRES_FOR_SUB can be a predefined value. For example, MIN_SIZE_AREA_THRES_FOR_SUB can be 8, 16, 32, 64, 128, or a positive integer.

[0916] As an alternative or additional embodiment, the extraction of motion information in sub-block units can only be performed if the result value of W×H is less than the fourth threshold (MAX_SIZE_AREA_THRES_FOR_SUB). For example, the indicator indicating whether to extract motion information in sub-block units can only be encoded / decoded if the result value of W×H is less than MAX_SIZE_AREA_THRES_FOR_SUB.

[0917] MAX_SIZE_AREA_THRES_FOR_SUB can be a predefined value, such as 8, 16, 32, 64, 128, or a positive integer.

[0918] Furthermore, when bidirectional matching is performed at the sub-block unit level, the size information of at least one sub-block that performs bidirectional matching (i.e., performs motion information derivation) can be transmitted / encoded / decoded using signals.

[0919] Optionally, at least one of the width and height of the sub-block can be determined based on at least one of the target block size, the target block encoding parameters, the target block motion information, the neighboring block encoding parameters, and the neighboring block motion information.

[0920] In some embodiments, at least one of the width and height of the sub-block may be one of the dimensions belonging to a predetermined list.

[0921] For example, a reservation list could be a list of size information.

[0922] For example, the predefined list can be predefined. In some embodiments, motion information can be derived at least once in sub-block units (i.e., bidirectional matching, where at least one of the width and height of the sub-block has a size of SUBBLOCK_SIZE_ALWAYS).

[0923] SUBBLOCK_SIZE_ALWAYS can be a predefined value. For example, SUBBLOCK_SIZE_ALWAYS can be 1, 2, 4, 8, 16, 32, 64, 128, or a positive integer.

[0924] Information about SUBBLOCK_SIZE_ALWAYS can be sent / encoded / decoded using signals.

[0925] In some embodiments, if the width of the target block is greater than a specific threshold width (e.g., SUBBLOCK_SIZE_ALWAYS), motion information can be derived at least once from sub-block units, where the sub-block units have the specific threshold width.

[0926] In some embodiments, if the height of the target block is greater than a specific threshold height (e.g., SUBBLOCK_SIZE_ALWAYS), motion information can be derived from sub-block units at least once, wherein the sub-block units have a specific threshold height.

[0927] In some embodiments, if the width of the target block is less than a certain threshold width (e.g., SUBBLOCK_SIZE_ALWAYS), motion information can be derived at least once using sub-block units, wherein the sub-blocks have the same width as the target block.

[0928] In some embodiments, if the height of the target block is less than a certain threshold height (e.g., SUBBLOCK_SIZE_ALWAYS), motion information can be derived at least once from sub-block units, wherein the sub-blocks have the same height as the target block.

[0929] In some embodiments, given a target block with a width of W and a height of H, if the larger of W and H is greater than MIN_SIZE_THRES_FOR_SUBBLOCK_SIZE, then at least one process for deriving motion information from sub-block units can be performed, where the size of the sub-block is SUBBLOCK_SIZE1. For example, information about SUBBLOCK_SIZE1 can be encoded / decoded. For example, SUBBLOCK_SIZE1 can be 1, 2, 4, 8, 16, 32, 64, 128, or a positive integer. MIN_SIZE_THRES_FOR_SUBBLOCK_SIZE can be a predefined value. For example, MIN_SIZE_THRES_FOR_SUBBLOCK_SIZE can be 8, 16, 32, 64, 128, or a positive integer.

[0930] As an additional or alternative embodiment, if the larger of W and H is less than MAX_SIZE_THRES_FOR_SUBBLOCK_SIZE, then at least one process for deriving motion information from sub-block units, where the size of the sub-block is SUBBLOCK_SIZE2, can be performed. For example, information about SUBBLOCK_SIZE2 can be encoded / decoded. For example, SUBBLOCK_SIZE2 can be 1, 2, 4, 8, 16, 32, 64, 128, or a positive integer. MAX_SIZE_THRES_FOR_SUBBLOCK_SIZE can be a predefined value. For example, MAX_SIZE_THRES_FOR_SUBBLOCK_SIZE can be 8, 16, 32, 64, 128, or a positive integer.

[0931] In some other embodiments, if the smaller of W and H is greater than MIN_SIZE_THRES_FOR_SUBBLOCK_SIZE, then at least one process of deriving motion information from sub-block units, wherein the size of the sub-block is SUBBLOCK_SIZE1, can be performed.

[0932] In an additional or optional embodiment, if the smaller of W and H is less than MAX_SIZE_THRES_FOR_SUBBLOCK_SIZE, then at least one process of deriving motion information from a sub-block unit, wherein the size of the sub-block is SUBBLOCK_SIZE2, may be performed.

[0933] In some other embodiments, if the resulting value of W×H (i.e., the area of ​​the target block or the total number of samples in the target block) is greater than MIN_SIZE_AREA_THRES_FOR_SUBBLOCK_SIZE, then at least one process of deriving motion information in sub-block units, wherein the sub-block has a size of SU...

Claims

1. An image decoding method using intra prediction, the method comprising: configuring at least one candidate list including one or more merge candidates based on reconstructed blocks reconstructed before a target block, wherein each merge candidate includes intra coding method information; selecting at least one merge candidate for the target block from the at least one candidate list; and generating a prediction block of the target block by performing intra prediction for the target block based on the selected at least one merge candidate.

2. The method of claim 1, further comprising: decoding merge candidate selection information from a bitstream, wherein the at least one merge candidate is selected from the at least one candidate list based on the merge candidate selection information.

3. The method of claim 1, wherein, The step of selecting the at least one merge candidate comprises: generating a target template consisting of reconstructed reference samples around the target block; generating a prediction template for the target template using each of at least a portion of the merge candidates included in the at least one candidate list; selecting the at least one merge candidate from the at least one candidate list by calculating a matching cost between the target template and the prediction template.

4. The method of claim 1, wherein, The number of the at least one candidate list is more than one, and Each of the plurality of candidate lists is differently configured according to a mode type indicating a type of intra coding method for a merge candidate.

5. The method of claim 4, wherein, The plurality of candidate lists includes at least two or more of the following candidate lists: a first candidate list consisting of merge candidates whose mode type is a template-based intra mode derivation method; a second candidate list consisting of merge candidates whose mode type is a gradient-based intra mode derivation method; a third candidate list consisting of merge candidates whose mode type is a block vector-based intra prediction mode; and a fourth candidate list consisting of merge candidates whose mode type is a frequency-of-occurrence-based intra prediction mode. The step of selecting the at least one merge candidate comprises:

6. The method of claim 4, wherein, selecting a candidate list to be applied to the target block from the plurality of candidate lists based on candidate list selection information included in a bitstream. The step of generating a prediction block of the target block comprises:

7. The method of claim 1, wherein, determining one or more pieces of intra coding method information to be applied to the target block based on the selected at least one merge candidate; and generating the prediction block of the target block based on the one or more pieces of intra coding method information. When the number of pieces of intra coding method information to be applied to the target block is determined to be more than one, the step of generating a prediction block of the target block comprises:

8. The method of claim 7, wherein, generating a preliminary prediction block of the target block based on each piece of intra coding method information; and generating the prediction block of the target block by applying a weight to the preliminary prediction block. The step of applying the weight to the preliminary prediction block comprises:

9. The method of claim 8, wherein, calculating a matching cost corresponding to each piece of intra coding method information by applying template matching based on each piece of intra coding method information to a target template consisting of reconstructed reference samples around the target block; and determining the weight to be applied to the preliminary prediction block based on the matching cost corresponding to each piece of intra coding method information. ​ 10. An image encoding method using intra prediction, the method comprising: configuring at least one candidate list comprising one or more merge candidates based on pre-encoded blocks coded before a target block, wherein each merge candidate comprises intra coding method information; selecting at least one merge candidate for the target block from the at least one candidate list; and generating a prediction block for the target block based on the selected at least one merge candidate.

11. A method for transmitting a bitstream comprising encoded image data, the method comprising: generating a bitstream by encoding a target block using intra prediction; and transmitting the bitstream to an image decoding device, wherein the step of generating a bitstream comprises: configuring at least one candidate list comprising one or more merge candidates based on pre-encoded blocks coded before a target block, wherein each merge candidate comprises intra coding method information; selecting at least one merge candidate for the target block from the at least one candidate list; and generating a prediction block for the target block based on the selected at least one merge candidate.

Citation Information

Patent Citations

  • Cylinderical battery winding equipment exhaust system

    KR1020230086135A