Method, device, and recording medium for image encoding / decoding using inter-component prediction model

WO2026197816A1PCT designated stage Publication Date: 2026-09-24ELECTRONICS & TELECOMM RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2026/004461
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-19
Filing Date
2026-03-19
Publication Date
2026-09-24

Smart Images

  • Figure KR2026004461_24092026_PF_FP_ABST
    Figure KR2026004461_24092026_PF_FP_ABST
Patent Text Reader

Abstract

Provided is an image decoding method for decoding a target block by using an inter-component prediction model. This image decoding method comprises the steps of: deriving at least one inter-component prediction model of a current block; and predicting color difference samples of the current block from luminance samples of the current block by using the at least one inter-component prediction model of the current block.
Need to check novelty before this filing date? Find Prior Art

Description

Method, apparatus, and recording medium for image encoding / decoding using an inter-component prediction model

[0001] The present disclosure relates to a method, apparatus, and recording medium for image encoding / decoding. Specifically, the present invention relates to a method, apparatus, and recording medium for image encoding / decoding using a component-to-component prediction model.

[0002] With the continuous development of the information and communication industry, services providing video through broadcasting and the Internet have spread globally.

[0003] Users demand videos with higher resolution and quality. To meet these user demands, video encoding and decoding technologies suitable for such videos are required. Video encoding technology can generate compressed video by compressing the video representing the images to have a smaller amount of data. Video decoding technology can generate reconstructed images using the compressed video.

[0004] Regarding video encoding and decoding technologies, various techniques exist, such as segmentation, prediction, transformation, quantization, filtering, and entropy encoding and decoding. By introducing, modifying, improving, and combining these various techniques, video and images can be compressed, transmitted, and stored more effectively.

[0005] One embodiment of the present disclosure provides a method and apparatus for image encoding / decoding using a component-to-component prediction model to improve encoding efficiency.

[0006] In addition, one embodiment of the present disclosure provides a method for transmitting or storing a bitstream generated by an image encoding method, or a recording medium for storing a bitstream.

[0007] One aspect of the present disclosure provides an image decoding method for decoding a target block using a component-to-component prediction model. The method may include the step of deriving at least one component-to-component prediction model of a current block; and the step of predicting color difference samples of a current block from luminance samples of a current block using at least one component-to-component prediction model of a current block.

[0008] Another aspect of the present disclosure provides an image encoding method for encoding a target block using a component-to-component prediction model. The method may include the steps of: deriving at least one component-to-component prediction model of a current block; predicting color difference samples of a current block from luminance samples of a current block using at least one component-to-component prediction model of a current block; and encoding information about at least one component-to-component prediction model of a current block.

[0009] Another aspect of the present disclosure provides a method for providing a bitstream generated by the image encoding method to an image decoding device or storing it in a computer-readable recording medium.

[0010] Another aspect of the present disclosure provides a computer-readable recording medium that stores a bitstream generated by the image encoding method.

[0011] FIG. 1 shows a system for video coding according to one embodiment.

[0012] Figure 2 shows a segmentation structure of an image according to one embodiment.

[0013] Figure 3 shows the structure of an intra prediction according to one embodiment.

[0014] FIG. 4 shows the structure of an inter prediction to explain an inter prediction process according to one embodiment.

[0015] FIG. 5 shows the order of addition of spatial candidates to the candidate list according to one embodiment.

[0016] Figure 6 shows a plurality of in-loop filters according to one example.

[0017] Figure 7 shows the structure of entropy encoding and entropy decoding according to one example.

[0018] FIG. 8 is a flowchart of an image encoding / decoding method that generates a color difference prediction signal using the component-to-component prediction model of the present disclosure.

[0019] Figure 9 is a diagram illustrating the number of reference lines and reference samples according to the block shape.

[0020] FIG. 10 is a diagram illustrating how to adjust the area of ​​reference samples so that a location found using Intra Template Matching (Intra TMP) / IBC is included in the reference area.

[0021] FIG. 11 is a diagram illustrating the determination of a reference region using block vectors derived from collocated blocks or surrounding blocks.

[0022] FIG. 12 is a diagram illustrating the derivation of a block vector from a block at a specific location within a corresponding block.

[0023] Figures 13a and 13b illustrate examples of template configurations.

[0024] Figures 14a, 14b, 14C, and 14D are examples of grouping samples within a range of sample values ​​between a minimum sample value and a maximum sample value.

[0025] Figure 15 illustrates an example of surrounding blocks adjusted (moved) using block vectors.

[0026] FIG. 16 illustrates an embodiment of deriving a new candidate (vector) using a block vector or motion vector stored in a block at a location indicated by a previously derived block vector or motion vector in a list.

[0027] Figure 17 illustrates an example of adjacent surrounding restoration samples for filtering prediction samples within a prediction block.

[0028] FIG. 18 illustrates an example of a downsampled luminance sample L'(i,j) and surrounding samples corresponding to a color difference sample C(i,j).

[0029] Figure 19 illustrates the positions of the downsampled luminance sample and the corresponding color difference sample.

[0030] FIG. 20 illustrates an example of using multiple downsampling filters to generate luminance samples corresponding to color difference samples.

[0031] Figure 21 illustrates a model that generates a color difference prediction signal and the information used by the model.

[0032] FIGS. 22a and FIGS. 22b are examples of multiple pattern models.

[0033] Figure 23 illustrates an example of chrominance signal prediction based on a component-to-component residual model in inter-frame prediction.

[0034] FIGS. 24a and FIGS. 24b illustrate examples of a luminance block and a color difference prediction block that directly correspond to a color difference block.

[0035] Figure 25 illustrates an example of a color difference prediction block.

[0036] Figure 26 illustrates an example of a reference area used to derive the model and an area required to calculate the error cost between the color difference prediction signal and the color difference signal.

[0037] Figures 27a, 27b, and 27c illustrate an example of calculating the error cost between the color difference prediction template and the color difference template generated for each model.

[0038] Figure 28 illustrates an example of reconstructing a model-based color difference prediction mode table.

[0039] FIGS. 29a and FIGS. 29b illustrate examples of filter or filter set configurations according to block size.

[0040] FIGS. 30a, FIGS. 30b, FIGS. 30c, and FIGS. 30d illustrate examples of filter set configurations according to the direction of the in-screen prediction mode.

[0041] Figure 31 illustrates an example of varying the filter depending on the prediction direction of the prediction mode within the screen or the position of the samples used to create the prediction mode.

[0042] Figure 32 illustrates classifying the list of candidates for component-to-component prediction models by group.

[0043] FIGS. 33a and FIGS. 33b illustrate surrounding adjacent / non-adjacent blocks from which model information and encoding information within a reference picture will be retrieved.

[0044] Figures 34 and 35 illustrate an example of sorting candidates within a candidate list based on color difference prediction costs.

[0045] Figure 36 illustrates the model and template matching costs according to the type of component-to-component prediction model within the candidate list.

[0046] FIG. 37 illustrates an example of creating a new model by interpolating / extending model parameters between at least two candidates in a candidate list.

[0047] FIG. 38 illustrates an example of weighted summing of color difference prediction signals / blocks / samples generated using multiple models.

[0048] Figure 39 illustrates an example of using different models for each part to obtain predicted color difference blocks.

[0049] FIG. 40 illustrates an example of applying filtering when blending color difference prediction blocks generated using multiple models.

[0050] FIG. 41 illustrates an example of obtaining a final predicted color difference block using multiple models.

[0051] FIG. 42 illustrates an example of applying filtering when obtaining a final predicted color difference block using multiple models.

[0052] Various modifications may be applied to the present invention. Additionally, the present invention may have various embodiments. Specific embodiments are described by the drawings and the detailed description.

[0053] Specific embodiments are not intended to limit the invention to specific embodiments, and it should be understood that all modifications, equivalents, and substitutions that fall within the spirit and scope of the invention are included as embodiments of the invention.

[0054] The embodiments are described in sufficient detail to enable those skilled in the art to practice the embodiments. It should be understood that the various embodiments are different but need not be mutually exclusive. For example, it should be understood that the shapes, structures, and characteristics described in relation to one embodiment may be applied to or implemented in other embodiments without departing from the spirit and scope of the invention. It should also be understood that the location or arrangement of components within one embodiment may be changed without departing from the spirit and scope of the invention. Accordingly, the following detailed description is not intended to be limiting, and the scope of the exemplary embodiments is limited only by the appended claims and all equivalents to the scope claimed by such claims, provided that they are appropriately described.

[0055] The detailed description of the embodiments described below may refer to the drawings relating to the embodiments. Descriptions described in the drawings or descriptions represented by the drawings may be considered part of the detailed description. In the drawings, similar reference numerals may refer to the same or similar functions for various aspects. Dependencies between components may not be limited to those depicted in the drawings.

[0056] In the embodiments, singular expressions may include plural expressions and may be limited to and / or limited to plural expressions unless the context clearly excludes plural expressions. That is to say, in the embodiments, expressions such as 'at least one' and 'one or more' may be replaced with 'plural'. Terms such as ' / ', 'and / or', 'at least one of' and 'one or more of' described for plural items may mean 1) one of the plural items, 2) some of the plural items, 3) a combination of some of the plural items, or 4) a combination of the plural items. Additionally, plural expressions may be replaced with singular expressions. Plural may mean an integer of 1, 2, 3, 4, or 5 or more.

[0057] In the embodiments, numbered terms such as 'first' and 'second' may be used to describe various components. These terms are used solely for the purpose of distinguishing one component from another and do not limit the components. For example, without departing from the scope of the present invention, the first component may be named the second component, and similarly, the second component may be named the first component.

[0058] The statement that a first component transmits (or provides) information to a second component may mean that the first component directly transmits information to the second component, or it may mean that the first component transmits information to the second component through another third component. Here, the information received (or acquired) by the second component may be information transmitted by the first component, or information generated by applying a specific processing to information transmitted by the first component.

[0059] The components of the embodiments may be illustrated independently to represent different characteristic functions, and this does not imply that each component corresponds to a separate hardware or a single software unit. That is, the components of the embodiments may be classified and enumerated for convenience of description. Two or more components described in the embodiments may be regarded as a single component. Furthermore, a single component described in the embodiments may be separated into multiple components that perform the functions of the said component separately. Embodiments in which such components are integrated and embodiments in which components are separated are also included within the scope of the present invention, provided that they do not depart from the essence of the invention.

[0060] The terms used in the embodiments are used merely to describe specific embodiments and are not intended to limit the invention. In the embodiments, terms such as "comprising" or "having" indicate the presence of features, numbers, steps, actions, components, parts, or combinations thereof described in the embodiments. The existence or addition of other features, numbers, steps, actions, components, parts, or combinations thereof not explicitly described in the embodiments is not excluded by these terms. That is, the description of a specific component of an embodiment as "comprising" does not exclude components other than the specific component, and means that additional components may also be included within the scope of the embodiments or the technical concept of the invention.

[0061] Some of the components of the embodiments may be optional components that are not essential for performing the essential functions of the invention. Such optional components may be used to enhance performance. The embodiments may be implemented as a structure comprising only the essential components required to realize the essence of the embodiments, excluding the optional components. Such a structure is also included within the scope of the embodiments.

[0062] In the following, embodiments are described in detail with reference to the attached drawings so that a person skilled in the art can easily implement the embodiments. In describing the embodiments, if it is determined that a detailed description of related known configurations or known functions could obscure the gist of this specification, such detailed description is omitted. Additionally, the same reference numerals are used for identical components within the drawings, and redundant descriptions of identical components are omitted.

[0063]

[0064] Replacement of terms in the examples

[0065] Below, terms listed in a single line may be used with the same meaning in the examples and may be used interchangeably in the examples.

[0066] - 'one or more', 'at least one'

[0067] - 'two or more', 'a plurality of', 'multiple', 'multiple'. (In the examples, 'one or more' or 'at least one' may be further limited to 'two or more', 'multiple', or 'multiple'.)

[0068] - 'Information', 'Signal'

[0069] - 'value', 'predefined value', 'specific value', 'threshold', 'threshold value', 'baseline value', 'reference value'

[0070] - 'statistical value', 'statistics value'

[0071] - 'indicator', 'index', 'index', 'flag', 'information'

[0072] - 'encoder', 'encoding apparatus'

[0073] - 'decoder', 'decoding apparatus'

[0074] - 'Entropy encoding', 'encoding', 'encoding'

[0075] - 'Entropy decoding', 'decoding', 'decoding'

[0076] - 'Coding', 'Encoding and / or decoding'

[0077] - 'video', 'moving picture', 'image', 'picture', 'picture', 'frame', 'screen'

[0078] - 'Reference picture', 'Reference video'

[0079] - 'Reference Picture List (RPL), 'Reference Image List'

[0080] - 'original', 'input', 'source'

[0081] - 'Block', 'Unit', 'Signal'

[0082] - 'square', 'square shape'

[0083] - 'pixel', 'pixel', 'sample', 'pel'

[0084] - 'region', 'area', 'part', 'segment'

[0085] - 'partition', 'split', 'divide'

[0086] - 'quad', 'quadronary'

[0087] - 'Luma component', 'Luma', 'luminance component', 'luminance', 'Y'

[0088] - 'Chroma component', 'Chroma', 'chrominance', 'chrominance component', 'Cb and Cr', 'Cb or Cr', 'Cb', 'Cr', 'U and V', 'U or V', 'U', 'V'

[0089] - 'target', 'current' (e.g., target block and current block, or target image and current image)

[0090] - 'neighbor', 'neighboring', 'adjacent', 'neighbor / neighboring' (e.g., neighbor block, adjacent block, and neighboring block)

[0091] - 'collocated', 'COL'

[0092] - 'reconstruction', 'reconstruction', 'decoding'

[0093] - 'reconstructed', 'reconstructed', 'decoded'

[0094] - 'Difference', 'Difference', 'Difference', 'Error', 'Residual', 'Residual'

[0095] - Largest Coding Unit (LCU), Coding Tree Unit (CTU)

[0096] - 'inter', 'inter-screen'

[0097] - 'Inter prediction', 'inter prediction', 'motion compensation'

[0098] - 'Inter Mode', 'Inter Prediction Mode', 'Inter-frame Mode', 'Inter-frame Prediction Mode'

[0099] - 'Motion Vector', 'Predicted Motion Vector', 'Advanced Motion Vector Prediction (AMVP)'

[0100] - 'List', 'Candidate List'

[0101] - 'spatial candidate', 'spatial merge candidate'

[0102] - 'temporal candidate', 'temporal merge candidate'

[0103] - 'prediction motion vector candidate', 'motion vector predictor'

[0104] - 'Prediction method', 'Prediction mode'

[0105] - 'Intra', 'Inside the screen'

[0106] - 'Intra prediction', 'Intra prediction'

[0107] - 'Intra Mode', 'Intra Prediction Mode'

[0108] - 'Dequantization', 'Scaling'

[0109] - 'Quantization matrix', 'Scaling list'

[0110] - 'Quantization matrix coefficients', 'Matrix coefficients'

[0111] - 'transform coefficient level', 'quantized level', 'quantized coefficient', 'quantized transform coefficient', 'quantized transform coefficient level'

[0112] - 'dequantized coefficient', 'dequantized transform coefficient'

[0113] - 'Scanning type', 'Scanning direction'

[0114] - 'directional mode', 'angle mode', 'angular mode', 'intra-prediction mode'

[0115] - 'Intra-prediction mode (mode) number', 'Intra-prediction mode (mode) index', 'Intra-prediction mode (mode) value', 'Intra-prediction mode (mode) angle', 'Intra-prediction mode (mode) direction', 'Intra-prediction direction (mode) number', 'Intra-prediction direction (mode) index', 'Intra-prediction direction (mode) value', 'Intra-prediction direction (mode) angle'

[0116] - 'Merge Mode', 'Motion Merge Mode'

[0117] - 'Geometric Partitioning Mode (GPM)', 'Triangle Partitioning Mode'

[0118] In addition to the terms exemplified above, terms having the same meaning according to the ordinary knowledge of the technical field may be used interchangeably in the embodiments.

[0119]

[0120] Information and range of values ​​of information described in the embodiments

[0121] In the embodiments, information may include a constant, a flag, an index, a variable, a coding parameter, an element, a syntax element, motion information, an attribute, an entity, an object, and data, etc. That is to say, the term 'information' may be interchangeable with 'data', 'flag', 'index', 'variable', 'element', 'syntax element', 'motion information', 'attribute', or 'entity'.

[0122] Information can have one of multiple values. 'The nth value' can mean the nth value among multiple values.

[0123] For example, the first value can represent '0' or (logical) false. The second value can represent '1' or (logical) true. Or, the first value can represent '1' or (logical) true. The second value can represent '0' or (logical) false.

[0124] A flag may be information having a value of either '0' or '1'. In the embodiments, the values ​​'0' and '1' of the flag may be replaced with '1' and '0', respectively. For example, information indicating whether a specific process is performed or information indicating whether a specific process is applied may be considered as a flag.

[0125] When a variable such as i or j is used to represent a row, column, or index, the variable may be an integer between 0 and n - 1 inclusive. Or, the variable may be an integer between 1 and n inclusive. Here, n may be the number of rows, the number of columns, or the number of entities pointed to by the index.

[0126]

[0127] Concepts related to coding

[0128] Concepts related to coding are explained below. The descriptions disclosed below may be applied to embodiments.

[0129] Predefined value: A predefined value may refer to a value commonly used by the encoding device and the decoder. For example, a predefined value may be interpreted as being limited to a fixed value. Alternatively, a predefined value may be a value shared by the encoding device and the decoder through signaling. Alternatively, a predefined value may be a value derived through the same procedure in the encoding device and the decoder so that the encoding device and the decoder have a common value. Alternatively, a predefined value may be a common value possessed by the encoding device and the decoder. The above description of a predefined value may also apply to predefined information. In the above descriptions, 'value' may be replaced with 'information'.

[0130] - Values ​​derived through the same procedure in the above-mentioned encoding device and decoding device may include values ​​derived through the same procedure for the same value and / or the same information in the encoding device and decoding device.

[0131] - Values ​​derived through the same procedure in the above-mentioned encoding device and decoding device may include values ​​derived using the same conditional statement for the same value and / or the same information in the encoding device and decoding device.

[0132] - The description of the predefined values ​​above may also apply to predefined information. In the descriptions above, 'value' may be replaced with 'information'.

[0133] Availability: The availability of specific modes for a specific target may mean that a selected mode among the specific modes is used for that specific target. Other modes belonging to the category of specific modes may be non-available modes. Non-available modes may not be used for a specific target. The above description of specific modes may also apply to other specific information. In the above descriptions, 'mode' may be replaced with 'information'.

[0134] Adjacency: 'Direction' for 'First Object'. 'Second Object' may refer to a 'Second Object' adjacent to the 'Direction' corner / face of the First Object. For example, the 'Top-left Block' for a 'Target Block' may be a block adjacent to the top-left of the Target Block. Here, the 'First Object' may be a Target Unit, Target Block, or Target Sample. 'Direction' may be one of left-above, above, right-above, left, right, left-below, below, and right-below. The 'Second Object' may be a Unit, Block, or Sample. For the directions of top-left, top-right, bottom-left, and bottom-right, the corner of the First Object and the corner of the Second Object may be diagonally adjacent. For the directions of top, left, right, and bottom, one face of the First Object and one face of the Second Object may be in contact with each other.

[0135] - For example, the block adjacent to the top-left of the target block may be the block adjacent to the top of the block adjacent to the left of the target block. The block adjacent to the top-right of the target block may be the block adjacent to the right of the block adjacent to the top of the target block. The block adjacent to the bottom-left of the target block may be the block adjacent to the bottom of the block adjacent to the left of the target block.

[0136] Coding: Coding can refer to encoding and / or decoding of an image.

[0137] Signal: A signal can represent information about an image, unit, or block. A specific signal can represent a specific image, a specific unit, or a specific block.

[0138] Image: An image can refer to a single picture constituting a video, or it can represent the video itself. For example, "encoding and / or decoding of an image" can mean "encoding and / or decoding of a video," or it can mean "encoding and / or decoding of one of the images constituting a video."

[0139] - An image can refer to the entirety of a picture, or it can refer to a part of a picture, such as a block.

[0140] Target image: The target image may be an encoding target image that is the subject of encoding and / or a decoding target image that is the subject of decoding. Additionally, the target image may be an input image processed by an encoding device and a restored image processed by a decoding device. The target image may be an image containing a target block.

[0141] Subpicture: A picture can be divided into one or more subpictures.

[0142] - A subpicture may be a square or rectangular area within the picture. A subpicture may include one or more CTUs.

[0143] - A subpicture may include one or more slices and / or one or more tiles. For example, a subpicture may consist of one or more slice rows and one or more slice columns. Alternatively, each subpicture may consist of one or more tile rows and one or more tile columns.

[0144] - A subpicture may include one or more slices that collectively cover a rectangular area within the picture. Accordingly, the boundary of each subpicture can always be the boundary of a slice. Additionally, each vertical subpicture boundary can always be the boundary of a vertical tile.

[0145] Slice: A slice may include one or more tiles within a picture. A slice may consist of one or more rows of tiles and one or more columns of tiles.

[0146] Tile: A tile can be a square or rectangular area within a picture. A tile can contain one or more CTUs. A picture can be divided into one or more tile rows and one or more tile columns.

[0147] CTU: An image can be divided into multiple Coding Tree Units (CTUs).

[0148] - A CTU may include one Y Coding Tree Block (CTB) and at least one of a Cb CTB and a Cr CTB associated with the Y CTB, and may include information for each CTB. The information may include syntax elements.

[0149] - Each CTU may be partitioned using one or more partitioning methods to form sub-units such as Coding Units (CU), Prediction Units (PU), and Transform Units (TU). One or more partitioning methods may include Quad Tree (QT) partitioning, Binary Tree (BT) partitioning, and Ternary Tree (TT) partitioning. Additionally, each CTU may be partitioned using Multi-Type Tree (MTT) partitioning, which uses a combination of multiple partitioning methods.

[0150] CTB: CTB can refer to one of Y CTB, Cb CTB, and Cr CTB.

[0151] Unit: A unit can be determined for specific processing in coding. A unit may be information about a specific region within an image. For specific processing in coding, the image may be recursively divided into multiple parts. A unit may represent the region to which the specific processing is applied and information about the aforementioned region.

[0152] - The unit type may represent a specific process applied to the unit. Depending on the unit type, a specific process may be applied to the unit. The 'specific' unit may be a unit for the process named 'specific' in the coding. For example, the unit may be at least one of the source unit, CTU, coding unit, prediction unit, residual unit, restored residual unit, transformation unit, and restored unit.

[0153] - A unit may include samples having a two-dimensional form or arrangement. In this respect, a 'unit' may mean a 'block'. For example, a block may be at least one of an original block, a CTB, a coding block (CB), a prediction block (PB), a residual block, a restored residual block, a transform block (TB), and a restored block. For example, a partition of a unit may mean a partition of a block corresponding to the unit.

[0154] - A unit may include syntactic elements. In other words, a block and the syntactic elements for the block can be combined and referred to as a unit.

[0155] - A block is an MxN array of samples. Here, M and N can represent positive integer values, and a block can commonly represent a two-dimensional array of samples. The current block can represent the encoding target block that is the subject of encoding during encoding, or the decoding target block that is the subject of decoding during decoding. Additionally, the current block can be at least one of a coding block, a prediction block, a residual block, a transformation block, or a restoration block. Blocks can have various sizes and shapes. For example, the shape of a block can be one or more of a tetragon, a rectangular, a square, a rectangle where the width differs from the height (i.e., an oblong), a trapezoid, a triangle, a right-angled triangle, and a pentagon. Here, the width and height of the rectangle can differ from each other. Additionally, the shape of a block may include other geometric figures that can be represented in two dimensions. For example, the shape of the block may be a square or a pentagon defined by subtracting the area of ​​a right triangle from the area of ​​a rectangle. Here, the right-angled vertex of the right triangle may be one of the vertices of the rectangle. Additionally, the shape of the block may be a combination of two or more of the aforementioned shapes. Additionally, the shape of the block may be the remainder of one of the aforementioned shapes after another shape has been subtracted.

[0156] - In the embodiments, the rectangle may be limited to a non-square rectangle. When the shape of a specific object in the embodiments is described as a rectangle, this description may additionally imply that the width and height of the specific object are different from each other.

[0157] - In the embodiments, the block may be limited to at least one of a vertically oriented block and a horizontally oriented block. A vertically oriented block may mean a block in which the vertical length is greater than the horizontal length. A horizontally oriented block may mean a block in which the horizontal length is greater than the vertical length.

[0158] - The unit may include a luma component block (i.e., a Y block) and two chroma component blocks (i.e., at least one of a Cb block and a Cr block), and may include information for each block. The information may include syntax elements.

[0159] - The unit information may include the unit type, unit size, unit depth, unit encoding order, and unit decoding order.

[0160] Target Unit: The target unit may be a block that is the target of encoding, an encoding target unit, and / or a decoding target unit that is the target of decoding. The target unit may be a specific region within the target picture to which one or more specific processing steps of coding are applied. By applying a specific processing step to the target unit, a unit of a specific type may be generated. Alternatively, the target unit may represent a unit having a specific type for a specific processing step of coding.

[0161] Depth: A block can be hierarchically divided into multiple sub-blocks with depth according to a tree structure. The multiple sub-blocks created by the division of a block can be referred to as partitions.

[0162] - The block depth can represent the level of the node corresponding to the block when the blocks constituting the image are represented as a tree structure. Alternatively, the block depth can represent the number of divisions applied until the block is determined. The block depth can increase by 1 as the block is further divided.

[0163] - In a tree structure, the root node can be considered to have the smallest level, and the leaf node the largest level. The root node may be the top node of the tree structure and may correspond to the first undivided block. The level of the root node may be 0 or 1. When the level of the root node is 0, a node with level 1 may represent the block determined by the first block being divided once. A node with level n may represent the block determined by the first block being divided n times. A leaf node may be the lowest node of the tree structure. A leaf node may be a node that cannot be further divided. The depth of a leaf node may be a predefined maximum depth. For example, the maximum depth may be a positive integer such as 3. The root node may represent a CTU. A leaf node may represent at least one of CU, PU, ​​or TU.

[0164] - Depth can have a type depending on the type of partition. QT depth can represent the depth for quadtree partitioning. BT depth can represent the depth for binary partitioning. TT depth can represent the depth for ternary partitioning.

[0165] Sample: A sample can be a base unit that constitutes a block. A sample can consist of one or more bits. Bit depth can be the number of bits that make up the sample. Samples range from 0 to 2 depending on the bit depth. Bd It can be expressed as values ​​up to -1.

[0166] PU: PU may refer to a base unit for processing related to prediction. For example, processing related to prediction may include inter-prediction, intra-prediction, intra-block copy (IBC) prediction, intra-compensation, and motion compensation.

[0167] A single PU can be divided into multiple sub-PUs that are smaller in size compared to the PU. These multiple sub-PUs can also serve as base units for processing related to prediction. In other words, a prediction unit partition generated by the division of the prediction unit can also be a prediction unit.

[0168] TU: A TU may be a base unit for processing related to a residual block. Processing related to a residual block may include at least one of transform, inverse transform, quantization, inverse quantization, transform coefficient encoding, transform coefficient decoding, entropy encoding, and entropy decoding. A single TU may be divided into a plurality of sub-transform units having a size smaller than that of the TU. The plurality of sub-TUs may also be base units for processing related to a residual block. That is to say, a transform unit partition generated by the division of the transform unit may also be a transform unit.

[0169] - The transformation may include one or more of a primary transformation and a secondary transformation, and the inverse transformation may include one or more of a primary inverse transformation and a secondary inverse transformation.

[0170] Parameter set: The parameter set can correspond to header information within the structure of the bitstream.

[0171] - The parameter set may include at least one of a Video Parameter Set (VPS), a Sequence Parameter Set (SPS), a Picture Parameter Set (PPS), an Adaptation Parameter Set (APS), and a Decoding Parameter Set (DPS).

[0172] Information signaled through a parameter set can be applied to pictures that reference the parameter set. For example, information within a VPS can be applied to pictures that reference the VPS. Information within an SPS can be applied to pictures that reference the SPS. Information within a PPS can be applied to pictures that reference the PPS. A parameter set can reference a higher-level parameter set. For example, a PPS can reference an SPS. An SPS can reference a VPS.

[0173] - Additionally, the parameter set may include tile group information, slice header information, and tile header information. A tile group may refer to a group or slice containing multiple tiles.

[0174] MPM (Most Probable Mode): MPM may represent an intra prediction mode that is likely to be used for intra prediction of a target block.

[0175] - One or more different MPMs can be determined based on coding parameters related to the target block and attributes of objects related to the target block.

[0176] - One or more MPMs may be determined based on the intra prediction mode of a reference block. There may be multiple reference blocks. One or more different MPMs may be determined depending on which intra prediction modes are used for one or more reference blocks. Reference blocks may include spatial neighbor blocks.

[0177] MPM List: An MPM list may be a list containing one or more MPMs. The number of one or more MPMs in an MPM list may be predefined.

[0178] MPM Index: The MPM index can indicate one or more MPMs in the MPM list that are used for intra prediction for the target block.

[0179] MPM Usage Indicator: The MPM Usage Indicator can indicate whether an MPM list is used for prediction regarding a target block.

[0180] Prediction mode: The prediction mode may be information indicating a prediction method for a target block, such as a mode used for intra-prediction or a mode used for inter-prediction. The prediction mode may refer to one of the prediction-related modes described in the embodiments. Additionally, the prediction mode may include at least one of an intra-mode, an inter-mode, and an intra-block copy mode.

[0181] Reference image list: The reference image list may be a list containing one or more reference images used for prediction of the target block.

[0182] - There may be multiple reference image lists. Multiple reference image lists may include List 0 (List 0; L0), List 1 (List 1; L1), etc.

[0183] - One or more reference image lists may be used for inter prediction for the target block. Parts such as 'L0' and 'L1' in the names of the information related to inter prediction may refer to the reference image lists associated with the information.

[0184] Reference picture: The reference picture may be an image referenced for prediction regarding the target block. Alternatively, the reference picture may be an image containing the reference block. The reference picture may include an image prior to the target image, the target image, and an image following the target image.

[0185] Reference image index: The reference image index may be an index indicating one reference image among one or more reference images in the reference image list that is used for prediction of the target block.

[0186] Reference Block: A reference block may be a block referenced for encoding / decoding of a target block, such as for prediction and filtering. For example, a reference block may include a reference sample referenced to derive a prediction sample, and may refer to a block that provides information used for decoding the target block.

[0187] Reference Sample: A reference sample may be a sample referenced for encoding / decoding of a target block, such as prediction and filtering.

[0188] Inter prediction indicator: The inter prediction indicator may indicate the direction of inter prediction for the target block. Inter prediction may be one of unidirectional prediction and bidirectional prediction. Alternatively, the inter prediction indicator may indicate the number of reference images used when generating prediction blocks for the target block. Alternatively, the inter prediction indicator may indicate the number of prediction blocks used for inter prediction for the target block. The reference direction may refer to the inter prediction indicator. For example, the inter prediction indicator may indicate either unidirectional or bidirectional. Alternatively, for an inter mode that uses only reference images within the L0 reference image list, the inter prediction indicator may have a first value of '0'; for an inter mode that uses only reference images within the L1 reference image list, the inter prediction indicator may have a second value of '1'; and for an inter mode that uses at least two of the reference images within the L0 reference image list and the L1 reference image list, the inter prediction indicator may have a third value of '2'.

[0189] Prediction List Utilization Flag: The prediction list utilization flag for a specific reference image list may indicate whether at least one reference image within that specific reference image list is used to generate the prediction block of the target block. For example, a value of '0' for the prediction list utilization flag for a specific reference image list may indicate that the prediction block is not generated using the reference images within that specific reference image list. A value of '1' for the prediction list utilization flag for a specific reference image list may indicate that the prediction block is generated using the reference images within that specific reference image list.

[0190] - An inter-prediction indicator can be derived using prediction list utilization flags. Conversely, an inter-prediction indicator can be derived using prediction list utilization flags. For example, an inter-prediction indicator can be derived using prediction list utilization flags for multiple reference image lists. If the inter-prediction indicator indicates that specific reference lists among the multiple reference image lists are being used, the prediction list utilization flags of the specific reference lists pointed to by the inter-prediction indicator among the prediction list utilization flags of the multiple reference image lists can be set to '1', and the prediction list utilization flags of the remaining reference image lists not pointed to by the inter-prediction indicator can be set to '0'.

[0191] Reference Direction: The reference direction may point to a list of reference images used for the prediction of the target block. For example, the reference direction may point to one or more of reference image list L0 and reference image list L1. The reference direction may be used interchangeably with "inter-frame prediction direction" and may be substituted for each other.

[0192] - The reference direction merely refers to the list of reference images used for prediction of the target block, and does not indicate that the directions of the reference images within the list are restricted to a forward direction or a backward direction. That is to say, each of the reference image list L0 and the reference image list L1 may include forward images and backward images, respectively. Here, the forward direction may indicate a direction from the target image to the image preceding the target image. Forward inter-prediction may be an inter-prediction that uses the image preceding the target image as a reference image. The backward direction may indicate a direction from the target image to the image following the target image. Backward inter-prediction may be an inter-prediction that uses the image following the target image as a reference image.

[0193] - A unidirectional reference direction may mean that a single reference image list is used. A bidirectional reference direction may mean that two reference image lists are used. For example, the reference direction may indicate one of the following: that only reference image list L0 is used, that only reference image list L1 is used, or that two reference image lists are used. Additionally, the reference direction may be indicated by an inter-predictor.

[0194] Picture Order Count (POC): The POC of a picture can represent the display order or output order of the picture.

[0195] Motion information: Motion information may be information used to specify a reference block. Motion information may include information used for inter prediction, such as a motion vector (MV), reference image index, reference image, inter prediction indicator, prediction list utilization flag, etc. Additionally, motion information may include information used in a specific inter prediction mode, such as an MV candidate, MV candidate index, merge candidate, and merge index. Additionally, motion information may include information related to the block vector described below. Information related to the block vector may mean information including at least one of a block vector, a block vector candidate, and a block vector candidate index.

[0196] - Multiple motion information for multiple reference image lists may be used for inter-prediction of the target block. Motion information for a specific reference image list may be used for prediction using that specific reference image list. Multiple (intermediate) prediction blocks may be derived from the multiple motion information. A (final) prediction block for the target block may be generated using statistical values ​​for the multiple (intermediate) prediction blocks.

[0197] MV: MV can be a 2-dimensional vector used in inter-prediction. MV can represent the offset between the target block and the reference block. Alternatively, MV can represent the difference between the location of the target block and the location of the reference block.

[0198] - For example, MV is (mv x , mv y It can be expressed in the form of ). mv x can represent a horizontal component, and mv yIt can represent a vertical component.

[0199] - The zero vector can be (0, 0) MV.

[0200] Block Vector (BV): A BV can be a two-dimensional vector used in intra-block copy prediction. A BV can represent the offset between a target block within a target image and a reference block within a target image. In other words, a BV can represent the displacement between a target block and a reference block within a target image.

[0201] - For example, BV is similar to MV (bv x , bv y It can be expressed in the form of ). bv x can represent a horizontal component, and bv y It can represent a vertical component.

[0202] - The zero vector can be (0, 0) BV.

[0203] Motion Information Candidates: In a specific prediction, motion information of the target block can be selected from motion information candidates determined by a specific method. A motion information candidate may refer to the motion information of a reference block, or it may refer to the reference block itself that possesses motion information. Here, the reference block may be a block determined by a specific method to select motion information candidates.

[0204] Candidate List: A candidate list may be a list containing one or more candidates. For example, a candidate list may include a motion information candidate list, a merge candidate list, an MV candidate list, an MPM list, etc. A candidate list may be generated in the same manner in both the encoding device and the decoder. That is to say, the candidate list used in the encoding device and the candidate list used in the decoder may be identical, and the same candidate list may be shared between the encoding device and the decoder. The encoding device may select a candidate from among the candidates in the candidate list to be used for processing the target block. An indicator pointing to the selected candidate may be signaled from the encoding device to the decoder. The decoder may use the indicator to identify the candidate from among the candidates in the candidate list to be used for processing the target block. Alternatively, the encoding device and the decoder may identify the candidate from among the candidates in the candidate list to be used for processing the target block by the same rule.

[0205] Motion Information Candidate List: A motion information candidate list may refer to a list constructed using one or more motion information candidates.

[0206] Motion Information Candidate Index: The motion information candidate index may be an identifier or indicator pointing to a motion information candidate among the motion information candidates in the motion information candidate list that is used for prediction regarding the target block.

[0207] - In a specific inter-prediction mode, motion information of other restored blocks may be used to derive motion information of the target block. Other blocks may include neighboring blocks. In this specific inter-prediction mode, the motion information for the target block itself is not signaled individually, but other information used to derive motion information of the target block based on motion information of other restored blocks may be signaled. In this case, the other information may include information indicating which of the other restored blocks' motion information is used to derive motion information of the target block, such as a motion information candidate index.

[0208] - For example, these inter-prediction modes may include AMVP mode, merge mode, and skip mode. The motion information candidate index may be a merge index or an MV candidate index.

[0209] - In the embodiments, MV may be part of the motion information. In the embodiments, information about motion information, such as motion information candidates, a list of motion information candidates, and an index of motion information candidates, may be replaced with information about MV, such as MV candidates, a list of MV candidates, and an index of MV candidates, and descriptions of motion information may also be applied to MV.

[0210] Merge: Merge can refer to the merging of motion information for multiple blocks, or it can refer to applying the motion information of one block to a target block as well. In other words, merge mode can refer to a mode where the motion information of a target block is derived from the motion information of a neighboring block.

[0211] Merge Candidate: A merge candidate may refer to a specific (restored) block used for merging with a target block, or it may refer to movement information of a specific block. Alternatively, a merge candidate may include movement information of a specific block.

[0212] - Merge candidates for the target block may include spatial merge candidates, temporal merge candidates, history-based candidates, average candidates based on the average of two merge candidates, and zero merge candidates.

[0213] Merge candidate list: The merge candidate list may be a list composed of one or more merge candidates.

[0214] Merge Index: The merge index may be an indicator pointing to a merge candidate among the merge candidates in the merge candidate list that is used for prediction regarding the target block. Among the merge candidates in the merge candidate list, the movement information of the merge candidate indicated by the merge index may be used as movement information for the target block.

[0215] Neighbor block: A neighbor block may refer to a block adjacent to the target block. Neighbor blocks may include spatial and temporal neighbor blocks. A neighbor block may also refer to a reconstructed neighbor block within the reference image. A neighbor block does not necessarily have to be in direct contact with the target block.

[0216] Spatial neighbor blocks: Spatial neighbor blocks can be blocks that are spatially adjacent to the target block.

[0217] - The target block and spatial neighbor blocks can be included within the target image.

[0218] - Spatial neighbor blocks may include blocks whose boundaries, at least a portion of which abuts at least a portion of the target block's boundary. Alternatively, spatial neighbor blocks may include blocks whose distance from the target block is less than or equal to a specific value.

[0219] - Spatial neighbor blocks may include blocks diagonally adjacent to the vertices of the target block.

[0220] - Spatial neighbor blocks may include a top-left block adjacent to the top-left of the target block, a top block adjacent to the top of the target block, a top-right block entered at the top-right of the target block, a left block adjacent to the left of the target block, a right block adjacent to the right of the target block, a bottom-left block adjacent to the bottom of the target block, and a bottom-right block adjacent to the bottom-right of the target block.

[0221] Temporal neighbor blocks: Temporal neighbor blocks can be blocks that are temporally adjacent to the target block.

[0222] - Temporal neighbor blocks may include a collocated block (COL block). A collocated block may be a block within a restored image in a reference image buffer. A collocated picture (col picture) may refer to an image containing a collocated block. A collocated picture may be an image included in a reference image list.

[0223] - Call blocks can be determined based on the location of target blocks within the target image. Two blocks being 'temporarily adjacent' may mean that the locations of the two blocks satisfy certain conditions.

[0224] - The position of the call block within the call image may be the same as the position of the target block within the target image. Alternatively, the position of the call block within the call image may correspond to the position of the target block within the target image. Here, the correspondence of the block positions may mean that the regions of the blocks are identical, that the region of one block is included within the region of another block, or that one block occupies a specific location within another block.

[0225] - For example, the location of a call block within a call image may be the same as the location of a target block within a target image. Alternatively, the call block may be a block containing call samples within a call image. A call sample may be a sample having coordinates identical to the coordinates of a specific sample in the target block.

[0226] - Temporal neighbor blocks may be blocks that are temporally adjacent to the spatial neighbor blocks of the target block.

[0227] Neighbor sample: A neighbor sample may refer to a sample within a neighbor block. Neighbor samples may include prediction samples, reconstructed samples, residual samples, and decoding samples.

[0228] Search range: The search range may refer to a two-dimensional area where a search for an MV is performed during inter-prediction. For example, when an optimal MV needs to be derived for processing a target block, the optimal MV can be selected from among the MVs pointing inside the search range.

[0229] Transform coefficient: The transform coefficient may be a coefficient generated by performing a transformation on the residual block. Alternatively, the transform coefficient may be a coefficient value generated by performing inverse quantization on the quantized level.

[0230] Quantized level: A quantized level can be an integer quantity used as an input for inverse quantization.

[0231] Quantization: Quantization can be a process that generates quantized levels for transform coefficients. Quantized levels can be generated by applying quantization to transform coefficients. Transformation can also be considered as part of quantization.

[0232] Inverse Quantization: Inverse quantization can be a process of multiplying a quantized level by a factor. By applying inverse quantization to the quantized level, (restored) transformation coefficients can be generated.

[0233] Quantization Parameter (QP): QP may refer to the argument used to generate quantized levels for transform coefficients in quantization. Additionally, QP may refer to the argument used to generate (restored) transform coefficients for quantized levels in inverse quantization. Alternatively, QP may be a value mapped to the quantization step size.

[0234] Delta QP: Delta QP can be the difference between the QP predicted by a specific process and the QP of the target block. In other words, the QP of the target block can be the sum of the predicted QP and Delta QP.

[0235] Quantization matrix: A quantization matrix may be a matrix used in quantization or inverse quantization to improve the subjective or objective image quality.

[0236] Quantization matrix coefficients: Quantization matrix coefficients can be each element within the quantization matrix.

[0237] Scan: Scan can refer to a method of arranging values ​​within a block or matrix. The values ​​can be coefficients. For example, a scan can mean arranging values ​​arranged in a 2D form into a 1D form, or rearranging values ​​arranged in a 1D form into a 2D form. An inverse scan can be the opposite arrangement (or rearrangement) of the arrangement performed in a scan.

[0238] Non-zero transformation coefficients: Non-zero transformation coefficients may refer to transformation coefficients that have a non-zero value or quantized levels that have a non-zero value.

[0239] Bitstream: A bitstream may refer to a series or sequence of bits containing encoded information generated by encoding of an image. A bitstream may contain information according to specific syntax elements. For example, the information may include syntax elements. An encoding device may generate a bitstream containing information according to specific syntax elements. A decoder may obtain information from the bitstream according to specific syntax elements.

[0240] Signaling: Signaling of information may indicate that information is transmitted from an encoding device to a decoding device via a bitstream. For example, the information may include syntactic elements. Alternatively, signaling may mean that the encoding device includes information within the bitstream. Information signaled by the encoding device may be used by the decoding device. In signaling, the bitstream may be transmitted over a network and may be contained within a recording medium. In embodiments, the description that information is signaled may include: 1) the encoding device determining and generating information for signaling of information; 2) the encoding device performing encoding on the information to generate encoded information; 3) the (encoded) information being transmitted from the encoding device to the decoding device via a bitstream; 4) the decoding device performing decoding on the encoded information to obtain information; and 5) the decoding device determining and generating information through signaling of information.

[0241] - An encoding device can generate encoded information by performing encoding on the information. The encoded information can be signaled through a bitstream. A decoding device can obtain information by performing decoding on the encoded information.

[0242] - The fact that information is signaled to a specific target may mean that the information is used for each specific target, and that the processing represented by the information is applied to each specific target. For example, the fact that information is signaled at a specific unit level may indicate that the information is used or processed for each specific unit.

[0243] - The signaled information may include one or more sub-information. That specific information is signaled may mean that each piece of information of the one or more sub-information included in the specific information is signaled.

[0244] Optional Signaling: Signaling for information may be performed optionally. Optional signaling for information may mean that an encoding device optionally includes information within a bitstream (depending on specific conditions). Optional signaling for information may mean that a decoder optionally obtains information from a bitstream (depending on specific conditions).

[0245] Omission of Signaling: Signaling for information may be omitted. Omission of signaling for information may mean that the encoding device does not include information in the bitstream (depending on specific conditions). Omission of signaling for information may mean that the decoding device does not obtain information from the bitstream (depending on specific conditions). The decoding device may derive information with omitted signaling using other information of the embodiments.

[0246] Symbol: May represent at least one piece of information of a target unit, such as syntactic elements, coding parameters, quantized levels, and transform coefficients of a target unit or target block. Additionally, the symbol may represent the target of entropy encoding or the result of entropy decoding.

[0247] Entropy encoding: Entropy encoding can allocate a small number of bits to symbols with a high probability of occurrence and a large number of bits to symbols with a low probability of occurrence. Through this allocation, the size of the bitstream representing the symbols can be reduced.

[0248] Entropy coding can utilize methods such as Variable Length Coding (VLC) and Context-Adaptive Binary Arithmetic Coding (CABAC). For example, in Variable Length Coding, entropy coding can be performed using variable-length tables. For instance, in CABAC, a binaryization method for symbols and a probabilistic model of symbols / bins can be derived for entropy coding, and context-based arithmetic coding can be performed.

[0249] Entropy Decoding: In entropy decoding, the processes performed in entropy encoding can be performed in reverse. Symbols can be generated by entropy decoding of a bitstream.

[0250] Parsing: Parsing can refer to determining the values ​​of syntactic elements by performing entropy decoding on the encoded information of a bitstream. Alternatively, parsing can refer to entropy decoding itself.

[0251] Statistical Value: The values ​​of information related to specific entity(s) described in the embodiments may be used as inputs for specific operations. The statistical value may be a value derived by a specific operation on the values ​​related to these specific entity(s). For example, the statistical value for specific information may be one or more of the following: an average value, a weighted average value (weighted average), a weighted sum (weighted sum), a minimum value, a maximum value, a mode, a median value, an interpolated value, a sum of products, and a product of sums. Additionally, information of the embodiment having specific values ​​determined by operations, such as constants, variables, and coding parameters, may have a specific statistical value according to the embodiment.

[0252]

[0253] Coding parameters

[0254] In the embodiments, the coding parameters may be information required for coding. The coding parameters may include information signaled from an encoding device to a decoder, information calculated / derived during the processing of coding described in the embodiments, and information used for the processing of coding described in the embodiments.

[0255] In the embodiments, the coding parameters include the size of the CTU, the size of the unit, the form of the unit, the shape of the unit, the depth of the unit, the minimum unit size, the maximum unit size, the maximum unit depth, the minimum unit depth, the unit splitting information, QT splitting information, BT splitting information, the splitting direction of the BT splitting, the splitting form of the BT splitting, TT splitting information, the splitting direction of the TT splitting, the splitting form of the TT splitting, MTT splitting information, the combination of MTT splittings, the splitting direction of the MTT splitting, the splitting form of the MTT splitting, the prediction mode, the intra prediction mode, the luminance intra prediction mode, the chroma intra prediction mode, the intra prediction mode, the inter splitting information, the coding block splitting information, the prediction block splitting information, the transformation block splitting information, the reference sample line index, the reference sample filtering method, the reference sample filter tab, the reference sample filter coefficients, the prediction block filtering method, the prediction block filter tab, the prediction block filter coefficients, the prediction block boundary filtering method, the prediction block boundary filter tab, the prediction block boundary filter coefficients, the inter prediction mode, motion information, MV, and Motion Vector Difference; MVD), MVD resolution, MV size, MV representation accuracy, reference image list, reference image, reference image index, inter prediction direction, inter prediction indicator, prediction list utilization flag, POC, MV candidate, MV candidate index, MV candidate list, AMVP mode usage information, merge candidate, merge index, merge candidate list, merge mode usage information, motion information correction information, skip mode usage information, intra-block copy mode usage information, BV (Block Vector), Block Vector Difference (BVD), BVD resolution, BV size, BV representation accuracy, BV candidate, BV candidate index, BV candidate list, interpolation filter filter tab, interpolation filter filter coefficients, transform type, transform size, transform selection information, primary transform usage information,Secondary transform usage information, primary transform selection information, secondary transform selection information, residual block presence information, coded block pattern, coded block flag, QP, delta QP, quantization matrix, deblocking filter usage information, deblocking filter coefficients, deblocking filter filter tab, deblocking filter strength, deblocking filter shape / form, adaptive sample offset usage information, adaptive sample offset value, adaptive sample offset category, adaptive sample offset type, adaptive loop filter usage information, adaptive loop filter coefficients, adaptive loop filter filter tab, adaptive loop filter shape / form, binarization / debinarization method, context model, context model determination method, context model update method, regular mode usage information, bypass mode usage information, significant coefficient flag, last significant coefficient flag, coefficient group coding flag, last significant coefficient position, flag indicating whether the coefficient value is greater than 1, whether the coefficient value is greater than 2 Flag indicating presence, flag indicating whether the coefficient value is greater than 3, remaining coefficient value information, sign information, context bin, bypass bin, restored sample, restored luminance sample, restored chroma sample, residual sample, residual luminance sample, residual chroma sample, transform coefficient, luminance transform coefficient, chroma transform coefficient, transform coefficient level, luminance transform coefficient level, chroma transform coefficient level, transform coefficient level scanning method, quantized level, luminance quantized level, chroma quantized level, size of the MV seek area on the decoder side, shape of the MV seek area on the decoder side, number of MV seeks on the decoder side, picture type, slice identification information, slice type, slice splitting information, tile group identification information, tile group type, tile group splitting information, tile identification information, tile type, tile splitting information, bit depth,It may include one or more of input sample bit depth, restored sample bit depth, residual sample bit depth, transform factor bit depth, quantized level bit depth, mapping availability information, information about the luminance signal, information about the chroma signal, the color space of the target block, the color space of the residual block, and temporal layer information.

[0256] In addition, the coding parameter may further include 1) a value of information that may be included in the coding parameter, 2) a combination of multiple pieces of information that may be included in the coding parameter, 3) a statistical value of information that may be included in the coding parameter, 4) information related to the coding parameter, 5) information used to calculate / derive the coding parameter, and 6) information calculated / derived using the coding parameter.

[0257] In the embodiments, "X usage information" may be "information indicating whether X is used / applied / executed." Alternatively, "X usage information" may be "information indicating whether X is available." For example, "specific mode usage information" may be information indicating whether a specific mode is used. Mode information may indicate a mode used for a target block among the modes described in the embodiments. In the embodiments, specific mode usage information may be replaced with mode information, and the description of specific mode usage information may also apply to mode information. "X usage information" and "X indicator" may be used interchangeably.

[0258] In the embodiments, coding parameters and syntax elements may correspond to each other. For example, a syntax element of the embodiment may be used as a coding parameter, and a coding parameter may be signaled as a syntax element.

[0259] In the embodiments, "X existence information" may be considered as "information indicating whether X exists" or "information indicating whether information indicating X exists within the bitstream".

[0260] In the embodiments, "X selection information" may be information indicating one of the candidates or methods for X. "X selection information" may be considered as an "X index".

[0261] In the embodiments, the splitting form of a specific tree may represent one of symmetric splitting and asymmetric splitting, and may represent one of QT, BT, TT, and non-split. The splitting direction of a specific tree may represent one of horizontal direction and vertical direction.

[0262] In the embodiments, when the coding parameter has one of a plurality of values, "coding parameter" may be replaced with "whether the coding parameter has a specific value among the plurality of values ​​available to the coding parameter".

[0263] In the embodiments, when the coding parameter refers to one of a plurality of targets, the "coding parameter" may be replaced with "whether the coding parameter refers to a specific target among the plurality of targets."

[0264] In the embodiments, the coding parameter may include at least one of the type of target picture and the type of target slice. The type of target picture may be one of an I-picture, a B-picture, and a P-picture. The type of target slice may be one of an I-slice, a B-slice, and a P-slice.

[0265] - If the target image to be encoded is an I-slice, the target image can be encoded using data within the image itself without inter-predicting that references other images. For example, an I-slice can be encoded using only intra-predicting.

[0266] - If the target image is a P slice, the target image can be encoded through inter-prediction using only the reference slice existing in a unidirectional direction. Here, the unidirectional direction can be forward or reverse.

[0267] - If the target image is a B slice, the target image can be encoded through inter-prediction using reference slices existing in both directions or through inter-prediction using a reference slice existing in one of the forward and backward directions. Here, both directions can be the forward and backward directions.

[0268] P slices and B slices encoded and / or decoded using a reference slice can be considered as images where inter-prediction is used.

[0269]

[0270] System for video coding

[0271] FIG. 1 shows a system for video coding according to one embodiment.

[0272] The system (100) may include at least one of an encoding device (110) and a decoding device (150).

[0273] Each of the encoding device (110) and the decoding device (150) may be a computer or an electronic apparatus.

[0274]

[0275] Structure of the encoding device

[0276] The encoding device (110) may include a processor (120), a storage (140), and a communicator (149).

[0277] The processor (120), storage (140), and communication device (149) can be connected via a bus.

[0278] The processor (120) may be a semiconductor device that executes instructions or computer-executable code, such as a Central Processing Unit (CPU). The processor (120) may be at least one hardware processor.

[0279] The processor (120) can perform generation and processing of information that is input to the encoding device (110) in the embodiments, output from the encoding device (110), or used inside the encoding device (110), and can perform comparison and judgment related to such information.

[0280] The processor (120) may include a plurality of components. The plurality of components may include a partitioner (122), a subtractor (124), a transformer (125), a quantizer (126), an inverse quantizer (127), an inverse transformer (128), an adder (129), a filter (130), and an entropy encoder (139).

[0281] At least some of the aforementioned multiple components may be program modules. Program modules may be included in the encoding device (110) in the form of an operating system, an application, and other program modules. Program modules may be instructions or computer-executable code stored in a storage (140) and executed by a processor (120).

[0282] The storage (140) may include various types of volatile storage media and non-volatile storage media. For example, the storage (140) may include memory such as ROM and RAM.

[0283] The storage (140) can store instructions and computer-executable code used for the operation of the encoding device (110), and can store information and bitstreams as described in the embodiments. The storage (140) may include a reference picture buffer (141).

[0284] The communication device (149) can perform functions related to the communication of information in the encoding device (110). For example, the communication device (149) can transmit a bitstream to the decoding device (150).

[0285] Among the names of the components of the encoding device (110), "-gi" ("-er" or "-or") may be replaced with "-bu" (- unit). The storage unit (140) may also be named a storage unit.

[0286]

[0287] Operation of the encoding device

[0288] The encoding device (110) can sequentially encode one or more images of the video.

[0289] The storage (140) can store the original image. In the encoding device (110), the original image can be used as the target image.

[0290] The processor (120) can generate a bitstream containing encoded information by performing encoding on the target image and can store the generated bitstream in a storage (140). The generated bitstream can be stored on a computer-readable recording medium and can be transmitted by the communication device (149) to the communication device (189) of the decoding device (150) via a wired and / or wireless transmission medium.

[0291] The splitter (122) can determine the target block by performing a split on the target image.

[0292] The predictor (123) can determine the prediction mode of the target block. The predictor (123) can generate a prediction block of the target block by performing a prediction according to the prediction mode.

[0293] The prediction mode of the target block may be one of the available prediction modes. For example, available prediction modes may include intra prediction, inter prediction, and IBC prediction.

[0294] For example, if the prediction mode is intra prediction, the predictor (123) can perform intra prediction on the target block to generate a prediction block of the target block.

[0295] For example, if the prediction mode is inter-prediction, the predictor (123) can perform inter-prediction on the target block to generate a prediction block of the target block.

[0296] For example, if the prediction mode is IBC, the predictor (123) can perform an IBC prediction for the target block to generate a prediction block of the target block.

[0297] The subtractor (124) can generate a residual block of the target block. The residual block may be the difference between the original block and the prediction block. The original block may be the region of the original image pointed to by the target block. Alternatively, the residual block may refer to a block generated by applying one or more of transformation and quantization to the difference between the original block and the prediction block.

[0298] The converter (125) can perform a conversion on the residual block to generate conversion coefficients.

[0299] The converter (125) can perform the conversion using one of a plurality of conversion methods.

[0300] For example, multiple transformation methods may include the Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Karhunen-Loeve Transform (KLT), and transformations based on each transformation.

[0301] The transform skip mode may be a mode that generates a restored block using the restored residual block and prediction block, for which transform and inverse transform have not been performed. When the transform skip mode is applied to a target block, the transform and inverse transform for the target block may be omitted, and only quantization and inverse quantization for the target block may be performed.

[0302] The quantizer (126) can generate quantized levels by applying quantization using quantization parameters to the conversion coefficients. In the embodiments, the quantized levels may also be referred to as conversion coefficients.

[0303] The entropy encoder (139) can generate encoded information by performing entropy encoding based on a probability distribution on information for decoding an image. The bitstream may contain encoded information.

[0304] Information for decoding the image may include quantized levels and syntax elements produced by the quantizer (126).

[0305] The probability distribution can be determined based on quantized levels and coding parameters.

[0306] The entropy encoder (139) can convert quantized levels, which have the form of a two-dimensional block, into the form of a one-dimensional vector by using scanning to perform encoding for the quantized levels. In the scanning, it can be determined which scan to use among an upper-right diagonal scan, a vertical scan, and a horizontal scan based on coding parameters such as the size of the block and the intra-prediction mode of the block.

[0307] When encoding is performed on a target image / block, the predictor (123) uses a reference image / block for prediction. The encoded target image / block can be used as a reference image / block for other images / blocks that are subsequently processed. Accordingly, the processor (120) can perform restoration on the encoded target block and can store the restored image containing the restored target block generated by the restoration in the reference picture buffer (141) as a reference image. Inverse quantization and inverse transform can be performed on the encoded target block for restoration.

[0308] The inverse quantizer (127) can generate inverse quantized conversion coefficients by performing inverse quantization on the quantized level.

[0309] The inverse converter (128) can generate inversely quantized and inversely converted coefficients by performing an inverse conversion on the inversely quantized conversion coefficients. In embodiments, the inversely quantized and / or inversely converted coefficients may refer to coefficients to which at least one of the inverse quantization and inverse conversion has been applied. The inversely quantized and inversely converted coefficients may be restored residual blocks.

[0310] The adder (129) can generate a recovery block by combining the prediction block and the recovered residual block.

[0311] The restoration block may pass through a filter (130). The filter (130) may apply one or more of a plurality of filters to the target. Each of the plurality of filters may be an in-loop filter. The target may be a restoration sample, a restoration block, or a restoration image.

[0312] The reference picture buffer (141) can store a restoration block / image provided from the filter (130). The restoration image may be an image containing the restoration block. Alternatively, the restoration image may be an image composed of restoration blocks.

[0313] The reference picture buffer (141) can provide the stored restored image to the predictor (123) as a reference image. In terms of storing the decoded (i.e., restored) picture, the reference picture buffer (141) may also be referred to as the Decoded Picture Buffer (DPB).

[0314]

[0315] Structure of the decoding device

[0316] The decoding device (150) may include a processor (160), a storage device (180), and a communication device (189).

[0317] The description of the processor (120), storage (140), and communication device (149) associated with the encoding device (110) may also apply to the processor (160), storage (180), and communication device (189) associated with the decoding device (150). Redundant descriptions are omitted.

[0318] The processor (160) may include a plurality of components. The plurality of components may include an entropy decoder (161), a splitter (162), a predictor (163), an inverse quantizer (167), an inverse converter (168), an adder (169), and a filter (170).

[0319] The storage (180) may include a reference picture buffer (181).

[0320] The communicator (189) can perform functions related to the communication of information in the decoding device (150). For example, the communicator (189) can receive a bitstream from the encoding device (110).

[0321] Among the names of the components of the decoding device (150), "-gi" ("-er" or "-or") may be replaced with "-bu" (- unit). The storage unit (180) may also be named a storage unit.

[0322]

[0323] Operation of the decoding device

[0324] The communication device (149) of the encoding device (110) can transmit the bitstream generated by the encoding device (100) to the decoding device (150). Alternatively, a computer-readable recording medium storing the bitstream can transmit the bitstream generated by the encoding device (100) to the decoding device (150).

[0325] The communication device (189) can receive a bitstream from the encoding device (110) via a wired and / or wireless transmission medium. The received bitstream can be stored in a storage device (180).

[0326] The processor (160) can obtain a bitstream from a storage (180) or a computer-readable recording medium.

[0327] A bitstream can contain encoded information.

[0328] The entropy decoder (161) can generate information for decoding an image by performing entropy decoding based on a probability distribution on the encoded information of the bitstream.

[0329] Information for decoding an image may include quantized levels and syntax elements, etc.

[0330] The entropy decoder (161) can convert quantized levels, which have the form of a one-dimensional vector, into the form of a two-dimensional block by using scanning to perform decoding on the quantized levels. In the scanning, it can be determined which scan to use among an upper-right diagonal scan, a vertical scan, and a horizontal scan based on coding parameters such as the size of the block and the intra-prediction mode of the block.

[0331] The entropy decoder (161) can provide syntax elements to other components of the processor (160), such as the splitter (162).

[0332]

[0333] Common explanation based on the relationship between the components of the encoding device and the components of the decoding device

[0334] The decoding device (150) performs decoding using the bitstream generated by the encoding device (110). The encoding device (110) may perform encoding for the target block using a restored image derived within the decoding device (150), rather than an original image that is not provided to the decoding device (150). Accordingly, the encoding device (110) and the decoding device (150) may need to generate the restored block / image in the same way. In this regard, the descriptions of the divider (122), predictor (123), inverse quantizer (127), inverse converter (128), adder (129), filter (130), and reference picture buffer (141) of the encoding device (110) disclosed in the embodiments may also be applied to the divider (162), predictor (163), inverse quantizer (167), inverse converter (168), adder (169), filter (170), and reference picture buffer (181) of the decoding device (150), respectively. Redundant descriptions are omitted.

[0335] Additionally, each of the divider (122), predictor (123), inverse quantizer (127), inverse converter (128), adder (129), and filter (130) of the encoding device (110) can generate syntactic element information that specifies processing for a target. Each of the divider (162), predictor (163), inverse quantizer (167), inverse converter (168), adder (169), and filter (170) of the decoding device (150) can perform processing for a target (such as that performed in the encoding device (110)) using the syntactic element information.

[0336] As described above, corresponding components of the encoding device (110) and the decoding device (150) may perform the same or corresponding functions. In embodiments, the processor may represent the processor (120) of the encoding device (110) and / or the processor (160) of the decoding device (150). For example, regarding the function of prediction, the processor may represent a predictor (123), a subtractor (124), and an adder (129), and may represent a predictor (163) and an adder (169). Regarding the function of conversion, the processor may represent a converter (125) and an inverse converter (128), and may represent an inverse converter (168). Regarding the function of quantization, the processor may represent a quantizer (126) and an inverse quantizer (127), and may represent an inverse quantizer (167). In terms of functions related to entropy encoding / decoding, the processing unit may represent an entropy encoder (139) and / or an entropy decoder (161). In terms of functions related to filtering, the processing unit may represent a filter (130) and / or a filter (170). The storage unit may represent a storage unit (140) of the encoding device (110) and / or a storage unit (180) of the decoding device (150). The reference picture buffer may represent a reference picture buffer (141) of the encoding device (110) and / or a reference picture buffer (181) of the decoding device (150). The communication unit may represent a communication unit (149) of the encoding device (110) and / or a communication unit (189) of the decoding device (150).

[0337]

[0338] Partitioning of the units that constitute the image

[0339] Figure 2 shows a segmentation structure of an image according to one embodiment.

[0340] Figure 2 schematically illustrates an example in which a single unit is divided into multiple sub-units.

[0341] CU can be used as a base unit for encoding and decoding of images. Additionally, CU can be a base unit for prediction, transformation, quantization, inverse quantization, inverse transformation, entropy encoding, and entropy decoding.

[0342] A CU can be used as a unit to which a prediction mode is applied. That is to say, in coding, it can be determined which of the available prediction modes will be applied to each CU. For example, available prediction modes may include intra prediction, inter prediction, and IBC intra block copy prediction.

[0343] The target image (200) can be sequentially divided into units of CTUs. A division structure can be determined for each CTU. The CTU can be divided into CUs according to the division structure. Alternatively, one CTU can be used as a CU. The size of the CTU can be the maximum size of the CU.

[0344] Each CU may have depth information. The depth information may represent the depth of the CU and the size of the CU. The depth of the CTU may be 0. The depth of the CU created by dividing the CTU may be 1. When a parent CU is divided into child CUs, the depth of the child CU may be 1 greater than the depth of the parent CU. The number of divided CUs may be a positive integer greater than or equal to 2, including 2, 4, 8, and 16. At least one of the width and height of the child CU created by dividing the parent CU may be smaller than at least one of the width and height of the parent CU, depending on the number of child CUs.

[0345] A partitioned CU can be recursively partitioned in the same way up to a predefined maximum depth or a predefined minimum size. The depth of a Smallest Coding Unit (SCU) can be the predefined maximum depth, and the size of an SCU can be the predefined minimum size. The size of an SCU can be the minimum CU size.

[0346] For example, the depth range of a CU can be values ​​from 0 to 3. Depending on the depth of the CU, the CU can have a size from 64x64 to 8x8. A CTU with a depth of 0 can be 64x64 blocks. 0 can be the minimum depth. An SCU with a depth of 3 can be 8x8 blocks. 3 can be the maximum depth. Depth 0 can represent a CTU that is 64x64 blocks. Depth 1 can represent a CU that is 32x32 blocks. Depth 2 can represent a CU that is 16x16 blocks. Depth 3 can represent an SCU that is 8x8 blocks.

[0347] The partition information of a CU may indicate whether the CU is partitioned. The partition information may be a 1-bit flag. All CUs except the SCU may include partition information. For example, the partition information of a CU that is not further partitioned may be a first value of '0', and the partition information of a CU that is partitioned may be a second value of '1'.

[0348] Quad Tree (QT) partitioning can mean that a single CU is partitioned into four CUs. When a parent CU is partitioned into four child CUs, the width and height of each child CU can be half the width and half the height of the parent CU, respectively.

[0349] A binary tree (BT) partition can mean that one CU is divided into two CUs. For example, if a parent CU is divided into two child CUs, the width or height of each child CU can be half the width or half the height of the parent CU.

[0350] Ternary tree (TT) partitioning can mean that a single CU is divided into three CUs. For example, when a parent CU is divided into three child CUs, the three child CUs can be created by dividing the width or height of the parent CU in a ratio of 1:2:1. The width or height of the child CUs can be 1 / 4, 1 / 2, and 1 / 4 of the width or height of the parent CU, respectively.

[0351] In FIG. 2, QT-type splitting was applied to the first CTU. QT splitting, BT splitting, and TT splitting were applied to the second CTU.

[0352] To split a CTU, at least one of different types of splits, such as QT splitting, BT splitting, and TT splitting, may be applied to the CTU. Different types of splits may be applied based on specific priorities.

[0353] For example, QT splitting may be applied preferentially to a CTU. A CU to which QT splitting can no longer be applied may correspond to a leaf node of QT. A CU that is a leaf node of QT may become a root node of BT and / or TT. A CU that is a leaf node of QT may be split into a BT form or a TT form, or may not be split further. In this case, QT splitting may not be applied again to a CU created by applying BT splitting or TT splitting to a CU that is a leaf node of QT.

[0354] The splitting of a CU corresponding to each node of QT can be signaled using QT splitting information. The QT splitting information may be a flag. The QT splitting information of a unit may be information indicating whether the unit is split into a QT form. A first value of the QT splitting information, '0', may indicate that the CU is not split into a QT form. QT splitting information having a first value may signify a Multi-Type Tree (MTT) split. MTT splitting may include BT splitting and TT splitting. A second value of the QT splitting information, '1', may indicate that the CU is split into a QT form.

[0355] There may be no priority between BT splitting and TT splitting. That is, CUs corresponding to the leaf nodes of QT can be split into BT form or TT form. Additionally, CUs generated by BT splitting or TT splitting can be split again into BT form or TT form, or they may not be split any further.

[0356] A CU corresponding to a leaf node of QT can be a root node of MTT. For a CU corresponding to each node of MTT, the CU may further include partition direction information and partition type information in the form of MTT.

[0357] The splitting direction information can indicate the splitting direction of the MTT split. The first value of the splitting direction information, '0', can indicate that the CU is split in the horizontal direction. The second value of the splitting direction information, '1', can indicate that the CU is split in the vertical direction.

[0358] The split type information may indicate the split type used for multi-type tree splitting. The first value of the split type information, '0', may indicate that CU is split into TT form. The second value of the split type information, '1', may indicate that CU is split into BT form.

[0359] Here, each of the aforementioned division direction information and division shape information may be a flag having a specified length (e.g., 1 bit).

[0360] The partitioning information of CU may also include QT partitioning information, partitioning direction information, and partitioning shape information.

[0361] A CU that is no longer divided by QT division, BT division, and TT division can be used as a unit for specific processing such as prediction, transformation, quantization, inverse quantization, inverse transformation, entropy encoding, and entropy decoding. That is, for a specific processing, the CU may no longer be divided. Therefore, division information for dividing such a CU into PU and / or TU, etc., may not exist within the bitstream.

[0362] On the other hand, if the size of a CU is larger than the maximum TU size, such a CU can be recursively partitioned until the size of the CU becomes less than or equal to the maximum TU size. For example, if the size of the CU is 64x64 and the maximum TU size is 32x32, the CU can be partitioned into 4 32x32 TUs for transformation. For example, if the size of the CU is 32x64 and the maximum TU size is 32x32, the CU can be partitioned into 2 32x32 TUs for transformation.

[0363] In such cases, information regarding whether the CU is split for transformation may not be signaled separately. Whether the CU is split may be determined without signaling by comparing the size of the CU (width / height) and the maximum TU size (width / height). For example, if the width of the CU is greater than the width of the maximum TU size, the CU may be split vertically into two. Additionally, if the height of the CU is greater than the height of the maximum TU size, the CU may be split horizontally into two.

[0364] For example, the minimum size of a CU can be 4x4. For example, the maximum size of a transformation block can be 64x64. For example, the minimum size of a transformation block can be 4x4. The minimum size of QT can be the minimum size of a CU corresponding to a leaf node of QT. The maximum depth of MTT can be the maximum depth of a path from the root node of MTT to a leaf node.

[0365] The BT maximum size may represent the maximum size of the CU corresponding to each node of the BT, and the TT maximum size may represent the maximum size of the CU corresponding to each node of the TT. The BT minimum size and / or the TT minimum size may be set as the minimum size of the CU.

[0366] If the depth of a CU within the MTT corresponding to a node of the MTT is equal to the maximum depth of the MTT, the CU may not be divided into BT form and / or TT form.

[0367] Based on the various sizes and depths of the aforementioned CU, each piece of information described in the embodiments may or may not be present in the bitstream.

[0368] Information regarding the maximum or minimum size described in the embodiments may be signaled at the upper level of the CU. In the embodiments, the upper level of the CU may include a video level, a sequence level, a picture level, a subpicture level, a tile group level, a tile level, and a slice level, etc.

[0369] The information described in the embodiments may be signaled separately for different types of slices. Different types of slices may include intra-slices and inter-slices.

[0370]

[0371] Processing of blocks based on block attributes

[0372] Whether a specific process described in the embodiments is applied or performed may be determined based on the attributes of the block associated with the specific process. Whether a specific process described in the embodiments is applied or performed may be determined based on whether the attributes of the block associated with the specific process satisfy specific conditions. For example, a block may include a target block, a neighbor block, and a reference block. A block may include other blocks described in the embodiments. A block may be one of the blocks and units described in the embodiments.

[0373] The block to which the specific treatment described in the embodiments is applied may have a square shape or a non-square shape.

[0374] In one embodiment, the attributes of the block may include the size of the block. The specific processing described in the embodiments may be applied / performed when specific conditions regarding the size of the block are met.

[0375] In one embodiment, specific conditions may include a minimum block size condition and a maximum block size condition. The block to which the minimum block size condition applies and the block to which the maximum block size condition applies may be different from each other.

[0376] In one embodiment, the minimum block size and / or maximum block size for a specific process may be predefined.

[0377] In one embodiment, the processing of the embodiment may be applied / performed when the block size is greater than or equal to the minimum block size and / or less than or equal to the maximum block size. Alternatively, in one embodiment, the processing of the embodiment may be applied / performed when the block size is greater than the minimum block size and / or less than the maximum block size.

[0378] In one embodiment, the processing of the embodiment may be applied / performed only when the block size is greater than or equal to the minimum block size and less than or equal to the maximum block size. Alternatively, the processing of the embodiment may be applied / performed only when the block size is greater than the minimum block size and less than or equal to the maximum block size. Alternatively, the processing of the embodiment may be applied / performed only when the block size is greater than the minimum block size and less than the maximum block size. The processing of the embodiment may be applied / performed only when the block size is greater than the minimum block size and less than the maximum block size.

[0379] In one embodiment, the processing of the embodiment may be applied / performed only when the block size is a predefined block size.

[0380] In the embodiments, the size of the block may be determined by various methods. For example, the size of the block may mean the width of the block or the height of the block. The size of the block may mean both the width and the height of the block. The size of the block may mean the area of ​​the block. The size of the block may mean 1) the result of a known formula using the width and height of the block, 2) the result of a formula of the embodiment, or 3) a statistical value.

[0381] Additionally, for the first size, the processing of the first embodiment among the embodiments may be applied / performed, and for the second size, the processing of the second embodiment among the embodiments may be applied / performed.

[0382] In the embodiments, the block size may be 2x2, 4x4, 8x8, 16x16, 32x32, 64x64, or 128x128, etc. Or, in the embodiments, the block size is (2*SIZE X )x(2*SIZE Y It may be ) etc. SIZE X is one of integers greater than or equal to 1. SIZE Y can be one of integers greater than or equal to 1.

[0383]

[0384] Predictive information for prediction

[0385] Predictive information can be used to generate a predicted block for a target block.

[0386] The encoding device (110) can generate prediction information required for prediction and can generate a bitstream containing the prediction information. The prediction information can be signaled from the encoding device (110) to the decoding device (150) through the bitstream. The decoding device (150) can obtain the prediction information from the bitstream and can generate a prediction block by performing a prediction on a target block using the prediction information.

[0387] Prediction information may include intra prediction information, inter prediction information, and IBC prediction information. In the embodiments, prediction information may be replaced with intra prediction information, inter prediction information, and / or IBC information. Intra prediction information may include information used for intra prediction as described in the embodiments. Inter prediction information may include information used for inter prediction as described in the embodiments. IBC information may include information used for IBC prediction as described in the embodiments.

[0388]

[0389] Intra prediction

[0390] Figure 3 shows the structure of an intra prediction according to one embodiment.

[0391]

[0392] Intra-prediction can be performed using reference samples and coding parameters of the target block. The reference sample may be a (restored) sample within the (restored) reference block. Alternatively, an intermediate prediction sample may be generated using a sample described in an example, such as the restored sample, and a reference sample may be generated again using the intermediate prediction sample. Processing described in an example, such as filtering, may be applied when generating the reference sample.

[0393] The reference block may be a (spatial) neighbor block of the target block. The coding parameter may be a coding parameter for the target block and / or a coding parameter for the reference block. In intra-prediction, the reference sample may refer to a neighbor sample.

[0394] A prediction block can be generated by performing intra prediction on a target block according to an intra prediction mode, based on a reference sample within the target image and information related to the reference sample. The size of the target block and the size of the prediction block may be the same.

[0395] In the embodiments, the prediction block may be a PU. Alternatively, the prediction block may correspond to the CU or TU described in the embodiments. The prediction block may have a square or rectangular shape.

[0396] An intra prediction mode can be represented by at least one of a mode number, a mode value, a mode angle, and a mode direction. The prediction directions of a plurality of intra prediction modes for a target block are illustrated in the lower right corner of FIG. 3. Among the plurality of intra prediction modes, the remaining intra prediction modes, excluding DC and planar modes, may be directional modes. A directional mode may be an intra prediction mode having a specific direction or a specific angle. An intra prediction mode for a target block may be selected from directional modes and non-directional modes.

[0397] In the bottom-right rectangle representing the target block, the number '0' may represent Planner mode, which is a non-directional intra prediction mode. The number '1' may represent DC mode, which is a non-directional intra prediction mode. In the bottom-right rectangle representing the target block, arrows extending from the center of the rectangle outwards may represent the prediction directions of directional intra prediction modes. Additionally, the number displayed near the arrow may represent an example of a mode value assigned to an intra prediction mode or a prediction direction of an intra prediction mode.

[0398] Intra prediction can be performed according to the intra prediction mode for the target block. One of the intra prediction modes available for the target block can be used as the intra prediction mode for the target block.

[0399] The number of intra prediction modes available to the target block may be a predefined value. Alternatively, the number of intra prediction modes available to the target block may be determined based on the attributes of the prediction block. For example, the attributes of the prediction block may include coding parameters such as shape, size, and color components.

[0400] For example, in FIG. 3, the directional modes illustrated by dashed lines (i.e., directional modes with numbers from -14 to -1 or numbers from 67 to 80) can be applied only to predictions for non-square blocks. Therefore, the number of available intra-prediction modes for predictions for square blocks may be 67. (Planner mode, DC mode, and 65 directional modes)

[0401] For example, the number of available intra prediction modes may vary depending on whether the color component of a block is a luminance signal or a chroma signal. The number of available intra prediction modes for a block with a luminance component may be greater than the number of available intra prediction modes for a block with a chroma component.

[0402] Intra-prediction modes may include a horizontal-below mode, a horizontal mode, a vertical mode, and a vertical-right mode. The horizontal-below mode may be an intra-prediction mode located at the bottom of the horizontal mode. The vertical-right mode may be a mode located to the right of the vertical mode. For example, in FIG. 3, the mode value of the horizontal mode may be 18. The mode value of the vertical mode may be 50. Intra-prediction modes with a mode value of 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, and 66 may be vertical-right modes. Intra prediction modes with a mode value of 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 and 17 may be horizontal bottom modes.

[0403] The number of the aforementioned intra-prediction modes and the mode number of each intra-prediction mode may be exemplary only. The number of the aforementioned intra-prediction modes and the mode number of each intra-prediction mode may be defined differently depending on the embodiment, implementation, and / or as necessary.

[0404] When the intra prediction mode is a planner mode, when generating a prediction block of a target block, the sample value of the prediction sample can be generated using a weighted sum (weighted sum) of the top reference sample of the target sample, the left reference sample of the target sample, the right top reference sample of the target block, and the left bottom reference sample of the target block, depending on the position of the prediction sample within the prediction block.

[0405] When the intra prediction mode is DC mode, a prediction block can be generated based on the average of the sample values ​​of multiple reference samples. The multiple reference samples may include top reference samples and left reference samples of the target block. The value of the prediction sample of the prediction block can be determined based on the average of the sample values ​​of the multiple reference samples. Additionally, filtering using the values ​​of the reference samples can be performed on specific rows and / or specific columns within the target block. The specific rows may be one or more top rows adjacent to the top reference samples. The specific columns may be one or more left columns adjacent to the left reference samples.

[0406] When the intra prediction mode is a directional mode, a prediction block can be generated using the top reference sample, left reference sample, right top reference sample, and / or left bottom reference sample of the target block.

[0407] The intra prediction mode of a target block can be determined based on the intra prediction mode of a neighboring block of the target block. Information for determining the intra prediction mode of the target block can be signaled.

[0408] For example, if the intra prediction modes of the target block and the neighbor block are the same, an indicator indicating that the intra prediction modes of the target block and the neighbor block are the same may be signaled.

[0409] For example, an indicator indicating an intra prediction mode such as the intra prediction mode of the target block among the intra prediction modes of multiple neighboring blocks may be signaled.

[0410] For example, if the intra prediction modes of the target block and neighboring blocks are different from each other, an indicator indicating the intra prediction mode of the target block may be signaled. Alternatively, information used to derive the intra prediction mode of the target block based on the intra prediction mode of the neighboring block may be signaled.

[0411] Reference samples used for intra-prediction of a target block may include bottom-left reference samples, left reference samples, top-left reference samples, top reference samples, and top-right reference samples, etc.

[0412] For example, left reference samples may be restored reference samples adjacent to the left side of the target block. Top reference samples may be restored reference samples adjacent to the top side of the target block. Top-left reference samples may be restored reference samples diagonally adjacent to the top-left side of the target block. Bottom-left reference samples may be reference samples located below the left reference samples among samples located on the same line as the left sample line composed of left reference samples. Top-right reference samples may be reference samples located to the right of the top reference samples among samples located on the same line as the top sample line composed of top reference samples.

[0413] Reference samples used for intra-prediction for a target block can be determined based on the intra-prediction mode of the target block. One or more reference samples may be used to determine the sample values ​​of the prediction samples of the prediction block. In FIG. 3, the direction of the intra-prediction mode indicated by the arrow may represent the direction from the prediction sample to the reference sample. The direction of the intra-prediction mode may represent the dependency relationship between the reference samples and the prediction samples. For example, depending on the intra-prediction mode, the sample value of a specific reference sample may be used as the sample value of at least one sample of the prediction block. Here, the specific reference sample and at least one sample of the prediction block may be samples designated by a straight line of the direction of the intra-prediction mode. That is to say, the sample value of the specific reference sample may be copied to the sample value of the prediction sample located in the reverse direction of the direction of the intra-prediction mode. Alternatively, the sample value of the prediction sample of the prediction block may be the sample value of the reference sample located in the direction of the intra-prediction mode relative to the location of the prediction sample.

[0414] Reference samples used for intra-prediction may not be limited to samples immediately adjacent to the target block. As illustrated in FIG. 3, at least one of reference sample line 0 to reference sample line 3 may be used for intra-prediction of the target block.

[0415] Each reference sample line in FIG. 3 may contain one or more reference samples. The smaller the number of the reference sample line, the closer the line of reference samples may be to the target block. Reference sample line 0 may be a line of reference samples immediately adjacent to the target block. When the top-left coordinates of the target block are (X, Y), the horizontal length is W, and the vertical length is H, the reference samples of reference sample line 0 may be samples with an x-coordinate of X-1 or a y-coordinate of Y-1. Here, the y-coordinates of the reference samples with an x-coordinate of X-1 may be Y-1 to Y+2H. The x-coordinates of the reference samples with a y-coordinate of Y-1 may be X-1 to X+2W. The reference samples of reference sample line A may be samples with an x-coordinate of XA-1 or a y-coordinate of YA-1. Here, the y-coordinates of the reference samples with an x-coordinate of XA-1 may be YA-1 to Y+2H+A. The x-coordinates of reference samples with y-coordinate YA-1 can be XA-1 to X+2W+A. A can be 1, 2, or 3.

[0416] Samples of segments A and F can be derived using padding that uses the nearest samples of segments B and E, respectively, instead of being obtained from restored neighbor blocks.

[0417] The reference sample line index may indicate a reference sample line among multiple reference sample lines used for intra-prediction of a target block. For example, the reference sample line index may have a value from 0 to 3. The reference sample line index may be signaled.

[0418] When intra-color component prediction is used for a target block, a prediction block for a second color component can be generated based on a reconstruction block of a first color component for the target block. For example, the first color component may be a luminance component, and the second color component may be a chroma component.

[0419] For intra-prediction between color components, parameters between the first and second color components can be derived based on a template. For example, the parameters can be parameters of a linear model.

[0420] For example, the template may include a top reference sample and / or a left reference sample of the target block, and may include a top reference sample and / or a left reference sample of the restoration block of the first color component corresponding to these reference samples.

[0421] Once the parameters are derived, a prediction block for a second color component for a target block can be generated by applying the reconstruction block of the first color component to a linear model. Depending on the image format or the type of intra-prediction between color components, subsampling or downsampling may be performed on the surrounding samples of the reconstruction block of the first color component and on the reconstruction block of the first color component. If subsampling is performed, the derivation of parameters and the intra-prediction between color components may be performed using corresponding samples derived by subsampling.

[0422] Intra Sub-Partitions (ISP) prediction may refer to sequential intra prediction for multiple subblocks generated by partitioning a target block. In ISP prediction, the target block may be partitioned into two or four subblocks in the horizontal and / or vertical directions. The partitioned subblocks may be restored sequentially. As intra prediction is performed on the subblocks, sub-prediction blocks for the subblocks may be generated. Additionally, as inverse quantization and / or inverse transformation is performed on the subblocks, sub-residual blocks for the subblocks may be generated. A restored subblock may be generated by adding the sub-prediction blocks to the sub-residual blocks. The restored subblocks may be used as reference samples for intra predictions for other subblocks to be processed next.

[0423] In performing a prediction for a target block, it may be determined whether samples included in a restored neighbor block can be used as reference samples for the target block. If there are non-available samples among the samples in the neighbor block that cannot be used as reference samples for the target block, a value generated by copying and / or interpolation using the sample value of at least one sample among the samples included in the restored neighbor block may replace the sample value of the non-available sample. If the value generated by copying and / or interpolation replaces the sample value of the sample, the sample may be used as a reference sample for the target block.

[0424] In intra-prediction, the sample value of a prediction sample in a prediction block can be determined by the sample value of a reference sample. The location of the reference sample can be specified by the location of the prediction sample and the direction of the intra-prediction mode. If the location specified by the location of the prediction sample and the direction of the intra-prediction mode is an integer location, the sample value of one reference sample pointed to by the integer location can be used to determine the sample value of the prediction sample in the prediction block. If the location specified by the location of the prediction sample and the direction of the intra-prediction mode is not an integer location, an interpolated reference sample can be generated based on the two reference samples closest to the specified location. The sample value of the interpolated reference sample can be used to determine the sample value of the prediction sample. That is to say, when the location specified by the location of the prediction sample and the direction of the intra-prediction mode represents the space between two reference samples, an interpolated sample value can be generated based on the sample values ​​of the two samples.

[0425]

[0426] Inter prediction

[0427] FIG. 4 shows the structure of an inter prediction to explain an inter prediction process according to one embodiment.

[0428] The rectangle shown in Fig. 4 can represent an image. Additionally, the arrow in Fig. 4 can represent the predicted direction.

[0429] Each image constituting a video can be classified into I-pictures (i.e., intra-pictures), P-pictures (i.e., uni-prediction pictures), and B-pictures (i.e., bi-prediction pictures) according to their coding type. Coding can be performed for each picture according to its coding type.

[0430] If the target picture is an I picture, coding for the target picture can be performed using information within the target picture without inter-prediction referencing other images. For example, coding for the I picture can be performed using intra-prediction and / or IBC prediction.

[0431] Coding for P picture and B picture can be performed by at least one of intra prediction, IBC prediction, and inter prediction using a reference image.

[0432] If the target picture is a P picture, coding for the target picture can be performed using unidirectional inter-prediction using a single reference image list.

[0433] When the target picture is picture B, coding for the target picture can be performed using unidirectional inter-prediction or bidirectional inter-prediction using two reference image lists.

[0434] Below, the inter prediction for the target block in the inter mode according to the embodiment is described in detail.

[0435] When the prediction mode of the target block is inter mode, inter prediction can be performed on the target block. The target block can be a prediction block or a partitioned prediction block.

[0436] Inter prediction can be performed using reference images and motion information. In inter prediction, a reference image can be selected using a reference image index, and a reference block corresponding to a target block within the reference image can be determined using motion information. A prediction block for the target block can be generated using the determined reference block.

[0437] Motion information can be derived using coding parameters, etc. For example, motion information can be derived using motion information of restored neighbor blocks, motion information of call blocks, and / or motion information of blocks adjacent to call blocks.

[0438] In the embodiments, a candidate list may be used for inter prediction. The candidate list may include multiple candidates. An index pointing to a candidate among the candidates in the candidate list that is used for inter prediction for a target block may be signaled. The candidate list may be derived in the same manner based on the same information in the encoding device (110) and the decoding device (150). Here, the same information may include a restored image and a restored block. Additionally, in order to specify a candidate by an index, the order of candidates within the candidate list may be constant.

[0439] In one embodiment, a prediction for a target block can be performed by using motion information of a spatial candidate or a temporal candidate as motion information of the target block. Motion information of a spatial candidate may be referred to as spatial motion information. Motion information of a temporal candidate may be referred to as temporal motion information.

[0440] Spatial candidates may be restored spatial neighbor blocks that are spatially adjacent to the target block.

[0441] Spatial candidates may be blocks that 1) exist within the target image, 2) have already been restored through decoding, and 3) are adjacent to the target block.

[0442] Spatial candidates may include the left block, top block, bottom-left block, top-right block, and top-left block of the target block.

[0443] Temporal candidates may be restored temporal neighbor blocks corresponding to the target block within the restored call (COL) image.

[0444] In the embodiments, the motion information of the spatial candidate may be the motion information of a block containing the spatial candidate. The motion information of the temporal candidate may be the motion information of a block containing the temporal candidate.

[0445] In inter prediction, a call block for a target block can be identified. The region of the target block within the target image and the region of the call block within the call image may be the same. That is to say, the call block may be a block that occupies a specific region within the call image. The specific region may be a region corresponding to the region of the target block within the call image.

[0446] Temporal candidates may be locations inside and / or outside the call block within the call image.

[0447] For example, a call block may include a first call block and a second call block. When the top-left coordinates of the call block are (xP, yP) and the size of the call block is (nPSW, nPSH), the first call block may be a block occupying the coordinates (xP + nPSW, yP + nPSH). The second call block may be a block occupying the coordinates (xP + (nPSW >> 1), yP + (nPSH >> 1)). The second call block may optionally be used as a call block when the first call block is unavailable.

[0448] The MV of the target block can be determined based on the MV of the call block. Scaling can be performed on the MV of the call block. The scaled MV of the call block can be used as the MV of the target block or the prediction MV. Alternatively, the MV of the temporal candidate stored in the candidate list associated with the inter-prediction can be the scaled MV.

[0449] The ratio of the scaled MV and the MV of the call block may be equal to the ratio of the first temporal distance and the second temporal distance. The first temporal distance may be the distance between the reference image of the target block and the target image. The second temporal distance may be the distance between the reference image of the call block and the call image.

[0450] The method by which motion information is derived can be determined by the inter prediction mode of the target block. For example, as an inter prediction mode, AMVP mode, merge mode, skip mode, merge mode with MVD, sub-block merge mode, GPM, Combined Inter Intra Prediction (CIIP) mode, and affine inter mode may be used. In the following embodiments, each of the inter prediction modes is described.

[0451]

[0452] AMVP mode

[0453] When AMVP mode is used as a prediction mode, a list of MV candidates including one or more MV candidates can be generated using spatial candidate MVs, temporal candidate MVs, history-based MV candidates, and zero vectors. At least one of the spatial candidate MVs, temporal candidate MVs, and zero vectors can be determined and used as an MV candidate.

[0454] Spatial candidates may include restored spatial neighbor blocks. The MV of a restored spatial neighbor block may be referred to as a spatial motion vector candidate. Temporal candidates may include a call block and a block adjacent to the call block. The MV of a call block or the MV of a block adjacent to the call block may be referred to as a temporal motion vector candidate. History-based MV candidates may be MVs in a list containing MVs of other blocks that were encoded / decoded before the encoding / decoding of the target block.

[0455] The encoding device (110) can determine the MV to be used for encoding a target block within a search range using an MV candidate list. The maximum number of MV candidates in the MV candidate list may be predefined. N may represent the predefined maximum number. For example, N may be 2. Alternatively, the maximum number of such candidates may be signaled from the encoding device to the decoding device or derived from the decoding device. The encoding device (110) can determine an MV candidate to be used as the predicted MV of the target block among the MV candidates in the MV candidate list. The MV to be used for encoding the target block may be an MV that can be encoded at the minimum cost. The encoding device (110) may determine whether to use the AMVP mode in encoding the target block and may generate AMVP mode usage information indicating whether the AMVP mode is used.

[0456] Inter prediction information may include 1) AMVP mode usage information, 2) MV candidate index, 3) MVD, 4) MVD resolution information, 5) reference direction and 6) reference image index, and may include residual blocks. Inter prediction information may be signaled from the encoding device (110) to the decoding device (150) in the form of a bitstream.

[0457] The decoding device (150) can obtain AMVP mode usage information from the bitstream. If the AMVP mode usage information indicates that the AMVP mode is being used, the decoding device (150) can obtain an MV candidate index, an MVD, MVD resolution information, a reference direction, and a reference image index from the bitstream. Among the MV candidates included in the MV candidate list, the MV candidate pointed to by the MV candidate index can be selected as the predicted MV of the target block.

[0458] The MVD may represent the difference between the MV that will actually be used for inter-prediction of the target block and the predicted MV. The encoding device (110) may derive a predicted MV that is close to the MV that will actually be used for inter-prediction of the target block in order to use an MVD of the smallest possible size. The decoding device (150) may derive the MV of the target block by summing the MVD and the predicted MV. That is to say, the MV of the target block derived by the decoding device (150) may be the sum of the MVD and the predicted MV candidates.

[0459] Additionally, the encoding device (110) can generate MVD resolution information. The MVD resolution information may be information used to adjust the resolution of the MVD. The decoding device (150) can adjust the resolution of the MVD using the MVD resolution information.

[0460] Meanwhile, the encoding device (110) can calculate the MVD based on an affine model. The affine control point MV of the target block can be derived based on the sum of the affine control point MV candidates and the MVD. Using the affine control point MV, the MV of each sub-block within the target block can be derived.

[0461]

[0462] Merge Mode

[0463] When merge mode is used, a merge candidate list containing multiple merge candidates can be generated using motion information of spatial candidates and motion information of temporal candidates, etc. Motion information may include 1) MV, 2) reference image index and 3) reference direction, etc. A merge candidate may be motion information.

[0464] Merge candidates may include 1) spatial merge candidates generated based on spatial candidates, 2) temporal merge candidates generated based on temporal candidates, 3) history-based merge candidates, 4) average merge candidates, and 5) zero merge candidates.

[0465] A history-based merge candidate may be movement information within a list containing movement information of other blocks that were encoded / decoded earlier than the encoding / decoding of the target block.

[0466] The average merge candidate may be a merge candidate generated based on the average of two merge candidates within the merge candidate list.

[0467] Zero merge candidates can be zero vector motion information. Zero vector motion information can be motion information where MV is a zero vector.

[0468] Merge candidates can be added to the merge candidate list according to a predefined method and a predefined order so that the merge candidate list has a set number of merge candidates. The same merge candidate list can be configured in the encoding device (110) and the decoding device (150) through the predefined method and a predefined order.

[0469] The encoding device (110) can select a merge candidate to be used for encoding a target block from among the merge candidates in the merge candidate list. The encoding device (110) can determine whether to use a merge mode in encoding the target block and can generate merge mode usage information indicating whether the merge mode is used.

[0470] Inter prediction information may include 1) merge mode usage information, 2) merge index and 3) correction information, etc., and may include residual blocks. Inter prediction information may be signaled in bitstream form from the encoding device (110) to the decoding device (150).

[0471] The decoding device (150) can obtain merge mode usage information from the bitstream. If the merge mode usage information indicates that the merge mode is being used, the decoding device (150) can obtain merge mode-related information, such as a merge index, from the bitstream.

[0472] The encoding device (110) can select the optimal merge candidate among the merge candidates included in the merge candidate list and can set the value of the merge index to point to the selected merge candidate.

[0473] Correction information may be information used for correcting the MV. The encoding device (110) may generate correction information. The decoding device (150) may derive a corrected MV by performing correction on the MV of a merge candidate selected by a merge index based on the correction information. The corrected MV may be used as the MV of the target block.

[0474] In one embodiment, the correction information may include an MVD. The correction information may include one or more of correction usage information, correction direction information, and correction magnitude information. The correction usage information may indicate whether to use correction for the MV. A merge mode that performs correction for the MV based on the correction information may be referred to as a merge mode having an MVD.

[0475] In merge mode, a prediction for the target block can be performed using the merge candidate pointed to by the merge index among the merge candidates included in the merge candidate list.

[0476] Movement information of the target block can be derived from 1) MV, 2) reference image index and 3) reference direction of the merge candidate pointed to by the merge index.

[0477] In one embodiment, the merge candidates in the merge candidate list may be specific modes that induce inter-prediction information. A merge candidate may be information pointing to a specific mode that induces inter-prediction information. Inter-prediction information of a target block may be induced according to the specific mode pointed to by the merge candidate. In this regard, the specific mode may be regarded as a specific inter-prediction information inducing mode or a specific movement information inducing mode. The specific mode may include a series of processes that induce inter-prediction information.

[0478] Inter-prediction information of the target block can be derived according to a specific mode pointed to by a merge candidate selected by a merge index among the merge candidates in the merge candidate list. For example, specific modes may include a mode for deriving motion information at the sub-block level and a mode for deriving motion information at the affine level, and may include other modes for deriving motion information as described in the embodiments.

[0479] Skip mode may be a mode that does not use residual blocks. That is to say, when skip mode is used, the restoration block may be identical to the prediction block. The description of the merge mode in the embodiments may also apply to skip mode. The difference between merge mode and skip mode may be whether or not residual blocks are signaled and used. That is to say, skip mode may be similar to merge mode except that residual blocks are not transmitted / used, and the description of merge mode may also apply to skip mode.

[0480] The subblock merge mode may be a mode in which motion information of a target subblock is induced for a target subblock within a target block. When the subblock merge mode is applied, a list of subblock merge candidates may be generated using affine control point motion vector merge candidates and / or subblock-based temporal merge candidates. The subblock-based temporal merge candidates may be motion information of the call subblock of the target subblock.

[0481] In GPM, a first prediction block and a second prediction block can be generated using two sets of motion information for a target block. For each coordinate of the target block, a final prediction sample of the final prediction block can be generated using the weighted sum of the first prediction sample of the first prediction block and the second prediction sample of the second prediction block.

[0482] Here, the first weight for the first prediction sample of the weighted consensus and the second weight for the second prediction sample can be determined based on the boundaries of the GPM. The boundaries may represent dividing lines that divide the target block. Depending on the boundaries, the target block may be divided into a first divided region and a second divided region.

[0483] If the distance between the final prediction sample and the boundary is less than or equal to a reference value, the value of the final prediction sample of the final prediction block may be determined using the weighted sum of the first prediction sample of the first prediction block and the second prediction sample of the second prediction block. If the distance between the final prediction sample and the boundary is greater than the reference value, one of the first weight and the second weight may be 1 and the other may be 0.

[0484] The Combined Inter-Intra Prediction (CIIP) mode may be a mode that derives a prediction sample of a target block using a weighted sum of a prediction sample generated by inter-prediction and a prediction sample generated by intra-prediction.

[0485] In the aforementioned modes, self-improvement of the derived motion information may be performed, and the improved motion information may be used as motion information for the target block. For example, blocks within a specific area determined based on the derived motion information may be searched, and the motion information of the block having the smallest Sum of Absolute Differences (SAD) value among the searched blocks may be used as the improved motion information for the target block. The specific area may be a square area within the reference image specified by the motion information. The point indicated by the motion information may be the center of the specific area.

[0486] In the aforementioned modes, compensation for prediction samples derived through inter-prediction can be performed using optical flow.

[0487]

[0488] FIG. 5 shows the order of addition of spatial candidates to the candidate list according to one embodiment.

[0489] In Fig. 5, the locations of the spatial candidates are shown.

[0490] The large block in the center can represent the target block. The five small blocks adjacent to the target block can represent spatial candidates.

[0491] The coordinates of the target block can be (xP, yP), and the size of the target block can be (nPSW, nPSH).

[0492] Spatial candidate A0 may be a block adjacent to the bottom-left of the target block. A0 may be a block occupying a sample of coordinates (xP - 1, yP + nPSH).

[0493] Spatial candidate A1 may be a block adjacent to the left of the target block. A1 may be the bottommost block among the blocks adjacent to the left of the target block. Or, A1 may be a block adjacent to the top of A0. A1 may be a block occupying a sample of coordinates (xP - 1, yP + nPSH - 1).

[0494] Spatial candidate B0 may be a block adjacent to the top right of the target block. B0 may be a block occupying a sample of coordinates (xP + nPSW, yP - 1).

[0495] Spatial candidate B1 may be a block adjacent to the top of the target block. B1 may be the rightmost block among the blocks adjacent to the top of the target block. Or, B1 may be a block adjacent to the left of B0. B1 may be a block occupying a sample of coordinates (xP + nPSW - 1, yP - 1).

[0496] Spatial candidate B2 may be a block adjacent to the top-left corner of the target block. B2 may be a block occupying a sample of coordinates (xP - 1, yP - 1).

[0497] As illustrated in Fig. 5, in adding spatial candidates to the candidate list, B1, A1, The order B0, A0, and B2 can be used. That is, B1, A1, Available spatial candidates can be added to the candidate list in the order of B0, A0, and B2. The order in which spatial candidates illustrated in FIG. 5 are added to the merge candidate list may be just one example.

[0498] The above candidate list may include a motion information candidate list, a merge candidate list, an MV candidate list, a BV candidate list, and an MPM list, etc.

[0499] To include a spatial or temporal candidate in the candidate list, it may be determined whether the spatial or temporal candidate is available. If a candidate block is outside the boundaries of an image, slice, or tile, the availability of the candidate block may be set to false. The description "availability is set to false" may mean "it is set to non-available."

[0500] The maximum number of candidates in the candidate list can be set. N can represent the set maximum number. The set maximum number can be signaled through a parameter set or header, etc. For example, the maximum number of candidates in the candidate list for a target block within a slice can be set by the slice header. For example, the value of N can be 5 by default.

[0501]

[0502] IBC mode

[0503] The IBC mode may be an intra-block copy prediction mode that generates a prediction block for a target block by referencing an already restored region within the target image. In this respect, the IBC mode may also be referred to as a current image reference mode. A block vector (BV) may be used to identify the already restored region.

[0504] Whether the target block is encoded / decoded in IBC mode can be determined using IBC mode usage information. The encoding device (110) can determine whether to use IBC mode in encoding the target block and can generate IBC mode usage information indicating whether IBC mode is used. The decoding device (150) can obtain IBC mode usage information from the bitstream.

[0505] In IBC mode, the predicted block of the target block can be generated based on the BV. The BV can specify the reference block. The BV can indicate the displacement between the target block and the reference block. The reference block can be a block within the target image. The description of the MV of the embodiments can also be applied to the BV.

[0506] The IBC mode may include a skip mode, a merge mode, and an AMVP mode, etc. The descriptions of the AMVP mode, merge mode, and skip mode of the embodiments may be similarly applied to the AMVP mode, merge mode, and skip mode of the IBC mode, respectively.

[0507] In skip mode or merge mode, a merge candidate list may be configured, and a merge index may specify one merge candidate from among the merge candidates in the merge candidate list. The BV of the specified merge candidate may be used as the BV of the target block.

[0508] In AMVP mode, BVD can be used. The description of MVD in the embodiments can also be applied to BVD.

[0509] The reference block in IBC mode may be limited to a block within an already restored region of the target image. Alternatively, the reference block may be contained within at least one of the target CTU or the left CTUs. For example, the value of BV may be restricted so that the reference block is located within a specific region. The specific region may be an area of ​​three blocks of a specific size that are encoded / decoded before the block of a specific size containing the target block. The specific size may be 64x64.

[0510]

[0511] Transformation and Quantization

[0512] Quantized levels can be generated by performing a transformation and / or quantization on the residual block. The residual block can represent the difference between the original block and the prediction block. A restored residual block can be generated by performing inverse quantization and / or inverse transformation on the quantized levels. The restored residual block can represent the difference between the restored block and the prediction block.

[0513] When a transformation or inverse transformation is performed, a separable transform or a 2D non-separable transform may be performed on the residual block. A separable transform may be a transformation that performs 1D transformations on the residual block in the horizontal and vertical directions, respectively.

[0514] The transformation kernels used for the transformation may include various DCT kernels such as DCT type 2 (DCT-II), 2) DST kernels, and 3) kernels derived by training. For 1D transformation, DCT type and DST type may include DCT-V, DCT-VIII, DST-I, and DST-VII in addition to DCT-II.

[0515] A transformation set may be used to determine the DCT type, DST type, or learning-derived kernel to be used for the transformation. Each transformation set may include multiple transformation candidates. Each transformation candidate may be a DCT type, a DST type, or a learning-derived kernel, etc.

[0516] The encoding device (110) can perform conversion and inverse conversion using conversion candidates included in the conversion set. The decoding device (150) can perform inverse conversion using conversion candidates included in the conversion set. Conversion selection information indicating which conversion candidate is used among the plurality of conversion candidates included in the conversion set applied to the residual block may be signaled. The conversion selection information may include vertical conversion selection information and horizontal conversion selection information. The vertical conversion selection information may indicate which conversion among the conversions belonging to the conversion set is used for the vertical conversion. The horizontal conversion selection information may indicate which conversion among the conversions belonging to the conversion set is used for the horizontal conversion.

[0517] The transformation may include at least one of a primary transformation and a secondary transformation. A primary transformation coefficient may be generated by performing a primary transformation on a residual block, and a secondary transformation coefficient may be generated by performing a secondary transformation on the transformation coefficient. Here, the transformation coefficient may include a primary transformation coefficient and a secondary transformation coefficient.

[0518] A first-order transformation may mean a Multiple Transform Selection (MTS) that applies different transformations to each of the 1D directions (i.e., vertical and horizontal directions).

[0519] A second-order transformation may be a transformation intended to improve the energy concentration of the transformation factors generated by a first-order transformation. A second-order transformation may be 1) a separable transformation like the first-order transformation, or 2) a 2D non-separable transformation. A 2D non-separable transformation may refer to a Low Frequency Non-Separable Transform (LFNST) or a Non-Separable Primary Transform (NSPT).

[0520] NSPT can be applied to specific block sizes such as 4x4, 4x8, 8x4, 4x16, 16x4, 8x8, 8x16, and 16x8 for intra-coding.

[0521] A first-order transformation may be performed using at least one of a plurality of predefined transformation methods. For example, the plurality of predefined transformation methods may include DCT, DST, and KLT, etc. Additionally, the first-order transformation may be a transformation having various transformation types according to transformation kernel functions that define DCT and DST. For example, the first-order transformation may include a plurality of transformations such as DCT-2, DCT-4, DCT-5, DCT-7, DCT-8, DST-1, DST-2, DST-4, DST-7, and DST-8 according to a plurality of transformation kernels.

[0522] In one embodiment, the transformation type may be determined based on coding parameters related to the target block. For example, the transformation type may be determined based on one or more of 1) the prediction mode of the target block (e.g., one of intra prediction and inter prediction), 2) the size of the target block, 3) the shape of the target block, 4) the intra prediction mode of the target block, 5) the components of the target block (e.g., one of luminance components and chroma components), and 6) the splitting type applied to the target block (e.g., one of QT, BT, TT, and non-split).

[0523] As in the first transformation, a transformation set can be defined in the second transformation as well. Methods for deriving and / or determining the transformation set of the embodiments can be applied to the second transformation as well as the first transformation.

[0524] In one embodiment, the first transformation and / or second transformation may be determined for a specific target. The transformation selection information may include transformation target information. The transformation target information may indicate the target to which the first transformation and / or second transformation is applied.

[0525] For example, first-order transformation and / or second-order transformation may be applied to one or more signal components among the luminance component and the chroma component.

[0526] In one embodiment, the transformation selection information may include first transformation usage information and second transformation usage information. The first transformation usage information may indicate whether a first transformation is applied to the residual block of the target block. The second transformation usage information may indicate whether a second transformation is applied to the residual block of the target block.

[0527] In one embodiment, whether a first transformation and / or a second transformation is applied may be determined based on coding parameters for the target / neighbor block, such as the size and shape of the target / neighbor block.

[0528] In one embodiment, the transformation selection information may include first transformation selection information and second transformation selection information. The first transformation selection information may indicate a transformation method applied to a residual block among a plurality of transformation methods that can be used in the first transformation. The first transformation selection information may be a first transformation index. The second transformation selection information may indicate a transformation method applied to a transformation coefficient among a plurality of transformation methods that can be used in the second transformation. The second transformation selection information may be a second transformation index.

[0529] In one embodiment, the transformation methods of the first transformation and the second transformation can each be derived based on specific information such as coding parameters. For example, the coding parameters may include coding parameters for target / neighbor blocks.

[0530] In the embodiments, information related to a transformation, such as transformation selection information, and sub-information of the transformation selection information may be signaled to a specific target. For example, the specific target may be a CU.

[0531] Information related to transformations, such as transformation selection information, and sub-information of transformation selection information can be derived for a specific target. For example, the specific target may be a CU.

[0532] Quantized levels can be generated by performing quantization on the result or residual block generated by performing a first-order transformation and / or a second-order transformation.

[0533] The description of the transformation described above may also be applied to the inverse transformation. In such application, the inverse processing of the processing described for the transformation may be performed in the inverse transformation. "Transformation" within the name related to the transformation may be changed to "inverse transformation." Additionally, the input of the transformation may be considered as the output of the inverse transformation. The output of the transformation may be considered as the input of the inverse transformation. The decoding device (150) may obtain information related to the transformation, such as transformation selection information, and may use the information related to the transformation to perform the inverse processing of the processing related to the transformation indicated by the information related to the transformation.

[0534] The target block may include multiple subblocks. Each subblock may be defined according to a minimum block size or minimum block shape. The target block may be divided into multiple subblocks, and each subblock may include coefficients such as 4x4, 2x8, and 8x2. The target block may be a transformation block. Transform coefficients or quantized levels may be represented in the form of a block. Transform coefficients may be quantized transformation coefficients.

[0535] Transform coefficients or quantized levels may be scanned according to at least one scanning type among diagonal scanning, vertical scanning, and horizontal scanning. Diagonal scanning may be top-right diagonal scanning or bottom-left diagonal scanning.

[0536] For example, by scanning the coefficients of a block using diagonal scanning, the coefficients can be changed or arranged into a one-dimensional vector form. Vertical scanning may be scanning the coefficients in the form of a two-dimensional block in a column direction. Horizontal scanning may be scanning the coefficients in the form of a two-dimensional block in a row direction.

[0537] The scanning type for the coefficients can be determined based on coding parameters such as intra prediction mode, block size, and block shape. For example, based on coding parameters such as intra prediction mode, block size, and block shape, it can be determined which scanning method—diagonal scanning, vertical scanning, and horizontal scanning—will be used. A block may be a transformation unit.

[0538] Scanning according to each scanning type can start at a specific starting point and end at a specific ending point.

[0539] In scanning, the scanning order according to the scanning type can first be applied between subblocks. Next, the scanning order according to the scanning type can be applied to the transformation coefficients or quantized levels within the subblocks.

[0540] The encoding device (110) can perform entropy encoding on the conversion coefficients or quantized levels to generate a bitstream containing entropy-encoded conversion coefficients or entropy-encoded quantized levels.

[0541] The decoding device (150) can generate the transform coefficients or quantized levels by obtaining entropy-encoded transform coefficients or entropy-encoded quantized levels from the bitstream and performing entropy decoding. The coefficients can be arranged in the form of two-dimensional blocks through inverse scanning. The arrangement of inverse scanning may be a rearrangement opposite to the arrangement of scanning.

[0542] Backscanned transform coefficients or backscanned quantized levels can be generated through backscanning of the coefficients. In this case, the backscanning types of backscanning may include diagonal scans, vertical scans, and horizontal scans, and a backscanning type of the inverse transform corresponding to the scanning type of the transform may be selected.

[0543] In the decoding device (150), inverse quantization can be performed on the (backscanned) coefficients. Depending on whether a second inverse transform is performed, a second inverse transform can be performed on the result generated by the performance of inverse quantization. Also, depending on whether a first inverse transform is performed, a first inverse transform can be performed on the result generated by the performance of the second inverse transform. By selectively performing a second inverse transform and a first inverse transform on the coefficients, a restored residual block can be generated.

[0544]

[0545] Filtering

[0546] To improve the image quality, filtering may be performed on the blocks. The value of the target sample may be determined or updated by the filtering.

[0547] The target sample may be one of the samples described in the embodiments. For example, the target sample may be one or more of the samples described in the embodiments, such as a prediction sample, a reference sample, a residual sample, a reconstructed sample, and a reconstructed sample to which filtering has been applied.

[0548] The target sample may be a sample within one or more of the target picture, target slice, target CTB, target block, reference sample line, and template. The target block may be one of the blocks described in the embodiments. For example, the target block may be one or more of the blocks described in the embodiments, such as a transformation block, prediction block, reference block, residual block, and restoration block.

[0549] In the embodiments, the filtering process described as being applied to one target may also be applied to other targets. For example, the filtering process described in a specific in-loop filtering may also be applied to a transformation block, a prediction block, a reference block, and a residual block, etc.

[0550] For the filtering of the embodiments, a specific type of filtering may be used. The type of filtering may include a filter tap (or filter tap length), a filter shape, a filter strength, filter coefficients (or weights), and an offset.

[0551] The filter tab may indicate the number of input samples used for the filter. The input samples may include target samples. Alternatively, the input samples may include specific values ​​determined for the target samples. The input samples may include one or more reference samples. One or more reference samples may be determined based on the attributes of the target block described in the embodiments. The attributes may include coding parameters. For example, the attributes of the target sample may include the location of the target sample. One or more reference samples may be specified based on their relative position to the location of the target sample.

[0552] The filter shape can represent the shape formed by input samples. A specific value determined for a target sample can be considered as the target sample. In other words, if a specific value determined for a target sample is used as an input sample for a filter, the target sample can also be considered as constituting the filter shape.

[0553] There may be multiple samples whose values ​​are determined by filtering. Filter strength may represent the range of samples whose values ​​are determined by filtering. Filter strength may be either strong filtering strength or weak filtering strength. The number of samples whose values ​​are determined by strong filtering strength may be greater than the number of samples whose values ​​are determined by weak filtering strength. Alternatively, filter strength may represent the range of values ​​that are changed by filtering. The range of sample values ​​changed by strong filtering strength may be wider than the range of sample values ​​changed by weak filtering strength.

[0554] Filter coefficients can be coefficients or weights of the input samples.

[0555] The offset can be a specific value added to the result calculated using the values ​​and coefficients of the input samples, such as a weighted sum.

[0556] Filtering, interpolation, and sampling may be common in that they update the values ​​of samples. Accordingly, the description of any one of filtering, interpolation, and sampling in the embodiments may also apply to the other one of filtering, interpolation, and sampling. Here, sampling may include at least one of upsampling, downsampling, and subsampling.

[0557] Filtering may include filtering performed by a predictor (123) and a predictor (163), etc.

[0558] In encoding for a target block, a prediction error may exist between the original sample of the original block and the prediction sample of the prediction block. To reduce the prediction error, filtering may be performed on at least one of the prediction sample of the prediction block and the reference sample referenced for prediction.

[0559] For example, in intra-prediction, the reference samples may include one or more of the top-left reference sample, top reference sample, top-right reference sample, left reference sample, and bottom-left reference sample. Filtering of the prediction samples may be performed by applying specific weights to the prediction samples, left reference samples, top reference samples, and / or top-left reference samples, respectively.

[0560] Filtering for at least one of the prediction sample and the reference sample may be performed based on the attributes of the target block and the attributes of the prediction sample. For example, whether filtering is performed, the type of filter, the area to which filtering is applied, the weights of the filtering, the reference sample, the range of the reference sample, and the location of the reference sample may each be determined based on the attributes of the target block and the attributes of the prediction sample.

[0561] For example, the attributes of the target block may include information related to the target block described in the embodiments, such as 1) size, 2) prediction mode, 3) intra prediction mode, 4) reference sample line, 5) sample value, and 6) coding parameter.

[0562] For example, the attributes of the prediction sample may include information related to the prediction sample described in the embodiments, such as 1) the sample value and 2) the location within the target block of the prediction sample, and may include coding parameters regarding the prediction sample.

[0563] Filtering may include in-loop filtering performed by a filter (130) and a filter (170), etc.

[0564]

[0565] Figure 6 shows a plurality of in-loop filters according to one example.

[0566] Multiple in-loop filters of in-loop filtering may include one or more of Luma Mapping with Chroma Scaling (LMCS), deblocking filter, Sample Adaptive Offset (SAO), and Adaptive Loop Filter (ALF).

[0567] Multiple in-loop filters can be connected sequentially. For example, multiple in-loop filters can be connected in the order of LMCS, deblocking filter, SAO, and ALF. Additionally, multiple in-loop filters can be connected in any order of all available permutations of the multiple in-loop filters. The output from one of the multiple in-loop filters can be used as an input to the next filter.

[0568] As illustrated in FIG. 6, an input image may be input to the first filter. The input image may be a block as described in the embodiments. For example, the input image may be a restored block generated by an adder (129) or an adder (169). The output from one filter may be input to the next filter. An output image may be generated by the last filter. The output image may be a filtered block as described in the embodiments. For example, the output image may be a filtered restored image generated by a filter (130) or a filter (170).

[0569] The target block can represent the image input to the filter. The filtered target block can represent the image output from the filter.

[0570] LMCS may include luminance signal mapping for the luminance signal of the target block and chroma signal scaling for the chroma signal of the target block.

[0571] Luminous signal mapping can perform codeword redistribution for the luminous signal.

[0572] Luma signal mapping may include forward mapping and inverse mapping. In forward mapping, the existing dynamic range may be divided into multiple intervals. The mapped dynamic range can be determined by performing codeword redistribution on the input image using a linear model for each interval. In inverse mapping, inverse mapping from the mapped dynamic range to the existing dynamic range is performed.

[0573] Chroma scaling can correct the chroma signal based on the interrelationship between the luminance signal and the corresponding chroma signal.

[0574] Forward mapping can be performed between inter-prediction of the luminance signal and restoration of the luminance signal, and between inter-prediction of the luminance signal and chroma scaling. Inverse mapping can be performed between restoration of the luminance signal and in-loop filtering of the luminance signal. Chroma scaling can be performed between inverse transform and restoration of the chroma signal.

[0575] According to this structure, inverse quantizations for the luminance and chroma signals, inverse transforms for the luminance and chroma signals, prediction for the luminance signal, and restoration for the luminance signal can be performed within a mapped dynamic domain. In-loop filtering for the luminance and chroma signals, inter-predictions for the luminance and chroma signals, intra-predictions for the chroma signal, and restoration for the chroma signal can be performed within the existing dynamic domain.

[0576] A deblocking filter can remove block distortion occurring at the boundaries between blocks within the reconstructed image. For example, the blocks may be transformed blocks. Additionally, the blocks may be subblocks of a specific block described in the embodiments. Here, the boundaries between blocks may refer to samples adjacent to the boundaries between blocks.

[0577] A deblocking filter can be applied to the vertical and horizontal boundaries between blocks. After filtering is performed on the vertical boundaries of the blocks, filtering can be performed again on the horizontal boundaries of the filtered blocks.

[0578] A deblocking filter may be applied optionally. Whether to apply a deblocking filter to a target block may be determined based on at least one of sample(s) contained within a specific number of columns or rows within the target block and sample(s) contained within a specific number of columns or rows within a neighboring block adjacent to a specific boundary.

[0579] When a deblocking filter is applied to a target block, the filter to be applied may be determined according to the required strength of deblocking filtering. In other words, among a plurality of different filters, the filter determined according to the strength of deblocking filtering may be applied to the target block. The plurality of filters may include one of a long-tap filter, a strong filter, a weak filter, and a Gaussian filter.

[0580] The maximum length of the deblocking filter can be determined based on attributes of the target block, such as the size of the target block, the components of the target block, and coding parameters.

[0581] SAO can compensate for distortion between the original image and the reconstructed image on a sample basis. For compensation, SAO can apply an appropriate offset to the sample values. In other words, the offset can be added to the sample values.

[0582] An offset can be determined for the target block. For example, the offset can be determined for each component of the CTB. The determined offset can be applied to samples within a specific component of the CTB.

[0583] SAO may include an SAO using an edge offset (EO) and an SAO using a band offset (BO). Depending on the characteristics of samples within a specific block, such as a CTU, whether to perform an SAO using EO and whether to perform an SAO using BO may be determined, respectively.

[0584] In an SAO using EO, correction for distortion of samples can be performed based on the direction of edges within the target block. The pattern classes of the EO may include horizontal patterns, vertical patterns, 135-degree diagonal patterns, and 45-degree diagonal patterns. For the target block, information indicating the pattern class applied to the target block and multiple offsets of the said pattern class may be signaled. There may be four offsets. For a target sample within the target block, adjacent samples of the target sample may be determined according to the direction of the pattern class. An offset to be applied to the target sample may be determined by the pattern of the adjacent samples.

[0585] In an offset using BO, correction for sample distortion can be performed by classifying the brightness values ​​of samples within the target block into specific bands. The bit depth of the input image can be divided into m intervals. For example, m can be 32. The specific bands can be n consecutive intervals among the m intervals. For example, n can be 4. n offsets for the n intervals can be signaled. Additionally, information indicating the first interval selected as the n intervals among the m intervals can be signaled. The offset of the interval corresponding to the target sample can be added to the sample value of the target sample of the target unit.

[0586] ALF can compensate for distortion between the restored image and the original image.

[0587] The filter coefficients of the ALF can be signaled through the bitstream.

[0588] The filter shape of the ALF can be determined by the components of the target block. For example, a 7x7 diamond-shaped filter can be used for the luminance component. A 5x5 diamond-shaped filter can be used for the chroma component.

[0589] In ALF, the characteristics of a specific block can be determined for that block, and the class of that specific block can be determined based on those characteristics. In other words, the determination of characteristics and the determination of the class in ALF can be performed in units of 4x4 blocks. Filter coefficients can be calculated according to the class. A specific block can be a block with a size of 4x4.

[0590] One of 25 classes can be determined as the class of a specific block based on the direction and activity determined using the gradient of the specific block. Depending on the gradient of the specific block, a rotation transformation, a vertical reflection transformation, and / or a diagonal reflection transformation may be applied to the filter.

[0591] Information regarding whether ALF is applied can be signaled to specific units such as CTBs.

[0592] An index indicating a filter to be applied to a specific unit among the available filters may be signaled. Here, the available filters may include fixed filters and filters configured using a parameter set. For example, the parameter set may be an Adaptive Parameter Set (APS). The fixed filters may be predefined identically in the encoding device (110) and the decoder (150). The filter coefficients of the filters configured using the parameter set may be determined based on the coding parameters.

[0593]

[0594] Entropy Encoding and Entropy Decoding

[0595] Figure 7 shows entropy encoding and entropy decoding according to one example.

[0596] The processes of entropy encoding by the entropy encoder (139) are illustrated at the top of Fig. 7.

[0597]

[0598] * The entropy encoder (139) may include a context modeler, a binarization unit, and an entropy encoding unit. The context modeler may include a context selection unit and a context memory.

[0599] The binarization unit can generate binaries for syntactic elements by performing binarization on the syntactic elements of the target block. Binarization may be a process of converting syntactic elements into the form of binaries.

[0600] Information about syntax elements and beans can be provided from the binarization unit to the context selection unit.

[0601] The context modeler can perform context updates.

[0602] Context can refer to occurrence probability information for each bin regarding syntactic elements that have already been encoded.

[0603] The context modeler may perform a context update to apply current probability information to the entropy encoding of the bins of the syntactic elements of the target block. The updated context may be stored in context memory. At this time, the updated context corresponding to the syntactic elements of the target block (or the bins within the syntactic elements of the target block) may be derived by the context modeler.

[0604] The context selector can select a context corresponding to a bin of a syntactic element of a target block. The selected context can be loaded from context memory and used as an updated context for entropy encoding of the bins of the syntactic element of the target block.

[0605] The updated context can be used for entropy encoding of syntactic elements of the target block.

[0606] The entropy encoding unit can generate encoded information for syntactic elements of a target block by performing entropy encoding using generated bins and an updated context, and can generate a bitstream containing the encoded information. The entropy encoding unit may use at least one of an arithmetic encoding method and a bypass encoding method.

[0607] At the bottom of Fig. 7, the processes of entropy decoding by the entropy decoder (161) are illustrated.

[0608] The entropy decoder (161) may include a context modeler, an entropy decoder, and an inverse binary converter. The context modeler may include a context selection unit and a context memory.

[0609] The context modeler can perform context updates.

[0610] The context can refer to the probability information of occurrence for each bin regarding syntactic elements that have already been decoded.

[0611] The context modeler may perform a context update to apply the currently decoded probability information to the entropy decoding for the bins of the syntactic elements of the target block. The updated context may be stored in context memory. At this time, the updated context corresponding to the syntactic elements of the target block (or the bins within the syntactic elements of the target block) may be derived by the context modeler.

[0612] The context selection unit can select a context corresponding to a bin of a syntactic element of a target block. The selected context can be loaded from context memory and can be used as an updated context for entropy decoding of the syntactic element of the target block.

[0613] The updated context can be used for entropy decoding of the syntactic elements of the target block.

[0614] The entropy decoding unit can generate bins for the segmentation elements of the target block by performing entropy decoding on the encoded information of the bitstream based on the updated context. The entropy decoding unit may use at least one of an arithmetic decoding method and a bypass decoding method.

[0615] The debinaryization unit can obtain syntactic elements of a target block by performing debinaryization on at least one of the generated beans. Debinaryization may be a process of converting at least one of the beans into the form of a syntactic element.

[0616] Information about syntax elements and beans can be provided from the inverse binary unit to the context selector.

[0617] The syntax element may be one of the coding parameters described in the examples.

[0618]

[0619] Methods for binarization, inbinarization, entropy encoding, and entropy decoding

[0620] In the embodiments, to perform signaling for specific information, one or more of the binarization method, inverse binarization method, entropy encoding method and entropy decoding method listed below may be used.

[0621] - Signed 0-th order Exponential Golomb binarization / debinarization method (abbreviated as se(v))

[0622] - Signed k-order exponential-Golomb binarization / debinarization method (abbreviated as sek(v))

[0623] - 0-order exponentiation-Golomb binarization / debinarization method for unsigned positive integers (abbreviated as ue(v))

[0624] - k-order exponential-Golomb binarization / debinarization method for unsigned positive integers (abbreviated as uek(v))

[0625] - Fixed-length binarization / debinarization method (abbreviated as f(n))

[0626] - Truncated Rice binarization / debinarization method or truncated unary binarization / debinarization method (abbreviated as tu(v))

[0627] - Truncated binary binarization / debinarization method (abbreviated as tb(v))

[0628] - Context-adaptive arithmetic encoding / decoding method (abbreviated as ae(v))

[0629] - Bit string in bytes (abbreviated as b(8))

[0630] - Signed integer binary / debinary conversion method (abbreviated as i(n))

[0631] - Unsigned positive integer binarization / debinarization method (abbreviated as u(n)) ('u(n)' may also refer to a fixed-length binarization / debinarization method.)

[0632] - Unary Binary / Debinary Method

[0633]

[0634] Adaptive execution of the processing of the examples

[0635] The processing of the embodiments may be performed in the same and / or corresponding way in the encoding device (110) and the decoding device (150). Additionally, a combination of one or more of the above embodiments may be used for encoding and / or decoding of the image.

[0636] The order in which the embodiments are applied may differ from one another in the encoding device (110) and the decoding device (150). Alternatively, the order in which the above embodiments are applied may be the same (at least partially) in the encoding device (110) and the decoding device (150).

[0637] The processing of the embodiments may be performed for each of the specific targets. The processing of the embodiments may be performed identically for the specific targets. For example, the specific targets may include a luminance signal and a chroma signal.

[0638] The processes of the embodiments may be selectively applied / performed based on specific conditions or specific targets.

[0639] In one embodiment, the processing of the embodiment may be selectively applied / performed according to a temporal layer. Temporal layer information for a specific processing may be information indicating the temporal layer where the processing can be applied / performed. Temporal layer information may be signaled for a specific processing. Temporal layer information may indicate the lowest layer and / or the highest layer where the specific processing can be applied, and may indicate the specific layer where the specific processing is applied / performed. Alternatively, a fixed temporal layer where the processing of the embodiment is applied / performed may be defined.

[0640] In one embodiment, a type to which the processing of the embodiments is applied / performed may be defined, and whether the processing of the embodiments is applied / performed may be determined based on the defined type. The type may include a picture type, a slice type, and a tile group type, etc.

[0641] According to the description of the embodiments, when applying / performing a specific process on a specific target, specific conditions may be required, and the specific process may be performed under a specific decision. Where it is determined whether a specific condition is satisfied based on a specific coding parameter, or where a specific decision is made based on a specific coding parameter, such specific coding parameter may be interpreted as being replaceable with another coding parameter. That is to say, the coding parameter affecting the specific condition or specific decision described in the embodiments may be considered merely exemplary, and in addition to the specified coding parameter, one or more other coding parameters or a combination of one or more other coding parameters may be understood to perform the role of the specified coding parameter.

[0642] The processing of the embodiments may be applied / performed based on the size of at least one of the blocks described in the embodiments. For example, the blocks may include a coding block, a prediction block, a transformation block, a reference block, a current block, and a target block. Alternatively, the blocks may include adjacent blocks of the embodiments. Here, the size may be defined as a minimum size and / or a maximum size for the processing of the embodiments, or as a fixed size for the processing of the embodiments. Additionally, for the processing of the embodiments, a first embodiment may be applied at a first size, and a second embodiment may be applied at a second size. That is, the processing of the embodiments may be applied in combination depending on the size. Additionally, the processing of the embodiments may be applied only when the block size is greater than or equal to the minimum size and less than or equal to the maximum size. That is, the processing of the embodiments may be applied only when the block size falls within a specific range.

[0643] FIG. 8 is a flowchart of an image encoding / decoding method that generates a color difference prediction signal using the component-to-component prediction model of the present disclosure.

[0644] In the present disclosure, steps S810 to S830 may be performed in each of the image encoding device and the image decoding device.

[0645] In step S810, reference samples are determined for deriving the Cross-Component Prediction Model (CCP model). The region containing the reference samples may be referred to as the reference region.

[0646] To determine a reference sample or a representative value of the reference sample for deriving a prediction model (or filter coefficients) between components, at least one of the surrounding restoration samples, prediction samples, and residual samples of the current block may be used.

[0647] Here, a sample can mean a pixel or a signal.

[0648] Representative values ​​can refer to statistical values ​​of reference samples. For example, representative values ​​can be the mean, weighted mean, weighted sum, etc.

[0649] Nearby restoration samples may include adjacent samples of the current block, non-adjacent samples, or samples within a reference picture.

[0650] The number of reference lines or reference samples required to determine reference samples or representative values ​​of reference samples can be determined using encoding information or decoding information.

[0651] In one embodiment, the number of reference lines or reference samples may be determined by considering CTU boundaries, VPDU boundaries, etc.

[0652] For example, when some variants of a block are adjacent to a CTU boundary or a VPDU boundary, the number of reference lines or reference samples derived from the area adjacent to the boundary may differ from the number of reference lines or reference samples derived from the area not adjacent to the boundary.

[0653] For example, the number of reference lines or reference samples derived from a region adjacent to the boundary may be less than the number of reference lines or reference samples derived from a region not adjacent to the boundary.

[0654] In addition, the number of reference lines and reference samples referenced to derive representative values ​​within a region adjacent to the boundary may be N. Here, N is a value greater than or equal to 0, which may be a pre-set value in the encoder / decoder and may be a value signaled from the encoder to the decoder.

[0655] Figure 9 is a diagram illustrating the number of reference lines and reference samples according to the block shape.

[0656] In one embodiment, the number of reference lines and reference samples used to derive the representative value may vary depending on the size of the surrounding block and the current block.

[0657] For example, if the size of an adjacent block is smaller than the size of the current block, the number of reference lines and reference samples used to derive the representative value of the filter in the reference region including the boundary abutting the adjacent block may be reduced.

[0658] For example, if the size of an adjacent block is larger than the size of the current block, the number of reference lines and reference samples used to derive the representative value of the filter in the reference region including the boundary abutting the adjacent block may increase.

[0659] For example, as the size of the current block decreases, the number of reference lines and reference samples used to derive the representative value of the filter in the reference area can be reduced.

[0660] For example, as the size of the current block increases, the number of reference lines and reference samples used to derive the representative value of the filter in the reference area that includes the face touching the adjacent block can be increased.

[0661] In this case, whether a modified reference line / sample is used can be signaled from the encoder to the decoder. Whether an adaptive reference line / sample is used can be signaled from the encoder to the decoder.

[0662] At this time, the number of changed reference lines / samples may be a pre-set value in the encoder / decoder, or it may be signaled from the encoder to the decoder.

[0663] In one embodiment, the number of reference lines and reference samples used to derive the representative value may vary depending on the shape or division type of the surrounding block or the current block.

[0664] For example, if the current block or surrounding blocks are divided so that the horizontal size is larger than the vertical size, the number of reference lines and reference samples used to derive the representative value of the filter in the reference area including the face adjacent to the horizontal face may increase.

[0665] For example, if the current block or surrounding blocks are divided so that the horizontal size is larger than the vertical size, the number of reference lines and reference samples used to derive the representative value of the filter in the reference area including the face adjacent to the vertical face may increase.

[0666] For example, if the current block or surrounding blocks are divided so that the horizontal size is larger than the vertical size, the number of reference lines and reference samples used to derive the representative value of the filter in the reference area including the face adjacent to the horizontal face may be reduced.

[0667] For example, if the current block or surrounding blocks are divided so that the horizontal size is larger than the vertical size, the number of reference lines and reference samples used to derive the representative value of the filter in the reference area including the face adjacent to the vertical face may be reduced.

[0668] In this case, whether a modified reference line / sample is used can be signaled from the encoder to the decoder. Whether an adaptive reference line / sample is used can be signaled from the encoder to the decoder.

[0669] At this time, the number of changed reference lines / samples may be a pre-set value in the encoder / decoder, or it may be signaled from the encoder to the decoder.

[0670] In one embodiment, the number of reference lines and the number of reference samples may vary depending on the prediction mode of the surrounding / current block.

[0671] For example, the number of reference lines and reference samples can be varied depending on the direction of the in-screen prediction mode.

[0672] For example, if the in-screen prediction mode is vertical mode, the number of reference lines / samples above the current block can be increased.

[0673] For example, if the in-screen prediction mode is vertical mode, the number of reference lines / samples to the left of the current block can be reduced.

[0674] For example, if the in-screen prediction mode is horizontal mode, the number of reference lines / samples above the current block can be reduced.

[0675] For example, if the in-screen prediction mode is horizontal mode, the number of reference lines / samples to the left of the current block can be increased.

[0676] At this time, whether the changed reference line / sample is used can be signaled from the encoder to the decoder.

[0677] At this time, the number of changed reference lines / samples may be a pre-set value in the encoder / decoder, or it may be signaled from the encoder to the decoder.

[0678] For example, if a neighboring block is encoded in an inter-frame prediction mode, samples included in that block can be excluded.

[0679] In the above example of increasing or decreasing the number of reference lines / samples, the increase or decrease may be applied to an example opposite to the above example. For example, as the opposite of an example in which the number of reference lines / samples is increased, the number of reference lines / samples may be set to decrease.

[0680] In one embodiment, the number of reference lines and the number of reference samples may vary by using the motion vector (MV) / block vector (BV) of surrounding blocks.

[0681] For example, the number of reference lines or reference samples for the direction or position pointed to by the motion vector / block vector can be increased.

[0682] In one embodiment, when determining a reference line and a reference sample, a region found through a matching method may be included in the reference region.

[0683] FIG. 10 is a diagram illustrating how to adjust the area of ​​reference samples so that a location found using Intra Template Matching (Intra TMP) / IBC is included in the reference area.

[0684] For example, as shown in Fig. 10, the reference area can be adjusted so that the location found using Intra Template Matching (Intra TMP) / IBC mode is included in the reference area.

[0685] For example, as shown in Fig. 10, a new reference region can be set based on the block vector location derived through IntraTMP / IBC.

[0686] For example, if the existing reference area was 6 lines from the face adjacent to the current block, you can set a new 6 lines as the reference area starting from the template area found through IBC / IntraTMP.

[0687] For example, a new reference region can be established based on the location pointed to by the block vector (BV) derived via IntraTMP / IBC. In this case, the established reference region may be an area non-adjacent to the current block.

[0688] The above block vector may be derived from a referenceable adjacent / non-adjacent / collocated block or a block within a previously encoded reference picture.

[0689] FIG. 11 is a diagram illustrating the determination of a reference region using block vectors derived from collocated blocks or surrounding blocks.

[0690] FIG. 12 is a diagram illustrating the derivation of a block vector from a block at a specific location within a corresponding block.

[0691] For example, as shown in FIG. 11, a collated block is an IntraTMP mode or IBC predicted block, the collated block has a block vector, the block vector of the current block is derived from the collated block, and a reference region for deriving model parameters can be determined using the derived block vector. In this case, the luminance / chrominance reference region can be determined as the block at the location indicated by the block vector within the current picture, the reference picture, or the collated picture. The reference block can be downsampled or subsampled, and the sampled reference block can be used as the reference region.

[0692] At this time, as shown in FIG. 12, a block vector can be derived from a block at a specific location within a luminance block corresponding to a color difference block. For example, five locations may be used. Alternatively, a block vector can be derived from blocks at specific locations within a juxtaposition block corresponding to a color difference block.

[0693] In determining reference samples for deriving a prediction model between components, multiple reference regions may be derived using at least one of surrounding reconstructed samples, prediction samples, and residual samples.

[0694] Figures 13a and 13b illustrate examples of template configurations.

[0695] For example, as shown in FIGS. 13a and FIGS. 13b, a plurality of templates can be derived using surrounding samples, and at least one reference region can be determined using the templates. Each template includes at least one line.

[0696] In determining reference samples for deriving a prediction model between components, multiple reference regions can be derived from luminance samples at different locations. Multiple reference regions can be constructed based on the locations of the luminance samples.

[0697] For example, an area containing luminance samples of even-numbered lines among the reference lines may be determined as a first reference area, and an area containing luminance samples of odd-numbered lines may be determined as a second reference area.

[0698] In determining reference samples for deriving a prediction model between components, samples used for deriving reference regions are classified based on a predefined threshold, and different reference regions can be constructed using the classified samples.

[0699] For example, first samples having sample values ​​below a threshold having a specific value and second samples having sample values ​​exceeding the threshold are distinguished, a first reference region is derived from the first samples, and a second reference region different from the first reference region can be derived from the second samples.

[0700] The above threshold may be a preset value in the encoder / decoder, or a value signaled from the encoder to the decoder.

[0701] In deriving multiple reference regions, reference regions can be derived based on the statistical values ​​of luminance samples.

[0702] For example, by using the average value of luminance samples as a threshold, samples with values ​​below the threshold and samples with values ​​exceeding the threshold can be distinguished, and different reference regions can be derived from each distinguished sample.

[0703] For example, the sample values ​​of the luminance samples can be sorted in order of proximity to the average value, N samples from the top and S samples from the bottom can be selected, and different reference regions can be derived using the selected sample groups. In this case, N and S may be pre-set values ​​in the encoder / decoder or values ​​signaled from the encoder to the decoder.

[0704] Figures 14a, 14b, 14C, and 14D are examples of grouping samples within a range of sample values ​​between a minimum sample value and a maximum sample value.

[0705] For example, in FIG. 14, the range of sample values ​​between the maximum and minimum values ​​of the samples can be divided into N groups, and different models can be derived for each group using the samples of each divided group. In this case, N may be a pre-set value in the encoder / decoder and may be a value signaled from the encoder to the decoder.

[0706] For example, in FIG. 14b, samples can be classified into groups. All samples can be grouped into two groups using the average of all sample values. For each group, samples with sample values ​​below the average value and samples with sample values ​​exceeding the average value can be separated and grouped. Different reference regions can be constructed for each group using the statistical values ​​of the samples. In this case, different reference regions can be constructed by calculating the average value of the samples for each group and creating a new sample group based on the average value.

[0707] For example, in FIG. 14C, samples can be divided into the above group units. All samples are grouped into Group 1 and Group 2 using the average of all sample values, and Group 1 is grouped into Group 1-1 and Group 1-2 using the average value of Group 1. Group 2 is grouped into Group 2-1 and Group 2-2 using the average value of Group 2.

[0708] For example, in FIG. 14D, a new reference region can be constructed by forming a new group between the average values ​​of samples calculated by group. Referring to FIG. 14C and FIG. 14D, samples between the average of sample values ​​of group 1 and the average of sample values ​​of group 2 can be regrouped into group 3.

[0709] In deriving multiple reference regions, samples for model derivation can be distinguished based on the centroid of the samples.

[0710] For example, a centroid for dividing samples into two groups can be derived in the following way.

[0711] STEP 1) Calculate the average value of all pixel values ​​(AVG_OLD).

[0712] STEP 2) Samples can be classified into two groups based on the average value (AVG_OLD), and the sample average value for each group can be recalculated (AVG_G1, AVG_G2) using the samples belonging to each group.

[0713] STEP 3) STEP 2 can be repeated with the median of the average values ​​calculated from the two groups as the new average value (AVG_NEW = (AVG_G1 + AVG_G2) / 2).

[0714] In this regard, at the start of STEP 2, AVG_OLD can be updated to AVG_NEW. At this time, if AVG_NEW is the same as AVG_OLD, the loop can be stopped.

[0715] STEP 4) The AVG_NEW calculated in STEP 3 can be determined as the centroid.

[0716] Whether to use the method of distinguishing samples using the above centroid is determined according to a pre-set convention in the encoder / decoder, or can be signaled from the encoder to the decoder. The above centroid value may be a pre-set value in the encoder / decoder, or a value signaled from the encoder to the decoder.

[0717] In determining the above reference area, a block at least one of the locations indicated by a motion vector or block vector derived from a surrounding block or a juxtaposed block, or an adjusted location of said location, may be included in the reference area for deriving an inter-component prediction model.

[0718] At this time, the block positions for deriving the block vector or motion vector may be five positions in FIG. 12.

[0719] For example, in FIG. 15, at least one of the blocks moved from an existing adjacent / non-adjacent block can be included in a reference area for model derivation using the block vector of the juxtaposed block.

[0720] Figure 15 illustrates an example of surrounding blocks adjusted (moved) using block vectors.

[0721] In this case, an area moved using a block vector or motion vector from a reference area configured to derive a prediction model between the components of the current block can be used as a new reference area. The area indicated by the vector may overlap with or be located outside the existing reference area.

[0722] In this case, if the block at the position indicated by the vector has a block vector or a motion vector, a new motion vector or block vector can be derived through a vector accumulation method as shown in FIG. 16. The area indicated by the first vector and the area indicated by the accumulation vector can be used as reference areas. In other words, if a second reference vector is available within the area indicated by the first reference vector, the block area or template area indicated by the second reference vector can be used as a reference area.

[0723] FIG. 16 illustrates an embodiment of deriving a new candidate (vector) using a block vector or motion vector stored in a block at a location indicated by a previously derived block vector or motion vector in a list.

[0724] In deriving multiple reference regions, at least one of the aforementioned methods may be selected and used.

[0725] Reference regions are derived using the aforementioned methods, inter-component prediction models are derived using the reference regions, and the optimal prediction model and optimal reference region can be determined based on the coding cost calculated during the prediction coding process.

[0726] The optimal method among the methods for distinguishing the above samples can be determined by the encoder / decoder using a predefined method shared with each other, or it can be signaled from the encoder to the decoder.

[0727] In color difference signal prediction based on a component-to-component prediction model, at least one reference sample among the predicted / residual / reconstructed samples within the reference region can be modified or reshaped.

[0728] In this case, transformation (modification) may refer not only to modifying the value by performing arithmetic operations using a specific value on the corresponding sample / pixel, but also to changing the sample / pixel value through filtering using a specific filter.

[0729] In one embodiment, samples belonging to a reference region can be modified using statistical values ​​of samples within a reference region or a prediction block.

[0730] For example, the value of a sample at a specific location within a reference area can be added to or subtracted from the samples included in the reference area.

[0731] For example, the average value of samples within a reference area or prediction block can be calculated, and the average value can be added to or subtracted from the samples within the reference area or prediction block.

[0732] For example, the values ​​of samples having maximum and minimum values ​​can all be transformed into a specific value. In this case, the specific value may be a pre-set value in the encoder / decoder, or it may be signaled from the encoder to the decoder.

[0733] For example, it can be set to 0.

[0734] For example, all sample values ​​that differ from the average by a threshold can be changed to a specific value.

[0735] In this case, the threshold can be the variance or standard deviation of the samples.

[0736] If the above specific value is the variance or standard deviation, and the sample values ​​within the reference area are greater than (mean + variance / standard deviation) or less than (mean - variance / standard deviation), the values ​​of those samples can be designated as the specific value.

[0737] In this case, the threshold and specific value may be pre-set values ​​in the encoder / decoder, or may be signaled from the encoder to the decoder.

[0738] In one embodiment, samples within a reference region or prediction block can be modified (adjusted) through filtering.

[0739] For example, reference samples or prediction samples can be modified through smoothing based on a Gaussian filter or a low-pass filter.

[0740] For example, reference samples or prediction samples can be modified using a sharpening filter.

[0741] For example, reference samples or prediction samples can be modified using edge-preserving filters such as bilateral filters.

[0742] For example, reference samples or prediction samples can be modified using edge detection filters such as Sobel / Canny filters.

[0743] For example, at least one of a plurality of downsampling filters can be selected to modify the reference sample or the prediction sample.

[0744] For example, samples within a reference region or prediction block can be modified using a component-to-component prediction model derived from adjacent / non-adjacent blocks within a referenceable current / reference picture.

[0745] When modifying samples within a reference region or prediction block through filtering, at least one of the aforementioned filters may be used.

[0746] In filtering prediction samples using the above filter, not only samples within the prediction block but also reference samples (surrounding samples) can be used.

[0747] Figure 17 illustrates an example of adjacent surrounding restoration samples for filtering prediction samples within a prediction block.

[0748] For example, as shown in FIG. 17, adjacent surrounding reconstructed samples can be used to filter the prediction samples of the top-left boundary of the current block. Alternatively, samples within the reference region can be filtered using the surrounding samples of each sample. Filtering may be intended to reduce the prediction error.

[0749] For example, when filtering prediction samples at the bottom right boundary of the current block as shown in Fig. 17, if there are no adjacent reconstructed samples available for filtering, padding samples can be derived using samples within the block, and filtering can be performed using the padded samples.

[0750] At this time, a new component-to-component prediction model (or color-to-component prediction) can be derived using reference region samples modified through the filtering or transformation method described above.

[0751] At this time, in each color-component-between prediction, whether the sample used for the prediction mode derivation is filtered and the filter information used can be signaled to the encoder for decoding.

[0752] At this point, in addition to the previously defined inter-component prediction model, a new inter-component prediction model can be defined using modified (or filtered) reference samples.

[0753] For example, a model can be defined and derived by varying the number of model parameters and the coefficients (gradient, location) that constitute the parameters. Additionally, a model can be defined and derived by varying the number of model parameters and the coefficients (gradient, location) that constitute the parameters depending on the type of filter used.

[0754] In the example above using a downsampling filter, the component-to-component prediction models can be defined differently depending on the type of downsampling filter used to sample the luminance sample.

[0755] For example, depending on the type of filter used, the model can be defined and derived by varying the number of model parameters and the model inputs (luminance samples, gradient, location).

[0756] The application status of the above-derived model, model information, and downsampling filter information used for model derivation may be pre-configured in the encoder and decoder or signaled from the encoder to the decoder.

[0757] At this time, samples within the reference region or prediction block can be modified using a pre-derived component-to-component prediction model.

[0758] For example, samples within a reference region or prediction block can be modified using a component-to-component prediction model derived from the current block or surrounding blocks (a pre-calculated or stored model), and a new component-to-component prediction model can be derived using the modified samples within the reference region or prediction block.

[0759] The above-derived model can be derived from the current block or surrounding blocks.

[0760] The use of the above-mentioned modified reference sample or prediction sample may be pre-configured in the encoder / decoder or signaled from the encoder to the decoder.

[0761] The method of modifying the above reference region or prediction sample can be equally applied to methods that use prediction samples (or blocks) or templates during the encoding / decoding process.

[0762] At this time, the sample value can be changed (or modified) by applying the reference area or prediction sample modification (or modification) method to at least one sample among the samples within the prediction block and the samples within the template.

[0763] For example, when performing IBC, samples within the IBC prediction block can be modified (or changed) and used.

[0764] At this time, using the statistical values ​​of samples within the reference area or prediction block, samples belonging to the reference area can be modified, or samples within the reference area or prediction block can be modified (adjusted) through filtering.

[0765] At this time, a prediction model between components can be derived using surrounding reconstructed samples within the current / reference picture, and then samples within the IBC prediction block can be filtered using the model.

[0766] For example, after deriving a CCCM model using surrounding reconstructed samples, the IBC prediction block can be filtered using it.

[0767] At this time, filtering can be performed on at least one sample within the prediction block using a component-to-prediction model derived (or stored) from adjacent / non-adjacent blocks within the referenceable current / reference picture.

[0768] Whether the method using the above-mentioned modified prediction sample is used may be pre-configured in the encoder / decoder or signaled from the encoder to the decoder.

[0769] For example, in IntraTMP (Intra template matching), IntraTMP can be performed by modifying (or changing) the samples within the template.

[0770] At this time, using the statistical values ​​of the samples within the template area, samples belonging to the template area can be modified, or samples within the template area can be modified (adjusted) through filtering.

[0771] At this time, a prediction model between components is derived using restored samples within the current / reference picture, and then at least one sample within the template region in the Intra TMP can be filtered using the model.

[0772] For example, after deriving a CCCM model using surrounding reconstructed samples, at least one sample within an Intra TMP template can be filtered using this.

[0773] At this time, at least one sample within the Intra TMP template can be filtered using a component-to-prediction model derived (or stored) from adjacent / non-adjacent blocks within the referenceable current / reference picture.

[0774] Whether to use the Intra TMP with the above-mentioned modified template may be pre-configured in the encoder / decoder or signaled from the encoder to the decoder.

[0775] Among multiple reference regions, the optimal reference region can be determined based on the encoding cost.

[0776] Specifically, a component-to-component prediction model is derived using samples within each reference region, and the encoding cost of the predictive encoding process is calculated using the component-to-component prediction model. Based on the encoding cost calculated from each reference region, the optimal reference region among multiple reference regions can be determined. Information regarding the selected reference region can be signaled from the encoder to the decoder.

[0777] For example, inter-component prediction models are derived from a top template, a left template, and an L-shaped template, the encoding cost according to the inter-component prediction models is calculated, and the inter-component prediction model and template having the lowest encoding cost are selected as the optimal model and optimal template and can be signaled.

[0778] Among multiple reference regions, the optimal reference region can be determined based on the template matching cost.

[0779] For example, a template is constructed that includes some of the reconstructed samples or predicted samples, a prediction template is derived by applying a component-to-component prediction model derived from a certain reference region to the template, and a template matching cost for that reference region can be calculated based on the difference between the prediction template and the reconstructed template. In other words, the template matching cost can be calculated based on the difference between the predicted samples and the reconstructed samples within the template region. The reference region with the lowest template matching cost can be determined as the optimal reference region.

[0780] The optimal reference region among multiple reference regions can be determined by comparing the coefficient values ​​of the component-to-component prediction models derived from the multiple reference regions.

[0781] For example, an upper reference region, a left reference region, and an L-shaped reference region are derived, and if the difference in coefficient values ​​between prediction models derived from the upper reference region and the left reference region is below a threshold, the upper reference region and the left reference region are considered similar, and the L-shaped reference region can be determined as the optimal reference region.

[0782] The use of downsampling or subsampling for a reference region can be determined based on encoding costs or template costs.

[0783] It can be applied in the same way as the example for determining the final reference region in multiple reference regions. When the encoding cost or template cost calculated from the downsampled reference region or the subsampled reference region is low, it may be determined that downsampling or subsampling is performed.

[0784] The use of downsampling or subsampling for a reference region can be determined using encoding information.

[0785] For example, depending on the value of the Quantization Parameter (QP), the use of downsampling or subsampling may be determined. If the QP value of the reference region is above a threshold, downsampling for the reference region may be omitted. If the QP value is below the threshold, downsampling for the reference region may be performed.

[0786] For example, the use of downsampling or subsampling may be determined based on gradient values ​​or histograms of gradient (HoG) representing the slope between samples within a surrounding block or reference region. If the gradient values ​​of samples within the reference region are below a specific threshold—that is, if there is little directionality between samples within the reference region—downsampling for the reference region may be omitted. If the gradient values ​​for similar directions exceed the threshold, downsampling for the reference region may be performed.

[0787] For example, the use of downsampling or subsampling may be determined based on the size of the current block or adjacent blocks. When the size of the current block or adjacent blocks is greater than a threshold, downsampling of the reference region may be performed. When the size of the current block or adjacent blocks is smaller than a threshold, downsampling of the reference region may be performed.

[0788] Downsampling or subsampling can be performed using multiple filters. The type of filter may vary depending on the encoding information. The filters may include downsampling filters or subsampling filters.

[0789] For example, as multiple filters, a 2x2 filter and a 3x3 filter may be used together. As a 2x2 filter, [[1 / 4, 1 / 4], [1 / 4, 1 / 4]] may be used. For example, as a 3x3 filter, [[1 / 16, 2 / 16, 1 / 16], [2 / 16, 4 / 16, 2 / 16], [1 / 16, 2 / 16, 1 / 16]] may be used.

[0790] Meanwhile, step S810 may be omitted based on the method for deriving the prediction model between components.

[0791] In step S820, an inter-component prediction model for predicting color difference signals is derived.

[0792] In one embodiment, the prediction model between components can be derived using surrounding samples of the current block.

[0793] In color difference signal prediction based on color component relationships, a prediction model between components and a predicted signal can be derived based on a Cross-component Linear Model (CCLM).

[0794] The color difference signal prediction method based on the above cross-component linear model can generate a color difference prediction block for color difference prediction using the sample values ​​of restored color difference samples adjacent to the current color difference block and restored luminance samples at corresponding locations.

[0795] Here, the sample may include a pixel, a statistical value of pixel values ​​of at least one pixel, or encoding information derived from at least one pixel.

[0796] At this time, a linear regression model can be derived as shown in Equation 1 below using statistical values ​​such as the correlation between the surrounding pixels of the luminance block and the surrounding pixels of the chrominance block, or maximum and minimum values, and the model coefficients a and b can be calculated. Then, the chrominance block to be currently predicted can be derived using these two values ​​and the pixels within the luminance block.

[0797] [Mathematical Equation 1] C'(i, j) = a * L'(i, j) + b or C'(i, j) = a * L(x, y) + b

[0798] In this case, C(i, j) may represent the current color difference prediction block or a sample within the color difference prediction block, and L'(i, j) may represent a restored luminance block / sample corresponding to the position of the color difference block to be predicted. However, L'(i, j) may represent a luminance sample of a luminance block in which the size of the luminance block has been adjusted to be equal to the size of the color difference block through subsampling, downsampling, etc. L(x, y) may represent a luminance sample of a luminance block that is not equal to the size of the color difference block. That is, L(x, y) may represent a luminance sample to which downsampling or subsampling has not been applied. Therefore, (x, y) of the luminance sample L(x, y) corresponding to the color difference sample C'(i, j) may be different from the color difference sample position (i, j).

[0799] At this time, the luminance sample position (x, y) corresponding to the color difference sample position (i, j) can be expressed as (x, y) = (2 * i, 2 * j).

[0800] In deriving model coefficients according to the above mathematical formula 1, not only can they be derived for each of the U and V signals, but model coefficients applicable to both U and V signals can also be derived.

[0801] Whether the above linear model is applied can be determined for U and V respectively, and this can be signaled.

[0802] In deriving a color difference prediction signal based on a cross-component linear model, a cross-component linear model can be derived using the offset (lumaOffset / chromaOffset) value of the luminance signal or the color difference signal as shown in Equation 2 below.

[0803] For example, offset values ​​of luminance signals or chrominance signals can be added to or subtracted from the coefficients of the cross-component linear model as shown in the formula below.

[0804] [Mathematical Formula 2] C' (i, j) (or, pred_C(i, j)) = a * (L'(i, j) - lumaOffset)+ chromaOffset

[0805] In this case, the luminance offset (lumaOffset) and chroma offset (chromaOffset) may be statistical values ​​such as the average value of samples within each reference area, or values ​​of samples at a specific location (e.g., a sample adjacent to the top-left sample of the current block).

[0806] In this case, if there are no samples needed to derive the offset value, the mid-valued value of the pixel can be designated as the offset value.

[0807] The mid-valued value of a pixel = (bit depth - 1) << 2. For example, if the bit depth is 10 bits, the mid-valued value is 512.

[0808] In deriving a cross-component linear model and generating color difference prediction samples using the derived model, different mathematical formulas may be used for model derivation and model application.

[0809] For example, when deriving model coefficients, the model coefficients are derived using Equation 2, to which an offset value is applied. When generating color difference prediction samples, model coefficients without an offset value applied, such as Equation 1, may be used.

[0810] For example, in the YUV = 4:4:4 or YUV = 4:2:2 format, adjacent sample lines of the luminance block and corresponding adjacent surrounding sample lines of the chrominance block at the same location can be used.

[0811] For example, in the YUV = 4:2:0 format, the adjacent sample line index (IntraChromaRefLineIdx) of the color difference block to be used can be derived as the floor(IntraLumaRefLineIdx / 2) value.

[0812] In addition, the number of luminance samples and the number of color difference samples can be matched by using only the even or odd-numbered samples from the luminance reference line restored based on a sub-sampling method. (The luminance sample corresponding to the color difference sample can be determined.)

[0813] In addition, a luminance sample located at a specific sample position can be determined as a sample corresponding to a chrominance sample. For example, when the size of the luminance block is 2Mx2N and the size of the chrominance block is MxN, samples located at (M+N)*X1 / 8, (M+N)*X2 / 8, (M+N)*X3 / 8, and (M+N)*X4 / 8 within the luminance block can be determined as luminance samples corresponding to chrominance samples. In this case, Xi (i = 1, 2, 3) is an arbitrary constant for determining the sample position and can increase as the number of samples increases. For example, when using 4 samples as above, the constants assigned to X1 through X4 can be arbitrarily determined as (1, 3, 5, 7), (2, 4, 6, 8), etc., and the values ​​may be different depending on the encoding information of the surrounding blocks.

[0814] In addition, N luminance samples can be used to derive a luminance sample corresponding to one color difference sample. In this case, N can be any positive integer.

[0815] For example, a luminance sample corresponding to a color difference sample can be determined using downsampling.

[0816] At this time, one representative luminance sample can be derived by using statistical values ​​such as sample values, average values, maximum values, minimum values, or median values ​​of N luminance samples within the entire or part of the reference area.

[0817] For example, when the luminance samples corresponding to the adjacent sample c(0, -1) of the current color difference block are denoted as L(-1, -2), L(0, -2), L(1, -2), L(-1, -1), L(0, -1), and L(1, -1), candidate values ​​for deriving a representative value can be derived using at least one of these samples through the mean, maximum, minimum, and median values. Here, the candidate values ​​for obtaining the representative value may refer to statistical values. Here, the representative value may refer to the sample values ​​of the luminance samples corresponding to the color difference samples used to derive model coefficients. Specifically, the reference line of the color difference block can be derived through the following equation. c(0, -1)= (L(-1, -2) + 2 * L(0, -2) + L(1, -2) + L(-1, -1) + 2 * L(0, -1) + L(1, -1) + 4) >> 3. Here, >> can mean the right shift operator.

[0818] The luminance samples corresponding to the above color difference samples may vary, and the formula for deriving the color difference prediction samples may also vary. Information regarding luminance samples corresponding to adjacent samples of the current color difference block or information regarding the formula for deriving adjacent samples of the current color difference block may be signaled by an encoder.

[0819] The above sample location may be a pre-set value in the encoder / decoder, or a value signaled from the encoder to the decoder. It may be transmitted from the encoder. Alternatively, it may be derived through encoding information of surrounding samples or blocks.

[0820] In predicting color difference signals based on the correlation between color component encoding information, a prediction signal can be derived based on a Gradient Linear Model (GLM), which represents the amount of change in sample values ​​or statistical values ​​of sample values ​​among various encoding information.

[0821] A color difference signal prediction method using a linear model between color components based on longitude values ​​can generate a color difference prediction block / signal using longitude (slope) values ​​obtained using restored color difference samples adjacent to the current color difference block and restored luminance samples at corresponding locations.

[0822] At this time, a linear regression model can be derived as shown in Equation 3 below by utilizing the correlation between the surrounding pixels of the luminance block and the surrounding pixels of the chrominance block, or the correlation between statistical values ​​such as maximum and minimum values, and the model coefficients a and b can be calculated. Then, the chrominance signal / block to be currently predicted can be derived using these two values ​​and the hardness values ​​of the pixels within the luminance block.

[0823] [Mathematical Equation 3]: C'(I, j) = a * G(I, j) + b

[0824] In this case, C(I, j) may represent the current color difference block to be predicted or the pixel value within the color difference block, and G(I, j) may represent a hardness sample derived from the restored luminance block corresponding to the location of the current color difference block to be predicted.

[0825] In deriving the model coefficients of the above mathematical formula 3, not only can they be derived for each of the U and V signals, but model coefficients applicable to both U and V signals can also be derived.

[0826] Whether the above linear model is applied can be determined for U and V respectively, and this can be signaled.

[0827] In detecting the above hardness value, multiple hardness detection filters can be used to derive a hardness sample corresponding to the color difference sample. In this case, the number of filters used may be at least one, and the number of filters may be a value pre-set in the encoder / decoder or a value signaled from the encoder to the decoder.

[0828] At this time, a hardness value can be detected for each filter that detects the hardness value, and a corresponding hardness-based linear model can be derived.

[0829] For example, let's assume that multiple hardness-based linear models as shown below were derived using hardness values ​​detected through multiple hardness detection filters. (The example uses four hardness detection filters to derive four models.)

[0830] Model 1: C' (i, j) = a * G(i, j) + b → Derivation of hardness value using Filter 1

[0831] Model 2: C'' (i, j) = c * G(i, j) + d → Derivation of hardness value using Filter 2

[0832] Model 3: C''' (i, j) = e * G(i, j) + f → Derivation of hardness value using Filter 3

[0833] Model 4: C'''' (i, j) = g * G(i, j) + h → Derivation of hardness value using Filter 4

[0834] In predicting color difference signals using the above-mentioned longitude-based linear model, color difference signals / samples / blocks can be derived by using multiple longitude-based models.

[0835] In deriving multiple models, hardness samples can be derived from luminance samples at different locations.

[0836] For example, hardness samples can be generated from luminance samples located at even positions on a reference line and used to derive one model. At the same time, hardness samples can be generated from luminance samples located at odd positions and used to derive another model.

[0837] For example, samples located at the top and left of the luminance block can be distinguished, hardness samples can be generated using each distinguished sample, and different models can be derived using the generated hardness samples.

[0838] In addition, when deriving multiple models, models can be derived based on the statistical values ​​of hardness samples.

[0839] For example, by using the average value of the hardness samples as a threshold, samples with values ​​below the threshold and samples with values ​​exceeding the threshold can be distinguished, and different models can be derived from each distinguished hardness sample.

[0840] For example, by sorting the hardness sample values ​​in order of proximity to the mean, N samples from the top and S samples from the bottom can be selected, and different models can be derived using these. In this case, N and S may be pre-set values ​​in the encoder / decoder, or values ​​signaled from the encoder to the decoder.

[0841] Alternatively, when deriving multiple models, models can be derived based on defined thresholds.

[0842] For example, by setting a specific value as a threshold, samples with values ​​below the threshold and samples with values ​​exceeding the threshold can be distinguished, and different models can be derived from each of the distinguished samples.

[0843] The above threshold may be a pre-set value in the encoder / decoder, or a value signaled from the encoder to the decoder.

[0844] Multiple models can be derived using hardness samples derived through multiple hardness detection filters.

[0845] At this time, multiple models can be derived using hardness samples through a single hardness detection filter.

[0846] In deriving multiple models, models derived from don- and decrypted blocks can be used.

[0847] When deriving multiple models, models derived from blocks of a different size from the current block can be used.

[0848] In deriving multiple models, models derived from blocks located at different positions within the partitioned block can be used.

[0849] In deriving multiple models, a new model can be derived by modifying the coefficient values ​​of a previously derived (calculated) model.

[0850] In deriving multiple models, a new model can be derived through the weighted sum of the coefficient values ​​of previously derived (calculated) models.

[0851] The method for deriving the above model can be used not only for multiple models but also for deriving a single model for the current block.

[0852] In color difference signal prediction based on the relationship between color components, a prediction signal can be generated based on the Convolutional Cross-Component Model (CCCM).

[0853] The above convolutional cross-component model can derive a model as shown in Equations 4 to 6 below by utilizing the correlation between surrounding samples of the luminance block and surrounding samples of the chrominance block, or the correlation between statistical values ​​such as maximum and minimum values, and calculate the filter coefficient α_ value and the offset (or bias) b value. Then, the chrominance block to be currently predicted can be derived using the derived filter-based model and samples within the luminance block.

[0854] [Mathematical Equation 4] pred_C(i, j) = α1· L' (i, j) + α2· L' (i-1, j) + α3· L' (i+1,j) +α4· L' (i, j-1) + α5· L' (i, j+1) + α6· P + α7· B

[0855] [Mathematical Equation 5] predChromaVal = α1· C + α2· N + α3· S + α4· E + α5· W + α6· P + α7· B

[0856] [Mathematical Equation 6] pred_C(i, j) = α1·L (x, y) +α2·L (x-1, y) +α3·L (x+1, y) +α4·L (x, y-1) +α5·L (x, y+1) +α6·P+α7·B

[0857] Here, pred_C(or pred_C(i, j) ) may mean the current color difference prediction block or a sample value within the color difference prediction block, and L(x, y) may mean the restored luminance block / sample at position (x, y).

[0858] FIG. 18 illustrates an example of a downsampled luminance sample L'(i, j) and surrounding samples corresponding to a color difference sample C(i, j).

[0859] At this time, C, N, S, E, and W can be expressed as shown in Fig. 18.

[0860] In this case, P is a non-linear term and can be defined as follows.

[0861] P(C) = (LC*LC + midVal) >> bitDepth, P = (LC*LC + 512) >> 10 (for 10-bit), where LC may mean the downsampled luminance sample L' at the position corresponding to the chrominance signal position C.

[0862] At this time, B (bias term) is a scalar offset value between input and output, and in the case of 10 bits, it can generally be set to 512, which is the midpoint of the samples.

[0863] However, L'(x, y) refers to a luminance sample of a luminance block in which the size of the luminance block is adjusted to be the same as the size of the chrominance block through subsampling, downsampling, etc., and L(x, y) may refer to a luminance sample of a luminance block that is not the same as the size of the chrominance block. Therefore, (x, y) of the luminance sample L(x, y) corresponding to the chrominance prediction sample pred_C(i, j) may be different from the chrominance sample location (i, j).

[0864] As in the case of Equation 6 above, a CCCM model can be defined and model parameters derived using a luminance sample (L) rather than a downsampled luminance sample (L').

[0865] When constructing reference samples for model derivation, specific values ​​can be added to or subtracted from the reference samples as shown in Equation 7 below.

[0866] For example, as in Equation 7, the offset (lumaOffset / chromaOffset) values ​​of the luminance signal or the color difference signal can be added or subtracted.

[0867] [Mathematical Equation 7] pred_C(i, j) = α1· (L' (i, j) -lumaOffset) + α2·(L' (i-1, j) - lumaOffest) + α3·(L' (i+1,j)- lumaOffset) + α4·(L' (i, j-1) -lumaOffset) + α5· (L' (i, j+1) -lumaOffset) + α6·P + α7·B + chromaOffset

[0868] In this case, the luminance offset (lumaOffset) and chroma offset (chromaOffset) may be values ​​derived using statistical values ​​such as the average value of samples within each reference area, or values ​​of samples at a specific location (e.g., a sample adjacent to the top-left sample of the current block).

[0869] In this case, if there are no reference samples necessary to derive the offset value, the pixel's mid-valued value can be used as the offset value.

[0870] Pixel mid-valued = (bit depth - 1) << 2 (For example, if the bit depth is 10 bits, the mid-valued is 512)

[0871] When deriving a convolution cross-component model using mathematical formula 7 and generating color difference prediction samples using the derived model, different formulas may be used in the derivation and application of the convolution cross-component model.

[0872] For example, when deriving model coefficients, the model coefficients are derived using Equation 7, to which an offset value is applied. When generating color difference prediction samples, model coefficients without an offset value applied, such as Equations 4 to 6, may be used.

[0873] In deriving the model coefficients of the above mathematical formulas 4 to 7, not only can they be derived for each of the U and V signals, but model coefficients applicable to both U and V signals can also be derived.

[0874] Whether the above model is applied can be determined for both U and V or for each, and this can be signaled from the encoder to the decoder.

[0875] In the above convolutional cross-component model, the shape of the filter can be determined according to the position (x, y) of the luminance sample to which the filter is applied and the number of filters applied, wherein the position of the sample to which the filter is applied and the number of filters may be pre-set values ​​in the encoder / decoder and may be values ​​signaled from the encoder to the decoder.

[0876] In the above filter-based linear model, the shape of the filter can be determined according to the position (x, y) of the luminance sample to which the filter is applied and the number of filters applied, and the position of the sample to which the filter is applied and the number of filters may vary depending on the encoding information of the current block and surrounding blocks.

[0877] At this time, the shape of the filter can be configured in various forms, such as a cross, radial, hexagonal, diamond, or triangular pattern.

[0878] In configuring the shape of the above filter, the filter coefficients constituting the filter can be positioned at a specific distance from each other.

[0879] For example, the distance between the filter coefficients included in the filter can be equal.

[0880] For example, the distances between filter coefficients included in the filter can be different.

[0881] The distance between filter coefficients included in the filter can be selected from at least one of the above methods.

[0882] At this time, the distance between the filter coefficients may be a preset value in the encoder / decoder, or it may be signaled from the encoder to the decoder.

[0883] In addition, the shape of the filter can be configured not only symmetrically but also asymmetrically.

[0884] In addition, if a sample at a filter coefficient location included in the filter is not available for reference, it can be replaced with a nearby sample that is available for reference.

[0885] For example, it can be replaced with a sample located at the same spatial position within the reference image.

[0886] In this case, the shape and size of the filter can be varied depending on the size of the surrounding / current block.

[0887] For example, if the size of a neighboring adjacent block is smaller than the size of the current block, the size of the filter can be reduced, or if the sample to which the filter is applied is outside the adjacent block, that sample can be excluded from filtering.

[0888] For example, if the size of the surrounding adjacent blocks is larger than the size of the current block, the size of the filter can be increased, or the number of coefficients or samples included in the filter can be increased.

[0889] For example, as the current block size increases, the size of the filter can be increased or the number of filter coefficients can be increased. In addition, the spacing between filter coefficients can be increased.

[0890] For example, as the current block size increases, the size of the filter can be reduced or the number of filter coefficients can be decreased. In addition, the spacing between filter coefficients can be narrowed.

[0891] For example, as the current block size decreases, the size of the filter can be increased or the number of filter coefficients can be increased. Additionally, the spacing between filter coefficients can be increased.

[0892] For example, as the current block size decreases, the size of the filter can be reduced or the number of filter coefficients can be reduced. In addition, the spacing between filter coefficients can be narrowed.

[0893] At this time, the shape and size of the filter can be varied depending on the shape (divided form) of the surrounding / current block.

[0894] For example, if the current block or surrounding blocks are divided so that the horizontal size is larger than the vertical size, the filter size can be increased in the horizontal direction or, conversely, tilted in the vertical direction.

[0895] For example, if the current block or surrounding blocks are divided so that the horizontal size is smaller than the vertical size, the filter size can be increased in the vertical direction or, conversely, tilted in the horizontal direction.

[0896] In the above case, if the number of filter coefficients is fixed, the filter coefficients at the position corresponding to the fixed direction can be increased or decreased.

[0897] At this time, the shape and size of the filter can be varied depending on the prediction mode of the surrounding / current block.

[0898] For example, the shape and size of the filter can be varied according to the in-screen prediction mode direction. In this case, the size of the filter can be increased or the number of filter coefficients can be increased in the direction of the in-screen prediction mode from the current sample position. Additionally, the filter spacing can be adjusted. Alternatively, the size of the filter can be decreased or the number of filter coefficients can be reduced in the direction of the in-screen prediction mode from the current sample position. Additionally, the filter spacing can be adjusted.

[0899] The shape and size of the above filter can be determined using at least one of the filter shape and size determination methods in the aforementioned 'Filter based Linear Model (FLM)'.

[0900]

[0901] In predicting color difference signals based on relationships between color components, a prediction signal can be generated based on a Gradient and Location-Based Convolutional Cross-Component Model (GL-CCCM) utilizing gradient and location information.

[0902] The above GL-CCCM can derive a model as shown in Equations 8 and 9 below by utilizing the correlation between the surrounding samples of the luminance block and the surrounding samples of the chrominance block, in particular the gradient information (, ) between the surrounding samples and the position information (X, Y) of the corresponding sample within the block, and can calculate the filter coefficient α value and the offset (or bias) b value. Then, the chrominance block to be currently predicted can be derived using the derived filter-based model and the samples within the luminance block.

[0903] [Mathematical Equation 8] pred_C (i, j) = α1·L' (i, j) +α2·G y +α3·G x +α4·Y+α5·X+ α6·P+α7·B

[0904] [Mathematical Equation 9] predChromaVal = α1·C+α2·G+α3·G+α4·Y+α5·X+α6·P+α7·B

[0905] Here, pred_C(pred_C(i, j) ) may mean the current color difference prediction block or a sample value within the color difference prediction block, and L(x, y) may mean the restored luminance block / sample at position (x, y).

[0906] The position information Y and X of the sample within the block may represent a position (coordinates) moved horizontally (X) and vertically (Y) from the top-left position of the block.

[0907] At this time, the above Gy and Gx can be expressed as follows.

[0908] Gy = (2N + NW + NE) - (2S + SW + SE)

[0909] Gx = (2W + NW + SW) - (2E + NE + SE)

[0910] At this time, C, N, S, E, and W can be expressed as shown in Fig. 18.

[0911] FIG. 18 illustrates an example of a downsampled luminance sample L'(i, j) and surrounding samples corresponding to a color difference sample C(i, j).

[0912] In this case, P is a non-linear term and can be defined as follows.

[0913] P(C) = (LC*LC + midVal) >> bitDepth, P = (LC*LC + 512) >> 10 (for 10-bit), where LC may mean the downsampled luminance sample L' at the position corresponding to the chrominance signal position C.

[0914] At this time, B (bias term) is a scalar offset value between input and output, and in the case of 10 bits, it can generally be set to 512, which is the midpoint of the samples.

[0915] However, L'(x, y) refers to a luminance sample of a luminance block in which the size of the luminance block is adjusted to be the same as the size of the chrominance block through subsampling, downsampling, etc., and L(x, y) may refer to a luminance sample of a luminance block that is not the same as the size of the chrominance block. Therefore, (x, y) of the luminance sample L(x, y) corresponding to the chrominance prediction sample pred_C(i, j) may be different from the chrominance sample location (i, j).

[0916] When constructing reference samples for model derivation, specific values ​​can be added to or subtracted from the reference samples, similar to Equation 2 or Equation 7.

[0917] For example, in mathematical formulas 8 to 9, the offset (lumaOffset / chromaOffset) values ​​of the luminance signal or the chroma difference signal can be added or subtracted.

[0918] In deriving a model and generating color difference prediction samples using the derived model, different formulas may be used in the derivation and application of the model.

[0919] For example, when deriving model coefficients, the model coefficients are derived using a formula to which an offset value is applied. When generating color difference prediction samples, model coefficients to which an offset value is not applied may be used, such as in Equation 8 or Equation 9.

[0920] Model coefficients according to the above mathematical formula 8 or mathematical formula 9 can be derived for each of the U and V signals, as well as model coefficients applicable to both U and V signals.

[0921] Whether the above model is applied can be determined for both U and V or for each, and this can be signaled from the encoder to the decoder.

[0922] In predicting color difference signals based on relationships between color components, a prediction signal can be generated based on a convolutional cross-component model (MDF-CCCM: CCCM using multiple downsampling filters) based on multiple downsampling filters.

[0923] Figure 19 illustrates the positions of the downsampled luminance sample and the corresponding color difference sample.

[0924] In FIG. 19, the positions of the color difference sample C(i, j) corresponding to the downsampled luminance sample L'(i, j) are shown.

[0925] The spatial positional relationship between samples may vary depending on the chrominance signal sampling format, or at least one of the vertical or horizontal alignment relationship between the luminance sample and the chrominance sample. The chrominance signal sampling format may be one of YUV = 4:4:4, YUV = 4:2:2, or YUV = 4:2:0. The vertical alignment relationship may indicate whether the chrominance sample is moved horizontally by a distance of 0.5 samples from the luminance sample. The horizontal alignment relationship may indicate whether the chrominance sample is moved vertically by a distance of 0.5 samples from the luminance sample. Information regarding the chrominance signal sampling format, or the vertical or horizontal alignment relationship between the luminance sample and the chrominance sample, may be signaled from the encoder to the decoder. As an example, sample C(i, j) may be located between L2 and L3 instead of between L2, L3, L4, and L5.

[0926] FIG. 20 illustrates an example of using multiple downsampling filters to generate luminance samples corresponding to color difference samples.

[0927] In FIG. 20, the downsampled luminance sample L'(i, j) can be F1(C(i, j)). L'(i, j) is the downsampled luminance sample at the position corresponding to the chrominance sample C(i, j). That is, the downsampling filter can generate L'(i, j) corresponding to the position of the chrominance sample C(i, j) using the restored luminance sample.

[0928] Figure 21 illustrates a model that generates a color difference prediction signal and the information used by the model.

[0929] In FIG. 21, ai may be a filter coefficient (model parameter) (derived by LDL decomposition). C may not be a chrominance sample, but may represent a downsampled luminance sample L' corresponding to the chrominance sample location. P() is a non-linear term -> P(C) = (C*C + midVal) >> bitDepth -> P = (C*C + 512) >> 10 (for 10-bit), where C may be a downsampled luminance sample L' at the location corresponding to the chrominance signal location C. B (bias term) may be a scalar offset value between input and output (generally set to 512, the midpoint of the sample, for 10-bit). Y and X may be positions (coordinates) moved horizontally (X) and vertically (Y) from the top-left position of the block. Gy = (2N + NW + NE) - (2S + SW + SE) and Gx = (2W + NW + SW) - (2E + NE + SE).

[0930] In deriving a component-to-component prediction model based on the correlation between surrounding samples of a luminance block and surrounding samples of a chrominance block, the above MDC-CCCM uses multiple downsampling filters as shown in FIG. 19 and FIG. 20 to generate luminance samples corresponding to chrominance samples, defines the model as shown in FIG. 21 (Example 1), and filter coefficients α k It was made possible to calculate the value and the offset (or bias) b value. Additionally, the chrominance block to be currently predicted can be derived using the derived filter-based model and samples within the luminance block.

[0931] In performing the above downsampling, various filters can be used in addition to the filters of FIGS. 19 and FIGS. 20.

[0932] In defining the model above, as shown in FIG. 21, not only luminance pixels (L, L') but also gradient information between pixels and location information of luminance pixels within a luminance block can be used. (Examples 2 and 3 in FIG. 21)

[0933] The above-mentioned component-to-component prediction models can be defined differently depending on the downsampling filter information used to generate the downsampled luminance sample L'(i, j).

[0934] For example, depending on the type of filter used, the model can be defined and model coefficients derived by varying the number of model parameters and the model inputs (luminance samples, gradient, location).

[0935] The application status of the above-derived model, model information, and downsampling filter information used for model derivation may be pre-configured in the encoder and decoder or signaled from the encoder to the decoder.

[0936] At this time, when defining a single component-to-component prediction model, the downsampled (L'(i, j)) luminance samples required for model derivation can be generated using at least one downsampling filter.

[0937] For example, as shown in FIGS. 19 and 20, when generating luminance samples corresponding to each color difference sample, multiple luminance samples may be generated by applying multiple filters to a single sample location, or downsampled luminance samples corresponding to the color difference sample location may be generated by applying different filters depending on the sample location. At this time, a component-to-component prediction model can be defined and model parameters of the model can be derived using samples downsampled by at least one filter.

[0938] When constructing reference samples for model derivation, specific values ​​may be added to or subtracted from the reference samples, similar to Equation 2 or Equation 7 below.

[0939] For example, the offset (lumaOffset / chromaOffset) value of the luminance signal or the color difference signal can be added or subtracted.

[0940] In deriving a model and generating color difference prediction samples using the derived model, different formulas may be used in the derivation and application of the model.

[0941] For example, when deriving model coefficients, the model coefficients are derived using a formula to which an offset value is applied. When generating color difference prediction samples, model coefficients without an offset value applied may be used.

[0942] In deriving the model parameters of the above model, not only can they be derived for each of the U and V signals, but model coefficients applicable to both U and V signals can also be derived.

[0943] Whether the above model is applied can be determined for both U and V or for each, and this can be signaled from the encoder to the decoder.

[0944] In predicting color difference signals based on relationships between color components, a prediction signal can be generated based on a convolutional cross-component model (NS-CCCM: CCCM using non-downsampled luma samples) based on multiple downsampling filters.

[0945] In FIG. 19, the positions of the chrominance sample C(i, j) corresponding to the downsampled luminance sample L'(i, j) are shown. These positions may change depending on the chrominance signal sampling format.

[0946] While the above NS-CCCM derives a color difference prediction signal by modeling the correlation between the luminance sample (L') of the surrounding restored area downsampled by the previously defined CCCM / MDF-CCCM / GL-CCCM and the corresponding color difference sample (C), the NS-CCCM can define a component-to-component prediction model to predict a color difference sample as shown in the following mathematical formula using adjacent luminance samples corresponding to the color difference sample location to be predicted as in FIG. 19.

[0947] [Mathematical Formula 10]

[0948]

[0949] [Mathematical Equation 11] pre_C = α1·L (x, y) +α2·L (x-1, y) +α3·L (x+1, y) +α4·L (x, y-1) +α5·L (x, y+1) +α6·P+α7·B

[0950] [Mathematical Formula 12]

[0951]

[0952] The above mathematical formulas are identical to the model in CCCM that takes a luminance sample (L) as input, rather than a downsampled luminance sample (L').

[0953] Here, pre_C represents the predicted color difference sample (color difference prediction sample), Li represents the luminance sample around the color difference signal, ai represents the model coefficient (parameter), and β represents the offset (or bias).

[0954] In this case, P is a non-linear term and can be defined as follows.

[0955] P(C) = (LC*LC + midVal) >> bitDepth, P = (C*C + 512) >> 10 (for 10-bit), where LC represents the downsampled luminance sample L' at the position corresponding to the chrominance signal position C.

[0956] As shown in Fig. 20, MDF-CCCM uses a downsampling filter to create downsampled luminance samples (L'0 ~ L'5) corresponding to color difference sample C using luminance samples (L0~L1) around color difference sample C, and uses these to derive a model and predict color difference prediction samples. In contrast, NS-CCC, as shown in Fig. 19, directly uses luminance samples (L0~L1) around color difference sample C without filtering for downsampling to define a model and derive model parameters. Then, the color difference signal is predicted through the derived model.

[0957] At this time, the equation can be constructed using the “offsetLuma” and “offsetChroma” values ​​as in Equation 12.

[0958] Here, the offset value can be defined as the average value of the reconstructed samples surrounding the luminance block and the chrominance block.

[0959] At this time, at least one sample among the surrounding restored samples can be used to calculate the offset value.

[0960] In this case, for calculating the offset value, the surrounding reconstructed samples can be limited to samples within the reference region used to derive the model parameters.

[0961] In this case, offsetChroma can be calculated for U and V individually or together.

[0962] In defining the above model, not only luminance pixels (L, L') but also gradient information between pixels and location information of luminance pixels within a luminance block can be used.

[0963] In deriving the model parameters of the above model, not only can they be derived for each of the U and V signals, but model coefficients applicable to both U and V signals can also be derived.

[0964] Whether the above model is applied can be determined for both U and V or for each, and this can be signaled from the encoder to the decoder.

[0965] In predicting color difference signals based on relationships between color components, multiple pattern models can be used.

[0966] In this context, a pattern model may refer to a model having various forms. Multiple pattern models may refer to models that have the same number of model coefficients (or model parameters) but different shapes. The input sample location may differ for each pattern model. Here, the entire luminance area to which the pattern models are applied for deriving model coefficient values ​​or predicting color difference samples may be the same.

[0967] FIGS. 22a and FIGS. 22b are examples of multiple pattern models.

[0968] At this time, referring to FIGS. 22a and FIGS. 22b, a pattern model having five model parameters can be defined as follows:

[0969] Vertical Model: predChromaVal = α0·C + α1·N + α2·S +α3·P + α4·B

[0970] Horizontal Model: predChromaVal = α0·C + α1·E + α2·W +α3·P + α4·B

[0971] In the above equation, C, N, S, E, and W represent the locations of the luminance samples to which the pattern model is applied. The parameter values ​​of each pattern model can be derived through regression analysis within the same reference region.

[0972] Pattern models with different filter application locations may be used within the same reference area. In other words, models with different filter application locations may be determined to be different pattern models. Referring to FIGS. 18 and FIGS. 22a, a first pattern vertical model applied to sample locations N, C, and S and a second vertical model applied to sample locations NW, W, and SW may be defined as different pattern models.

[0973] Color difference signals can be derived using different models for the model derivation and model application, respectively. For example, a model with offsets applied to the model coefficients may be used in the model parameter derivation step, while a model with offsets excluded may be applied in the model application step.

[0974] In the color difference signal prediction using the plurality of pattern models described above, a prediction for the current color difference block is performed using the plurality of pattern models, the color difference block is restored using the color difference prediction block, and an optimal model can be determined from the plurality of pattern models based on the encoding cost for the restored color difference block. A plurality of optimal models can be determined. At this time, information on the optimal pattern model can be signaled from the encoder to the decoder. For example, a first pattern model applied to the location of a luminance restoration sample (i.e., a collocated luminance sample) corresponding to the location of the color difference prediction sample, and a second pattern model applied to the surrounding samples of the corresponding luminance restoration sample are determined as optimal models, and information regarding the optimal models can be signaled. Here, the surrounding samples of the luminance restoration sample may be samples located at one of the same location, the left location, or the top location of the luminance restoration sample. Information regarding the optimal models may include information regarding which location the second pattern model is applied to among the location, the left location, or the top location of the corresponding luminance restoration sample.

[0975] In the color difference signal prediction using the plurality of pattern models described above, color difference prediction samples are generated by applying the plurality of pattern models to luminance samples within a template area, a template matching cost is calculated based on the difference between the color difference prediction samples and the color difference reconstruction samples within the template area, and an optimal model among the plurality of pattern models can be determined based on the template matching cost. The model with the smallest template matching cost can be determined as the optimal model.

[0976] In color difference signal prediction based on color component-interrelationships, a new color difference prediction signal can be generated through component-interrelation between previously derived prediction blocks.

[0977] At this time, a luminance prediction block and a chrominance prediction block are derived using a prediction method other than the prediction between components, a prediction model between components is derived using the prediction between the two prediction blocks, and a new chrominance signal can be derived by applying the prediction model between components to the luminance prediction block or the luminance restoration block.

[0978] Figure 23 illustrates an example of chrominance signal prediction based on a component-to-component residual model in inter-frame prediction.

[0979] For example, as shown in FIG. 23, an inter-component prediction model is derived using the inter-component correlation between the luminance prediction signal and the color difference prediction signal derived by inter-frame prediction, an inter-component prediction signal is derived from the luminance restoration signal using the inter-component prediction model, and a final color difference prediction signal can be generated by weighting the inter-component prediction signal and the color difference prediction signal.

[0980] The aforementioned prediction method between components can be applied not only to the prediction block but also to the residual block.

[0981] For example, luminance residual blocks and chrominance residual blocks can be derived, and after deriving a component-to-component prediction model through component-to-component prediction between the two blocks, a new chrominance residual signal can be derived using the derived model.

[0982] At this time, for model derivation, an L-shaped reference area or a block shape of FIG. 11 may also be used as a reference area.

[0983] At this time, when configuring the block-shaped reference area of ​​FIG. 11, the luminance offset (offsetLuma) or chrominance offset (offsetChroma) value can be the average value of the four edge samples of the luminance block or chrominance block. For example, the average value of the top-left sample, top-right sample, bottom-left sample, and bottom-right sample can be used as the offset.

[0984] At this time, at least one sample among the prediction samples within the block can be used to calculate the offset value.

[0985] In this case, to calculate the offset value, the surrounding prediction samples can be limited to samples within the reference region used to derive the model parameters.

[0986] In this case, the offering can be obtained by U and V separately or together.

[0987] In this case, the reference area can be downsampled or subsampled for use.

[0988] For the derivation and application of the various models described above, motion vectors or block vectors are derived from the surroundings (adjacent / non-adjacent, juxtaposed blocks, blocks within a reference picture), a reference region is determined at a moved position using the motion vectors or block vectors, and a model can be derived using the reference region.

[0989] At this time, for model derivation, an L-shaped reference region may be used, or a block-shaped region may be used by considering a referenceable restored color difference sample region.

[0990] In the aforementioned model, the filter coefficient α kTo derive the value and offset (bias) b value, at least one of the following information may be used: adjacent surrounding samples of the color difference block currently to be predicted, adjacent surrounding samples of the luminance block corresponding to the position of the color difference block currently to be predicted, and the reference line index of the corresponding luminance block.

[0991] At this time, when determining the reference lines for deriving the filter coefficients, at least N lines may be used as reference lines. Here, N is a positive number greater than or equal to 0, which may be a preset value in the encoder / decoder, or may be a value signaled from the encoder to the decoder. It may be transmitted from the encoder.

[0992] At this time, among N luminance samples, at least one statistical value such as the average, maximum value, minimum value, median value of at least one sample can be calculated to derive and use one or more representative luminance samples, and filter coefficients can be derived from said representative samples.

[0993] For example, among the adjacent samples of the current color difference block, only those samples whose sample values ​​are greater than the average value of the samples can be selected and used to derive filter coefficients.

[0994] For example, among the adjacent samples of the current color difference block, only those samples whose sample values ​​are smaller than the average value of the samples can be selected and used to derive filter coefficients.

[0995] The luminance samples for deriving the above filter coefficients may vary, and the formula for deriving the color difference prediction samples may also vary. Information regarding the luminance samples or the formula for deriving adjacent samples of the current color difference block may be signaled in the encoder.

[0996] The above sample location may be a preset value in the encoder / decoder, or a value signaled from the encoder to the decoder. It may be transmitted from the encoder. Alternatively, it may be derived through encoding information of surrounding samples or blocks.

[0997] The reference area (or line) for deriving the above filter coefficients and offset can be selected not only from a referenceable area within the same picture but also from a reference area selected from different referenceable pictures.

[0998] In this case, sample values ​​within the reference area taken from different pictures can be scaled and used.

[0999] In the aforementioned inter-component prediction model, the shape of the filter can be determined according to the position (x, y) of the luminance sample to which the filter is applied and the number of filters applied, where the position of the sample to which the filter is applied and the number of filters may be values ​​pre-set in the encoder / decoder and may be values ​​signaled from the encoder to the decoder.

[1000] In the above-mentioned prediction of color difference signals based on inter-component prediction models, color difference signals / samples / blocks can be derived by using multiple inter-component prediction models.

[1001] In deriving the above multiple prediction models between the components, the multiple reference regions described in step S810 may be used.

[1002] In the prediction of color difference signals based on the plurality of models described above, among the plurality of models, the model information used for generating the color difference signal / block may be signaled or derived using surrounding encoding information. In this case, the model information includes mathematical formulas necessary to derive the color difference signal and coefficient and constant information included therein. Furthermore, the model information may further include not only coefficient and constant values ​​but also sign information of the model coefficients or sign information of the model constants. Furthermore, the model information may include information indicating one of the luminance sample corresponding to the color difference sample, the position to the left of the luminance sample, or the position to the top of the luminance sample as the sample location to which the model is applied.

[1003] In this case, let's assume that, using CCLM as an example, four models were derived as follows.

[1004] Model 1: C' (i, j) = a * L'(I, j) + b

[1005]

[1006] Model 2: C'' (i, j) = c * L'(I, j) + d

[1007] Model 3: C''' (i, j) = e * L'(I, j) + f

[1008] Model 4: C'''' (i, j) = g * L'(I, j) + h

[1009] (Model Information Signaling Method) In the generation of color difference signals / blocks based on multiple models, the model information used in the generation of the final color difference signal / block can be signaled.

[1010] FIGS. 24a and FIGS. 24b illustrate examples of a luminance block and a color difference prediction block that directly correspond to a color difference block.

[1011] Figure 25 illustrates an example of a color difference prediction block.

[1012] Specifically, FIG. 25 can 1) generate prediction samples (C', C'', C''', C'''') or color difference prediction blocks (BLK_C', BLK_C'', BLK_C'''', BLK_C'''') for the corresponding region using models, 2) select a color difference prediction block and model having the minimum error cost through color difference prediction, and 3) signal the corresponding model information from the encoder to the decoder.

[1013] At this time, color difference prediction blocks are generated using the above-mentioned multiple models as shown in FIG. 24a, FIG. 24b, and FIG. 25, and after performing color difference prediction, the error cost can be calculated.

[1014] At this point, the final model can be determined based on the error cost, and the corresponding model information can be signaled.

[1015] For example, the model with the smallest error cost can be selected as the final prediction model.

[1016] In this case, the error cost can be calculated by calculating U and V separately or by summing them.

[1017] As a method for signaling model information used to generate a final color difference signal / block, a table in which sets of model coefficients or sets of parameters are recorded is defined in an image encoding device and an image decoder, and an index pointing to a specific set can be signaled from the image encoding device to the image decoder. Code information of the coefficients can be transmitted through index signaling for a predefined code table, included in the table recording the sets of coefficients, or included in each set of coefficients.

[1018] (Method for deriving model information) In the generation of color difference signals / blocks based on multiple models, the final model used for generating the final color difference block can be derived by performing color difference prediction through model derivation and model application in a reference area.

[1019] For example, to derive the final model information mentioned above, a region for deriving the model in the reference region and a region for calculating the error cost by applying the derived model can be distinguished and used.

[1020] Figure 26 illustrates an example of a reference area used to derive the model and an area required to calculate the error cost between the color difference prediction signal and the color difference signal.

[1021] At this time, as shown in Fig. 26, the area used to derive the model among the reference areas is called R1, and the reference area R2 required to generate a color difference prediction signal by applying the derived model and to calculate the error cost with respect to the color difference signal is called R2.

[1022] At this time, one or more models are derived using the luminance samples in R1 and their corresponding chrominance samples, and each derived model is applied to the luminance samples in the R2 region to generate chrominance prediction samples (C', C'', C''', C'''') for each model. At this time, the chrominance prediction samples can also be combined to form a template. (PRED_TMP_C', PRED_TMP_C'', PRED_TMP_C''', PRED_TMP_C'''')

[1023] Figures 27a, 27b, and 27c illustrate an example of calculating the error cost between the color difference prediction template and the color difference template generated for each model.

[1024] At this time, as shown in FIGS. 27a, 27b, and 27c, the error cost between the color difference prediction samples (or color difference prediction templates) generated for each model and the restored color difference samples (or color difference templates) in the R2 region can be calculated, and the final prediction model can be derived based on the error cost.

[1025] For example, the model with the smallest error cost can be selected as the final prediction model.

[1026] In this case, the error cost can be calculated by calculating U and V separately or by summing them.

[1027] The size, location, and shape of the above R1 and R2 regions may be variable. In this case, the information may be a value pre-set in the encoder / decoder, or a value signaled from the encoder to the decoder.

[1028] The above R1 and R2 regions can be determined based on template matching. For example, as in TIMD, a reference region (R1, template) can be determined using a template to generate a prediction block (template), and the error cost can be calculated by comparing it with an adjacent template (R2) of the block.

[1029] In the above-mentioned multiple model-based color difference signal prediction, the amount of bits required for model information signaling can be reduced by using the model information derivation method.

[1030] For example, when there are multiple models, a prediction mode index (model index) can be defined in the form of a table for each model. In this case, the table can be named the 'Color Difference Prediction Mode Table'.

[1031] At this time, in the model information derivation method, color difference prediction samples (C', C'', C''', C'''') around the color difference block are created for each model, and the error cost between the generated color difference prediction samples and the reconstructed color difference samples (C) around the corresponding color difference block can be calculated.

[1032] At this time, the prediction mode index can be sorted and used based on the magnitude of the calculated error cost value.

[1033] Figure 28 illustrates an example of reconstructing a model-based color difference prediction mode table.

[1034] At this time, for example, if models with a small error cost are sorted to have upper index values ​​on the 'model-based color difference prediction mode table' as shown in FIG. 28, then color difference prediction blocks are generated only for the top N models, and then signaling the final determined model information after color difference prediction can save signaling bits.

[1035] At this time, when generating color difference prediction blocks and selecting models to be included in the model-based color difference prediction mode table, the decision can be made based on the statistical value of the error cost.

[1036] For example, a color difference prediction mode table and a color difference prediction block can be generated by selecting only the models with error costs less than or equal to the average error cost.

[1037] For example, S models with small variance values ​​can be selected to generate a color difference prediction mode table and a color difference prediction block.

[1038] Here, N and S are values ​​between 1 and 4 inclusive, which may be preset values ​​in the encoder / decoder, or may be values ​​signaled from the encoder to the decoder.

[1039] In the above-mentioned multiple model-based color difference signal prediction, the type and size of the prediction models between components may vary depending on the block size.

[1040] FIGS. 29a and FIGS. 29b illustrate examples of filter or filter set configurations according to block size.

[1041] Specifically, referring to FIGS. 29a and 29b, filter set 1 may be used when the width and height of the block are greater than 16, and filter set 2 may be used when the width or height of the block is less than or equal to 16. At least one filter may be selected and used from each filter set.

[1042] When multiple filters are selected and used from the above filter set, the information of the final determined filter can be signaled from the encoder to the decoder.

[1043] For example, as shown in FIGS. 29a and 29b, multiple model sets are defined according to block size, and at least one model can be selected from the model sets. Information about the selected model can be signaled from the encoder to the decoder.

[1044] In the above-mentioned multiple model-based color difference signal prediction, the model used can be determined based on the direction of the prediction mode.

[1045] FIGS. 30a, FIGS. 30b, FIGS. 30c, and FIGS. 30d illustrate examples of filter set configurations according to the direction of the in-screen prediction mode.

[1046] You can select and use at least one filter from each of the filter sets below.

[1047] When multiple filters are selected and used from the above filter set, the information of the finally determined filter can be signaled from the encoder to the decoder.

[1048] Figure 31 illustrates an example of varying the filter depending on the prediction direction of the prediction mode within the screen or the position of the samples used to create the prediction mode.

[1049] Depending on the prediction direction of the prediction mode within the screen or the position of the samples used to create the prediction mode, the filters used can be different as shown in Fig. 32.

[1050] For example, as shown in FIG. 30a, FIG. 30b, FIG. 30c, FIG. 30d and FIG. 31, multiple models are defined for each in-frame prediction mode, and a color difference signal can be derived using at least one model in the model set based on the in-frame prediction mode.

[1051] In one embodiment, the prediction model between components can be derived by using a method of combining or modifying previously derived models.

[1052] At this point, a new model can be derived by combining different models.

[1053] At this time, a new inter-component prediction model is derived by combining multiple models among the aforementioned inter-component prediction models, and a color difference prediction signal / sample / block can be generated using the derived inter-component prediction model.

[1054] For example, as shown in Equation 7 below, a cross-component linear model (CCLM) and a gradient linear model (GLM) can be combined to generate color difference prediction signals / samples / blocks.

[1055] Mathematical Equation 7: C' (i, j) = a1 * G(i, j) + a2 * L'(i, j) + b

[1056] Here, a1 and a2 are model coefficient values, b is the offset (or bias), L'(I, j) is the luminance sample corresponding to the color difference sample, and G(i, j) is the hardness sample corresponding to the color difference sample.

[1057] In deriving the above-mentioned inter-component prediction model for the target block, information on the inter-component prediction model to be used for the target block can be derived from the previously encoded / decoded block.

[1058] For example, the derived model can be saved for each block, and the model saved in the current block can be retrieved and used.

[1059] In this case, the model may be saved for every block, or it may be saved only when the color difference mode of the corresponding block is determined to be a component-to-component model-based color difference mode.

[1060] In determining the reference block for model derivation (refer to stored model coefficients) of the above current block, at least one block from adjacent or non-adjacent blocks may be selected and used.

[1061] In determining the reference block for the above model derivation (refer to the stored model coefficients), at least one block among the blocks in the reference picture may be selected and used.

[1062] At this time, pre-calculated or stored model information from surrounding blocks (pre-encoded / decoded) can be retrieved and used for color difference prediction for the target block.

[1063] In deriving a prediction model between components, a new model can be derived through the weighted sum of the coefficient values ​​of previously derived (calculated) models.

[1064] At this time, at least one of the coefficient values, biases, or non-linear terms constituting the new model can be derived through the weighted sum of the coefficient values, biases, or non-linear terms of the previously derived (calculated) models.

[1065] For example, let's assume that the previously derived model is as follows.

[1066] Model 1: C' (i, j) = a * L'(i, j) + b

[1067] Model 2: C' (i, j) = c * L'(i, j) + d

[1068] At this time, let's assume that the newly generated Model 3 using the above Model 1 and Model 2 is as follows.

[1069] Model 3: C' (i, j) = e * L'(i, j) + f

[1070] At this time, the coefficients e and bias f of Model 3 can be generated through the weighted sum of the model coefficients and biases of Models 1 and 2.

[1071] e = w*a + (1-w)*c, f = w*b + (1-w)*d (w: weight)

[1072] In this case, the weight w can be a preset value in the encoder / decoder, or a value signaled from the encoder to the decoder.

[1073] As a prediction model between components, a model derived from a block having a different size from the current block can be used.

[1074] For example, if different models are generated because the location of the reference region or reference sample varies depending on the block size, a model generated from a large-sized block can be used as a model for a small-sized block. Conversely, a model derived for a small-sized block can be used as a model for a large-sized block. For example, if the current block is a block partitioned from a higher-level block, an inter-component prediction model derived from the higher-level block can be used as an inter-component prediction model for the current block, or an inter-component prediction model for the current block can be used as an inter-component prediction model for the higher-level block or another block included within the higher-level block. As an inter-component prediction model, a model derived from a block located at a different position among the partitioned blocks can be used.

[1075] For example, if different models are generated because the location of the reference region or reference sample varies depending on the location of the sub-block within the partition block, the model derived from any sub-block within the partition block can be used in the current block.

[1076] At this time, when deriving a prediction model between components, a new model can be derived by modifying the coefficient values ​​of the previously derived model.

[1077] At this time, at least one of the coefficients, biases, or non-linear terms constituting the new model can be derived by modifying the coefficients, biases, or non-linear terms of the previously derived (calculated) models.

[1078] For example, when the model derived from the surrounding blocks is CCLM, a new model can be created by modifying the values ​​of a or b for the model coefficients a and the bias value b.

[1079] Here, when the values ​​obtained by transforming a and b are denoted as a' and b', a' and b' can be constructed by adding, subtracting, multiplying, or dividing a and b by a specific value v.

[1080] a' = a + v, a' = a - v, a' = a * v, a' = a / v

[1081] The coefficients v, t, s, etc. used to derive the above a', b' may be preset values ​​in the encoder / decoder, or values ​​signaled from the encoder to the decoder. For example, a table in which multiple offset values ​​are recorded may be defined, and an index pointing to a specific offset value may be signaled from the video encoding device to the video decoder.

[1082] In this case, when the range of a' and b' values ​​is denoted as W, X, Y, Z (W < a' < X, Y < b' < Z), W, X, Y, Z may be preset values ​​in the encoder / decoder, or may be values ​​signaled from the encoder to the decoder.

[1083] When a new model is derived by modifying the above model coefficients (or parameters), the difference in coefficient (or parameter) values ​​between the original model and the modified model can be signaled from the encoder to the decoder.

[1084] The above model coefficients can be corrected through template matching based on the modification method.

[1085] For example, the model coefficients can be corrected by applying the modified model to the template area to calculate the template matching cost, and then modifying the model coefficients to minimize the template matching cost.

[1086] At this time, the template matching cost can be calculated by varying the model coefficients by a fixed value range (e.g., ±5), and the model coefficients can be adjusted to have the smallest template matching cost.

[1087] At this time, a new model can be derived through the weighted sum of the coefficient values ​​of models derived from surrounding blocks.

[1088] At this time, after performing color difference signal prediction for the target (or current) block using at least one model derived from surrounding blocks, the optimal component-to-component prediction model among the models to be used for prediction can be determined. The model information used at this time (model index, parameters, block location from which the model was obtained, etc.) can be signaled from the encoder to the decoder.

[1089] In one embodiment, the inter-component prediction model can be derived using inter-component prediction merge.

[1090] When predicting between components for the current block, in addition to the model derived from the current block, there is a method to retrieve and use component-to-component prediction model information stored from pre-encoded / decoded surrounding blocks for color difference prediction for the current block, and to create new models by modifying the pre-encoded models and use them for color difference prediction for the current block.

[1091] At this time, to perform inter-component prediction merging, a chrominance signal prediction for the target (or current) block is performed using at least one of the above-derived models, and then the optimal inter-component prediction model among the models to be used for prediction can be determined. The model information used at this time (model index, parameters, block location from which the model was obtained, etc.) can be signaled from the encoder to the decoder.

[1092] Here, the candidate list may refer to a list used for sorting, adding, or removing candidates to construct a merge list or a final merge list for adding merge candidates.

[1093] The above component-to-component prediction merge can proceed as follows.

[1094] When performing inter-component prediction merging, a candidate list for storing inter-component prediction parameters derived from neighboring blocks can be constructed.

[1095] In this case, the size of the candidate list may vary depending on the encoding information of the current block (block size, whether it is split, reference block location, prediction mode information, etc.).

[1096] In this case, the candidate list may include not only model information but also encoding information of the block that derived the model.

[1097] In this case, the size of the corresponding candidate list is a pre-defined value in the encoder and decoder, or the value can be signaled from the encoder to the decoder.

[1098] Figure 32 illustrates classifying the list of candidates for component-to-component prediction models by group.

[1099] In constructing the above candidate list, the candidates can be classified based on the component-to-component prediction method of the candidates as shown in FIG. 32, and then each candidate list can be constructed.

[1100] For example, the candidate list can be constructed by separating into linear prediction methods (CCLM, GLM) and non-linear prediction methods (CCCM, GL-CCCM, MDF-CCCM, NS-CCCM).

[1101] For example, the candidate list may be configured differently depending on the method of configuring the reference area. In this case, the candidate list may be configured differently depending on whether downsampling or subsampling is performed on the reference area.

[1102] For example, a candidate list can be constructed for each component-to-component prediction mode.

[1103] In constructing the above candidate list, at least one candidate list can be constructed by combining multiple candidate lists.

[1104] In constructing the above candidate list, the candidate list can be formed by selecting only specific candidates from multiple groups. For example, a candidate list can be formed by selecting only candidates from the candidate lists of each group whose candidate index is greater than or equal to N (a defined number).

[1105] In constructing the above candidate list, one or more candidate lists can be separated to create multiple candidate lists.

[1106] The order of the STEP defined below is the same as the order in which candidates are added to the candidate list (starting with the block corresponding to STEP 1), and the order may vary depending on the predefined or current block's encoding information (block size, whether it is split, reference block location, prediction mode information, etc.). Additionally, this information may be signaled from the encoder to the decoder.

[1107] At this point, if the candidate list becomes full in the corresponding STEP, adding candidates can be stopped.

[1108] When adding candidates to the candidate list, similar or identical candidates may not be added.

[1109] For example, if at least one model parameter is the same, it may not be added.

[1110] For example, when using multiple models such as MMLM, if at least one model is identical, it may not be added to the candidate list.

[1111] When adding a candidate to the candidate list, you can determine the information to be inherited from surrounding blocks.

[1112] Here, only some model parameters of the component-to-component prediction model are taken, and the remaining parameters can be derived using information from the current block. For example, in the component-to-component prediction model of the previous example, the non-linear parameter (P) and offset B (or lumaOffset, chromaOffset) are not inherited and can be recalculated using surrounding samples of the current block.

[1113] STEP 1) Add candidates (model information and encoding information) stored in neighboring adjacent and non-adjacent blocks to the list

[1114] At this time, the neighboring adjacent / non-neighboring blocks from which model information and encoding information are to be retrieved may be any blocks among the neighboring referenceable blocks that have been encoded / decoded in a component-to-component prediction mode.

[1115] In this case, the surrounding adjacent / non-adjacent blocks from which model information and encoding information are to be obtained may be specific predefined locations or specific pattern shapes defined as crosses, radials, hexagons, diamonds, or triangles.

[1116] FIGS. 33a and FIGS. 33b illustrate surrounding adjacent / non-adjacent blocks from which model information and encoding information within a reference picture will be retrieved.

[1117] At this time, the surrounding adjacent / non-adjacent blocks from which model information and encoding information are to be retrieved may include blocks in the reference picture as well as blocks in the current picture, as shown in FIGS. 33a and 33b. In FIGS. 33a and 33b, the numbers may represent the order of samples or the order of blocks containing samples.

[1118] At this time, the block at the position moved using the motion vector or block vector in Fig. 15 may be included in the position for inducing the candidate.

[1119] In this case, the locations of neighboring adjacent / non-adjacent blocks from which model information and encoding information are obtained may vary depending on the encoding information of the current block (block size, partitioning status, reference block location, prediction mode information, etc.).

[1120] At this time, the order of adding neighboring adjacent and non-adjacent blocks to the candidate list may vary depending on predefined or current block encoding information (block size, partitioning status, reference block location, prediction mode information, etc.).

[1121] For example, items can be added to the candidate list first in order of proximity to the target block.

[1122] For example, you can add items to the candidate list starting from the inside in a spiral order centered on the target block.

[1123] For example, you can add surrounding blocks within the current picture to the candidate list first, and then add surrounding blocks within the reference picture later.

[1124] STEP 2) Add history-based candidates (model information and encoding information) to the list

[1125] At this time, a FIFO-type list such as HMVP can be defined and candidates stored in the list in the order of encoding / decoding, so that when performing component-to-component merge prediction in the target block, they can be added to the candidate list for component-to-component prediction merge.

[1126] STEP 3) Add modified candidates to the list

[1127] At this time, at least one of the models derived from the current block or models in the candidate list can be selected, modified, and added.

[1128] The above variation, using the previously mentioned method, for example, can create a new model by scaling model parameters or offsets and add it to the candidate list.

[1129] For example, in the case of CCLM, each model can be created by multiplying the CCLM model coefficients (or parameters) by the values ​​of [0, 1 / 8, -1 / 8, 2 / 8, -2 / 8, 3 / 8, -3 / 8, 4 / 8, -4 / 8, 5 / 8, -5 / 8, 6 / 8] and added to the candidate list.

[1130] At this time, at least one parameter in the derived model can be modified.

[1131] For example, in the case of CCCM, a new model can be created by modifying only a1 to a5 out of a total of 7 model coefficients (a1 to a7).

[1132] STEP 4) Sort candidates within the candidate list

[1133] In this case, the candidates within the candidate list can be sorted based on template cost.

[1134] In the above template cost-based alignment, each model in the candidate list is applied to samples surrounding the luminance block to generate color difference prediction samples corresponding to each model, and the error cost between these and the surrounding samples of the corresponding color difference block can be calculated and used as the template cost.

[1135] When sorting candidates within a candidate list, at least one candidate can be sorted.

[1136] For example, only candidates within index 5 of the candidate list can be sorted.

[1137] For example, only linear prediction series candidates can be sorted.

[1138] For example, you can sort only the candidates corresponding to the most component-to-prediction methods within the candidate list.

[1139] For example, only CCLM candidates can be sorted.

[1140] After sorting the candidates in the candidate list, N candidates can be removed from the candidate list. Similarly, only S candidates can be added to the candidate list based on the above template cost. In this case, N and S may represent integers greater than or equal to 0. Also, N and S may be pre-set values ​​in the encoder / decoder or signaled from the encoder to the decoder.

[1141] For example, candidates with a template cost above a certain threshold can be removed from the candidate list by comparing them with the candidate with the smallest template cost.

[1142] STEP 5) Perform color difference signal prediction using candidates in the candidate list

[1143] In performing the above color difference prediction, each model within the candidate list is applied to generate a color difference prediction block corresponding to each model, and the error cost between the derived color difference prediction blocks and the original color difference block can be calculated. Furthermore, the error cost between the color difference prediction blocks derived by the intra-chroma modes and the original color difference block can also be calculated.

[1144] At this time, the error cost in the prediction process after prediction can be the SAD or SATD value. Here, computational complexity can be reduced by omitting some of the transformation (primary / secondary transform), quantization, and entropy encoding processes when calculating the error cost.

[1145] Figures 34 and 35 illustrate an example of sorting candidates within a candidate list based on color difference prediction costs.

[1146] At this time, when sorting the candidates within the candidate list, they can be sorted based on the color difference prediction error cost as shown in FIGS. 34 and FIGS. 35.

[1147] When sorting the above candidate list, the process of adding and deleting candidates within the candidate list can be performed in the same way as the method in STEP 4.

[1148] STEP 6) Perform rate-distortion optimization-based chroma signal encoding on candidates in the candidate list

[1149] Among the candidates in the candidate list of STEP 5 above, M can be selected to perform rate-distortion optimization-based coding including transform, quantization, and entropy coding. Here, the complexity of the coding process can be reduced by selecting only arbitrary candidates to go through the STEP 6 process. For example, M candidates with low error costs are selected, and rate-distortion optimization can be performed on the M selected candidates. At this time, the M candidates can be selected from the candidates in the candidate list and from intra-chroma modes. Rate-distortion optimization for some candidates is performed in the encoder, and information regarding the finally selected candidates can be signaled from the encoder to the decoder.

[1150] Among the above M candidates, the index of the candidate with the minimum error cost can be signaled from the encoder to the decoder. At this time, since the original color difference signal cannot be used in the decoder, the index of the candidate to be signaled can be an index within the sorted list based on the template cost of STEP 4.

[1151] The decoder receives the index of the corresponding candidate and can perform color difference prediction based on component-to-component prediction merge.

[1152] In performing the above-mentioned component-to-component prediction merge, a new model can be created using information on candidates (component-to-component prediction models) within the candidate list and added to the candidate list.

[1153] Figure 36 illustrates the model and template matching costs according to the type of component-to-component prediction model within the candidate list.

[1154] FIG. 37 illustrates an example of creating a new model by interpolating / extending model parameters between at least two candidates in a candidate list.

[1155] In the above-mentioned creation of a new model based on candidate information within a candidate list, as shown in FIGS. 36 and 37, new model parameters can be derived by interpolating / extrapolating model parameters between at least two candidates within the candidate list, and thereby a new model can be created.

[1156] In generating a new model based on candidate information within the candidate list, a new model may be generated by weighting the information (parameters) of the models within the candidate list and then added to the candidate list. For the new model, at least two candidates within the candidate list sorted based on template matching costs may be used. In other words, a combination candidate may be generated by fusing at least two candidates among multiple inter-component prediction model candidates, and the combination candidate may be added to the candidate list. The combination candidate may include the inter-component prediction models of at least two candidates and weights used to blend the color difference blocks predicted using the inter-component prediction models of the candidates.

[1157] At this time, a new candidate can be generated by weighted summing the model calculated in the current block and the models in the candidate list.

[1158] In this case, weighted sums can be performed between identical component-to-component prediction methods.

[1159] In this case, the weighted sum of the model parameters can be applied not only among the candidates in the candidate list but also between the prediction model derived in the current block and the models in the candidate list.

[1160] At this time, at least one of the new candidates generated through the above weighted sum can be added to the candidate list.

[1161] The addition of the above candidates can be performed in the process from STEP 1 to STEP 6.

[1162] However, since the processes of STEP 5 / STEP 6 are not performed in the decoder, if a candidate is added during this process, a new flag and index within the candidate list must be added for that candidate.

[1163] For example, a flag such as “ccp_pairwise_model” can be used to signal to the decoder whether to use the corresponding model.

[1164] For example, after incrementing the index of the candidate list, the corresponding mode can be added at the end.

[1165] In the above candidate addition, if a new candidate is added during the STEP 4 candidate sorting process—that is, if the candidates in the candidate list are sorted based on template matching costs—and the template matching cost of the new candidate is greater than that of other candidates in the list, the candidate addition may be omitted. Additionally, when a new candidate is added, the indices of candidates with a template matching cost greater than that of the new candidate may increase by the number of newly added candidates. In this case, the size of the candidate list may be increased by the number of increased candidates, and candidates whose sorted candidate indices exceed the maximum size of the candidate list may be removed from the candidate list.

[1166] In the process of creating a new model through the above weighted sum, model creation can be restricted.

[1167] For example, in STEP 4, when selecting candidates to perform a weighted sum after sorting the candidates in the existing candidate list, candidates whose error cost is greater than a certain threshold may be excluded from the candidates to perform the weighted sum.

[1168] For example, if the difference in indices within the candidate list between candidates to be merged exceeds a threshold, the candidate with the larger index may be excluded from the candidates for performing the weighted sum.

[1169] For example, if the difference in error costs between candidates to be merged exceeds a threshold, the candidate with the larger index in the candidate list may be excluded from the candidates for performing the weighted sum.

[1170] In determining the weights for generating the new model above, the weights can be determined based on the error value.

[1171] For example, based on the template matching costs of the two candidates to perform a weighted sum, a higher weight can be assigned to the candidate with the lower matching cost.

[1172] For example, weights can be determined using regression analysis based on template matching costs. The weights can be determined based on the template matching cost, which represents the ...

Claims

In an image decoding method using a component-to-component prediction model, A step of deriving a prediction model between at least one component of the current block; and A step of predicting color difference samples of the current block from luminance samples of the current block using a prediction model between at least one component of the current block. A video decoding method including In paragraph 1, The step of deriving a prediction model between at least one component of the above-mentioned current block is, A step of generating a candidate list including multiple candidate prediction models between components; A step of adding a combination candidate generated by combining at least two candidates among the plurality of predicted model candidates between components to the candidate list; and A step of deriving prediction models between multiple components of the current block from candidates within the above candidate list. A video decoding method including In paragraph 2, The above combination candidate includes weights used to blend color difference blocks predicted using component-to-component prediction models of the above at least two candidates, and An image decoding method in which the above weights are determined based on a template matching cost representing the difference between a predicted template predicted using component-to-component prediction models of at least two candidates and a previously restored template. In paragraph 2, The above candidate list is, An image decoding method that is aligned based on the template matching costs of the above-mentioned multiple prediction model candidates. In paragraph 1, The step of deriving a prediction model between at least one component of the above-mentioned current block is, A step of decoding at least one of coefficient information or candidate information of a prediction model between at least one component of the current block from a bitstream; and A step of deriving a prediction model between at least one component of the current block based on at least one of coefficient information or candidate information of a prediction model between at least one component of the current block. A video decoding method including In paragraph 5, The coefficient information of the prediction model between at least one component of the above current block is, An image decoding method that includes index information pointing to one of a predefined set of coefficients. In paragraph 1, The step of deriving a prediction model between at least one component of the above-mentioned current block is, A step of deriving a first component prediction model applied to a luminance sample corresponding to each color difference sample; and A step of deriving a second component prediction model applied to surrounding samples of a luminance sample corresponding to each color difference sample. A video decoding method including In Paragraph 7, An image decoding method in which the above surrounding samples have a position determined based on information decoded from a bitstream. In paragraph 1, The step of deriving a prediction model between at least one component of the above-mentioned current block is, A step of deriving a plurality of reference regions of the current block based on predefined templates or sample value thresholds; and Step of deriving a component-to-component prediction model from each of the above plurality of reference regions A video decoding method including In paragraph 1, The step of deriving a prediction model between at least one component of the above-mentioned current block is, A step of inducing a reference area of ​​the current block; and Step of deriving a prediction model between components from the reference region of the current block above A video decoding method including In Paragraph 10, The step of deriving the reference area of ​​the current block above is, If a second reference vector is available within the region indicated by the first reference vector of the current block, the step of deriving the region indicated by the second reference vector as the reference region. A video decoding method including In Paragraph 10, The step of deriving the reference area of ​​the current block above is, A step of determining at least one downsampling filter based on color difference format information and spatial alignment information between the color difference samples and the luminance samples; and Step of applying the at least one downsampling filter to the reference area of ​​the current block. A video decoding method including In Paragraph 10, A step of adjusting the prediction model between the components based on sample statistical information of the reference area of ​​the current block. A video decoding method including further In an image encoding method using a component-to-component prediction model, A step of deriving a prediction model between at least one component of the current block; A step of predicting color difference samples of the current block from luminance samples of the current block using at least one component-to-component prediction model of the current block; and A step of encoding information about a prediction model between at least one component of the above current block. A video encoding method including A method for providing image data to an image decoding device, A step of generating a bitstream by encoding a target block using a prediction mode that utilizes a prediction model between components; and The method includes the step of transmitting the bitstream to the image decoder, and The step of generating the above bitstream is, A step of deriving a prediction model between at least one component of the current block; A step of predicting color difference samples of the current block from luminance samples of the current block using at least one component-to-component prediction model of the current block; and A step of encoding information about a prediction model between at least one component of the above current block. A method including