Image encoding / decoding method, apparatus, and recording medium

By employing an inter-component prediction method in image encoding/decoding to perform chroma block prediction using luminance block information, the reference area is expanded and the number of reference samples is reduced, thus solving the problem of low encoding/decoding efficiency in traditional methods and achieving efficient compression of intra-frame chroma prediction.

CN121970333APending Publication Date: 2026-05-01ELECTRONICS & TELECOMM RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ELECTRONICS & TELECOMM RES INST
Filing Date
2024-10-04
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Traditional image encoding/decoding methods have limitations in encoding/decoding efficiency, especially when using prediction methods in chroma signals, they fail to effectively utilize information in luminance signals, and the use of reference regions in intra-frame prediction is limited.

Method used

Inter-component prediction methods such as cross-component linear model (CCLM), gradient linear model (GLM), filter-based linear model (FLM), and convolutional cross-component model (CCCM) are employed to predict chroma blocks by utilizing information in the luminance block, thereby expanding the reference region and reducing the number of reference samples, thus improving coding efficiency.

Benefits of technology

By utilizing luminance block information to improve the signal transmission and coding efficiency of chrominance blocks, the number of bits required for prediction methods in chrominance blocks is reduced, achieving effective compression of intra-frame chrominance prediction and improving coding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121970333A_ABST
    Figure CN121970333A_ABST
Patent Text Reader

Abstract

An image encoding / decoding method, apparatus, and recording medium of the present disclosure may comprise the steps of: determining an inter-component prediction model of a current block; and generating a chroma prediction signal by using the determined inter-component prediction model.
Need to check novelty before this filing date? Find Prior Art

Description

Image encoding / decoding methods, devices and recording media Technical Field

[0001] This disclosure relates to image encoding / decoding methods and apparatus, and recording media for storing bit streams; more specifically, it relates to methods and apparatus for encoding / decoding images using efficient prediction methods in chroma signals, and recording media for storing bit streams.

[0002] This disclosure relates to image encoding / decoding methods and apparatus, and recording media for storing bit streams; more specifically, it relates to methods and apparatus for encoding / decoding images based on intra-frame prediction, and recording media for storing bit streams. Background Technology

[0003] Recently, the demand for high-resolution and high-quality images, such as HD (High Definition) and UHD (Ultra High Definition) images, has increased across various application areas. As image data becomes higher resolution and higher quality, the relative data volume increases compared to existing image data, thus increasing the costs associated with transmission and storage when the image data is transmitted using media such as existing wired and wireless broadband circuits or stored using existing storage media. This necessitates efficient image encoding / decoding technologies for images with higher resolution and image quality to address these issues arising from the increasing resolution and quality of image data.

[0004] Various techniques exist, such as inter-frame prediction techniques that use image compression technology to predict pixel values ​​included in the current frame based on previous or subsequent frames of the current frame, intra-frame prediction techniques that use pixel information in the current frame to predict pixel values ​​included in the current frame, transformation and quantization techniques for compressing residual signal energy, and entropy coding techniques that assign short symbols to values ​​with high occurrence frequency and long symbols to values ​​with low occurrence frequency. Image data can be effectively compressed, transmitted, or stored by using these image compression techniques. Summary of the Invention

[0005] Technical Issues: Traditional image encoding / decoding methods and apparatuses that use prediction methods from chroma signals have limitations in encoding / decoding because they cannot consider information in the luminance signal in a variety of ways. Furthermore, image encoding / decoding methods and apparatuses that use prediction methods from chroma signals may be limited in terms of effective encoding / decoding because they only use a limited amount of information from the chroma signal.

[0006] Because the use of reference regions in traditional image encoding / decoding is limited by intra-frame prediction, there are limitations in improving coding efficiency.

[0007] This disclosure provides encoding / decoding methods and apparatus by utilizing information in the luminance block, thereby improving the efficiency of signal transmission and encoding in the chrominance block.

[0008] This disclosure provides an image encoding / decoding method and apparatus for performing chroma prediction by using a cross-component linear model (CCLM), a gradient linear model (GLM), a filter-based linear model (FLM), a convolutional cross-component model (CCCM), and another inter-component prediction model derived from the convolutional cross-component model, as one of the types of inter-component prediction methods, in order to reduce the number of bits required for encoding / decoding in the prediction methods in the chroma block.

[0009] This disclosure provides a method and apparatus for performing at least one of expanding a reference region and reducing the number of reference samples in the reference region during intra-frame prediction, as well as a recording medium for storing a bit stream to improve coding efficiency.

[0010] The present disclosure may be an image encoding / decoding method and apparatus, used to perform the steps of determining a chroma signal prediction method, determining a reference region and representative value for deriving an inter-component prediction model, and using the inter-component prediction model to generate a chroma prediction signal when determining intra-frame prediction in a chroma block.

[0011] This disclosure may be an image encoding / decoding method and apparatus for performing at least one of expanding a reference region and reducing the number of reference samples in the reference region during intra-frame prediction, as well as a recording medium for storing bit streams.

[0012] Technical Effects: This disclosure improves the efficiency of signal transmission and encoding in chroma blocks by providing encoding / decoding methods and apparatus that utilize information in luminance blocks.

[0013] This disclosure can reduce the number of bits required for encoding / decoding in prediction methods in chroma blocks by providing an image encoding / decoding method and apparatus based on inter-component methods.

[0014] This disclosure provides an image encoding / decoding method and apparatus for efficient compression in chroma intra-frame prediction methods.

[0015] This disclosure can improve coding efficiency by providing a method and apparatus for performing at least one of expanding a reference region and reducing the number of reference samples in the reference region during intra-frame prediction, as well as a recording medium for storing a bitstream. Attached Figure Description

[0016] Figure 1 is a block diagram illustrating the configuration of an embodiment of the encoding device to which the present disclosure is applied.

[0017] Figure 2 is a block diagram illustrating the configuration of an embodiment of the decoding device to which the present disclosure is applied.

[0018] Figure 3 is a schematic diagram illustrating the partitioning structure of an image as it is encoded and decoded.

[0019] Figure 4 is a diagram showing the form of prediction units (PUs) that a coding unit (CU) can include.

[0020] Figure 5 is a diagram showing the form of a transformation unit (TU) that can be included in a CU.

[0021] Figure 6 shows the block division based on the example.

[0022] Figure 7 is a diagram illustrating an embodiment for explaining the intra-frame prediction process.

[0023] Figure 8 is a diagram showing the reference samples used in the intra-frame prediction process.

[0024] Figure 9 is a diagram illustrating an embodiment used to explain the inter-frame prediction process.

[0025] Figure 10 illustrates a spatial candidate according to an embodiment.

[0026] Figure 11 illustrates the order in which motion information of spatial candidates is added to the merging list according to an embodiment.

[0027] Figure 12 illustrates the transformation and quantization process based on the example.

[0028] Figure 13 shows a diagonal scan according to the example.

[0029] Figure 14 shows a horizontal scan based on an example.

[0030] Figure 15 shows a vertical scan based on the example.

[0031] Figure 16 is a configuration diagram of an encoding device according to an embodiment.

[0032] Figure 17 is a configuration diagram of a decoding device according to an embodiment.

[0033] Figure 18 illustrates a template configuration method for template matching based on intra / inter-frame prediction.

[0034] Figure 19 illustrates a template configuration method for two-sided matching.

[0035] Figure 20 illustrates a method for configuring a template in template matching and two-sided matching by clearing at least one pixel or at least one line.

[0036] Figure 21 illustrates a method for configuring a template by considering at least one of the prediction patterns of the current block or neighboring blocks.

[0037] Figure 22 shows a flowchart of an image encoding / decoding method for generating chroma prediction signals using the inter-component prediction model of the present invention.

[0038] Figure 23 shows examples of downsampling filter coefficients and DNN-based downsampling.

[0039] Figure 24 is an example of a sample point within a range of sample point values ​​between the minimum and maximum sample point values, divided into groups.

[0040] Figure 25 is an example of generating new groups by using the average of the samples calculated for each group.

[0041] Figure 26 shows an example of a chroma prediction block and a luminance block that directly corresponds to the chroma block.

[0042] Figure 27 shows an example of a chromaticity prediction block.

[0043] Figure 28 shows an example of the region in the reference area used to derive the model and the region required to calculate the error cost between the chromaticity prediction signal and the chromaticity signal.

[0044] Figure 29 shows an example of calculating the error cost between the chromaticity template and the chromaticity prediction template generated for each model.

[0045] Figure 30 shows an example of reconfiguring a model-based chromaticity prediction mode table.

[0046] Figure 31 shows an example of a gradient detection filter or gradient detection style.

[0047] Figure 32 shows the gradient block derived through four gradient detection filters.

[0048] Figure 33 shows the reference region for deriving the gradient block derived through four gradient detection filters and the reference region required for calculating the error cost.

[0049] Figure 34 shows an example of calculating the error cost of the chromaticity prediction template and the chromaticity template derived through four gradient-based linear models.

[0050] Figure 35 shows an example of a gradient detection filter (style).

[0051] Figures 36 and 37 show examples of reconfiguring a model-based chromaticity prediction mode table.

[0052] Figure 38 shows the form of applying luminance samples and filters corresponding to chrominance samples.

[0053] Figure 39 shows the form of the cross-shaped filter and the location of the sample points where the corresponding filter is applied.

[0054] Figure 40 shows various examples of filter forms and the locations of sample points where the filter is applied.

[0055] Figure 41 shows an example of a filter or filter bank configuration based on block size.

[0056] Figure 42 shows an example of filter bank configuration based on the direction of the intra-prediction mode.

[0057] Figure 43 shows an example of changing the filter based on the prediction direction of the intra-frame prediction mode or the location of the sample points used to create the prediction mode.

[0058] Figure 44 shows an example of downsampled luminance sample L'(i, j) and neighboring samples corresponding to the position of chrominance sample C(i, j).

[0059] Figure 45 shows an example of applying luminance samples and filters corresponding to chrominance samples.

[0060] Figure 46 shows an example of using multiple downsampling filters to generate luminance samples corresponding to chrominance samples.

[0061] Figure 47 shows the model used to generate the chromaticity prediction signal and the information used by the corresponding model.

[0062] Figure 48 shows the positions of the chromaticity samples corresponding to the downsampled luminance samples.

[0063] Figure 49 shows an example of changing the number of reference lines and the number of reference points according to the block partitioning format.

[0064] Figure 50 shows an example of changing the number of reference lines and reference samples based on the prediction patterns of neighboring blocks.

[0065] Figure 51 shows an example of changing the number of reference lines and the number of reference samples according to the intra-frame prediction mode.

[0066] Figure 52 shows an example of adjusting the reference region so that the location found by intra-frame template matching (intra-frame TMP) / IBC is included in the reference region.

[0067] Figure 53 shows an example of determining the reference region for deriving model parameters by referencing the block vector of the co-located block.

[0068] Figure 54 shows an example of deriving a block vector from a block at a specific location within a co-located block.

[0069] Figure 55 shows the neighboring reference area of ​​the current block.

[0070] Figure 56 shows the neighboring reference blocks at any location around the current block.

[0071] Figure 57 shows an example of combining reference regions based on the neighboring reference regions of the current block.

[0072] Figure 58 shows the location of non-adjacent blocks used to determine the reference area.

[0073] Figure 59 shows the size and form of the blocks actually encoded / decoded at non-adjacent block locations.

[0074] Figure 60 shows an example of the location of non-adjacent blocks.

[0075] Figure 61 shows an example of selecting the range of a reference block.

[0076] Figure 62 shows an example of a neighboring block in a non-adjacent block, located at a position where it is moved from a neighboring block within a reference image used to configure the reference area by using a block vector.

[0077] Figure 63 shows an example of the position of the block vector used when a co-occurring block exists.

[0078] Figure 64 shows an example of a Gaussian filter or low-pass filter used to modify reference or prediction samples.

[0079] Figure 65 shows an example of neighboring reconstructed samples used to filter predicted samples within a prediction block.

[0080] Figure 66 shows an example of blending chromaticity prediction blocks generated by using multiple models.

[0081] Figure 67 shows an example of applying filtering when mixing chromaticity prediction blocks generated by using multiple models.

[0082] Figure 68 shows an example of obtaining the final predicted chroma block by using multiple models.

[0083] Figure 69 shows an example of applying filtering when obtaining the final predicted chroma patch using multiple models.

[0084] Figure 70 shows the classification of candidate lists for prediction models between group pairs of components.

[0085] Figure 71 shows an example of neighboring / non-neighboring blocks that will bring model information and encoding information.

[0086] Figure 72 shows the adjacent / non-adjacent blocks that will bring model information and encoding information within the reference frame.

[0087] Figure 73 shows an example of arranging candidates in a candidate list based on template cost.

[0088] Figure 74 shows an example of arranging candidates in the candidate list based on chromaticity prediction cost.

[0089] Figure 75 shows the model and template matching costs in the candidate list based on the prediction model type between components.

[0090] Figure 76 illustrates an example of generating a new model by interpolating / extrapolating model parameters between at least two or more candidates in the candidate list.

[0091] Figure 77 shows an example of removing a candidate from the candidate list.

[0092] Figure 78 shows an example of a configuration list based on candidate derivation methods.

[0093] Figure 79 shows an example of a configuration list based on inter-component prediction methods.

[0094] Figure 80 shows an example of configuring a merge list (candidate list) based on a classification method based on candidate derivation.

[0095] Figure 81 shows an example of performing a weighted sum on chromaticity prediction signals / blocks / samples generated using multiple models.

[0096] Figure 82 shows an example of using different models for each part to obtain predicted chromaticity blocks.

[0097] Figure 83 shows an example of a predicted signal corrected using an inter-component prediction model between chromaticity signals (U / V).

[0098] Figure 84 shows an example of chromaticity signal prediction based on the inter-component residual model in inter-frame prediction.

[0099] Figure 85 shows an example of a linear model pattern between colors.

[0100] Figure 86 shows an example of deriving two linear models.

[0101] Figure 87 shows an example of a linear model before / after gradient adjustment.

[0102] Figure 88 shows an example of a 5-tap spatial filter.

[0103] Figure 89 shows an example of a reference region used to derive filter coefficients.

[0104] Figure 90 shows four gradient patterns.

[0105] Figure 91 shows the template region used to derive a template-based intra-prediction mode.

[0106] Figure 92 illustrates the process of deriving intra-prediction modes in the method for deriving decoder-side intra-prediction modes.

[0107] Figure 93 shows an example of a neighboring reconstruction sample point.

[0108] Figure 94 shows an example of candidates for spatial geometric partitioning patterns.

[0109] Figure 95 shows an example of a template region for a spatial geometric partitioning pattern.

[0110] Figure 96 shows an example of mixed width in a spatial geometry partitioning pattern.

[0111] Figure 97 shows the neighboring blocks used when deriving the MPM list. Detailed Implementation

[0112] This invention can be modified in various ways and can have various embodiments, which will be described in detail below with reference to the accompanying drawings. However, it should be understood that these embodiments are not intended to limit the invention to the specific forms disclosed, and they include all variations, equivalents, or modifications included within the spirit and scope of the invention.

[0113] The following exemplary embodiments will be described in detail with reference to the accompanying drawings, which illustrate specific embodiments. These embodiments are described to enable those skilled in the art to readily implement them. It should be noted that the various embodiments differ from one another but are not necessarily mutually exclusive. For example, specific shapes, structures, and characteristics described herein may be implemented as other embodiments without departing from the spirit and scope of other embodiments associated with one embodiment. Furthermore, it should be understood that the position or arrangement of various components in each disclosed embodiment can be changed without departing from the spirit and scope of the embodiments. Therefore, the appended detailed description is not intended to limit the scope of this disclosure, and the scope of the exemplary embodiments is defined only by the appended claims and their equivalents (provided they are properly described).

[0114] In the accompanying drawings, similar reference numerals are used to designate the same or similar functions in various respects. The shape, size, etc., of the components in the drawings may be exaggerated for clarity of description.

[0115] Terms such as “first” and “second” may be used to describe various components, but components are not limited by these terms. These terms are used only to distinguish one component from another. For example, without departing from the scope of this specification, a first component may be referred to as a second component. Similarly, a second component may be referred to as a first component. The term “and / or” may include a combination of multiple related descriptive terms or any one of multiple related descriptive terms.

[0116] It will be understood that when a component is referred to as "connected" or "combined" to another component, the two components may be directly connected or combined with each other, or there may be an intermediate component between the two components. On the other hand, it will be understood that when a component is referred to as "directly connected or combined," there is no intermediate component between the two components.

[0117] Furthermore, the components described in the embodiments are shown independently to indicate different functional characteristics, but this does not mean that each component is formed by a single piece of hardware or software. That is, for ease of description, multiple components are arranged and included separately. For example, at least two of the multiple components may be integrated into a single component. Conversely, a component may be divided into multiple components. As long as it does not depart from the spirit of this specification, embodiments in which multiple components are integrated or embodiments in which some components are separated are included within the scope of this specification.

[0118] The terminology used in the embodiments is for describing particular embodiments only and is not intended to limit the invention. Singular expressions include plural expressions unless the context specifically indicates the contrary. In the embodiments, it should be understood that terms such as “comprising” or “having” are intended only to indicate the presence of features, numbers, steps, operations, components, parts, or combinations thereof, and are not intended to exclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof. That is, in the embodiments, the expression describing a component as “comprising” a specific component means that additional components may be included within the practice or technical spirit of the invention, but do not exclude the presence of components other than the specific component stated therein.

[0119] In embodiments, the term "at least one" may mean one of one or more quantities (such as 1, 2, 3, and 4). In embodiments, the term "a plurality of" may mean one of two or more quantities (such as 2, 3, and 4).

[0120] Some components of the embodiments are not essential components for performing the necessary functions, but may be optional components used only to improve performance. Embodiments may be implemented using only the essential components necessary to achieve the essence of the embodiments. For example, a structure that includes only the essential components (excluding optional components used only to improve performance) is also included within the scope of the embodiments.

[0121] The embodiments will now be described in detail with reference to the accompanying drawings, enabling those skilled in the art to readily implement the embodiments. In the following description of the embodiments, detailed descriptions of well-known functions or configurations that are considered to obscure key points of this specification will be omitted. Furthermore, the same reference numerals are used throughout the drawings to designate the same components, and repeated descriptions of the same components will be omitted.

[0122] In the following text, "image" may refer to a single frame that constitutes a video, or it may refer to the video itself. For example, "encoding and / or decoding of an image" may mean "encoding and / or decoding of a video," and may also mean "encoding and / or decoding of any one of the multiple images that constitute a video."

[0123] In the following text, the terms “video” and “moving footage” may be used to have the same meaning and may be used interchangeably.

[0124] In the following text, the target image can be an encoded target image that is the target to be encoded and / or a decoded target image that is the target to be decoded. Furthermore, the target image can be an input image input to an encoding device or an input image input to a decoding device. Also, the target image can be the current image, i.e., the target currently to be encoded and / or decoded. For example, the terms "target image" and "current image" can be used to have the same meaning and can be used interchangeably.

[0125] In the following text, the terms “image,” “picture,” “frame,” and “screen” may be used to have the same meaning and may be used interchangeably with each other.

[0126] In the following text, a target block can be an encoding target block (i.e., the target to be encoded) and / or a decoding target block (i.e., the target to be decoded). Furthermore, a target block can be the current block, i.e., the target currently to be encoded and / or decoded. Here, the terms "target block" and "current block" can be used to have the same meaning and are interchangeable. A current block can represent an encoding target block that is the encoding target during encoding and / or a decoding target block that is the decoding target during decoding. Furthermore, a current block can be at least one of an encoding block, a prediction block, a residual block, and a transform block.

[0127] In the following text, the terms “block” and “unit” may be used to have the same meaning and may be used interchangeably. Optionally, “block” may refer to a specific unit.

[0128] In the following text, the terms “region” and “fragment” are used interchangeably.

[0129] In the following embodiments, specific information, data, flags, indexes, elements, and attributes may have their own values. The value "0" corresponding to each of the information, data, flags, indexes, elements, and attributes may indicate false, logical false, or a first predefined value. In other words, the values ​​"0," false, logical false, and the first predefined value are interchangeable. The value "1" corresponding to each of the information, data, flags, indexes, elements, and attributes may indicate true, logical true, or a second predefined value. In other words, the values ​​"1," true, logical true, and the second predefined value are interchangeable.

[0130] When variables such as i or j are used to indicate rows, columns, or indices, the value i can be an integer 0 or a greater than 0, or an integer 1 or a greater than 1. In other words, in the embodiments, each of the rows, columns, and indices can be counted starting from 0 or starting from 1.

[0131] In the embodiments, the term "one or more" or the term "at least one" may mean the term "multiple". The terms "one or more" or the term "at least one" may be used interchangeably with "multiple".

[0132] The terminology used in the embodiments will be described below.

[0133] Encoder: An encoder represents a device used to perform encoding. In other words, an encoder can represent an encoding device.

[0134] Decoder: A decoder refers to a device used to perform decoding. In other words, a decoder can represent a decoding device.

[0135] Unit: A unit can represent a component of image encoding and decoding. The terms "unit" and "block" can be used interchangeably and have the same meaning.

[0136] – A cell can be an M×N sample array. Each of M and N can be a positive integer. Cells can typically represent sample arrays in two-dimensional form.

[0137] In the process of image encoding and decoding, a "unit" can be a region generated by partitioning an image. In other words, a "unit" can be a specified region within an image. A single image can be partitioned into multiple units. Optionally, an image can be partitioned into sub-parts, and a unit can represent each sub-part created when encoding or decoding is performed on the partitioned sub-parts.

[0138] – During the encoding and decoding of an image, predefined processing can be performed on each unit according to its type.

[0139] Based on function, unit types can be classified as macrounits, coding units (CUs), prediction units (PUs), residual units, transform units (TUs), etc. Optionally, based on function, units can represent blocks, macroblocks, coding tree units, coding tree blocks, coding units, coding blocks, prediction units, prediction blocks, residual units, residual blocks, transform units, transform blocks, etc. For example, the target unit that serves as the object of encoding and / or decoding can be at least one of CUs, PUs, residual units, and TUs.

[0140] – The term “unit” can refer to a block of luma components, a block of chroma components corresponding to the luma components, and information about the syntax elements for each block, such that the unit is specified to be distinct from the block.

[0141] The size and shape of the unit can be implemented differently. In addition, the unit can have any of a variety of sizes and shapes. Specifically, the shape of the unit can include not only squares, but also geometric shapes that can be represented in two dimensions (2D) (such as rectangles, trapezoids, triangles and pentagons).

[0142] – In addition, cell information may include one or more of the following: cell type, cell size, cell depth, cell encoding order, and cell decoding order. For example, the cell type may indicate one of CU, PU, ​​residual cell, and TU.

[0143] – A cell can be divided into sub-cells, each sub-cell having a smaller size than the related cell.

[0144] Depth: Depth represents the degree to which a cell is partitioned. Furthermore, cell depth indicates the level at which a cell exists when represented by a tree structure.

[0145] – Cell partitioning information may include depth, which indicates the depth of the cell. Depth may indicate the number of times the cell is partitioned and / or the degree to which the cell is partitioned.

[0146] In a tree structure, the root node can be considered to have the smallest depth and the leaf nodes have the largest depth. The root node can be the highest (top) node. The leaf nodes can be the lowest nodes.

[0147] A single unit can be hierarchically partitioned into multiple sub-units, with each unit possessing depth information based on a tree structure. In other words, a unit and the sub-units generated by partitioning that unit can correspond to a node and its child nodes, respectively. Each partitioned sub-unit can have a unit depth. Since depth indicates the number of times a unit is partitioned and / or the degree to which a unit is partitioned, the partitioning information of a sub-unit can include information about the size of that sub-unit.

[0148] In a tree structure, the top node can correspond to the initial node before partitioning. The top node can be called the "root node". Furthermore, the root node can have a minimum depth value. Here, the depth of the top node can be level "0".

[0149] A node with a depth of level "1" can represent a cell generated when the initial cell is partitioned once. A node with a depth of level "2" can represent a cell generated when the initial cell is partitioned twice.

[0150] – A leaf node of depth “n” can represent a cell generated when the initial cell is partitioned n times.

[0151] – A leaf node can be a bottom node that cannot be further partitioned. The depth of a leaf node can be the maximum level. For example, a predefined value for the maximum level could be 3.

[0152] –QT depth can represent the depth for a four-partition drive. BT depth can represent the depth for a two-partition drive. TT depth can represent the depth for a three-partition drive.

[0153] Samples: Samples can be the basic units that make up a block. They can be 0 to 2 based on the bit depth (Bd). Bd The value of -1 is used to represent the sample point.

[0154] – A sample point can be a pixel or a pixel value.

[0155] – In the following text, the terms “pixel” and “sample” may be used to have the same meaning and may be used interchangeably.

[0156] Code Tree Unit (CTU): A CTU can consist of a single luma component (Y) code tree block and two chroma component (i.e., Cb, Cr) code tree blocks associated with the luma component code tree block. Furthermore, a CTU can represent information including the aforementioned blocks and the syntax elements used for each block.

[0157] – Each coding tree unit (CTU) can be partitioned using one or more partitioning methods (such as quadtree (QT), binary tree (BT), and ternary tree (TT)) to configure sub-units, such as coding units, prediction units, and transform units. A quadtree can represent a quaternion tree. Additionally, one or more partitioning methods can be used to partition each coding tree unit using multi-type tree (MTT).

[0158] – “CTU” can be used as a term to specify a pixel block as a processing unit in image decoding and encoding processes, such as in the case of partitioning an input image.

[0159] Coding Tree Block (CTB): "CTB" can be used as a term to specify any one of the Y coding tree block, Cb coding tree block, and Cr coding tree block.

[0160] Neighboring Blocks: Neighboring blocks (or adjacent blocks) can represent blocks that are adjacent to the target block. Neighboring blocks can also represent neighboring blocks that are being reconstructed.

[0161] – In the following text, the terms “neighboring block” and “adjacent block” may be used to have the same meaning and may be used interchangeably with each other.

[0162] – Neighboring blocks can represent reconstructed neighboring blocks.

[0163] Spatial neighbor block: A spatial neighbor block can be a block that is spatially adjacent to the target block. Neighbor blocks can include spatial neighbor blocks.

[0164] – Target blocks and spatially adjacent blocks can be included in the target frame.

[0165] – A spatially adjacent block can represent a block whose boundary is in contact with the target block or a block located within a predetermined distance from the target block.

[0166] – A spatial neighbor block can represent a block that is adjacent to the vertex of the target block. Here, a block that is adjacent to the vertex of the target block can be a block that is vertically adjacent to a neighbor block that is horizontally adjacent to the target block, or a block that is horizontally adjacent to a neighbor block that is vertically adjacent to the target block.

[0167] Temporally neighboring blocks: Temporally neighboring blocks can be blocks that are temporally adjacent to the target block. Neighboring blocks can include temporally neighboring blocks.

[0168] –Time-proximity blocks can include col blocks.

[0169] The col block can be a block in a previously reconstructed co-location frame (col frame). The position of the col block in the col frame can correspond to the position of the target block in the target frame. Optionally, the position of the col block in the col frame can be equal to the position of the target block in the target frame. The col frame can be a frame included in the list of reference frames.

[0170] - A temporally neighboring block can be a block that is spatially adjacent to the target block in time.

[0171] Prediction mode: Prediction mode can be information indicating the mode used for intra-frame prediction or the mode used for inter-frame prediction.

[0172] Prediction Unit: A prediction unit can be a basic unit used for prediction (such as inter-frame prediction, intra-frame prediction, inter-frame compensation, intra-frame compensation, and motion compensation).

[0173] A single prediction unit can be divided into multiple partitions or sub-prediction units with smaller sizes. These partitions can also be the basic units used in performing prediction or compensation. Partitions generated by dividing the prediction unit can also be prediction units.

[0174] Prediction cell partitioning: Prediction cell partitioning can be the shape in which prediction cells are divided.

[0175] Reconstructed neighboring cells: The reconstructed neighboring cells can be cells that are adjacent to the target cell and have already been decoded and reconstructed.

[0176] - The reconstructed neighboring units can be those that are spatially adjacent to the target unit or temporally adjacent to the target unit.

[0177] – The reconstructed spatial neighboring unit can be a unit included in the target image that has already been reconstructed through encoding and / or decoding.

[0178] – The reconstructed temporal neighbor unit can be a unit included in the reference image that has already been reconstructed through encoding and / or decoding. The position of the reconstructed temporal neighbor unit in the reference image can be the same as the position of the target unit in the target image, or it can correspond to the position of the target unit in the target image. Furthermore, the reconstructed temporal neighbor unit can be a block adjacent to a corresponding block in the reference image. Here, the position of the corresponding block in the reference image can correspond to the position of the target block in the target image. The fact that the positions of the blocks correspond to each other can indicate that the positions of the blocks are the same, that one block is included in another block, or that one block occupies a specific position in another block.

[0179] Sub-screen: A screen can be divided into one or more sub-screens. A sub-screen can consist of one or more parallel block rows and one or more parallel block columns.

[0180] – A sub-screen can be an area within a screen that has a square or rectangular shape (i.e., a non-square rectangle). Furthermore, a sub-screen can include one or more CTUs.

[0181] – A sub-picture can be a rectangular area of ​​one or more strips in the picture.

[0182] A sub-picture may include one or more parallel blocks, one or more bricks, and / or one or more stripes.

[0183] Parallel blocks: Parallel blocks can be areas in the image that have a square or rectangular shape (i.e., a non-square rectangle).

[0184] – A parallel block may include one or more CTUs.

[0185] – Parallel blocks can be partitioned into one or more blocks.

[0186] Blocking: Blocks can represent one or more CTU lines in a parallel block.

[0187] – A parallel block can be partitioned into one or more blocks. Each block may contain one or more CTU rows.

[0188] – Parallel blocks that are not partitioned into two parts can also be represented as blocks.

[0189] Strip: A strip may include one or more parallel blocks of a frame. Optionally, a strip may include one or more sub-blocks of a parallel block.

[0190] A sub-picture can contain one or more stripes that share a rectangular area covering the picture. Therefore, each sub-picture boundary is always a stripe boundary, and each vertical sub-picture boundary is always a vertical parallel block boundary.

[0191] Parameter set: The parameter set corresponds to the header information in the internal structure of the bitstream.

[0192] – The parameter set may include at least one of the following: Video Parameter Set (VPS), Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Adaptive Parameter Set (APS), Decoding Parameter Set (DPS).

[0193] Information transmitted via signals for each parameter set can be applied to the screen referencing the corresponding parameter set. For example, information in VPS can be applied to the screen referencing VPS. Information in SPS can be applied to the screen referencing SPS. Information in PPS can be applied to the screen referencing PPS.

[0194] Each parameter set can reference a higher-level parameter set. For example, PPS can reference SPS, and SPS can reference VPS.

[0195] Additionally, the parameter set may include parallel block groups, stripe header information, and parallel block header information. A parallel block group can be a group comprising multiple parallel blocks. Furthermore, the meaning of "parallel block group" can be the same as that of "strip".

[0196] Rate-distortion optimization: Coding devices can use rate-distortion optimization to provide high coding efficiency by utilizing a combination of the following: coding unit (CU) size, prediction mode, prediction unit (PU) size, motion information, and transform unit (TU) size.

[0197] – Rate distortion optimization schemes can calculate the rate distortion cost of each combination to select the optimal combination from these combinations. This can be done using the equation " To calculate the rate distortion cost, the combination that minimizes the rate distortion cost is typically chosen as the optimal combination under the rate distortion optimization scheme.

[0198] –D can represent distortion. D can be the average of the squares of the differences between the original transform coefficients and the reconstructed transform coefficients in the transform unit (i.e., mean square error).

[0199] –R can represent the rate, which can use relevant context information to represent the bit rate.

[0200] – R represents the Lagrange multiplier. It can include not only coding parameter information (such as prediction mode, motion information, and coding block flags), but also bits generated from encoding the transform coefficients.

[0201] The coding device can perform processes such as inter-frame prediction and / or intra-frame prediction, transform, quantization, entropy coding, inverse quantization (dequantization), and / or inverse transform to compute accurate D and R. These processes greatly increase the complexity of the coding device.

[0202] Bitstream: A bitstream can represent a stream of bits that encode image information.

[0203] Parsing: Parsing can be the determination of the value of a syntax element by performing entropy decoding on a bitstream. Alternatively, the term "parsing" can refer to this entropy decoding itself.

[0204] Symbols: A symbol can be at least one of the syntax elements, encoding parameters, and transform coefficients of the encoding target unit and / or the decoding target unit. Furthermore, a symbol can be the target of entropy encoding or the result of entropy decoding.

[0205] Reference frame: The reference frame can be an image referenced by the cell to perform inter-frame prediction or motion compensation. Optionally, the reference frame can be an image that includes reference cells referenced by the target cell to perform inter-frame prediction or motion compensation.

[0206] In the following text, the terms “reference screen” and “reference image” may be used to have the same meaning and may be used interchangeably.

[0207] Reference frame list: The reference frame list can be a list of one or more reference images that are used for inter-frame prediction or motion compensation.

[0208] – The types of reference screen lists can include combined list (LC), list 0 (L0), list 1 (L1), list 2 (L2), list 3 (L3), etc.

[0209] – For inter-frame prediction, one or more reference frame lists can be used.

[0210] Inter-frame prediction indicator: The inter-frame prediction indicator indicates the direction of inter-frame prediction for the target cell. Inter-frame prediction can be either unidirectional or bidirectional. Optionally, the inter-frame prediction indicator can represent the number of reference frames used to generate the prediction cell for the target cell. Optionally, the inter-frame prediction indicator can represent the number of prediction blocks used for inter-frame prediction or motion compensation of the target cell.

[0211] Prediction list utilization flags: Prediction list utilization flags indicate whether to use at least one reference screen from a specific reference screen list to generate prediction cells.

[0212] – The prediction list utilization flag can be used to derive the inter-frame prediction indicator. Conversely, the inter-frame prediction indicator can be used to derive the prediction list utilization flag. For example, a prediction list utilization flag indicating "0" (as a first value) indicates that for the target cell, reference frames from the reference frame list are not used to generate the prediction block. A prediction list utilization flag indicating "1" (as a second value) indicates that for the target cell, the reference frame list is used to generate the prediction cell.

[0213] Reference screen index: The reference screen index can be an index that indicates a specific reference screen in the list of reference screens.

[0214] Screen Order Count (POC): The POC value of a screen indicates the order in which the corresponding screens are displayed.

[0215] Motion Vector (MV): A motion vector can be a 2D vector used for inter-frame prediction or motion compensation. A motion vector represents the offset between the target image and a reference image.

[0216] – For example, MV can be represented in the form of (mvx, mvy). mvx indicates the horizontal component, and mvy indicates the vertical component.

[0217] Search Range: The search range can be a 2D region where a search for the MV is performed during inter-frame prediction. For example, the size of the search range can be M×N. M and N can both be positive integers.

[0218] Motion vector candidates: Motion vector candidates can be blocks that are used as prediction candidates when motion vectors are predicted, or motion vectors that are used as prediction candidates.

[0219] – Motion vector candidates can be included in the motion vector candidate list.

[0220] Motion vector candidate list: The motion vector candidate list can be a list using one or more motion vector candidate configurations.

[0221] Motion vector candidate index: The motion vector candidate index can be an indicator used to indicate motion vector candidates in the motion vector candidate list. Optionally, the motion vector candidate index can be an index of motion vector predictors.

[0222] Motion information: Motion information may include at least one of the following: a list of reference frames, a reference image, motion vector candidates, a motion vector candidate index, a merge candidate and a merge index, as well as information on motion vectors, reference frame indexes and inter-frame prediction indicators.

[0223] Merge candidate list: The merge candidate list can be a list that uses one or more merge candidate configurations.

[0224] Merging Candidates: Merging candidates can be spatial merging candidates, temporal merging candidates, combined merging candidates, combined dual-prediction merging candidates, history-based candidates, candidates based on the average of two candidates, zero merging candidates, etc. Merging candidates may include inter-frame prediction indicators and may include motion information such as prediction type information, reference frame index for each list, motion vectors, prediction list utilization flags, and inter-frame prediction indicators.

[0225] Merge index: A merge index can be an indicator used to indicate merge candidates in a merge candidate list.

[0226] - The merge index can indicate the reconstructed cell used to derive merge candidates among the reconstructed cells that are spatially adjacent to the target cell and temporally adjacent to the target cell.

[0227] – The merge index can indicate at least one of the multiple motion information candidates to be merged.

[0228] Transform Unit: A transform unit can be the basic unit for residual signal encoding and / or residual signal decoding (such as transform, inverse transform, quantization, dequantization, transform coefficient encoding, and transform coefficient decoding). A single transform unit can be partitioned into multiple sub-transform units with smaller sizes. Here, the transform may include one or more primary transforms and secondary transforms, and the inverse transform may include one or more primary inverse transforms and secondary inverse transforms.

[0229] Scaling: Scaling can represent the process of multiplying a factor by a transformation coefficient.

[0230] – As a result of scaling the levels of the transform coefficients, transform coefficients can be generated. Scaling can also be referred to as “inverse quantization”.

[0231] Quantization parameter (QP): The quantization parameter can be a value used to generate a transform coefficient level for the transform coefficients during quantization. Optionally, the quantization parameter can also be a value used to generate transform coefficient values ​​by scaling the transform coefficient level during dequantization. Optionally, the quantization parameter can be a value mapped to the quantization step size.

[0232] Delta quantization parameter: The Delta quantization parameter represents the difference between the quantization parameter of the target cell and the predicted quantization parameter.

[0233] Scan: A scan can refer to a method of arranging the coefficients in a cell, block, or matrix in order. For example, a method for arranging a 2D array in the form of a one-dimensional (1D) array can be called a "scan". Alternatively, a method for arranging a 1D array in the form of a 2D array can also be called a "scan" or "inverse scan".

[0234] Transform coefficients: Transform coefficients can be coefficient values ​​generated when the encoding device performs a transform. Optionally, transform coefficients can be coefficient values ​​generated when the decoding device performs at least one of entropy decoding and dequantization.

[0235] – The level of quantization or the level of quantized transform coefficients generated by applying quantization to transform coefficients or residual signals can also be included in the meaning of the term “transform coefficients”.

[0236] Quantization level: The quantization level can be a value generated when the encoding device performs quantization on the transform coefficients or residual signal. Optionally, the quantization level can be a value that serves as the target for dequantization when the decoding device performs dequantization.

[0237] – The transformation coefficient levels of quantization, as a result of transformation and quantization, can also be included in the meaning of the quantization levels.

[0238] Non-zero transform coefficients: Non-zero transform coefficients can be transform coefficients with values ​​other than 0, or transform coefficient levels with values ​​other than 0. Optionally, non-zero transform coefficients can be transform coefficients with a value amplitude that is not zero, or transform coefficient levels with a value amplitude that is not zero.

[0239] Quantization matrix: A quantization matrix is ​​a matrix used during the quantization or dequantization process to improve the subjective or objective image quality of an image. A quantization matrix can also be referred to as a "scaling list".

[0240] Quantization matrix coefficients: Quantization matrix coefficients can be each element in the quantization matrix. Quantization matrix coefficients are also referred to as "matrix coefficients".

[0241] Default matrix: The default matrix can be a quantization matrix predefined by the encoding and decoding devices.

[0242] Non-default matrix: A non-default matrix can be a quantization matrix that is not predefined by the encoding and decoding devices. A non-default matrix can represent a quantization matrix sent by the user from the encoding device to the decoding device.

[0243] Most Probable Mode (MPM): MPM can represent an intra-prediction mode that is highly likely to be used for intra-prediction of the target block.

[0244] – Encoding and decoding devices can determine one or more MPMs based on encoding parameters associated with the target block and attributes of entities associated with the target block.

[0245] – Encoding and decoding devices can determine one or more MPMs based on the intra-prediction modes of reference blocks. A reference block may include multiple reference blocks. These multiple reference blocks may include spatially neighboring blocks to the left of the target block and spatially neighboring blocks above the target block. In other words, one or more distinct MPMs can be determined based on which intra-prediction modes have been used for the reference blocks.

[0246] – One or more MPMs can be identified in the same way in both the encoding and decoding devices. That is, the encoding and decoding devices can share the same list of MPMs, which includes one or more MPMs.

[0247] MPM List: The MPM list can be a list that includes one or more MPMs. The number of one or more MPMs in the MPM list can be predefined.

[0248] MPM Indicator: The MPM indicator can indicate one or more MPMs in the MPM list that will be used for intra-prediction against the target block. For example, the MPM indicator can be an index used for the MPM list.

[0249] Since the MPM list is determined in the same way in both the encoding and decoding devices, it is not necessary to send the MPM list itself from the encoding device to the decoding device.

[0250] – The MPM indicator can be signaled from the encoding device to the decoding device. Because the MPM indicator is signaled, the decoding device can determine which MPM in the MPM list will be used for intra-frame prediction against the target block.

[0251] MPM Usage Indicator: The MPM usage indicator indicates whether an MPM usage mode will be used for prediction of the target block. The MPM usage mode can be determined using a list of MPMs to identify the MPMs that will be used for intra-frame prediction of the target block.

[0252] –MPM uses indicators that can be sent from the encoding device to the decoding device using signals.

[0253] Signaling: "Signaling" can indicate that information is sent from an encoding device to a decoding device. Optionally, "signaling" can indicate that information is included in a bitstream or recording medium by the encoding device. Information sent by the encoding device using signals can be used by the decoding device.

[0254] An encoding device generates encoded information by encoding the information to be transmitted as a signal. The encoded information is then sent from the encoding device to a decoding device. The decoding device obtains the information by decoding the transmitted encoded information. Here, the encoding can be entropy encoding, and the decoding can be entropy decoding.

[0255] Selective signaling: Information can be selectively transmitted using signals. Selective signaling for information can mean that an encoding device (based on specific conditions) selectively includes information in a bitstream or recording medium. Selective signaling for information can mean that a decoding device (based on specific conditions) selectively extracts information from a bitstream.

[0256] Omission of signaling: Signaling used for information can be omitted. Omission of signaling for information may mean that the encoding device (under certain conditions) does not include the information in the bitstream or recording medium. Omission of signaling for information may mean that the decoding device (under certain conditions) does not extract the information from the bitstream.

[0257] Statistical values: Variables, coding parameters, constants, etc., can have computable values. Statistical values ​​can be values ​​generated by performing calculations (operations) on the values ​​of a specified target. For example, a statistical value can indicate one or more of the following: the mean, weighted average, weighted sum, minimum, maximum, mode, median, and interpolation of the values ​​of a specific variable, a specific coding parameter, a specific constant, etc.

[0258] Figure 1 is a block diagram illustrating the configuration of an embodiment of the encoding device to which the present disclosure is applied.

[0259] Encoding device 100 can be an encoder, a video encoding device, or an image encoding device. Video may include one or more images (frames). Encoding device 100 can sequentially encode one or more images of the video.

[0260] Referring to Figure 1, the encoding device 100 includes an inter-frame prediction unit 110, an intra-frame prediction unit 120, a switcher 115, a subtractor 125, a transform unit 130, a quantization unit 140, an entropy coding unit 150, an inverse quantization unit 160, an inverse transform unit 170, an adder 175, a filter unit 180, and a reference frame buffer 190.

[0261] The encoding device 100 can perform encoding on the target image using intra-frame mode and / or inter-frame mode. In other words, the prediction mode of the target block can be one of intra-frame mode and inter-frame mode.

[0262] In the following text, the terms "intra-frame mode", "intra-frame prediction mode", "in-frame mode" and "in-frame prediction mode" may be used to have the same meaning and may be used interchangeably.

[0263] In the following text, the terms "inter-frame mode", "inter-frame prediction mode", "inter-picture mode" and "inter-picture prediction mode" may be used to have the same meaning and may be used interchangeably.

[0264] In the following text, the term "image" may refer to only a portion of an image or to a block. Furthermore, the processing of an "image" may refer to the sequential processing of multiple blocks.

[0265] Furthermore, the encoding device 100 can generate a bitstream including encoded information by encoding the target image, and can output and store the generated bitstream. The generated bitstream can be stored in a computer-readable storage medium and can be streamed via wired and / or wireless transmission media.

[0266] When intra-frame mode is used as prediction mode, switcher 115 can switch to intra-frame mode. When inter-frame mode is used as prediction mode, switcher 115 can switch to inter-frame mode.

[0267] The encoding device 100 can generate a prediction block for the target block. Furthermore, after the prediction block has been generated, the encoding device 100 can use the residual between the target block and the prediction block to encode the residual block for the target block.

[0268] When the prediction mode is intra-frame mode, the intra-frame prediction unit 120 can use pixels from previously encoded / decoded neighboring blocks adjacent to the target block as reference samples. The intra-frame prediction unit 120 can use the reference samples to perform spatial prediction on the target block and can generate prediction samples for the target block via spatial prediction. The prediction samples can represent samples in the prediction block.

[0269] The inter-frame prediction unit 110 may include a motion prediction unit and a motion compensation unit.

[0270] When the prediction mode is inter-frame mode, the motion prediction unit can search for the region in the reference image that best matches the target block during motion prediction, and can derive motion vectors for the target block and the found region based on the found region. Here, the motion prediction unit can use the search range as the target region for the search.

[0271] A reference image may be stored in a reference frame buffer 190. More specifically, when the encoding and / or decoding of a reference image has been processed, the encoded and / or decoded reference image may be stored in the reference frame buffer 190.

[0272] Since it stores the decoded screen, the reference screen buffer 190 can be a decoded screen buffer (DPB).

[0273] The motion compensation unit can generate a predicted block for the target block by performing motion compensation using motion vectors. Here, the motion vector can be a two-dimensional (2D) vector used for inter-frame prediction. Furthermore, the motion vector can indicate the offset between the target image and the reference image.

[0274] When the motion vector has values ​​other than integers, the motion prediction unit and motion compensation unit can generate prediction blocks by applying interpolation filters to a portion of the reference image. To perform inter-frame prediction or motion compensation, it can be determined which of the following modes—skip mode, merge mode, advanced motion vector prediction (AMVP) mode, and current frame reference mode—corresponds to the method for predicting and compensating for the motion of the PU included in the CU based on the CU, and inter-frame prediction or motion compensation can be performed according to that mode.

[0275] Subtractor 125 generates a residual block, which is the difference between the target block and the prediction block. The residual block can also be referred to as the "residual signal".

[0276] The residual signal can be the difference between the original signal and the predicted signal. Optionally, the residual signal can be a signal generated by transforming or quantizing the difference between the original signal and the predicted signal, or a signal generated by transforming and quantizing the difference. The residual block can be the residual signal for a block unit.

[0277] The transformation unit 130 can generate transformation coefficients by transforming the residual block, and can output the generated transformation coefficients. Here, the transformation coefficients can be coefficient values ​​generated by transforming the residual block.

[0278] Transformation unit 130 may use one of a number of predefined transformation methods when performing a transformation.

[0279] The predefined transformation methods may include Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Karhunen-Loeve Transform (KLT), etc.

[0280] The transformation method for transforming the residual block can be determined based on at least one of the coding parameters for the target block and / or neighboring blocks. For example, the transformation method can be determined based on at least one of the inter-frame prediction mode for the PU, the intra-frame prediction mode for the PU, the size of the TU, and the shape of the TU. Optionally, transformation information indicating the transformation method can be transmitted from the encoding device 100 to the decoding device 200 by a signal.

[0281] When using the transform skip mode, the transform unit 130 can omit the operation of transforming the residual block.

[0282] By quantizing the transform coefficients, a quantized transform coefficient level or a quantized level can be generated. In the following examples, each of the quantized transform coefficient level and the quantized level may also be referred to as a "transform coefficient".

[0283] Quantization unit 140 can generate quantized transform coefficient levels (i.e., quantized levels or quantized coefficients) by quantizing the transform coefficients according to quantization parameters. Quantization unit 140 can output the generated quantized transform coefficient levels. In this case, quantization unit 140 can use a quantization matrix to quantize the transform coefficients.

[0284] Entropy coding unit 150 can generate a bitstream by performing probability distribution-based entropy coding based on values ​​calculated by quantization unit 140 and / or coding parameter values ​​calculated during the encoding process. Entropy coding unit 150 can output the generated bitstream.

[0285] The entropy coding unit 150 can perform entropy coding on information about the pixels of the image and information required to decode the image. For example, the information required to decode the image may include syntax elements, etc.

[0286] When applying entropy coding, fewer bits can be allocated to more frequently occurring symbols, and more bits can be allocated to less frequently occurring symbols. Because symbols are represented through this allocation, the size of the bit string used to encode the target symbol can be reduced. Therefore, entropy coding can improve the compression performance of video coding.

[0287] Furthermore, for entropy coding, the entropy coding unit 150 can use coding methods such as Exponential Golomb, Context Adaptive Variable Length Coding (CAVLC), or Context Adaptive Binary Arithmetic Coding (CABAC). For example, the entropy coding unit 150 can use a variable-length code / code (VLC) table to perform entropy coding. For example, the entropy coding unit 150 can derive a binarization method for the target symbol. Furthermore, the entropy coding unit 150 can derive a probabilistic model for the target symbol / bit. The entropy coding unit 150 can use the derived binarization method, probabilistic model, and context model to perform arithmetic coding.

[0288] The entropy coding unit 150 can transform the coefficients in 2D block form into 1D vector form through the transform coefficient scanning method, so as to encode the quantized transform coefficient levels.

[0289] Encoding parameters can be information required for encoding and / or decoding. Encoding parameters may include information encoded by encoding device 100 and transmitted from encoding device 100 to decoding device, and may also include information that can be deduced during encoding or decoding. For example, the information transmitted to the decoding device may include syntax elements.

[0290] Encoding parameters can include not only information such as syntax elements (or flags or indexes) encoded by the encoding device and transmitted by the encoding device to the decoding device via signals, but also information derived during the encoding or decoding process. Furthermore, encoding parameters can include information required for encoding or decoding an image. For example, encoding parameters can include at least one value of the following, a combination of the following, or statistics of the following: cell / block size, cell / block shape / form, cell / block depth, cell / block partitioning information, cell / block partitioning structure, information indicating whether a cell / block is partitioned in a quadtree structure, information indicating whether a cell / block is partitioned in a binary tree structure, partitioning direction of the binary tree structure (horizontal or vertical), partitioning form of the binary tree structure (symmetric or asymmetric), information indicating whether a cell / block is partitioned in a ternary tree structure, partitioning direction of the ternary tree structure (horizontal or vertical), partitioning form of the ternary tree structure (symmetric or asymmetric), and partitioning direction of the ternary tree structure (horizontal or vertical). Information including: symmetric partitioning, whether the indicator unit / block is partitioned using a multi-type tree structure, the combination and direction of partitions in the multi-type tree structure (horizontal or vertical, etc.), partition form of the multi-type tree structure (symmetric or asymmetric partitioning, etc.), partition tree form of the multi-type tree (binary or ternary tree), prediction type (intra-frame prediction or inter-frame prediction), intra-frame prediction mode / direction, intra-frame luma prediction mode / direction, intra-frame chroma prediction mode / direction, intra-frame partition information, inter-frame partition information, coded block partition flag, prediction block partition flag, transform block partition flag, reference sample filtering method, reference sample filter taps, reference sample filter coefficients, and prediction block filtering method. Prediction block filter taps, prediction block filter coefficients, prediction block boundary filtering method, prediction block boundary filter taps, prediction block boundary filter coefficients, inter-frame prediction mode, motion information, motion vector, motion vector difference, reference frame index, inter-frame prediction direction, inter-frame prediction indicator, prediction list utilization flag, reference frame list, reference image, POC, motion vector prediction factor, motion vector prediction index, motion vector prediction candidate, motion vector candidate list, information indicating whether merging mode is used, merging index, merging candidate, merging candidate list, information indicating whether skip mode is used, interpolation filter type, interpolation filter taps, interpolation filter... The filter coefficients, magnitude of the motion vector, accuracy of the motion vector representation, transform type, transform magnitude, information indicating whether the first transform is used, information indicating whether the additional (second) transform is used, first transform selection information (or first transform index), second transform selection information (or second transform index), information indicating the presence or absence of the residual signal, code block style, code block flag, quantization parameters, residual quantization parameters, quantization matrix, information about the in-loop filter, information indicating whether the in-loop filter is applied, coefficients of the in-loop filter, taps of the in-loop filter, shape / form of the in-loop filter, information indicating whether the deblocking filter is applied.Deblocking filter coefficients, deblocking filter taps, deblocking filter strength, deblocking filter shape / form, information indicating whether adaptive sample offset is applied, adaptive sample offset value, adaptive sample offset category, adaptive sample offset type, information indicating whether adaptive loop filter is applied, adaptive loop filter coefficients, adaptive loop filter taps, adaptive loop filter shape / form, binarization / debinarization method, context model, context model determination method, context model update method, information indicating whether normal mode is executed, information indicating whether bypass mode is executed, valid coefficient flag, last valid coefficient flag, coefficient group encoding flag, last valid coefficient position, information indicating whether the coefficient value is greater than 1, information indicating whether the coefficient value is greater than 2, information indicating whether the coefficient value is greater than 3, remaining coefficient value information, positive / negative sign information, reconstructed luminance sample, reconstructed chrominance sample, context binary bits. The following parameters are included: bypass binary bits, residual luminance samples, residual chrominance samples, transform coefficients, luminance transform coefficients, chrominance transform coefficients, quantization level, luminance quantization level, chrominance quantization level, transform coefficient level, transform coefficient level scanning method, size of the motion vector search area on the decoding device side, shape / form of the motion vector search area on the decoding device side, number of motion vector searches on the decoding device side, CTU size, minimum block size, maximum block size, maximum block depth, minimum block depth, image display / output order, stripe identification information, stripe type, stripe partition information, parallel block group identification information, parallel block group type, parallel block group partition information, parallel block identification information, parallel block type, parallel block partition information, screen type, bit depth, input sample bit depth, reconstructed sample bit depth, residual sample bit depth, transform coefficient bit depth, quantization level bit depth, information about the luminance signal, information about the chrominance signal, color space of the target block, and color space of the residual block. Furthermore, information related to the above encoding parameters may also be included in the encoding parameters. Information used to calculate and / or derive the above encoding parameters may also be included in the encoding parameters. Information calculated or derived using the above encoding parameters can also be included in the encoding parameters.

[0291] The first transformation selection information can indicate the first transformation to be applied to the target block.

[0292] The second transformation selection information can indicate the second transformation to be applied to the target block.

[0293] The residual signal can represent the difference between the original signal and the predicted signal. Optionally, the residual signal can be a signal generated by transforming the difference between the original signal and the predicted signal. Optionally, the residual signal can be a signal generated by transforming and quantizing the difference between the original signal and the predicted signal. The residual block can be the residual signal for a block.

[0294] Here, sending information via a signal can indicate that the encoding device 100 includes entropy-encoded information generated by performing entropy encoding on flags or indices in the bitstream, and can also indicate that the decoding device 200 obtains information by performing entropy decoding on the entropy-encoded information extracted from the bitstream. This information may include flags, indices, etc.

[0295] A signal can refer to information that will be transmitted using a signal. In the following text, information used for images and blocks may be referred to as a "signal". Furthermore, in the following text, the terms "information" and "signal" may be used to have the same meaning and may be used interchangeably. For example, a specific signal may be a signal representing a specific block. A raw signal may be a signal representing a target block. A prediction signal may be a signal representing a predicted block. A residual signal may be a signal representing a residual block.

[0296] The bitstream may include information based on a specific syntax. Encoding device 100 may generate a bitstream that includes information according to the specific syntax. Decoding device 200 may obtain information from the bitstream according to the specific syntax.

[0297] Since the encoding device 100 performs encoding via inter-frame prediction, the encoded target image can be used as a reference image for other images to be processed subsequently. Therefore, the encoding device 100 can reconstruct or decode the encoded target image and store the reconstructed or decoded image as a reference image in the reference frame buffer 190. For decoding, inverse quantization and inverse transform of the encoded target image can be performed.

[0298] The quantization levels can be dequantized by the dequantization unit 160 and inversely transformed by the inverse transform unit 170. The dequantization unit 160 can generate dequantized coefficients by performing an inverse transform on the quantization levels. The inverse transform unit 170 can generate coefficients that have undergone both dequantization and inverse transform by performing an inverse transform on the dequantized coefficients.

[0299] The coefficients that have undergone dequantization and inverse transform can be added to the prediction block by adder 175. Adding the coefficients that have undergone dequantization and inverse transform to the prediction block generates a reconstructed block. Here, the coefficients that have undergone dequantization and / or inverse transform can represent one or more coefficients that have undergone dequantization and inverse transform, and can also represent the reconstructed residual block. Here, the reconstructed block can represent either the recovered block or the decoded block.

[0300] The reconstructed blocks can be filtered by filter unit 180. Filter unit 180 can apply one or more filters, including deblocking filter, sample adaptive offset (SAO) filter, adaptive loop filter (ALF), and nonlocal filter (NLF), to the reconstructed samples, reconstructed blocks, or reconstructed images. Filter unit 180 may also be referred to as an "in-loop filter".

[0301] Deblocking filters eliminate block distortion that occurs at the boundaries between blocks in a reconstructed image. To determine whether to apply a deblocking filter, the number of columns or rows of pixels included in the block and on which the determination of whether to apply the deblocking filter to the target block is based can be determined.

[0302] When a deblocking filter is applied to a target block, the applied filter can vary depending on the required deblocking strength. In other words, among different filters, one that takes into account the strength of the deblocking filter can be applied to the target block. When a deblocking filter is applied to a target block, one or more filters, such as long-tap filters, strong filters, weak filters, and Gaussian filters, can be applied to the target block according to the required deblocking strength.

[0303] Furthermore, when performing vertical and horizontal filtering on the target block, horizontal and vertical filtering can be performed in parallel.

[0304] SAO can add an appropriate offset to the pixel value to compensate for coding errors. SAO can perform correction on the image to which deblocking is applied based on pixels, where the correction uses an offset of the difference between the original image and the image to which deblocking is applied. To perform offset correction on an image, methods can be used to divide the pixels included in the image into a specific number of regions, determine the region to be offset within the divided regions, and apply the offset to the determined region, or methods can be used to apply the offset taking into account the edge information of each pixel.

[0305] The ALF can perform filtering based on values ​​obtained by comparing the reconstructed image with the original image. After the pixels included in the image have been divided into a predetermined number of groups, the filter to be applied to each group can be determined, and filtering can be performed differently for each group. Information related to whether an adaptive loop filter is applied can be sent to each CU via a signal. This information can be sent via a signal for the luminance signal. The shape and filter coefficients of the ALF to be applied to each block can be different for each block. Alternatively, an ALF with a fixed form can be applied to the block regardless of its characteristics.

[0306] Nonlocal filters can perform filtering based on reconstructed blocks similar to the target block. Regions similar to the target block can be selected from the reconstructed image, and statistical properties of the selected similar regions can be used to perform filtering on the target block. Information regarding whether a nonlocal filter is applied can be sent to the coding unit (CU) via a signal. Furthermore, the shape and filter coefficients of the nonlocal filter applied to the block can vary depending on the block.

[0307] The reconstructed blocks or reconstructed image filtered by filter unit 180 can be stored as a reference frame in reference frame buffer 190. The reconstructed blocks filtered by filter unit 180 can be part of the reference frame. In other words, the reference frame can be a reconstructed frame composed of reconstructed blocks filtered by filter unit 180. The stored reference frame can then be used for inter-frame prediction or motion compensation.

[0308] Figure 2 is a block diagram illustrating the configuration of an embodiment of the decoding device to which the present disclosure is applied.

[0309] Decoding device 200 can be a decoder, video decoding device, or image decoding device.

[0310] Referring to Figure 2, the decoding device 200 may include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an intra-frame prediction unit 240, an inter-frame prediction unit 250, a switcher 245, an adder 255, a filter unit 260, and a reference frame buffer 270.

[0311] Decoding device 200 can receive bit streams output from encoding device 100. Decoding device 200 can receive bit streams stored in computer-readable storage media and can also receive bit streams transmitted via wired / wireless transmission media.

[0312] The decoding device 200 can perform decoding on the bitstream in intra-frame mode and / or inter-frame mode. Furthermore, the decoding device 200 can generate a reconstructed image or a decoded image via decoding, and can output the reconstructed image or the decoded image.

[0313] For example, switcher 245 can be used to switch between intra-frame mode and inter-frame mode based on the prediction mode used for decoding. When the prediction mode used for decoding is intra-frame mode, switcher 245 can be operated to switch to intra-frame mode. When the prediction mode used for decoding is inter-frame mode, switcher 245 can be operated to switch to inter-frame mode.

[0314] The decoding device 200 can obtain the reconstructed residual block by decoding the input bitstream and can generate the prediction block. When the reconstructed residual block and the prediction block are obtained, the decoding device 200 can generate the reconstructed block, which is the target to be decoded, by adding the reconstructed residual block and the prediction block.

[0315] The entropy decoding unit 210 can generate symbols by performing entropy decoding on the bitstream based on the probability distribution of the bitstream. The generated symbols may include symbols in the form of quantized transform coefficient levels (i.e., quantized levels or quantized coefficients). Here, the entropy decoding method may be similar to the entropy coding method described above. That is, the entropy decoding method may be the inverse process of the entropy coding method described above.

[0316] The entropy decoding unit 210 can transform coefficients in one-dimensional (1D) vector form into 2D block shapes by a transform coefficient scanning method in order to decode the quantized transform coefficient levels.

[0317] For example, the coefficients of a block can be transformed into a 2D block shape by scanning the block coefficients using a top-right diagonal scan. Optionally, which of the top-right diagonal scan, vertical scan, and horizontal scan will be used can be determined based on the size of the corresponding block and / or the intra-frame prediction mode.

[0318] The quantized coefficients can be dequantized by the dequantization unit 220. The dequantization unit 220 generates dequantized coefficients by performing dequantization on the quantized coefficients. Furthermore, the dequantized coefficients can be inversely transformed by the inverse transform unit 230. The inverse transform unit 230 generates a reconstructed residual block by performing an inverse transform on the dequantized coefficients. As a result of performing dequantization and inverse transform on the quantized coefficients, a reconstructed residual block can be generated. Here, when generating the reconstructed residual block, the dequantization unit 220 can apply the quantization matrix to the quantized coefficients.

[0319] When using intra-frame mode, intra-frame prediction unit 240 can generate a prediction block by performing spatial prediction on the target block, wherein the spatial prediction uses the pixel values ​​of previously decoded neighboring blocks adjacent to the target block.

[0320] The inter-frame prediction unit 250 may include a motion compensation unit. Optionally, the inter-frame prediction unit 250 may be designated as a "motion compensation unit".

[0321] When using inter-frame mode, the motion compensation unit can generate a prediction block by performing motion compensation on the target block, wherein the motion compensation uses motion vectors and a reference image stored in the reference frame buffer 270.

[0322] The motion compensation unit can apply an interpolation filter to a portion of the reference image when the motion vector has values ​​other than integers, and can use the reference image with the interpolation filter applied to generate prediction blocks. To perform motion compensation, the motion compensation unit can determine, based on the CU, which of the following modes—skip mode, merge mode, advanced motion vector prediction (AMVP) mode, and current frame reference mode—corresponds to the motion compensation method used by the PU included in the CU, and can perform motion compensation according to the determined mode.

[0323] The reconstructed residual block and the predicted block can be added to each other by adder 255. Adder 255 generates a reconstructed block by adding the reconstructed residual block and the predicted block.

[0324] The reconstructed blocks can be filtered by filter unit 260. Filter unit 260 can apply at least one of a deblocking filter, a SAO filter, an ALF filter, and an NLF filter to the reconstructed blocks or the reconstructed image. The reconstructed image can be a picture that includes the reconstructed blocks.

[0325] The filter unit can output a reconstructed image.

[0326] The reconstructed image and / or reconstructed blocks filtered by filter unit 260 can be stored as a reference image in reference image buffer 270. The reconstructed blocks filtered by filter unit 260 can be part of the reference image. In other words, the reference image can be an image composed of reconstructed blocks filtered by filter unit 260. The stored reference image can then be used for inter-frame prediction or motion compensation.

[0327] Figure 3 is a schematic diagram illustrating the partitioning structure of an image as it is encoded and decoded.

[0328] Figure 3 schematically illustrates an example of a single cell being divided into multiple sub-cells.

[0329] To effectively partition an image, coding units (CUs) can be used in encoding and decoding. The term "unit" can be used to collectively specify 1) a block comprising image samples and 2) a syntax element. For example, "partition of a unit" can mean "partition of a block corresponding to a unit".

[0330] A CU can be used as the basic unit for image encoding / decoding. A CU can be used as the unit to which one of the intra-frame and inter-frame modes is applied during image encoding / decoding. In other words, during image encoding / decoding, it can be determined which of the intra-frame and inter-frame modes will be applied to each CU.

[0331] Furthermore, the CU can be the basic unit for predicting, transforming, quantizing, inverse transforming, dequantizing, and encoding / decoding transform coefficients.

[0332] Referring to Figure 3, image 300 can be sequentially partitioned into units corresponding to maximum coding units (LCUs), and the partitioning structure can be determined for each LCU. Here, LCU can be used to have the same meaning as coding tree unit (CTU).

[0333] Partitioning a cell can represent partitioning the block corresponding to the cell. Block partitioning information may include depth information about the depth of the cell. The depth information may indicate the number of times the cell is partitioned and / or the degree to which the cell is partitioned. A single cell may be hierarchically partitioned into multiple sub-cells, and the single cell may have depth information based on a tree structure.

[0334] Each partitioned sub-unit can have depth information. The depth information can be information indicating the size of the CU. Depth information can be stored for each CU.

[0335] Each CU can have depth information. When a CU is partitioned, the depth of the CU generated from the partition can be increased by 1 from the depth of the partitioned CU.

[0336] The partitioning structure represents the distribution of coding units (CUs) in the LCU 310 used for effective encoding of an image. This distribution can be determined by whether a single CU will be partitioned into multiple CUs. The number of CUs generated by partitioning can be a positive integer of 2 or greater, including 2, 3, 4, 8, 16, etc.

[0337] Depending on the number of CUs generated through partitioning, the horizontal and vertical dimensions of each CU generated through partitioning can be smaller than the horizontal and vertical dimensions of the CU before partitioning. For example, the horizontal and vertical dimensions of each CU generated through partitioning can be half the horizontal and vertical dimensions of the CU before partitioning.

[0338] Each partitioned CU can be recursively partitioned into four CUs in the same manner. Through recursive partitioning, at least one of the horizontal and vertical dimensions of each partitioned CU can be reduced compared to at least one of the horizontal and vertical dimensions of the CU before partitioning.

[0339] The partitioning of a CU can be performed recursively until a predefined depth or predefined size is reached.

[0340] For example, the depth of a CU can range from 0 to 3. The size of a CU can range from 64×64 to 8×8, depending on its depth.

[0341] For example, the depth of LCU 310 can be 0, and the depth of the minimum coding unit (SCU) can be a predefined maximum depth. Here, as mentioned above, the LCU can be a CU with the maximum coding unit size, and the SCU can be a CU with the minimum coding unit size.

[0342] Partitioning can begin at LCU 310, and the depth of the CU can be increased by 1 whenever the horizontal and / or vertical dimensions of the CU are reduced by partitioning.

[0343] For example, for each depth, an unpartitioned CU can have a size of 2N×2N. Furthermore, when CUs are partitioned, a CU of size 2N×2N can be partitioned into four CUs, each with a size of N×N. The value of N is halved each time the depth increases by 1.

[0344] Referring to Figure 3, an LCU with a depth of 0 can have 64×64 pixels or 64×64 blocks. 0 can be the minimum depth. An SCU with a depth of 3 can have 8×8 pixels or 8×8 blocks. 3 can be the maximum depth. Here, a CU with 64×64 blocks as an LCU can be represented by depth 0. A CU with 32×32 blocks can be represented by depth 1. A CU with 16×16 blocks can be represented by depth 2. A CU with 8×8 blocks as an SCU can be represented by depth 3.

[0345] Information about whether a corresponding CU is partitioned can be represented by the CU's partition information. Partition information can be 1 bit. All CUs except the SCU can include partition information. For example, the partition information value for an unpartitioned CU can be a first value. The partition information value for a partitioned CU can be a second value. When the partition information indicates whether a CU is partitioned, the first value can be "0" and the second value can be "1".

[0346] For example, when a single CU is partitioned into four CUs, the horizontal and vertical dimensions of each of the four CUs created through partitioning can be half the horizontal and vertical dimensions of the CU before partitioning. When a 32×32 CU is partitioned into four CUs, the size of each of the four partitioned CUs can be 16×16. When a single CU is partitioned into four CUs, the CU can be considered to have been partitioned using a quadtree structure. In other words, quadtree partitioning can be considered to have been applied to the CU.

[0347] For example, when a single CU is partitioned into two CUs, the horizontal or vertical dimension of each of the two resulting CUs can be half the horizontal or vertical dimension of the CU before partitioning. When a 32×32 CU is vertically partitioned into two CUs, the size of each of the two resulting CUs can be 16×32. When a 32×32 CU is horizontally partitioned into two CUs, the size of each of the two resulting CUs can be 32×16. When a single CU is partitioned into two CUs, the CU can be considered to have been partitioned using a binary tree structure. In other words, binary tree partitioning can be considered to have been applied to the CU.

[0348] For example, when a single CU is partitioned (or divided) into three CUs, the original CU before partitioning is partitioned such that its horizontal or vertical dimensions are divided in a 1:2:1 ratio, thus enabling the generation of three sub-CUs. For example, when a 16×32 CU is horizontally partitioned into three sub-CUs, the three sub-CUs generated by the partitioning can have dimensions of 16×8, 16×16, and 16×8 respectively from top to bottom. For example, when a 32×32 CU is vertically partitioned into three sub-CUs, the three sub-CUs generated by the partitioning can have dimensions of 8×32, 16×32, and 8×32 respectively from left to right. When a single CU is partitioned into three CUs, the CU can be considered to be partitioned in the form of a ternary tree. In other words, ternary tree partitioning can be considered to have been applied to the CU.

[0349] Both quadtree partitioning and binary tree partitioning are applied to LCU 310 in Figure 3.

[0350] In the encoding device 100, a 64×64 coding tree unit (CTU) can be partitioned into multiple smaller CUs using a recursive quadtree structure. A single CU can be partitioned into four CUs of the same size. Each CU can be recursively partitioned and can have a quadtree structure.

[0351] By using recursive partitioning of the CU, the optimal partitioning method that causes the minimum rate distortion cost can be selected.

[0352] The coding tree unit (CTU) 320 in Figure 3 is an example of a CTU in which quadtree partitioning, binary tree partitioning, and ternary tree partitioning are all applied.

[0353] As described above, to partition a CTU, at least one of quadtree partitioning, binary tree partitioning, and ternary tree partitioning can be applied to the CTU. Partitioning can be applied based on specific priorities.

[0354] For example, quadtree partitioning can be preferentially applied to CTUs. CUs that cannot be further partitioned in quadtree form can correspond to the leaf nodes of a quadtree. CUs corresponding to the leaf nodes of a quadtree can be the root nodes of a binary tree and / or a ternary tree. That is, CUs corresponding to the leaf nodes of a quadtree can be partitioned in binary or ternary tree form, or may not be further partitioned. In this case, to prevent each CU generated by applying binary or ternary tree partitioning to the CUs corresponding to the leaf nodes of a quadtree from being quadtree partitioned again, the operations of block partitioning and / or signaling block partitioning information are effectively performed.

[0355] Four-partition information can be used to signal the partitions of a CU corresponding to each node of a quadtree. A four-partition message with a first value (e.g., "1") indicates that the corresponding CU is partitioned in quadtree form. A four-partition message with a second value (e.g., "0") indicates that the corresponding CU is not partitioned in quadtree form. The four-partition message can be a flag with a specific length (e.g., 1 bit).

[0356] There may be no priority between binary tree partitioning and ternary tree partitioning. That is, the CU corresponding to the leaf node of a quadtree can be partitioned in either binary or ternary tree form. Furthermore, CUs generated by binary or ternary tree partitioning can be further partitioned in either binary or ternary tree form, or they may not be further partitioned.

[0357] A partition executed when there is no priority between a binary tree partition and a ternary tree partition can be called a "multi-type tree partition". That is, the CU corresponding to a leaf node of a quadtree can be the root node of a multi-type tree. The partitioning of the CU corresponding to each node of the multi-type tree can be signaled using at least one of the following: information indicating whether the CU is partitioned according to the multi-type tree, partitioning direction information, and partitioning tree information. For the partitioning of the CU corresponding to each node of the multi-type tree, the information indicating whether the multi-type tree partitioning is executed, the partitioning direction information, and the partitioning tree information can be signaled sequentially.

[0358] For example, information indicating whether a CU is partitioned in a multi-type tree and having a first value (e.g., "1") can indicate that the corresponding CU is partitioned in a multi-type tree format. Information indicating whether a CU is partitioned in a multi-type tree and having a second value (e.g., "0") can indicate that the corresponding CU is not partitioned in a multi-type tree format.

[0359] When the CU corresponding to each node of the multi-type tree is partitioned in the form of a multi-type tree, the corresponding CU may further include partitioning direction information.

[0360] Partition direction information indicates the partitioning direction of a multi-type tree partition. Partition direction information with a first value (e.g., "1") indicates that the corresponding CU is partitioned in the vertical direction. Partition direction information with a second value (e.g., "0") indicates that the corresponding CU is partitioned in the horizontal direction.

[0361] When a CU corresponding to each node of a multi-type tree is partitioned in the form of a multi-type tree, the corresponding CU may further include partition tree information. The partition tree information can indicate the tree used for multi-type tree partitioning.

[0362] For example, partition tree information with a first value (e.g., "1") can indicate that the corresponding CU is partitioned in a binary tree format. Partition tree information with a second value (e.g., "0") can indicate that the corresponding CU is partitioned in a ternary tree format.

[0363] Here, each of the above-mentioned information indicating whether partitioning by multi-type tree is performed, partition tree information, and partition direction information can be a flag with a specific length (e.g., 1 bit).

[0364] Entropy encoding and / or entropy decoding can be performed on at least one of the above four partition information, information indicating whether partitioning according to multiple tree types has been executed, partition direction information, and partition tree information. To perform entropy encoding / entropy decoding of this information, information from neighboring CUs adjacent to the target CU can be used.

[0365] For example, it can be assumed that the partitioning patterns (i.e., partitioned / non-partitioned, partitioned tree, and / or partitioned direction) of the left and / or upper CUs are highly similar to the partitioning patterns of the target CU. Therefore, based on the information of neighboring CUs, contextual information for entropy encoding and / or entropy decoding of the information for the target CU can be derived. Here, the information of neighboring CUs may include at least one of the following: 1) four-partition information of the neighboring CUs, 2) information indicating whether the neighboring CUs are partitioned according to multiple types of trees, 3) partitioned direction information of the neighboring CUs, and 4) partitioned tree information of the neighboring CUs.

[0366] In another embodiment of binary tree partitioning and ternary tree partitioning, binary tree partitioning can be performed first. That is, binary tree partitioning can be applied first, and then the CUs corresponding to the leaf nodes of the binary tree can be set as the root nodes of the ternary tree. In this case, quadtree partitioning or binary tree partitioning may not be performed on the CUs corresponding to the nodes of the ternary tree.

[0367] A CU that is not further partitioned by quadtree partitioning, binary tree partitioning, and / or ternary tree partitioning can be a unit for encoding, prediction, and / or transformation. That is, a CU may not be further partitioned for use in prediction and / or transformation. Therefore, the partitioning structure used to partition CUs into prediction units (PUs) and / or transformation units (TUs), its partitioning information, etc., may not exist in the bitstream.

[0368] However, when the size of a CU (Computer Unit) used as a partitioning unit is larger than the size of the largest transform block, the CU can be recursively partitioned until the size of the CU becomes smaller than or equal to the size of the largest transform block. For example, when the size of the CU is 64×64 and the size of the largest transform block is 32×32, the CU can be partitioned into four 32×32 blocks to perform the transform. Similarly, when the size of the CU is 32×64 and the size of the largest transform block is 32×32, the CU can be partitioned into two 32×32 blocks.

[0369] In this case, it is not necessary to separately send a signal indicating whether the CU has been partitioned for transformation. Without signal transmission, partitioning of the CU can be determined by comparing its horizontal (and / or vertical) dimensions with the horizontal (and / or vertical) dimensions of the largest transform block. For example, when the horizontal dimension of the CU is greater than the horizontal dimension of the largest transform block, the CU can be vertically bisected. Furthermore, when the vertical dimension of the CU is greater than the vertical dimension of the largest transform block, the CU can be horizontally bisected.

[0370] Information regarding the maximum and / or minimum size of the CU and the maximum and / or minimum size of the transform block can be signaled or determined at a level higher than the CU level. For example, higher levels could be sequence level, frame level, parallel block level, parallel block group level, or stripe level. For example, the minimum size of the CU could be set to 4×4. For example, the maximum size of the transform block could be set to 64×64. For example, the maximum size of the transform block could be set to 4×4.

[0371] Information regarding the minimum size of the CU corresponding to the leaf node of the quadtree (i.e., the minimum size of the quadtree) and / or the maximum depth of the path from the root node to the leaf node of the multi-type tree (i.e., the maximum depth of the multi-type tree) can be signaled or determined at a level higher than the CU level. For example, higher levels could be sequence level, frame level, stripe level, parallel block group level, or parallel block level. Information regarding the minimum size of the quadtree and / or the maximum depth of the multi-type tree can be signaled or determined individually at each of the intra-strip and inter-strip levels.

[0372] Information regarding the difference between the size of the CTU and the maximum size of the transform block can be signaled or determined at a level higher than the CU level. For example, higher levels could be sequence level, frame level, stripe level, parallel block group level, or parallel block level. Information regarding the maximum size of the CU corresponding to each node of the binary tree (i.e., the maximum size of the binary tree) can be determined based on the size of the CTU and the aforementioned difference. The maximum size of the CU corresponding to each node of the ternary tree (i.e., the maximum size of the ternary tree) can have different values ​​depending on the stripe type. For example, the maximum size of the ternary tree in an intra-strip level could be 32×32. For example, the maximum size of the ternary tree in an inter-strip level could be 128×128. For example, the minimum size of the CU corresponding to each node of the binary tree (i.e., the minimum size of the binary tree) and / or the minimum size of the CU corresponding to each node of the ternary tree (i.e., the minimum size of the ternary tree) can be set to the minimum size of the CU.

[0373] In another example, the maximum size of a binary tree and / or the maximum size of a ternary tree can be signaled or determined at the stripe level. Furthermore, the minimum size of a binary tree and / or the minimum size of a ternary tree can be signaled or determined at the stripe level.

[0374] Based on the various block sizes and depths described above, the four partition information, information indicating whether partitioning by multiple tree types has been performed, partition tree information, and / or partition direction information may or may not exist in the bitstream.

[0375] For example, when the size of the CU is not greater than the minimum size of the quadtree, the CU may not include the four-partition information, and the four-partition information of the CU can be inferred as the second value.

[0376] For example, when the size (horizontal and vertical dimensions) of the CU corresponding to each node of the multi-type tree is greater than the maximum size (horizontal and vertical dimensions) of the binary tree and / or the maximum size (horizontal and vertical dimensions) of the ternary tree, the CU may not be partitioned in binary and / or ternary tree form. In this determination method, information indicating whether partitioning by multi-type tree is performed may not be sent by signal, but can be inferred as a second value.

[0377] Optionally, the CU may not be partitioned in binary and / or ternary tree form when the size (horizontal and vertical dimensions) of the CU corresponding to each node of the multi-type tree is equal to the minimum size (horizontal and vertical dimensions) of the binary tree, or when the size (horizontal and vertical dimensions) of the CU is twice the minimum size (horizontal and vertical dimensions) of the ternary tree. With this determination method, information indicating whether partitioning by multi-type tree is performed can be sent without signaling, but can be inferred as a second value. This is because when the CU is partitioned in binary and / or ternary tree form, it generates CUs smaller than the minimum size of the binary tree and / or the minimum size of the ternary tree.

[0378] Optionally, binary or ternary partitioning can be limited based on the size of the virtual pipeline data unit (i.e., the size of the pipeline buffer). For example, binary or ternary partitioning can be limited when a CU is partitioned into sub-CUs that do not fit the size of the pipeline buffer. The size of the pipeline buffer can be equal to the maximum size of the transform block (e.g., 64×64).

[0379] For example, when the size of the pipeline buffer is 64×64, the following partitions can be restricted.

[0380] - A ternary tree partition for an N×M CU (where N and / or M are 128) - A horizontal binary tree partition for a 128×N CU (where N<=64) - A vertical binary tree partition for an N×128 CU (where N<=64) Optionally, when the depth of the CU corresponding to each node of the multi-type tree is equal to the maximum depth of the multi-type tree, the CU may not be partitioned in binary and / or ternary tree form. With this determination method, information indicating whether partitioning by multi-type tree is performed can be sent without signaling, but can be inferred as a second value.

[0381] Optionally, information indicating whether partitioning by the multi-type tree has been performed may be signaled only if at least one of vertical binary tree partitioning, horizontal binary tree partitioning, vertical ternary tree partitioning, and horizontal ternary tree partitioning is possible for each CU corresponding to each node of the multi-type tree. Otherwise, the CU may not be partitioned in binary and / or ternary tree form. With this determination method, the information indicating whether partitioning by the multi-type tree has been performed may not be signaled, but rather inferred as a second value.

[0382] Optionally, for each CU corresponding to a node in a multi-type tree, partitioning direction information may be signaled only when both vertical binary tree partitioning and horizontal binary tree partitioning are feasible, or only when both vertical ternary tree partitioning and horizontal ternary tree partitioning are feasible. Otherwise, partitioning direction information may not be signaled, but may be inferred as a value indicating the direction in which the CU can be partitioned.

[0383] Optionally, for each CU corresponding to a node of a multi-type tree, partition tree information may be signaled only when both vertical binary tree partitioning and vertical ternary tree partitioning are feasible, or only when both horizontal binary tree partitioning and horizontal ternary tree partitioning are feasible. Otherwise, partition tree information may not be signaled, but may be inferred as the value of the tree indicating the partitions that can be applied to the CU.

[0384] Figure 4 is a diagram showing the form of prediction units that a coding unit can include.

[0385] Within the control units (CUs) partitioned from the control unit (LCU), CUs that are no longer partitioned can be divided into one or more prediction units (PUs). This partitioning is also known as "partitioning".

[0386] A PU can be the basic unit used for prediction. A PU can be encoded and decoded in any of the skip mode, inter-frame mode, and intra-frame mode. A PU can be partitioned into various shapes according to each mode. For example, the target block described above with reference to FIG1 and the target block described above with reference to FIG2 can both be PUs.

[0387] A CU may not be classified as a PU. When a CU is not classified as a PU, the dimensions of the CU and the PU can be equal.

[0388] In skip mode, partitioning may not be present in the CU. Skip mode also supports a 2N×2N mode 410 without partitioning, where the PU and CU have the same size.

[0389] In inter-frame mode, eight types of partition shapes can exist in the CU. For example, in inter-frame mode, 2N×2N mode 410, 2N×N mode 415, N×2N mode 420, N×N mode 425, 2N×nU mode 430, 2N×nD mode 435, nL×2N mode 440 and nR×2N mode 445 are supported.

[0390] In intra-frame mode, 2N×2N mode 410 and N×N mode 425 are supported.

[0391] In 2N×2N mode 410, a PU with a size of 2N×2N can be encoded. A PU with a size of 2N×2N can represent a PU with the same size as the CU. For example, a PU with a size of 2N×2N can have a size of 64×64, 32×32, 16×16, or 8×8.

[0392] In N×N mode 425, PUs with an N×N size can be encoded.

[0393] For example, in intra-frame prediction, when the PU size is 8×8, the PUs from four partitions can be encoded. The size of the PU from each partition can be 4×4.

[0394] When encoding a PU in intra-frame mode, the PU can be encoded using any of a number of intra-frame prediction modes. For example, HEVC technology provides 35 intra-frame prediction modes, and the PU can be encoded in any of these 35 intra-frame prediction modes.

[0395] The rate-distortion cost can be used to determine which of the 2N×2N modes 410 and N×N modes 425 will be used to encode the PU.

[0396] Encoding device 100 can perform encoding operations on a PU of size 2N×2N. Here, the encoding operation can be an operation of encoding the PU in each of a plurality of intra-prediction modes that can be used by encoding device 100. Through the encoding operation, the optimal intra-prediction mode for the 2N×2N PU can be derived. This optimal intra-prediction mode can be the intra-prediction mode among the plurality of intra-prediction modes that can be used by encoding device 100 that exhibits the minimum rate-distortion cost when encoding the 2N×2N PU.

[0397] Furthermore, the encoding device 100 can sequentially perform encoding operations on each PU obtained by performing N×N partitioning. Here, the encoding operation can be an operation of encoding the PU in each of a plurality of intra-prediction modes that can be used by the encoding device 100. Through the encoding operation, the optimal intra-prediction mode for a PU of size N×N can be derived. This optimal intra-prediction mode can be the intra-prediction mode that results in the minimum rate-distortion cost when encoding a PU of size N×N among the plurality of intra-prediction modes that can be used by the encoding device 100.

[0398] The encoding device 100 can determine which of the PUs, one of size 2N×2N and one of size N×N, will be encoded based on a comparison between the rate-distortion cost of a PU of size 2N×2N and the rate-distortion cost of a PU of size N×N.

[0399] A single CU can be partitioned into one or more PUs, and a PU can be partitioned into multiple PUs.

[0400] For example, when a single PU is partitioned into four PUs, the horizontal and vertical dimensions of each of the four PUs created through partitioning can be half the horizontal and vertical dimensions of the original PU. When a 32×32 PU is partitioned into four PUs, the size of each of the four partitioned PUs can be 16×16. When a single PU is partitioned into four PUs, the PU can be considered to have been partitioned in a quadtree structure.

[0401] For example, when a single PU is partitioned into two PUs, the horizontal or vertical dimension of each of the two resulting PUs can be half the horizontal or vertical dimension of the original PU. When a 32×32 PU is vertically partitioned into two PUs, the size of each of the two partitioned PUs can be 16×32. When a 32×32 PU is horizontally partitioned into two PUs, the size of each of the two partitioned PUs can be 32×16. When a single PU is partitioned into two PUs, the PU can be considered to have been partitioned in a binary tree structure.

[0402] Figure 5 is a diagram showing the form of a transform unit that can be included in an encoding unit.

[0403] A transform unit (TU) can be a basic unit in a CU used for processes such as transform, quantization, inverse transform, dequantization, entropy coding, and entropy decoding.

[0404] The TU can be square or rectangular. The shape of the TU can be determined based on the size and / or shape of the CU.

[0405] Within the CUs partitioned from the LCU, CUs that are no longer partitioned into CUs can be divided into one or more TUs. Here, the partitioning structure of the TUs can be a quadtree structure. For example, as shown in Figure 5, a single CU 510 can be partitioned once or more according to a quadtree structure. Through this partitioning, a single CU 510 can be composed of TUs of various sizes.

[0406] The CU can be considered to be recursively partitioned when a single CU is partitioned two or more times. Through partitioning, a single CU can be composed of transformation units (TUs) of various sizes.

[0407] Optionally, a single CU can be divided into one or more TUs based on the number of vertical and / or horizontal lines dividing the CU.

[0408] The CU can be divided into symmetrical TUs or asymmetrical TUs. To divide into asymmetrical TUs, information about the size and / or shape of each TU can be transmitted from the encoding device 100 to the decoding device 200 via signals. Alternatively, the size and / or shape of each TU can be derived from the information about the size and / or shape of the CU.

[0409] A CU may not be classified as a TU. When a CU is not classified as a TU, the dimensions of the CU and the TU may be equal.

[0410] A single CU can be partitioned into one or more TUs, and a TU can be partitioned into multiple TUs.

[0411] For example, when a single TU is partitioned into four TUs, the horizontal and vertical dimensions of each of the four TUs generated by the partitioning can be half the horizontal and vertical dimensions of the original TU. When a TU of size 32×32 is partitioned into four TUs, the size of each of the four partitioned TUs can be 16×16. When a single TU is partitioned into four TUs, the TU can be considered to have been partitioned in a quadtree structure.

[0412] For example, when a single TU is partitioned into two TUs, the horizontal or vertical dimension of each of the two resulting TUs can be half the horizontal or vertical dimension of the original TU. When a TU of size 32×32 is vertically partitioned into two TUs, the size of each of the two partitioned TUs can be 16×32. When a TU of size 32×32 is horizontally partitioned into two TUs, the size of each of the two partitioned TUs can be 32×16. When a single TU is partitioned into two TUs, the TU can be considered to have been partitioned in a binary tree structure.

[0413] The CU can be divided in a different way than that shown in Figure 5.

[0414] For example, a single CU can be divided into three CUs. The horizontal or vertical dimensions of the three CUs generated by the division can be 1 / 4, 1 / 2, and 1 / 4 of the horizontal or vertical dimensions of the original CU before the division, respectively.

[0415] For example, when a 32×32 CU is vertically divided into three CUs, the resulting three CUs can have sizes of 8×32, 16×32, and 8×32, respectively. In this way, when a single CU is divided into three CUs, the CU can be considered to be divided in the form of a ternary tree.

[0416] One of the exemplary partitioning forms (i.e., quadtree partitioning, binary tree partitioning, and ternary tree partitioning) can be applied to the partitioning of the CU, and multiple partitioning schemes can be combined together for the partitioning of the CU. Here, the combination of multiple partitioning schemes is referred to as "composite tree partitioning".

[0417] Figure 6 shows the block division based on the example.

[0418] In video encoding and / or decoding processes, as shown in Figure 6, target blocks can be divided. For example, a target block can be a CU (Computer Unit).

[0419] For the partitioning of the target block, an indicator indicating the partitioning information can be sent from the encoding device 100 to the decoding device 200 by a signal. The partitioning information can be information indicating how the target block is partitioned.

[0420] The partitioning information can be one or more of the following: a partitioning flag (hereinafter referred to as "split_flag"), a quad-binary flag (hereinafter referred to as "QB_flag"), a quadtree flag (hereinafter referred to as "quadtree_flag"), a binary tree flag (hereinafter referred to as "binarytree_flag"), and a binary type flag (hereinafter referred to as "Btype_flag").

[0421] The "split_flag" can be a flag indicating whether a block has been split. For example, a split_flag value of 1 indicates that the corresponding block has been split, while a split_flag value of 0 indicates that the corresponding block has not been split.

[0422] The `QB_flag` can be a flag indicating whether the block is partitioned into a quadtree or a binary tree. For example, a `QB_flag` value of 0 indicates that the block is partitioned into a quadtree, while a `QB_flag` value of 1 indicates that the block is partitioned into a binary tree. Alternatively, a `QB_flag` value of 0 indicates that the block is partitioned into a binary tree, while a `QB_flag` value of 1 indicates that the block is partitioned into a quadtree.

[0423] The "quadtree_flag" can be a flag indicating whether the block is partitioned as a quadtree. For example, a quadtree_flag value of 1 indicates that the block is partitioned as a quadtree, while a quadtree_flag value of 0 indicates that the block is not partitioned as a quadtree.

[0424] The "binarytree_flag" can be a flag indicating whether the block is partitioned as a binary tree. For example, a binarytree_flag value of 1 indicates that the block is partitioned as a binary tree, while a binarytree_flag value of 0 indicates that the block is not partitioned as a binary tree.

[0425] The `Btype_flag` can be a flag indicating which of the vertical or horizontal partitions corresponds to the partitioning direction when a block is divided in a binary tree format. For example, a `Btype_flag` value of 0 indicates that the block is partitioned horizontally, and a `Btype_flag` value of 1 indicates that the block is partitioned vertically. Alternatively, a `Btype_flag` value of 0 indicates that the block is partitioned vertically, and a `Btype_flag` value of 1 indicates that the block is partitioned horizontally.

[0426] For example, the partitioning information of the blocks in Figure 6 can be derived by sending at least one of quadtree_flag, binarytree_flag, and Btype_flag, as shown in Table 1 below.

[0427] Table 1

[0428] For example, the block partitioning information in Figure 6 can be derived by sending at least one of split_flag, QB_flag, and Btype_flag, as shown in Table 2 below.

[0429] Table 2

[0430] The partitioning method may be limited to quadtrees or binary trees depending on the size and / or shape of the blocks. When this restriction is applied, the `split_flag` may be a flag indicating whether the blocks are partitioned in a quadtree or a binary tree format. The size and shape of the blocks can be deduced from the block depth information, and the depth information can be signaled from the encoding device 100 to the decoding device 200.

[0431] When the block size falls within a certain range, it is possible to partition in quadtree form only. For example, the certain range can be defined by at least one of the maximum block size and the minimum block size that can be partitioned in quadtree form only.

[0432] Information indicating the maximum and minimum block sizes that can be partitioned only in quadtree form can be transmitted from the encoding device 100 to the decoding device 200 via a bitstream. Furthermore, this information can be transmitted via a signal for at least one of the units such as video, sequence, picture, parameter, parallel block group, and strip (or segment).

[0433] Optionally, the maximum block size and / or minimum block size can be fixed sizes predefined by the encoding device 100 and the decoding device 200. For example, when the block size is greater than 64×64 and less than 256×256, it is possible to partition in quadtree form only. In this case, split_flag can be a flag indicating whether to perform partitioning in quadtree form.

[0434] When the size of a block is larger than the maximum size of a transform block, it is possible to partition it only in the form of a quadtree. Here, the sub-blocks generated by partitioning can be at least one of CU and TU.

[0435] In this case, split_flag can be a flag indicating whether the CU is partitioned in the form of a quadtree.

[0436] When the size of a block falls within a certain range, it is possible to partition it using only a binary tree or a ternary tree. For example, the specific range can be defined by at least one of the maximum block size and the minimum block size that allows partitioning using only a binary tree or a ternary tree.

[0437] Information indicating the maximum and / or minimum block size, which can be partitioned in a binary tree or a ternary tree manner, can be transmitted from the encoding device 100 to the decoding device 200 via a bitstream. Furthermore, this information can be transmitted via a signal for at least one of the units such as sequences, frames, and stripes (or segments).

[0438] Optionally, the maximum block size and / or minimum block size can be fixed sizes predefined by the encoding device 100 and the decoding device 200. For example, when the block size is greater than 8×8 and less than 16×16, it is possible to partition in binary tree form only. In this case, split_flag can be a flag indicating whether to perform partitioning in binary tree or ternary tree form.

[0439] The above description of partitioning in the form of a quadtree can be applied equally to binary and / or ternary tree forms.

[0440] The partitioning of a block may be limited by previous partitioning. For example, when a block is partitioned in a specific binary tree form and multiple sub-blocks are generated from said partition, each sub-block may be further partitioned only in that specific tree form. Here, the specific tree form can be at least one of a binary tree form, a ternary tree form, and a quadtree form.

[0441] When the horizontal or vertical dimensions of a partition block are dimensions that cannot be further subdivided, the aforementioned indicator does not need to be sent.

[0442] Figure 7 is a diagram illustrating an embodiment of intra-frame prediction processing.

[0443] The arrows extending radially from the center of the graph in Figure 7 indicate the prediction direction of the intra-prediction mode. Furthermore, the numbers appearing near the arrows indicate examples of mode values ​​assigned to the intra-prediction mode or its prediction direction.

[0444] In Figure 7, number 0 can represent the planar mode as a non-directional intra-frame prediction mode. Number 1 can represent the DC mode as a non-directional intra-frame prediction mode.

[0445] Intra-frame coding and / or decoding can be performed using reference samples from neighboring blocks of the target block. A neighboring block can be a reconstructed neighboring block. Reference samples can represent neighboring samples.

[0446] For example, intra-frame encoding and / or decoding can be performed using the values ​​of reference samples included in the reconstructed neighboring blocks or the encoding parameters of the reconstructed neighboring blocks.

[0447] Encoding device 100 and / or decoding device 200 can generate a prediction block by performing intra-frame prediction on the target block based on information about samples in the target image. When intra-frame prediction is performed, encoding device 100 and / or decoding device 200 can generate a prediction block for the target block by performing intra-frame prediction based on information about samples in the target image. When intra-frame prediction is performed, encoding device 100 and / or decoding device 200 can perform directional prediction and / or non-directional prediction based on at least one reconstructed reference sample.

[0448] A prediction block can be a block generated as a result of performing intra-frame prediction. A prediction block can correspond to at least one of CU, PU, ​​and TU.

[0449] The cells of the prediction block may have a size corresponding to at least one of CU, PU, ​​and TU. The prediction block may have a square shape with a size of 2N×2N or N×N. The size N×N may include sizes such as 4×4, 8×8, 16×16, 32×32, 64×64, etc.

[0450] Optionally, the prediction block can be a square block with a size of 2×2, 4×4, 8×8, 16×16, 32×32, 64×64, etc., or a rectangular block with a size of 2×8, 4×8, 2×16, 4×16, 8×16, etc.

[0451] Intra-prediction can be performed using intra-prediction modes for the target block. The number of intra-prediction modes that a target block can have can be a predefined fixed value, or it can be a value determined differently based on the attributes of the prediction block. For example, the attributes of the prediction block can include the size of the prediction block, the type of the prediction block, etc. In addition, the attributes of the prediction block can indicate the coding parameters used for the prediction block.

[0452] For example, the number of intra-prediction modes can be fixed at N, regardless of the size of the prediction block. Alternatively, the number of intra-prediction modes can be, for example, 3, 5, 9, 17, 34, 35, 36, 65, 67, or 95.

[0453] Intra-frame prediction mode can be either non-directional or directional.

[0454] For example, intra-frame prediction modes may include two non-directional modes and 65 directional modes corresponding to numbers 0 to 66 shown in Figure 7.

[0455] For example, when using a specific intra-prediction method, the intra-prediction modes may include two non-directional modes and 93 directional modes corresponding to numbers -14 to 80 shown in Figure 7.

[0456] The two non-directional modes may include DC mode and planar mode.

[0457] Directional patterns can be prediction patterns with a specific direction or angle. Directional patterns can also be called "angle patterns".

[0458] An intra-prediction mode can be represented by at least one of a mode number, a mode value, a mode angle, and a mode direction. In other words, the terms "(mode) number of intra-prediction mode", "(mode) value of intra-prediction mode", "(mode) angle of intra-prediction mode", and "(mode) direction of intra-prediction mode" can be used to have the same meaning and can be used interchangeably with each other.

[0459] The number of intra-prediction modes can be M. The value of M can be 1 or greater. In other words, the number of intra-prediction modes can be M, where M includes the number of non-directional modes and the number of directional modes.

[0460] The number of intra-prediction modes can be fixed at M, regardless of the block size and / or color components. For example, the number of intra-prediction modes can be fixed at either 35 or 67, regardless of the block size.

[0461] Optionally, the number of intra-frame prediction modes may vary depending on the shape, size, and / or type of color components of the block.

[0462] For example, in Figure 7, the direction prediction mode shown by the dashed line can be applied only to the prediction of non-square blocks.

[0463] For example, the larger the block size, the more intra-prediction modes there are. Alternatively, the larger the block size, the fewer intra-prediction modes there are. When the block size is 4×4 or 8×8, the number of intra-prediction modes can be 67. When the block size is 16×16, the number of intra-prediction modes can be 35. When the block size is 32×32, the number of intra-prediction modes can be 19. When the block size is 64×64, the number of intra-prediction modes can be 7.

[0464] For example, the number of intra-prediction modes can vary depending on whether the color component is a luma signal or a chrominance signal. Optionally, the number of intra-prediction modes corresponding to the luma component block can be greater than the number of intra-prediction modes corresponding to the chrominance component block.

[0465] For example, in the vertical mode with a mode value of 50, prediction can be performed along the vertical direction based on the pixel values ​​of the reference sample. Similarly, in the horizontal mode with a mode value of 18, prediction can be performed along the horizontal direction based on the pixel values ​​of the reference sample.

[0466] Even in directional modes other than those described above, the encoding device 100 and the decoding device 200 can still perform intra-frame prediction on the target unit using reference samples based on the angle corresponding to the directional mode.

[0467] Intra-prediction modes located to the right of the vertical mode can be called "vertical-right mode". Intra-prediction modes located below the horizontal mode can be called "horizontal-bottom mode". For example, in Figure 7, intra-prediction modes with mode values ​​of 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, and 66 can be vertical-right mode. Intra-prediction modes with mode values ​​of 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, and 17 can be horizontal-bottom mode.

[0468] Non-directional modes can include DC mode and planar mode. For example, the value for DC mode can be 1, and the value for planar mode can be 0.

[0469] Orientation modes can include angle modes. Among the various intra-frame prediction modes, all modes except DC mode and planar mode can be orientation modes.

[0470] When the intra-frame prediction mode is DC mode, a prediction block can be generated based on the average pixel values ​​of multiple reference pixels. For example, the pixel values ​​of the prediction block can be determined based on the average pixel values ​​of multiple reference pixels.

[0471] The number of intra-prediction modes and the mode values ​​of each intra-prediction mode described above are merely exemplary. The number of intra-prediction modes and the mode values ​​of each intra-prediction mode described above may be defined differently depending on the embodiment, implementation, and / or requirements.

[0472] To perform intra-frame prediction on a target block, a step can be performed to check whether samples included in the reconstructed neighboring blocks can be used as reference samples for the target block. When there are samples among the samples in the neighboring blocks that cannot be used as reference samples for the target block, a value generated by interpolation and / or duplication using at least one sample value from the samples included in the reconstructed neighboring blocks can replace the sample value of the sample that cannot be used as a reference sample. When a value generated by duplication and / or interpolation replaces the sample value of an existing sample, that sample can be used as a reference sample for the target block.

[0473] When using intra-frame prediction, filters can be applied to at least one of the reference samples and the prediction samples based on the size of the target block and at least one of the intra-frame prediction modes.

[0474] The type of filter to be applied to at least one of the reference sample and the prediction sample can vary based on at least one of the intra-prediction mode of the target block, the size of the target block, and the shape of the target block. The filter type can be classified based on one or more of the filter tap length, the value of the filter coefficients, and the filter strength. The filter tap length can represent the number of filter taps. Furthermore, the number of filter taps can represent the length of the filter.

[0475] When the intra-frame prediction mode is planar mode, the sample value of the predicted target block can be generated by weighting the top reference sample, left reference sample, upper right reference sample, and lower left reference sample of the target block according to the position of the predicted target sample in the prediction block.

[0476] When the intra-frame prediction mode is DC mode, the average of reference samples above and to the left of the target block can be used when generating the prediction block for the target block. Furthermore, filtering using the values ​​of the reference samples can be performed on specific rows or columns within the target block. The specific row can be one or more upper rows adjacent to the reference samples. The specific column can be one or more left columns adjacent to the reference samples.

[0477] When the intra-frame prediction mode is directional mode, the top reference sample, left reference sample, top right reference sample, and / or bottom left reference sample of the target block can be used to generate the prediction block.

[0478] To generate the above predicted samples, real-number-based interpolation can be performed.

[0479] The intra-prediction mode of the target block can be predicted from the intra-prediction modes of neighboring blocks adjacent to the target block, and the information used for prediction can be entropy encoded / entropy decoded.

[0480] For example, when the intra-prediction modes of the target block and neighboring blocks are the same, a predefined flag can be used to signal that the intra-prediction modes of the target block and neighboring blocks are the same.

[0481] For example, an indicator can be sent to indicate an intra-prediction mode that is the same as the intra-prediction mode of the target block among the intra-prediction modes of multiple neighboring blocks.

[0482] When the intra-prediction modes of the target block and neighboring blocks are different from each other, entropy coding and / or entropy decoding can be used to encode and / or decode information about the intra-prediction mode of the target block.

[0483] Figure 8 is a diagram showing the reference samples used in the intra-frame prediction process.

[0484] The reconstruction reference points used for intra-frame prediction of the target block may include the lower left reference point, the left reference point, the upper left corner reference point, the upper reference point, and the upper right reference point.

[0485] For example, a left reference sample may represent a reconstructed reference pixel adjacent to the left side of the target block. A top reference sample may represent a reconstructed reference pixel adjacent to the top of the target block. A top-left reference sample may represent a reconstructed reference pixel located at the top-left corner of the target block. A bottom-left reference sample may represent a reference sample located below the left-side sample line, which is the same line as the left-side sample line formed by the left reference samples. A top-right reference sample may represent a reference sample located to the right of the upper sample line, which is the same line as the upper sample line formed by the upper reference samples.

[0486] When the size of the target block is N×N, the number of the lower left reference point, the left reference point, the upper reference point, and the upper right reference point can all be N.

[0487] A prediction block can be generated by performing intra-frame prediction on the target block. The process of generating a prediction block may include determining the values ​​of the pixels in the prediction block. The target block and the prediction block can have the same size.

[0488] The reference sample used for intra-prediction of the target block can be changed according to the intra-prediction mode of the target block. The direction of the intra-prediction mode can represent the dependency between the reference sample and the pixels of the prediction block. For example, the value of a specified reference sample can be used as the value of one or more specified pixels in the prediction block. In this case, the specified reference sample and the one or more specified pixels in the prediction block can be samples and pixels located on a straight line along the direction of the intra-prediction mode. In other words, the value of the specified reference sample can be copied as the value of a pixel located in the opposite direction to the direction of the intra-prediction mode. Optionally, the value of a pixel in the prediction block can be the value of a reference sample located in the direction of the intra-prediction mode relative to the pixel's position.

[0489] In the example, when the intra-prediction mode of the target block is vertical, the upper reference sample can be used for intra-prediction. When the intra-prediction mode is vertical, the value of a pixel in the prediction block can be the value of a reference sample vertically above that pixel. Therefore, the upper reference sample adjacent to the top of the target block can be used for intra-prediction. Furthermore, the value of a pixel in a row of the prediction block can be the same as the value of a pixel at the upper reference sample.

[0490] In the example, when the intra-prediction mode of the target block is horizontal, the left reference sample can be used for intra-prediction. When the intra-prediction mode is horizontal, the value of a pixel in the prediction block can be the value of a reference sample horizontally to the left of that pixel. Therefore, the left reference sample adjacent to the left side of the target block can be used for intra-prediction. Furthermore, the value of a pixel in a column of the prediction block can be the same as the value of the pixel in the left reference sample.

[0491] In the example, when the mode value of the intra-prediction mode for the current block is 34, at least some of the left reference samples, the top-left reference sample, and the top reference sample can be used for intra-prediction. When the mode value of the intra-prediction mode is 34, the value of a pixel in the prediction block can be the value of a reference sample located diagonally at the top-left corner of that pixel.

[0492] Furthermore, in the case of an intra-prediction mode with mode values ​​ranging from 52 to 66, at least a portion of the upper right reference samples can be used for intra-prediction.

[0493] Furthermore, in the case of intra-prediction modes with mode values ​​ranging from 2 to 17, at least a portion of the lower left reference samples can be used for intra-prediction.

[0494] Furthermore, in the case of intra-prediction modes with mode values ​​ranging from 19 to 49, the upper left reference sample can be used for intra-prediction.

[0495] The number of reference samples used to determine the pixel value of a pixel in the prediction block can be 1, 2, or more.

[0496] As described above, the pixel value of a pixel in a prediction block can be determined based on the pixel's position and the position of a reference sample indicated by the direction of the intra-prediction mode. When both the pixel's position and the position of the reference sample indicated by the direction of the intra-prediction mode are integer positions, the value of a reference sample indicated by the integer position can be used to determine the pixel value of the pixel in the prediction block.

[0497] When the pixel position and the position of the reference sample indicated by the direction of the intra-prediction mode are not integer positions, an interpolated reference sample can be generated based on the two reference samples closest to that reference sample position. The value of the interpolated reference sample can be used to determine the pixel value of the pixel in the prediction block. In other words, when the pixel position in the prediction block and the position of the reference sample indicated by the direction of the intra-prediction mode indicate the position between two reference samples, an interpolation based on the values ​​of those two samples can be generated.

[0498] The predicted block generated by prediction may differ from the original target block. In other words, there may be prediction errors, which are the differences between the target block and the predicted block, and there may also be prediction errors between pixels in the target block and pixels in the predicted block.

[0499] In the following text, the terms “difference,” “error,” and “residual” are used to have the same meaning and are interchangeable.

[0500] For example, in the case of intra-frame prediction, the greater the distance between the pixels of the predicted block and the reference sample, the greater the potential prediction error. This prediction error can lead to discontinuities between the generated predicted block and its neighboring blocks.

[0501] To reduce prediction error, filtering operations can be used for prediction blocks. These filtering operations can be configured to adaptively apply filters to regions within the prediction block that are considered to have large prediction errors. For example, regions considered to have large prediction errors could be the boundaries of the prediction block. Furthermore, the regions within the prediction block considered to have large prediction errors can vary depending on the intra-prediction mode, and the characteristics of the filters can also vary depending on the intra-prediction mode.

[0502] As shown in Figure 8, for intra-frame prediction of the target block, at least one of reference lines 0 to 3 can be used.

[0503] Each reference line in Figure 8 indicates a reference point line that includes one or more reference points. When the reference line number is smaller, it indicates a reference point line that is closer to the target block.

[0504] Samples in fragments A and F can be obtained by padding instead of from reconstructed neighboring blocks, wherein the padding uses the samples from fragments B and E that are closest to the target block.

[0505] An index information indicating the reference sample lines to be used for intra-frame prediction of the target block can be transmitted using a signal. The index information can indicate which of a plurality of reference sample lines will be used for intra-frame prediction of the target block. For example, the index information can have a value corresponding to any one of 0 to 3.

[0506] When the upper boundary of the target block is the boundary of the CTU, only reference sample line 0 can be available. Therefore, in this case, index information does not need to be sent. When additional reference sample lines besides reference sample line 0 are used, filtering of the prediction block, which will be described later, is not required.

[0507] In the case of intra-frame prediction between colors, the predicted block of the target block of the second color component can be generated based on the corresponding reconstructed block of the first color component.

[0508] For example, the first color component can be the luminance component, and the second color component can be the chromaticity component.

[0509] To perform inter-color intra-frame prediction, the parameters of a linear model between the first and second color components can be derived based on a template.

[0510] The template may include a reference point above the target block (upper reference point) and / or a reference point to the left of the target block (left reference point), and may include the upper reference point and / or left reference point of the reconstructed block of the first color component corresponding to the reference point.

[0511] For example, the following values ​​can be used to derive the parameters of a linear model: 1) the value of the sample point of the first color component with the maximum value among the samples in the template, 2) the value of the sample point of the second color component corresponding to the sample point of the first color component, 3) the value of the sample point of the first color component with the minimum value among the samples in the template, and 4) the value of the sample point of the second color component corresponding to the sample point of the first color component.

[0512] When deriving the parameters of a linear model, the predicted block of the target block can be generated by applying the corresponding reconstructed block to the linear model.

[0513] Depending on the image format, subsampling can be performed on samples adjacent to the reconstructed block of the first color component and on the corresponding reconstructed block of the first color component. For example, when one sample of the second color component corresponds to four samples of the first color component, a corresponding sample can be calculated by subsampling the four samples of the first color component. When subsampling is performed, the parameters of the linear model and intra-frame prediction between colors can be performed based on the subsampled corresponding sample.

[0514] In intra-frame prediction mode, information about whether to perform inter-color intra-frame prediction and / or the range of templates can be sent by signaling.

[0515] The target block can be divided into two or four sub-blocks in the horizontal and / or vertical directions.

[0516] Sub-blocks generated by partitioning can be reconstructed sequentially. That is, when intra-prediction is performed on each sub-block, a sub-prediction block for that sub-block can be generated. Furthermore, when inverse quantization and / or inverse transform is performed on each sub-block, a sub-residual block for the corresponding sub-block can be generated. Reconstructed sub-blocks can be generated by adding the sub-prediction blocks to the sub-residual blocks. The reconstructed sub-blocks can be used as reference samples for intra-prediction of sub-blocks with the next higher priority.

[0517] A sub-block can be a block containing a specific number (e.g., 16) or more samples. For example, when the target block is an 8×4 block or a 4×8 block, the target block can be divided into two sub-blocks. Furthermore, when the target block is a 4×4 block, it cannot be divided into sub-blocks. When the target block has another size, it can be divided into four sub-blocks.

[0518] Signals can be used to send information about whether to perform intra-frame prediction based on these sub-blocks and / or about the partitioning direction (horizontal or vertical).

[0519] This sub-block-based intra-prediction can be restricted so that it is performed only when reference sample line 0 is used. When performing sub-block-based intra-prediction, filtering of the prediction block, which will be described below, may not be performed.

[0520] The final prediction block can be generated by filtering the prediction block generated via intra-frame prediction.

[0521] Filtering can be performed by applying specific weights to the target sample, left reference sample, top reference sample, and / or top-left reference sample, which are the targets to be filtered.

[0522] The weights and / or reference samples (e.g., the range of reference samples, the location of reference samples, etc.) used for filtering can be determined based on at least one of the block size, intra-frame prediction mode, and the location of the target filter sample in the prediction block.

[0523] For example, filtering can be performed only in specific intra-frame prediction modes (e.g., DC mode, planar mode, vertical mode, horizontal mode, diagonal mode, and / or adjacent diagonal mode).

[0524] Adjacent diagonal patterns can be patterns with numbers obtained by adding k to the diagonal pattern's number, or patterns with numbers obtained by subtracting k from the diagonal pattern's number. In other words, the number of an adjacent diagonal pattern can be the sum of the diagonal pattern's number and k, or the difference between the diagonal pattern's number and k. For example, k can be a positive integer of 8 or less.

[0525] The intra prediction mode of the target block can be derived using the intra prediction modes of neighboring blocks that exist near the target block, and this derived intra prediction mode can be entropy encoded and / or entropy decoded.

[0526] For example, when the intra prediction mode of the target block is the same as that of the neighboring blocks, specific flag information can be used to signal information indicating that the intra prediction mode of the target block is the same as that of the neighboring blocks.

[0527] Furthermore, for example, an indicator information of neighboring blocks whose intra-prediction modes are the same as those of the target block can be transmitted using signals.

[0528] For example, when the intra prediction mode of the target block is different from that of the neighboring blocks, entropy coding and / or entropy decoding can be performed on the information about the intra prediction mode of the target block by performing entropy coding and / or entropy decoding based on the intra prediction modes of the neighboring blocks.

[0529] Figure 9 is a diagram illustrating an embodiment used to explain the inter-frame prediction process.

[0530] The rectangles shown in Figure 9 can represent images (or frames). Furthermore, the arrows in Figure 9 can indicate prediction directions. An arrow pointing from the first frame to the second frame indicates that the second frame references the first frame. That is, each image can be encoded and / or decoded according to the prediction direction.

[0531] Images can be classified into intra-frame frames (I-frames), single-predictive or predictive-coded frames (P-frames), and double-predictive or double-predictive-coded frames (B-frames) based on their encoding type. Each frame can be encoded and / or decoded according to its encoding type.

[0532] When the target image to be encoded is an I-frame, the target image can be encoded using the data contained in the image itself without inter-frame prediction referencing other images. For example, an I-frame can be encoded solely via intra-frame prediction.

[0533] When the target image is a P-frame, it can be encoded using inter-frame prediction with reference frames existing in one direction. Here, the one direction can be a forward direction or a backward direction.

[0534] When the target image is a B-frame, the image can be encoded via inter-frame prediction using reference frames present in both directions, or via inter-frame prediction using reference frames present in one of the forward and backward directions. Here, the two directions can be the forward and backward directions.

[0535] P-frames and B-frames that are encoded and / or decoded using reference frames can be considered as images using inter-frame prediction.

[0536] The following will describe in detail the inter-frame prediction in inter-frame mode according to the embodiments.

[0537] Reference images and motion information can be used to perform inter-frame prediction or motion compensation.

[0538] In inter-frame mode, encoding device 100 may perform inter-frame prediction and / or motion compensation on the target block. Decoding device 200 may perform inter-frame prediction and / or motion compensation on the target block corresponding to the inter-frame prediction and / or motion compensation performed by encoding device 100.

[0539] Motion information of the target block can be derived independently by the encoding device 100 and the decoding device 200 during inter-frame prediction. Motion information can be derived using the motion information of reconstructed neighboring blocks, the motion information of the col block, and / or the motion information of blocks adjacent to the col block.

[0540] For example, encoding device 100 or decoding device 200 can perform prediction and / or motion compensation by using motion information of spatial candidates and / or temporal candidates as motion information for a target block. The target block may represent a PU and / or a PU partition.

[0541] Spatial candidates can be reconstructed blocks that are spatially adjacent to the target block.

[0542] The time candidate can be a reconstructed block that corresponds to the target block in a previously reconstructed co-location frame (col frame).

[0543] In inter-frame prediction, the encoding device 100 and the decoding device 200 can improve encoding efficiency and decoding efficiency by utilizing motion information from spatial candidates and / or temporal candidates. The motion information from spatial candidates can be referred to as "spatial motion information." The motion information from temporal candidates can be referred to as "temporal motion information."

[0544] Below, the motion information of spatial candidates can be the motion information of PUs including spatial candidates. The motion information of temporal candidates can be the motion information of PUs including temporal candidates. The motion information of candidate blocks can be the motion information of PUs including candidate blocks.

[0545] Inter-frame prediction can be performed using a reference frame.

[0546] The reference image can be at least one of an image preceding or following the target image. The reference image can be an image used for prediction of the target block.

[0547] In inter-frame prediction, regions within a reference frame can be specified using a reference frame index (or refIdx) used to indicate the reference frame, motion vectors that will be described subsequently, and so on. Here, the region specified in the reference frame can indicate a reference block.

[0548] Inter-frame prediction can select a reference frame, and can also select a reference block corresponding to the target block from the reference frame. In addition, inter-frame prediction can use the selected reference block to generate a prediction block for the target block.

[0549] Motion information can be derived by each of the encoding device 100 and the decoding device 200 during inter-frame prediction.

[0550] Spatial candidates can be 1) blocks existing in the target frame, 2) blocks that have been previously reconstructed via encoding and / or decoding, and 3) blocks adjacent to or located at the corner of the target block. Here, a "block located at the corner of the target block" can be a block that is vertically adjacent to a horizontally adjacent neighboring block, or a block that is horizontally adjacent to a vertically adjacent neighboring block. Furthermore, "block located at the corner of the target block" can have the same meaning as "block adjacent to the corner of the target block." The meaning of "block located at the corner of the target block" can be included within the meaning of "block adjacent to the target block."

[0551] For example, a spatial candidate can be a reconstruction block located to the left of the target block, a reconstruction block located above the target block, a reconstruction block located at the lower left corner of the target block, a reconstruction block located at the upper right corner of the target block, or a reconstruction block located at the upper left corner of the target block.

[0552] Each of the encoding device 100 and the decoding device 200 can identify a block existing in the col frame at a spatial position corresponding to the target block. The position of the target block in the target frame and the position of the identified block in the col frame can correspond to each other.

[0553] Each of the encoding device 100 and the decoding device 200 can identify a col block existing at a predefined relevant location for the identified block as a time candidate. The predefined relevant location can be a location existing inside and / or outside the identified block.

[0554] For example, a col block can include a first col block and a second col block. When the coordinates of the identified block are (xP, yP) and the size of the identified block is represented by (nPSW, nPSH), the first col block can be the block located at coordinates (xP+nPSW, yP+nPSH). The second col block can be the block located at coordinates (xP+(nPSW>>1), yP+(nPSH>>1)). When the first col block is unavailable, the second col block can be used selectively.

[0555] The motion vector of the target block can be determined based on the motion vector of the col block. Each of the encoding device 100 and the decoding device 200 can scale the motion vector of the col block. The scaled motion vector of the col block can be used as the motion vector of the target block. Furthermore, the motion vectors of motion information for time candidates stored in a list can also be scaled motion vectors.

[0556] The ratio of the motion vector of the target block to the motion vector of the col block can be the same as the ratio of the first time distance to the second time distance. The first time distance can be the distance between the reference frame and the target frame of the target block. The second time distance can be the distance between the reference frame and the col frame of the col block.

[0557] The scheme used to derive motion information can be changed depending on the inter-frame prediction mode of the target block. For example, inter-frame prediction modes applied to inter-frame prediction may include Advanced Motion Vector Prediction Factor (AMVP) mode, merge mode, skip mode, merge mode with motion vector difference, sub-block merge mode, triangular partitioning mode, inter-frame / intra-frame combined prediction mode, affine inter-frame mode, and current frame reference mode. The merge mode can also be called "motion merge mode." Each mode will be described in detail below.

[0558] 1) AMVP mode When using AMVP mode, the encoding device 100 can search for similar blocks in the neighborhood of the target block. The encoding device 100 can obtain a predicted block by performing a prediction on the target block using the motion information of the found similar blocks. The encoding device 100 can encode the residual block, which is the difference between the target block and the predicted block.

[0559] 1-1) Create a list of candidate motion vectors for prediction.When the AMVP mode is used as the prediction mode, each of the encoding device 100 and the decoding device 200 can use spatial candidate motion vectors, temporal candidate motion vectors, and zero vectors to create a list of prediction motion vector candidates. The prediction motion vector candidate list may include one or more prediction motion vector candidates. At least one of the spatial candidate motion vectors, temporal candidate motion vectors, and zero vectors can be determined and used as a prediction motion vector candidate.

[0560] In the following text, the terms “predicted motion vector (candidate)” and “motion vector (candidate)” can be used to have the same meaning and can be used interchangeably.

[0561] In the following text, the terms “predicted motion vector candidate” and “AMVP candidate” can be used to have the same meaning and can be used interchangeably.

[0562] In the following text, the terms “predicted motion vector candidate list” and “AMVP candidate list” can be used to have the same meaning and can be used interchangeably.

[0563] Spatial candidates can include reconstructed spatial neighbor blocks. In other words, the motion vectors of the reconstructed neighbor blocks can be referred to as "spatial prediction motion vector candidates".

[0564] A temporal candidate can include the col block and the blocks adjacent to the col block. In other words, the motion vector of the col block or the motion vector of the blocks adjacent to the col block can be called a "temporal prediction motion vector candidate".

[0565] The zero vector can be a (0,0) motion vector.

[0566] The predicted motion vector candidate can be a motion vector predictor used to predict the motion vector. Furthermore, in the encoding device 100, each predicted motion vector candidate can be an initial search position for the motion vector.

[0567] 1-2) Search for motion vectors using the list of predicted motion vector candidates. Encoding device 100 can use a list of predicted motion vector candidates to determine, within a search range, the motion vectors that will be used to encode the target block. Furthermore, encoding device 100 can determine, from among the predicted motion vector candidates present in the list of predicted motion vector candidates, the predicted motion vector candidates that will be used as the target block.

[0568] The motion vector used to encode the target block can be a motion vector that can be encoded at the minimum cost.

[0569] In addition, the encoding device 100 can determine whether to use the AMVP mode to encode the target block.

[0570] 1-3) Transmission of inter-frame prediction informationEncoding device 100 can generate a bitstream that includes inter-frame prediction information required for inter-frame prediction. Decoding device 200 can use the inter-frame prediction information of the bitstream to perform inter-frame prediction on the target block.

[0571] Inter-frame prediction information may include 1) mode information indicating whether the AMVP mode is used, 2) predicted motion vector index, 3) motion vector difference (MVD), 4) reference direction and 5) reference frame index.

[0572] In the following text, the terms “predicted motion vector index” and “AMVP index” can be used interchangeably and have the same meaning.

[0573] In addition, inter-frame prediction information may include residual signals.

[0574] When the mode information indicates that the AMVP mode is used, the decoding device 200 can obtain the predicted motion vector index, MVD, reference direction, and reference frame index from the bitstream through entropy decoding.

[0575] The predicted motion vector index indicates which of the predicted motion vector candidates included in the predicted motion vector candidate list will be used to predict the target block.

[0576] 1-4) Inter-frame prediction in AMVP mode using inter-frame prediction information The decoding device 200 can use a list of predicted motion vector candidates to derive predicted motion vector candidates, and can determine the motion information of the target block based on the derived predicted motion vector candidates.

[0577] The decoding device 200 can use the predicted motion vector index to determine a motion vector candidate for the target block from among the predicted motion vector candidates included in the predicted motion vector candidate list. The decoding device 200 can select the predicted motion vector candidate indicated by the predicted motion vector index from among the predicted motion vector candidates included in the predicted motion vector candidate list as the predicted motion vector for the target block.

[0578] Encoding device 100 can generate an entropy-coded predicted motion vector index by applying entropy coding to the predicted motion vector index, and can generate a bitstream including the entropy-coded predicted motion vector index. The entropy-coded predicted motion vector index can be transmitted from encoding device 100 to decoding device 200 via the bitstream. Decoding device 200 can extract the entropy-coded predicted motion vector index from the bitstream, and can obtain the predicted motion vector index by applying entropy decoding to the entropy-coded predicted motion vector index.

[0579] The motion vector actually used for inter-frame prediction of the target block may not match the predicted motion vector. MVD (Motion Vector Difference) can be used to indicate the difference between the actual motion vector used for inter-frame prediction of the target block and the predicted motion vector. The encoding device 100 can derive a predicted motion vector similar to the actual motion vector used for inter-frame prediction of the target block in order to use the smallest possible MVD.

[0580] Motion Vector Difference (MVD) can be the difference between the motion vector of the target block and the predicted motion vector. Encoding device 100 can compute the MVD and generate an entropy-coded MVD by applying entropy coding to the MVD. Encoding device 100 can generate a bitstream including the entropy-coded MVD.

[0581] The MVD can be sent from the encoding device 100 to the decoding device 200 via a bitstream. The decoding device 200 can extract the entropy-encoded MVD from the bitstream and obtain the MVD by applying entropy decoding to the entropy-encoded MVD.

[0582] The decoding device 200 can derive the motion vector of the target block by summing the MVD and the predicted motion vector. In other words, the motion vector of the target block derived by the decoding device 200 can be the sum of the MVD and the motion vector candidate.

[0583] Furthermore, the encoding device 100 can generate entropy-coded MVD resolution information by applying entropy coding to the calculated MVD resolution information, and can generate a bitstream including the entropy-coded MVD resolution information. The decoding device 200 can extract the entropy-coded MVD resolution information from the bitstream, and can obtain the MVD resolution information by applying entropy decoding to the entropy-coded MVD resolution information. The decoding device 200 can use the MVD resolution information to adjust the MVD resolution.

[0584] Additionally, the encoding device 100 can calculate the MVD based on an affine model. The decoding device 200 can derive the affine control motion vector of the target block by the sum of the MVD and the affine control motion vector candidates, and can use the affine control motion vector to derive the motion vector of the sub-block.

[0585] The reference direction can indicate a list of reference frames that will be used to predict the target block. For example, the reference direction can indicate one of reference frame list L0 and reference frame list L1.

[0586] The reference direction only indicates the list of reference frames that will be used to predict the target block, and does not necessarily mean that the direction of the reference frames is limited to the forward or backward direction. In other words, each of the reference frame lists L0 and L1 can include frames in the forward and / or backward directions.

[0587] A unidirectional reference direction can mean using a single reference screen list. A bidirectional reference direction can mean using two reference screen lists. In other words, the reference direction can indicate one of the following: using only reference screen list L0, using only reference screen list L1, or using both reference screen lists.

[0588] A reference frame index indicates a reference frame in a list of reference frames used to predict a target block. Encoding device 100 can generate an entropy-coded reference frame index by applying entropy coding to the reference frame index, and can generate a bitstream including the entropy-coded reference frame index. The entropy-coded reference frame index can be signaled from encoding device 100 to decoding device 200 via the bitstream. Decoding device 200 can extract the entropy-coded reference frame index from the bitstream, and can obtain the reference frame index by applying entropy decoding to the entropy-coded reference frame index.

[0589] When two reference frame lists are used to predict a target block, a single reference frame index and a single motion vector can be used for each of the reference frame lists. Furthermore, when two reference frame lists are used to predict a target block, two prediction blocks can be specified for the target block. For example, the (final) prediction block for the target block can be generated using the average or weighted sum of the two prediction blocks for the target block.

[0590] The motion vector of the target block can be derived by predicting the motion vector index, MVD, reference direction, and reference screen index.

[0591] The decoding device 200 can generate a predicted block for a target block based on the derived motion vector and the reference frame index. For example, the predicted block can be a reference block indicated by the derived motion vector in a reference frame indicated by the reference frame index.

[0592] Since the predicted motion vector index and MVD are encoded, while the motion vector of the target block itself is not encoded, the number of bits sent from the encoding device 100 to the decoding device 200 can be reduced, and the encoding efficiency can be improved.

[0593] For the target block, motion information from reconstructed neighboring blocks can be used. In certain inter-frame prediction modes, the encoding device 100 may not encode the actual motion information of the target block separately. Instead of encoding the target block's motion information, additional information can be encoded, which enables the derivation of the target block's motion information using the reconstructed motion information from neighboring blocks. Because this additional information is encoded, the number of bits sent to the decoding device 200 can be reduced, and encoding efficiency can be improved.

[0594] For example, in inter-frame prediction modes where motion information of the target block is not directly encoded, skipping modes and / or merging modes may exist. Here, the motion information of each of the neighboring units in the encoding device 100 and decoding device 200 that can be used to indicate reconstruction will be used as the identifier and / or index of the unit's motion information for the target unit.

[0595] 2) Merge Mode Merging is a scheme used to derive motion information for a target block. The term "merging" can mean merging the motion of multiple blocks. Merging can also mean that the motion information of one block is applied to other blocks. In other words, a merging pattern can be a pattern for deriving the motion information of a target block from the motion information of neighboring blocks.

[0596] When using the merging mode, the encoding device 100 can use motion information from spatial candidates and / or temporal candidates to predict motion information of the target block. Spatial candidates may include reconstructed spatially adjacent blocks that are spatially adjacent to the target block. Spatially adjacent blocks may include left-side adjacent blocks and top-side adjacent blocks. Temporal candidates may include col blocks. The terms "spatial candidate" and "spatial merging candidate" are used interchangeably and have the same meaning. The terms "temporal candidate" and "temporal merging candidate" are used interchangeably and have the same meaning.

[0597] The encoding device 100 can obtain a prediction block through prediction. The encoding device 100 can encode a residual block, which is the difference between the target block and the prediction block.

[0598] 2-1) Create a list of candidate mergers When using the merging mode, each of the encoding device 100 and the decoding device 200 can create a merging candidate list using motion information from spatial candidates and / or motion information from temporal candidates. The motion information may include 1) a motion vector, 2) a reference frame index, and 3) a reference direction. The reference direction can be unidirectional or bidirectional. The reference direction may represent an inter-frame prediction indicator.

[0599] The merge candidate list can include merge candidates. Merge candidates can be motion information. In other words, the merge candidate list can be a list that stores multiple pieces of motion information.

[0600] The merged candidate can be the motion information of multiple temporal and / or spatial candidates. In other words, the merged candidate list can include the motion information of temporal and / or spatial candidates, etc.

[0601] Furthermore, the merge candidate list may include new merge candidates generated by combining merge candidates that already exist in the merge candidate list. In other words, the merge candidate list may include new motion information generated by combining multiple motion information items that previously existed in the merge candidate list.

[0602] In addition, the merge candidate list may include history-based merge candidates. History-based merge candidates may be motion information of blocks that were encoded and / or decoded before the target block.

[0603] In addition, the list of merge candidates may include merge candidates based on the average of two merge candidates.

[0604] Merging candidates can be specific patterns for deriving inter-frame prediction information. Merging candidates can be information indicating specific patterns for deriving inter-frame prediction information. Inter-frame prediction information for a target block can be derived based on the specific patterns indicated by the merging candidates. Furthermore, the specific pattern can include the processing of deriving a series of inter-frame prediction information. This specific pattern can be an inter-frame prediction information derivation pattern or a motion information derivation pattern.

[0605] Inter-frame prediction information for the target block can be derived based on the pattern indicated by the merge candidate selected in the merge candidate list via the merge index.

[0606] For example, the motion information derivation mode in the merge candidate list can be at least one of the following modes: 1) motion information derivation mode for sub-block units and 2) affine motion information derivation mode.

[0607] In addition, the list of merged candidates may include motion information for the zero vector. The zero vector may also be referred to as a "zero merged candidate".

[0608] In other words, the multiple motion information in the merged candidate list can be at least one of the following: 1) motion information of spatial candidates, 2) motion information of temporal candidates, 3) motion information generated by combining multiple motion information that previously existed in the merged candidate list, and 4) zero vector.

[0609] Motion information may include 1) motion vectors, 2) reference frame indexes, and 3) reference directions. The reference direction can also be referred to as an "inter-frame prediction indicator." The reference direction can be unidirectional or bidirectional. A unidirectional reference direction can indicate L0 prediction or L1 prediction.

[0610] A list of merge candidates can be created before performing predictions in merge mode.

[0611] The number of merge candidates in the merge candidate list can be predefined. Each of the encoding device 100 and the decoding device 200 can add merge candidates to the merge candidate list according to a predefined scheme and a predefined priority, so that the merge candidate list has a predefined number of merge candidates. The merge candidate list of the encoding device 100 and the merge candidate list of the decoding device 200 can be made identical to each other using a predefined scheme and a predefined priority.

[0612] Merging can be applied based on either CU or PU. When merging is performed based on CU or PU, the encoding device 100 can send a bitstream including predefined information to the decoding device 200. For example, the predefined information may include 1) information indicating whether merging is to be performed on each block partition, and 2) information about the blocks to be merged among the blocks that are spatial candidates and / or temporal candidates for the target block.

[0613] 2-2) Search for motion vectors using a merged candidate list Encoding device 100 can determine merge candidates to be used for encoding a target block. For example, encoding device 100 can perform prediction on the target block using merge candidates from the merge candidate list and can generate residual blocks for the merge candidates. Encoding device 100 can encode the target block using the merge candidate that generates the minimum cost in encoding both the prediction and the residual blocks.

[0614] In addition, the encoding device 100 can determine whether to use a merge mode to encode the target block.

[0615] 2-3) Transmission of inter-frame prediction information Encoding device 100 can generate a bitstream including inter-frame prediction information required for inter-frame prediction. Encoding device 100 can generate entropy-coded inter-frame prediction information by performing entropy coding on the inter-frame prediction information, and can send the bitstream including the entropy-coded inter-frame prediction information to decoding device 200. The entropy-coded inter-frame prediction information can be sent by encoding device 100 to decoding device 200 via a bitstream signal. Decoding device 200 can extract the entropy-coded inter-frame prediction information from the bitstream, and can obtain inter-frame prediction information by applying entropy decoding to the entropy-coded inter-frame prediction information.

[0616] The decoding device 200 can use the inter-frame prediction information of the bitstream to perform inter-frame prediction on the target block.

[0617] Inter-frame prediction information may include 1) mode information indicating whether a merging mode is used, 2) merging index, and 3) correction information.

[0618] In addition, inter-frame prediction information may include residual signals.

[0619] Decoding device 200 can obtain the merge index from the bitstream only when the mode information indicates that the merge mode is used.

[0620] Pattern information can be merge flags. The unit of pattern information can be a block. Information about a block can include pattern information, and the pattern information can indicate whether a merge pattern is applied to the block.

[0621] The merge index can indicate which merge candidate from the merge candidate list will be used to predict the target block. Optionally, the merge index can indicate which block from the spatially or temporally adjacent neighboring blocks will be merged with the target block.

[0622] Encoding device 100 can select the merge candidate with the highest encoding performance from the merge candidate list, and can set the value of the merge index to indicate the selected merge candidate.

[0623] The correction information can be information used to correct motion vectors. Encoding device 100 can generate the correction information. Decoding device 200 can correct the motion vectors of the merging candidates selected by the merging index based on the correction information.

[0624] The correction information may include at least one of information indicating whether correction will be performed, correction direction information, and correction size information. A prediction mode that corrects motion vectors based on correction information transmitted by a signal may be referred to as a "merging mode with motion vector difference".

[0625] 2-4) Inter-frame prediction using merging mode with inter-frame prediction information Decoding device 200 can perform prediction on target block using a merge candidate indicated by a merge index from among the merge candidates included in the merge candidate list.

[0626] The motion vector of the target block can be specified by the motion vector of the merge candidate indicated by the merge index, the reference screen index, and the reference direction.

[0627] 3) Skip Mode Skip mode can be a mode that applies spatial or temporal motion information to the target block without alteration. Furthermore, skip mode can be a mode that does not use the residual signal. In other words, when using skip mode, the reconstructed block can be identical to the predicted block.

[0628] The difference between merge mode and skip mode lies in whether or not residual signals are sent or used. In other words, skip mode is similar to merge mode except that residual signals are not sent or used.

[0629] When using skip mode, encoding device 100 can transmit information about blocks whose motion information will be used as motion information for target blocks via a bitstream to decoding device 200. Encoding device 100 can generate entropy-encoded information by performing entropy encoding on this information, and can transmit the entropy-encoded information as a signal to decoding device 200 via a bitstream. Decoding device 200 can extract the entropy-encoded information from the bitstream, and can obtain information by applying entropy decoding to the entropy-encoded information.

[0630] Furthermore, when using skip mode, encoding device 100 may not send other syntax information (such as MVD) to decoding device 200. For example, when using skip mode, encoding device 100 may not send the syntax elements associated with at least one of MVD, code block flag, and transform coefficient level to decoding device 200.

[0631] 3-1) Create a list of candidate mergers The merge candidate list can also be used in skip mode. In other words, the merge candidate list can be used in both merge mode and skip mode. In this respect, the merge candidate list can also be referred to as the "skip candidate list" or the "merge / skip candidate list".

[0632] Optionally, the skip mode may use an additional candidate list that differs from the candidate list used in the merge mode. In this case, in the following description, the merge candidate list and merge candidate may be replaced by the skip candidate list and the skip candidate, respectively.

[0633] A list of merged candidates can be created before performing predictions in skip mode.

[0634] 3-2) Search for motion vectors using a merged candidate list Encoding device 100 can determine merge candidates to be used for encoding the target block. For example, encoding device 100 can perform prediction on the target block using merge candidates from the merge candidate list. Encoding device 100 can encode the target block using the merge candidate that generates the minimum cost in the prediction.

[0635] In addition, the encoding device 100 can determine whether to use a skip mode to encode the target block.

[0636] 3-3) Transmission of inter-frame prediction information Encoding device 100 can generate a bitstream that includes inter-frame prediction information required for inter-frame prediction. Decoding device 200 can use the inter-frame prediction information of the bitstream to perform inter-frame prediction on the target block.

[0637] Inter-frame prediction information may include 1) mode information indicating whether a skip mode is used and 2) a skip index.

[0638] Skipping indexes is the same as merging indexes as described above.

[0639] When using skip mode, the target block can be encoded without using the residual signal. Inter-frame prediction information may not include the residual signal. Optionally, the bitstream may not include the residual signal.

[0640] Decoding device 200 may obtain the skip index from the bitstream only when the mode information indicates that a skip mode is used. As mentioned above, the merge index and the skip index may be the same. Decoding device 200 may obtain the skip index from the bitstream only when the mode information indicates that either a merge mode or a skip mode is used.

[0641] Skip index indicates which of the merge candidates included in the merge candidate list will be used to predict the target block.

[0642] 3-4) Inter-frame prediction in skip mode using inter-frame prediction information Decoding device 200 can perform prediction on the target block using a merge candidate indicated by a skip index from among the merge candidates included in the merge candidate list.

[0643] The motion vector of the target block can be specified by the motion vector of the merge candidate indicated by the skip index, the reference screen index, and the reference direction.

[0644] 4) Current screen reference mode The current frame reference mode can represent a prediction mode that uses the previously reconstructed area in the target frame to which the target block belongs.

[0645] Motion vectors can be used to specify previously reconstructed areas. The reference frame index of the target block can be used to determine whether the target block has been encoded in the current frame reference mode.

[0646] A flag or index indicating whether a target block is encoded in the current screen reference mode can be sent by the encoding device 100 to the decoding device 200. Optionally, whether a target block is encoded in the current screen reference mode can be inferred from the target block's reference screen index.

[0647] When a target block is encoded in the current frame reference mode, the current frame can exist in a fixed position or any position in the reference frame list for the target block.

[0648] For example, the fixed position could be the position where the reference screen index value is 0 or the last position.

[0649] When the target image exists at any position in the list of reference images, an additional reference image index indicating such an arbitrary position can be sent by the encoding device 100 to the decoding device 200 by a signal.

[0650] 5) Sub-block merging mode Sub-block merging mode can be a mode that derives motion information from sub-blocks of the CU.

[0651] When applying the sub-block merging mode, the motion information of col-sub-blocks of the target sub-block in the reference image (i.e., based on the temporal merging candidate of the sub-block) and / or affine control point motion vector merging candidates can be used to generate a list of sub-block merging candidates.

[0652] 6) Triangular partitioning modeIn the triangular partitioning mode, the target block can be partitioned diagonally, and sub-target blocks generated through partitioning can be produced. For each sub-target block, motion information of the corresponding sub-target block can be derived, and the derived motion information can be used to derive the predicted samples of each sub-target block. The predicted samples of the target block can be derived by weighted summing of the predicted samples of the sub-target blocks generated through partitioning.

[0653] 7) Combined inter-frame and intra-frame prediction modes The combined inter-frame-intra-frame prediction mode can be a mode that uses a weighted sum of prediction samples generated via inter-frame prediction and prediction samples generated via intra-frame prediction to derive prediction samples for the target block.

[0654] In the above mode, the decoding device 200 can autonomously correct the derived motion information. For example, the decoding device 200 can search for motion information with the minimum sum of absolute differences (SAD) in a specific region based on a reference block indicated by the derived motion information, and can derive the found motion information as corrected motion information.

[0655] In the above mode, the decoding device 200 can use optical flow to compensate for the prediction samples derived via inter-frame prediction.

[0656] In the AMVP mode, merge mode, skip mode, etc. described above, the index information of the list can be used to specify the motion information among multiple motion information in the list that will be used to predict the target block.

[0657] To improve coding efficiency, the encoding device 100 may use only the index of the element in the signal transmission list that generates the minimum cost in inter-frame prediction of the target block. The encoding device 100 may encode this index and may signal the encoded index.

[0658] Therefore, the encoding device 100 and the decoding device 200 must be able to derive the lists described above (i.e., the candidate list for predicted motion vectors and the candidate list for merging) using the same scheme and based on the same data. Here, the same data may include reconstructed frames and reconstructed blocks. Furthermore, in order to specify elements using indices, the order of elements in the lists must be fixed.

[0659] Figure 10 illustrates a spatial candidate according to an embodiment.

[0660] Figure 10 shows the locations of spatial candidates.

[0661] The large block in the center of the graph represents the target block. The five smaller blocks represent spatial candidates.

[0662] The coordinates of the target block can be (xP, yP), and the size of the target block can be represented by (nPSW, nPSH).

[0663] Spatial candidate A0 can be a block adjacent to the lower left corner of the target block. A0 can be a block that occupies the pixel located at coordinates (xP-1, yP+nPSH).

[0664] Spatial candidate A1 can be the block that is adjacent to the left side of the target block. A1 can be the bottommost block among the blocks that are adjacent to the left side of the target block. Alternatively, A1 can be the block that is adjacent to the top side of A0. A1 can be the block that occupies the pixel located at coordinates (xP-1, yP+nPSH-1).

[0665] Spatial candidate B0 can be the block adjacent to the top right corner of the target block. B0 can be a block that occupies the pixel located at coordinates (xP+nPSW, yP-1).

[0666] Spatial candidate B1 can be the block that is top-adjacent to the target block. B1 can be the rightmost block among the blocks that are top-adjacent to the target block. Alternatively, B1 can be the block that is left-adjacent to B0. B1 can be the block that occupies the pixel located at coordinates (xP+nPSW-1, yP-1).

[0667] Spatial candidate B2 can be a block adjacent to the top-left corner of the target block. B2 can be a block that occupies the pixel located at coordinates (xP-1, yP-1).

[0668] Determining the availability of spatial and temporal candidates In order to include spatial or temporal motion information in the list, it must be determined whether the spatial or temporal motion information is available.

[0669] In the following text, candidate blocks may include spatial candidates and temporal candidates.

[0670] For example, the determination can be performed by sequentially applying steps 1) through 4).

[0671] Step 1) When the PU including the candidate block is outside the boundary of the screen, the availability of the candidate block can be set to "false". The statement "availability is set to false" can have the same meaning as "set to unavailable".

[0672] Step 2) When the PU including the candidate block is outside the boundary of the stripe, the availability of the candidate block can be set to "false". When the target block and the candidate block are in different stripes, the availability of the candidate block can be set to "false".

[0673] Step 3) When the PU including the candidate block is outside the boundary of the parallel block, the availability of the candidate block can be set to "false". When the target block and the candidate block are in different parallel blocks, the availability of the candidate block can be set to "false".

[0674] Step 4) When the prediction mode of the PU including the candidate block is intra-frame prediction mode, the availability of the candidate block can be set to "false". When the PU including the candidate block does not use inter-frame prediction, the availability of the candidate block can be set to "false".

[0675] Figure 11 illustrates the order in which motion information of spatial candidates is added to the merging list according to an embodiment.

[0676] As shown in Figure 11, when multiple motion information entries from spatial candidates are added to the merging list, the order A1, B1, B0, A0, and B2 can be used. That is, multiple motion information entries from available spatial candidates can be added to the merging list in the order A1, B1, B0, A0, and B2.

[0677] Methods for deriving merge lists in merge mode and skip mode As described above, the maximum number of merge candidates in the merge list can be set. The maximum number can be indicated by "N". The set number can be sent from the encoding device 100 to the decoding device 200. The stripe header can include N. In other words, the maximum number of merge candidates in the merge list for the target block of the stripe can be set via the stripe header. For example, the value of N can essentially be 5.

[0678] Multiple motion information (i.e., merge candidates) can be added to the merge list in the order of steps 1) to 4).

[0679] Step 1) Among the spatial candidates, available spatial candidates can be added to the merge list. Multiple motion information entries from available spatial candidates can be added to the merge list in the order shown in Figure 11. Here, if the motion information of an available spatial candidate overlaps with other motion information already existing in the merge list, the motion information of the available spatial candidate may not be added to the merge list. The operation of checking whether the corresponding motion information overlaps with other motion information existing in the list can be simply referred to as "overlap check".

[0680] The maximum number of motion information entries that can be added is N.

[0681] Step 2) When the number of motion information entries in the merge list is less than N and time candidates are available, the motion information of the time candidates can be added to the merge list. However, if the available motion information of time candidates overlaps with other motion information already existing in the merge list, the available motion information of time candidates may not be added to the merge list.

[0682] Step 3) When the number of motion information entries in the merge list is less than N and the target strip type is "B", the combined motion information generated by combining bidirectional prediction (double prediction) can be added to the merge list.

[0683] The target strip can be a strip that includes the target block.

[0684] Combined motion information can be a combination of L0 motion information and L1 motion information. L0 motion information can be motion information that only refers to the L0 reference frame list. L1 motion information can be motion information that only refers to the L1 reference frame list.

[0685] The merged list may contain one or more L0 motion entries. Additionally, the merged list may contain one or more L1 motion entries.

[0686] Combined motion information may include one or more pieces of combined motion information. When generating combined motion information, the L0 motion information and L1 motion information that will be used in the step of generating combined motion information can be predefined from the one or more L0 motion information and the one or more L1 motion information. One or more pieces of combined motion information can be generated in a predefined order via bidirectional prediction using a pair of different motion information from a merge list. One of the different motion information in the pair can be L0 motion information, and the other of the different motion information in the pair can be L1 motion information.

[0687] For example, the combined motion information with the highest priority can be a combination of L0 motion information with a merge index of 0 and L1 motion information with a merge index of 1. When the motion information with a merge index of 0 is not L0 motion information, or when the motion information with a merge index of 1 is not L1 motion information, neither combined motion information is generated nor added. Next, the combined motion information with the next higher priority can be a combination of L0 motion information with a merge index of 1 and L1 motion information with a merge index of 0. Subsequent detailed combinations can conform to other combinations in the field of video encoding / decoding.

[0688] Here, when the combined motion information overlaps with other motion information already existing in the merge list, the combined motion information may not be added to the merge list.

[0689] Step 4) When the number of motion information entries in the merge list is less than N, the motion information of the zero vector can be added to the merge list.

[0690] Zero-vector motion information can be motion information where the motion vector is zero.

[0691] The number of zero-vector motion information entries can be one or more. The reference frame indices for one or more zero-vector motion information entries can be different from each other. For example, the reference frame index value for the first zero-vector motion information entry can be 0. The reference frame index value for the second zero-vector motion information entry can be 1.

[0692] The number of zero-vector motion information entries can be the same as the number of reference frames in the reference frame list.

[0693] The reference direction for zero-vector motion information can be bidirectional. Both motion vectors can be zero vectors. The number of zero-vector motion information entries can be the smaller of the number of reference frames in reference frame list L0 and the number of reference frames in reference frame list L1. Optionally, when the number of reference frames in reference frame list L0 and the number of reference frames in reference frame list L1 are different from each other, a unidirectional reference direction can be used for reference frame indexing that can be applied to only a single reference frame list.

[0694] Encoding device 100 and / or decoding device 200 may subsequently add zero-vector motion information to the merge list while changing the reference screen index.

[0695] When zero-vector motion information overlaps with other motion information already existing in the merge list, the zero-vector motion information may not be added to the merge list.

[0696] The order of steps 1) to 4) above is merely exemplary and can be changed. Furthermore, some steps in the above steps may be omitted based on predefined conditions.

[0697] A method for deriving a candidate list of predicted motion vectors in AMVP mode The maximum number of predicted motion vector candidates in the candidate list can be predefined. N can be used to indicate the predefined maximum number. For example, the predefined maximum number could be 2.

[0698] Multiple pieces of motion information (i.e., predicted motion vector candidates) can be added to the predicted motion vector candidate list in the order of steps 1) to 3).

[0699] Step 1) Available spatial candidates can be added to the list of predicted motion vector candidates. Spatial candidates may include a first spatial candidate and a second spatial candidate.

[0700] The first spatial candidate can be one of A0, A1, scaled A0, and scaled A1. The second spatial candidate can be one of B0, B1, B2, scaled B0, scaled B1, and scaled B2.

[0701] Multiple motion information entries from available spatial candidates can be added to the predicted motion vector candidate list in the order of first spatial candidates and second spatial candidates. In this case, if the motion information of an available spatial candidate overlaps with other motion information already existing in the predicted motion vector candidate list, the motion information of the available spatial candidate may not be added to the predicted motion vector candidate list. In other words, when the value of N is 2, if the motion information of the second spatial candidate is the same as the motion information of the first spatial candidate, then the motion information of the second spatial candidate may not be added to the predicted motion vector candidate list.

[0702] The maximum number of motion information entries that can be added is N.

[0703] Step 2) When the number of motion information entries in the predicted motion vector candidate list is less than N and a time candidate is available, the motion information of the time candidate can be added to the predicted motion vector candidate list. In this case, if the motion information of the available time candidate overlaps with other motion information already existing in the predicted motion vector candidate list, the available time candidate may not be added to the predicted motion vector candidate list.

[0704] Step 3) When the number of motion information entries in the candidate list of predicted motion vectors is less than N, zero-vector motion information can be added to the candidate list of predicted motion vectors.

[0705] Zero-vector motion information may include one or more zero-vector motion information pieces. The reference frame indices of the one or more zero-vector motion information pieces may be different from each other.

[0706] Encoding device 100 and / or decoding device 200 can sequentially add multiple zero-vector motion information to the candidate list of predicted motion vectors while changing the reference frame index.

[0707] When zero-vector motion information overlaps with other motion information already existing in the candidate list of predicted motion vectors, the zero-vector motion information may not be added to the candidate list of predicted motion vectors.

[0708] The description of zero-vector motion information presented above, combined with the merged list, can also be applied to zero-vector motion information. Repeated descriptions will be omitted.

[0709] The order of steps 1) to 3) described above is merely exemplary and can be changed. Furthermore, some steps may be omitted based on predefined conditions.

[0710] Figure 12 illustrates the transformation and quantization process based on the example.

[0711] As shown in Figure 12, quantization levels can be generated by performing transformation and / or quantization on the residual signal.

[0712] The residual signal can be generated as the difference between the original block and the predicted block. Here, the predicted block can be a block generated via intra-frame prediction or inter-frame prediction.

[0713] The residual signal can be transformed into a signal in the frequency domain through a transformation process that is part of the quantization process.

[0714] The transform kernel used for the transform can include various DCT kernels, such as Discrete Cosine Transform (DCT) Type 2 (DCT-II) kernel and Discrete Sine Transform (DST) kernel.

[0715] These transform kernels can perform separable or two-dimensional (2D) non-separable transforms on the residual signal. A separable transform can be a transform indicating that a one-dimensional (1D) transform is performed on the residual signal in each of the horizontal and vertical directions.

[0716] In addition to DCT-II, the DCT and DST types adaptively used for 1D transformations may also include DCT-V, DCT-VIII, DST-I, and DST-VII, as shown in each of Tables 3 and 4 below.

[0717] Table 3

[0718] Table 4

[0719] As shown in Tables 3 and 4, transform sets can be used when deriving the DCT or DST type to be used for the transform. Each transform set may include multiple transform candidates. Each transform candidate may be a DCT type or a DST type.

[0720] Table 5 below shows examples of the transform sets that will be applied in the horizontal direction and the transform sets that will be applied in the vertical direction according to the intra-frame prediction mode.

[0721] Table 5

[0722] Table 5 shows the numbers of the vertical transform set and horizontal transform set that will be applied to the horizontal direction of the residual signal according to the intra-frame prediction mode of the target block.

[0723] As illustrated in Table 5, the transform sets to be applied in the horizontal and vertical directions can be predefined based on the intra-prediction mode of the target block. Encoding device 100 can use transforms included in the transform set corresponding to the intra-prediction mode of the target block to perform transforms and inverse transforms on the residual signal. Furthermore, decoding device 200 can use transforms included in the transform set corresponding to the intra-prediction mode of the target block to perform an inverse transform on the residual signal.

[0724] In the transform and inverse transform, as illustrated in Tables 3, 4, and 5, the set of transforms to be applied to the residual signal can be determined and may not be transmitted by signal. Transform indication information can be transmitted from the encoding device 100 to the decoding device 200 by signal. The transform indication information may be information indicating which of the multiple transform candidates included in the set of transforms to be applied to the residual signal is used.

[0725] For example, when the target block size is 64×64 or smaller, transform sets with three transforms can be configured according to the intra-frame prediction mode. The optimal transform method can be selected from a total of nine multi-transform methods generated by combinations of three transforms in the horizontal direction and three transforms in the vertical direction. With such an optimal transform method, the residual signal can be encoded and / or decoded, thus improving coding efficiency.

[0726] Here, information indicating which of the multiple transformations belonging to each transform set has been used for at least one of the vertical and horizontal transformations can be entropy encoded and / or entropy decoded. Here, truncated univariate binarization can be used to encode and / or decode such information.

[0727] As mentioned above, various transformation methods can be applied to residual signals generated via intra-frame prediction or inter-frame prediction.

[0728] The transformation may include at least one of a first transformation and a second transformation. Transform coefficients can be generated by performing a first transformation on the residual signal, and second transformation coefficients can be generated by performing a second transformation on the transform coefficients.

[0729] The first transformation can be referred to as the "primary transformation". Furthermore, the first transformation can also be referred to as the "Adaptive Multitransformation (AMT) scheme". As mentioned above, AMT can represent applying different transformations to various 1D directions (i.e., the vertical and horizontal directions).

[0730] A secondary transformation can be a transformation used to increase the energy concentration of the transformation coefficients generated by the first transformation. Similar to the first transformation, a secondary transformation can be a separable transformation or a non-separable transformation. Such a non-separable transformation can be a non-separable secondary transformation (NSST).

[0731] The first transformation can be performed using at least one of a predefined plurality of transformation methods. For example, the predefined plurality of transformation methods may include the Discrete Cosine Transform (DCT), the Discrete Sine Transform (DST), the Karhunen-Loeve Transform (KLT), etc.

[0732] Furthermore, depending on the kernel function defined for the Discrete Cosine Transform (DCT) or Discrete Sine Transform (DST), the first transform can be of various types.

[0733] For example, the transform type can be determined based on at least one of the following: 1) the prediction mode of the target block (e.g., one of intra-frame prediction and inter-frame prediction), 2) the size of the target block, 3) the shape of the target block, 4) the intra-frame prediction mode of the target block, 5) the components of the target block (e.g., one of luma component and chroma component), and 6) the partition type applied to the target block (e.g., one of quadtree, binary tree and ternary tree).

[0734] For example, based on the transform kernels presented in Table 6 below, the first transform may include transforms such as DCT-2, DCT-5, DCT-7, DST-7, DST-1, DST-8, and DCT-8. Table 6 below illustrates various transform types and transform kernel functions used for Multiple Transform Selection (MTS).

[0735] MTS can refer to the selection of a combination of one or more DCT and / or DST cores to transform the residual signal in the horizontal and / or vertical directions.

[0736] Table 6

[0737] In Table 6, i and j can be integer values ​​that are equal to or greater than 0 and less than or equal to N-1.

[0738] A secondary transformation can be performed on the transformation coefficients generated by performing the first transformation.

[0739] For example, in the first transformation, a transformation set can also be defined in the secondary transformation. The methods used to derive and / or determine the above transformation set can be applied not only to the first transformation but also to the secondary transformation.

[0740] The first and second transformations can be determined for a specific target.

[0741] For example, the first and second transformations can be applied to one or more signal components corresponding to the luma and chroma components. Whether to apply the first and / or second transformations can be determined based on at least one of the coding parameters for the target block and / or neighboring blocks. For example, whether to apply the first and / or second transformations can be determined based on the size and / or shape of the target block.

[0742] In the encoding device 100 and the decoding device 200, transformation information indicating the transformation method to be used for the target can be derived by using specified information.

[0743] For example, the transformation information may include transformation indices that will be used for primary and / or secondary transformations. Optionally, the transformation information may indicate that primary and / or secondary transformations are not used.

[0744] For example, when the target of the primary and secondary transforms is a target block, the transform method to be applied to the primary and / or secondary transforms, as indicated by the transform information, can be determined based on at least one of the encoding parameters for the target block and / or blocks adjacent to the target block.

[0745] Optionally, the encoding device 100 may send transformation information indicating the transformation method for a specific target to the decoding device 200 via a signal.

[0746] For example, for a single CU, the decoding device 200 can deduce transformation information such as whether a primary transformation is used, the index indicating the primary transformation, whether a secondary transformation is used, and the index indicating the secondary transformation. Alternatively, for a single CU, transformation information indicating the following can be transmitted by signals: whether a primary transformation is used, the index indicating the primary transformation, whether a secondary transformation is used, and the index indicating the secondary transformation.

[0747] Quantized transform coefficients (i.e., quantization levels) can be generated by quantizing the result produced by performing a first transform and / or a secondary transform, or by quantizing the residual signal.

[0748] Figure 13 shows a diagonal scan according to the example.

[0749] Figure 14 shows a horizontal scan based on an example.

[0750] Figure 15 shows a vertical scan based on the example.

[0751] The quantized transform coefficients can be scanned via at least one of (top right) diagonal scan, vertical scan, and horizontal scan, based on at least one of intra-frame prediction mode, block size, and block shape. The block can be a transform unit (TU).

[0752] Each scan can be started at a specific start point and terminated at a specific end point.

[0753] For example, quantized transform coefficients can be converted to 1D vector form by scanning the coefficients of the block using a diagonal scan as shown in Figure 13. Alternatively, a horizontal scan as shown in Figure 14 or a vertical scan as shown in Figure 15 can be used instead of a diagonal scan, depending on the block size and / or intra-frame prediction mode.

[0754] A vertical scan can be an operation that scans 2D block-type coefficients in the column direction. A horizontal scan can be an operation that scans 2D block-type coefficients in the row direction.

[0755] In other words, the choice between diagonal, vertical, and horizontal scans can be determined based on the block size and / or inter-frame prediction mode.

[0756] As shown in Figures 13, 14 and 15, the quantized transform coefficients can be scanned along the diagonal, horizontal or vertical direction.

[0757] The quantized transformation coefficients can be represented by block shapes. Each block can include multiple sub-blocks. Each sub-block can be defined based on either the minimum block size or the minimum block shape.

[0758] During scanning, the scanning order, based on the type or direction of the scan, can be applied first to the sub-blocks. Furthermore, the scanning order, based on the direction of the scan, can be applied to the quantized transform coefficients within each sub-block.

[0759] For example, as shown in Figures 13, 14, and 15, when the target block size is 8×8, quantized transform coefficients can be generated by a first transform, a second transform, and quantization of the residual signal of the target block. Therefore, one of three types of scan sequences can be applied to four 4×4 sub-blocks, and the quantized transform coefficients can be scanned for each 4×4 sub-block according to the scan sequence.

[0760] Encoding device 100 can generate entropy-coded quantized transform coefficients by performing entropy coding on scanned quantized transform coefficients, and can generate a bit stream including the entropy-coded quantized transform coefficients.

[0761] The decoding device 200 can extract entropy-encoded quantized transform coefficients from the bitstream, and can generate quantized transform coefficients by performing entropy decoding on the entropy-encoded quantized transform coefficients. The quantized transform coefficients can be arranged in a 2D block format via inverse scanning. Here, as a method of inverse scanning, at least one of upper right diagonal scanning, vertical scanning, and horizontal scanning can be performed.

[0762] In the decoding device 200, inverse quantization can be performed on the quantized transform coefficients. A secondary inverse transform can be performed on the result generated by inverse quantization, depending on whether a secondary inverse transform is performed. Furthermore, a first inverse transform can be performed on the result generated by the secondary inverse transform, depending on whether a first inverse transform will be performed. The reconstructed residual signal can be generated by performing a first inverse transform on the result generated by the secondary inverse transform.

[0763] For the luminance component reconstructed via intra-frame prediction or inter-frame prediction, an inverse mapping with dynamic range can be performed before loop filtering.

[0764] The dynamic range can be divided into 16 equal segments, and the mapping functions for the corresponding segments can be signaled. These mapping functions can be signaled at the stripe level or the parallel block group level.

[0765] The inverse mapping function can be derived from the mapping function to perform the inverse mapping.

[0766] Loop filtering, reference frame storage, and motion compensation can be performed in the inverse mapping region.

[0767] Predicted blocks generated via inter-frame prediction can be transformed to a mapped region using a mapping function, and the transformed predicted blocks can be used to generate reconstructed blocks. However, since intra-frame prediction is performed in the mapped region, predicted blocks generated via intra-frame prediction can be used to generate reconstructed blocks without requiring mapping and / or inverse mapping.

[0768] For example, when the target block is a residual block of the chrominance component, the residual block can be transformed into the inverse mapping region by scaling the chrominance component of the mapping region.

[0769] Scaling availability can be signaled at the stripe level or the parallel block group level.

[0770] For example, scaling can be applied only when the mapping is available for the luminance component and the partitions of the luminance and chrominance components follow the same tree structure.

[0771] Scaling can be performed based on the average value of the samples in the luminance prediction block corresponding to the chrominance prediction block. Here, when the target block uses inter-frame prediction, the luminance prediction block can represent the mapped luminance prediction block.

[0772] The scaling values ​​can be derived by using an index-referenced lookup table of the segment to which the average value of the sample values ​​of the brightness prediction block belongs.

[0773] The residual block can be transformed into the inverse mapping region by scaling the residual block using the finally derived values. Subsequently, for blocks of the chroma components, reconstruction, intra-frame prediction, inter-frame prediction, loop filtering, and storage of reference frames can be performed in the inverse mapping region.

[0774] For example, information indicating whether the mapping and / or inverse mapping of the luminance and chrominance components is available can be sent via a sequence parameter set using signals.

[0775] A predicted block for the target block can be generated based on a block vector. The block vector indicates the displacement between the target block and a reference block. The reference block can be a block in the target image.

[0776] In this way, the prediction mode that generates prediction blocks by referencing the target image can be called the "intra-block copy (IBC) mode".

[0777] The IBC mode can be applied to CUs with specific dimensions. For example, the IBC mode can be applied to an M×N CU. Here, M and N can be less than or equal to 64.

[0778] IBC modes can include skip mode, merge mode, AMVP mode, etc. In skip mode or merge mode, a merge candidate list can be configured, and the merge index is signaled, allowing a single merge candidate to be specified from among the existing merge candidates in the merge candidate list. The block vector of the specified merge candidate can be used as the block vector of the target block.

[0779] In AMVP mode, differential block vectors can be signaled. Furthermore, predicted block vectors can be derived from the target block's left and top neighboring blocks. Additionally, the index of which neighboring block will be used can be signaled.

[0780] In IBC mode, the predicted block can be included in the target CTU or the left CTU, and can be limited to blocks within the previously reconstructed region. For example, the value of the block vector can be restricted such that the predicted block of the target block is located in a specific region. The specific region can be defined by three 64×64 blocks that are encoded and / or decoded before the 64×64 block including the target block. Restricting the value of the block vector in this way reduces memory consumption and device complexity caused by the implementation of IBC mode.

[0781] Figure 16 is a configuration diagram of an encoding device according to an embodiment.

[0782] Encoding device 1600 can correspond to encoding device 100 described above.

[0783] Encoding device 1600 may include a processing unit 1610, a memory 1630, a user interface (UI) input device 1650, a UI output device 1660, and a storage 1640 that communicate with each other via a bus 1690. Encoding device 1600 may also include a communication unit 1620 connected to a network 1699.

[0784] The processing unit 1610 may be a central processing unit (CPU) or a semiconductor device for executing processing instructions stored in memory 1630 or storage 1640. The processing unit 1610 may be at least one hardware processor.

[0785] The processing unit 1610 can generate and process signals, data, or information input to, output from, or used in the encoding device 1600, and can perform checks, comparisons, determinations, etc., related to the signals, data, or information. In other words, in this embodiment, the processing unit 1610 can perform the generation and processing of data or information, as well as the checks, comparisons, and determinations related to the data or information.

[0786] The processing unit 1610 may include an inter-frame prediction unit 110, an intra-frame prediction unit 120, a switcher 115, a subtractor 125, a transform unit 130, a quantization unit 140, an entropy coding unit 150, an inverse quantization unit 160, an inverse transform unit 170, an adder 175, a filter unit 180, and a reference frame buffer 190.

[0787] At least some of the following components—inter-frame prediction unit 110, intra-frame prediction unit 120, switcher 115, subtractor 125, transform unit 130, quantization unit 140, entropy coding unit 150, inverse quantization unit 160, inverse transform unit 170, adder 175, filter unit 180, and reference frame buffer 190—may be program modules and capable of communicating with external devices or systems. These program modules may be included in the encoding device 1600 in the form of an operating system, application module, or other program modules.

[0788] The program modules can be physically stored in various types of known storage devices. Furthermore, at least some of the program modules can also be stored in a remote storage device capable of communicating with the encoding device 1600.

[0789] Program modules may include, but are not limited to, routines, subroutines, programs, objects, components, and data structures for performing functions or operations according to an embodiment or for implementing abstract data types according to an embodiment.

[0790] The program module can be implemented using instructions or code that are executed by at least one processor of the encoding device 1600.

[0791] The processing unit 1610 can execute instructions or codes in the inter-frame prediction unit 110, intra-frame prediction unit 120, switcher 115, subtractor 125, transform unit 130, quantization unit 140, entropy coding unit 150, dequantization unit 160, inverse transform unit 170, adder 175, filter unit 180 and reference frame buffer 190.

[0792] The storage unit may represent memory 1630 and / or storage 1640. Each of memory 1630 and storage 1640 may be any of a variety of volatile or non-volatile storage media. For example, memory 1630 may include at least one of read-only memory (ROM) 1631 and random access memory (RAM) 1632.

[0793] The storage unit can store data or information used for the operation of the encoding device 1600. In an embodiment, the data or information of the encoding device 1600 can be stored in the storage unit.

[0794] For example, storage units can store images, blocks, lists, motion information, inter-frame prediction information, bitstreams, etc.

[0795] The encoding device 1600 can be implemented in a computer system that includes a computer-readable storage medium.

[0796] The storage medium may store at least one module required for the operation of the encoding device 1600. The memory 1630 may store at least one module and may be configured such that the at least one module is operated by the processing unit 1610.

[0797] The communication unit 1620 can be used to perform functions related to communication of data or information with the encoding device 1600.

[0798] For example, communication unit 1620 can send a bit stream to decoding device 1700, which will be described later.

[0799] Figure 17 is a configuration diagram of a decoding device according to an embodiment.

[0800] Decoding device 1700 can be used in conjunction with decoding device 200 as described above.

[0801] The decoding device 1700 may include a processing unit 1710, a memory 1730, a user interface (UI) input device 1750, a UI output device 1760, and a storage 1740 that communicate with each other via a bus 1790. The decoding device 1700 may also include a communication unit 1720 connected to a network 1799.

[0802] Processing unit 1710 may be a central processing unit (CPU) or semiconductor device for executing processing instructions stored in memory 1730 or storage 1740. Processing unit 1710 may be at least one hardware processor.

[0803] The processing unit 1710 can generate and process signals, data, or information input to, output from, or used in the decoding device 1700, and can perform checks, comparisons, determinations, etc., related to the signals, data, or information. In other words, in this embodiment, the processing unit 1710 can perform the generation and processing of data or information, as well as the checks, comparisons, and determinations related to the data or information.

[0804] The processing unit 1710 may include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an intra-frame prediction unit 240, an inter-frame prediction unit 250, a switcher 245, an adder 255, a filter unit 260, and a reference frame buffer 270.

[0805] At least some of the entropy decoding unit 210, inverse quantization unit 220, inverse transform unit 230, intra-frame prediction unit 240, inter-frame prediction unit 250, adder 255, switcher 245, filter unit 260, and reference frame buffer 270 of decoding device 200 may be program modules and are capable of communicating with external devices or systems. These program modules may be included in decoding device 1700 in the form of an operating system, application module, or other program modules.

[0806] The program modules can be physically stored in various types of known storage devices. Furthermore, at least some of the program modules can also be stored in a remote storage device capable of communicating with the decoding device 1700.

[0807] Program modules may include, but are not limited to, routines, subroutines, programs, objects, components, and data structures for performing functions or operations according to an embodiment or for implementing abstract data types according to an embodiment.

[0808] The program module can be implemented using instructions or code executed by at least one processor of the decoding device 1700.

[0809] The processing unit 1710 can run instructions or codes in the entropy decoding unit 210, the dequantization unit 220, the inverse transform unit 230, the intra-frame prediction unit 240, the inter-frame prediction unit 250, the switcher 245, the adder 255, the filter unit 260, and the reference frame buffer 270.

[0810] The storage unit may represent memory 1730 and / or storage 1740. Each of memory 1730 and storage 1740 may be any of a variety of volatile or non-volatile storage media. For example, memory 1730 may include at least one of ROM 1731 and RAM 1732.

[0811] The storage unit can store data or information used for the operation of the decoding device 1700. In an embodiment, the data or information of the decoding device 1700 can be stored in the storage unit.

[0812] For example, storage units can store images, blocks, lists, motion information, inter-frame prediction information, bitstreams, etc.

[0813] The decoding device 1700 can be implemented in a computer system that includes a computer-readable storage medium.

[0814] The storage medium may store at least one module required for the operation of the decoding device 1700. The memory 1730 may store at least one module and may be configured such that the at least one module is operated by the processing unit 1710.

[0815] The communication unit 1720 can be used to perform functions related to communication of data or information with the decoding device 1700.

[0816] For example, communication unit 1720 can receive bit streams from encoding device 1600.

[0817] In the following text, "processing unit" may refer to processing unit 1610 of encoding device 1600 and / or processing unit 1710 of decoding device 1700. For example, regarding prediction-related functions, the processing unit may represent switch 115 and / or switch 245. Regarding inter-frame prediction-related functions, the processing unit may represent inter-frame prediction unit 110, subtractor 125, and adder 175, and may also represent inter-frame prediction unit 250 and adder 255. Regarding intra-frame prediction-related functions, the processing unit may represent intra-frame prediction unit 120, subtractor 125, and adder 175, and may also represent intra-frame prediction unit 240 and adder 255. Regarding transform-related functions, the processing unit may represent transform unit 130 and inverse transform unit 170, and may also represent inverse transform unit 230. Regarding quantization-related functions, the processing unit may represent quantization unit 140 and inverse quantization unit 160, and may also indicate inverse quantization unit 220. Regarding functions related to entropy encoding and / or entropy decoding, the processing unit may represent entropy encoding unit 150 and / or entropy decoding unit 210. Regarding functions related to filtering, the processing unit may represent filter unit 180 and / or filter unit 260. Regarding functions related to reference frames, the processing unit may instruct reference frame buffer 190 and / or reference frame buffer 270.

[0818] The same method can be used by encoding device 1600 and decoding device 1700 to perform the embodiments. Furthermore, at least one of the embodiments or at least a combination thereof can be used to encode / decode the image.

[0819] The application order of the embodiments may differ from each other through the encoding device 1600 and the decoding device 1700, and the application order of the embodiments may be (at least partially) the same as each other through the encoding device 1600 and the decoding device 1700.

[0820] An embodiment can be executed for each of the luminance and chrominance signals, and an embodiment can be executed equivalently for both the luminance and chrominance signals.

[0821] The blocks in the application examples can be square or non-square in shape.

[0822] Whether to apply and / or execute at least one of the above embodiments can be determined based on conditions related to the block size. In other words, when conditions related to the block size are met, at least one of the above embodiments can be applied and / or executed. These conditions include a minimum block size and a maximum block size. The block can be one of the blocks described above in conjunction with the embodiments and the units described above in conjunction with the embodiments. The block applying the minimum block size and the block applying the maximum block size can be different from each other.

[0823] For example, the above embodiments can be applied and / or executed when the block size is equal to or greater than the minimum block size and / or less than or equal to the maximum block size.

[0824] For example, the above embodiments can be applied only when the block size is a predefined block size. The predefined block size can be 2×2, 4×4, 8×8, 16×16, 32×32, 64×64, or 128×128. The predefined block size can be (2... SIZE X )×(2 SIZE Y SIZE X It can be any integer 1 or larger. SIZE Y It can be one of the integers 1 or greater.

[0825] For example, the above embodiments can be applied only to cases where the block size is equal to or greater than the minimum block size. The above embodiments can be applied only to cases where the block size is greater than the minimum block size. The minimum block size can be 2×2, 4×4, 8×8, 16×16, 32×32, 64×64, or 128×128. Alternatively, the minimum block size can be (2... SIZE MIN_X )×(2 SIZE MIN_Y SIZE MIN_X It can be any integer 1 or larger. SIZE MIN_Y It can be one of the integers 1 or greater.

[0826] For example, the above embodiments can be applied only to cases where the block size is less than or equal to the maximum block size. The maximum block size can be 2×2, 4×4, 8×8, 16×16, 32×32, 64×64, or 128×128. Alternatively, the maximum block size can be (2... SIZE MAX_X )×(2 SIZE MAX_YSIZE MAX_X It can be any integer 1 or larger. SIZE MAX_Y It can be one of the integers 1 or greater.

[0827] For example, the above embodiments can be applied only to cases where the block size is equal to or greater than the minimum block size and less than or equal to the maximum block size. (This is repeated four times in the original text.)

[0828] In the above embodiments, the block size can be the horizontal dimension (width) or the vertical dimension (height) of the block. The block size can indicate both the horizontal and vertical dimensions of the block. The block size can indicate the area of ​​the block. Each of the area, minimum block size, and maximum block size can be an integer equal to or greater than 1. Alternatively, the block size can be the result (or value) of a known equation using the horizontal and vertical dimensions of the block, or the result (or value) of the equation in the embodiments.

[0829] Furthermore, in the embodiments, the first embodiment can be applied to a first size, and the second embodiment can be applied to a second size.

[0830] The embodiments of this disclosure can be applied according to time layers. To identify the time layer to which an embodiment is applicable, a separate identifier can be signaled, and the embodiment can be applied to the time layer specified by the corresponding identifier. Here, the identifier can be defined as the lowest (bottom) layer and / or the highest (top) layer to which the embodiment is applicable, and can be defined as indicating a specific layer to which the embodiment is applied. Furthermore, a fixed time layer for applying the embodiment can also be defined.

[0831] For example, the embodiment may only be applied when the time layer of the target image is the lowest layer. For example, the embodiment may only be applied when the time layer identifier of the target image is equal to or greater than 1. For example, the embodiment may only be applied when the time layer of the target image is the highest layer.

[0832] The stripe type or parallel block group type of the application embodiment can be defined, and the application embodiment can be applied according to the corresponding stripe type or parallel block group type.

[0833] The information about a block in this specification may refer to at least one of the following: information about neighboring blocks, information about reference blocks, and information about the current block.

[0834] In addition, the information of a block may include at least one of encoding information or encoding parameters.

[0835] Encoding information or encoding parameters can include not only information (flags, indexes, etc.) encoded by the encoder and sent to the decoder as signals like syntax elements, but also information deduced from encoding or decoding processes, and can refer to the information required when encoding or decoding an image.

[0836] The coding parameter information may include at least one of the information used in inter-frame prediction, intra-frame prediction, transform, inverse transform, quantization, dequantization, entropy coding / decoding, or loop filtering.

[0837] In other words, block information refers to a combination or value of at least one of the following: block size, block depth, block partitioning information, block form (square or non-square), whether partitioned in a quadtree format, whether partitioned in a binary tree format, partitioning direction in the binary tree format (horizontal or vertical), partitioning direction in the binary tree format (symmetric or asymmetric partitioning), block index, prediction mode (intra-frame prediction or inter-frame prediction), intra-frame luma prediction mode / direction, intra-frame chroma prediction mode / direction, intra-frame prediction mode candidate list, intra-frame prediction mode candidate index, intra-frame partitioning information, inter-frame partitioning information, coded block partitioning flag, prediction block partitioning flag, transform block partitioning flag, reference sample filter tap, reference sample filter coefficient, prediction block filter tap, prediction block filter coefficient, prediction block boundary filter tap, prediction block boundary filter coefficient, motion vector (motion vector for at least one of L0, L1, L2, L3, etc.), motion vector difference (motion vector difference for at least one of L0, L1, L2, L3, etc.). The parameters include: inter-frame prediction direction (for at least one of unidirectional prediction, bidirectional prediction, etc.), reference image index (for at least one of L0, L1, L2, L3, etc.), inter-frame prediction indicator, prediction list utilization flag, reference image list, motion vector prediction index, motion vector prediction candidate, motion vector candidate list, whether to use merge mode, merge index, merge candidate, merge candidate list, whether to use skip mode, interpolation filter type, interpolation filter taps, interpolation filter coefficients, motion vector size, and motion vector representation precision (represented as n times the number of samples or 1 / n times the number of samples, such as integer samples, 1 / 2 samples, 1 / 4 samples, 1 / 8 samples, 1 / 16 samples, 1 / 32 samples, etc. n can be a positive integer. Alternatively, n can be determined based on at least one of the coding parameters of the current block and the candidate coding parameters. Furthermore, U can be a value pre-configured in the encoder / decoder, or a value sent from the encoder to the decoder by a signal.Transform type, transform size, information on whether the first transform is used, information on whether the second transform is used, first transform index, second transform index, information on the presence of residual signal, code block style, code block flag, quantization parameters, residual quantization parameters, quantization matrix, whether an intra-loop filter is applied, intra-loop filter coefficients, intra-loop filter taps, intra-loop filter shape / form, whether a deblocking filter is applied, deblocking filter coefficients, deblocking filter taps, deblocking filter strength, deblocking filter shape / form, whether adaptive sample offset is applied, adaptive sample offset value, adaptive sample offset category, adaptive sample offset type, whether an adaptive loop filter is applied, adaptive loop filter coefficients, adaptive loop filter taps, adaptive loop filter shape / form, binarization / debinarization method, context model determination method, context model update method, whether to execute normal mode, whether to execute bypass mode, context binary bits, bypass binary bits, effective coefficients. The following information is included in the coding parameters: last valid coefficient flag, coding flag in coefficient group unit, last valid coefficient position, flag indicating whether the coefficient value is greater than 1, flag indicating whether the coefficient value is greater than 2, flag indicating whether the coefficient value is greater than 3, information about remaining coefficient values, symbol information, reconstructed luminance sample, reconstructed chrominance sample, residual luminance sample, residual chrominance sample, luminance transform coefficient, chrominance transform coefficient, luminance quantization level, chrominance quantization level, transform coefficient level scanning method, size of the decoder-side motion vector search area, form of the decoder-side motion vector search range, number of decoder-side motion vector searches, CTU size information, minimum block size information, maximum block size information, maximum block depth information, minimum block depth information, stripe identification information, stripe partition information, parallel block identification information, parallel block type, parallel block partition information, input sample bit depth, reconstructed sample bit depth, residual sample bit depth, transform coefficient bit depth or quantization level bit depth, and values ​​that can be derived from these values.

[0838] The matching techniques described in this specification are methods that calculate error costs by adjusting the position of templates defined between comparison targets. Matching techniques include template matching and two-sided matching.

[0839] Figure 18 illustrates a template configuration method for template matching based on intra / inter-frame prediction.

[0840] Figure 19 illustrates a template configuration method for two-sided matching.

[0841] Figure 20 illustrates a method for configuring a template by clearing at least one pixel or line in template matching and two-sided matching.

[0842] As shown in Figure 18, template matching can configure the template (current template) for the current block by using the neighboring pixels of the current block, and can configure the reference template that matches the current template by using the pixels in the search range of the reference image.

[0843] As shown in Figure 19, two-sided matching can configure the template using pixels from the reference image used. In this case, the template can be configured with at least one pixel from the pixels reconstructed in the current image or the reference image.

[0844] When configuring a template in a matching technique, as shown in Figure 20, at least one pixel or at least one line can be configured by leaving at least one pixel or at least one line blank.

[0845] In addition, the number of pixels, the number of lines, and the form of the template can be configured based on different encoding parameters.

[0846] As an example, the template format can be configured differently based on the partitioning form or size of the current block or neighboring blocks and their statistical values.

[0847] For example, a template can be configured with a surface whose dimensions are the same as those of the surface that contacts the current block or a neighboring block.

[0848] For example, a template can be configured with a surface whose dimensions are smaller than those of the surface that contacts the current block or a neighboring block.

[0849] For example, a template can be configured with a surface whose dimensions are larger than the dimensions of the surface that contacts the current block or a neighboring block.

[0850] For example, a template can be configured by setting the maximum, minimum, and median values ​​of the dimensions of each surface of the current block and neighboring blocks as the dimensions of the surface.

[0851] For example, when the size of the current block or neighboring blocks is less than a threshold, the template can be configured by reducing lines or pixels, or the size of the template can be reduced.

[0852] For example, when the size of the current block or a neighboring block is less than a threshold, the template can be configured by adding lines or pixels, or the size of the template can be increased.

[0853] For example, when the size of the current block or neighboring blocks is greater than a threshold, the template can be configured by reducing lines or pixels, or the size of the template can be reduced.

[0854] For example, when the size of the current block or a neighboring block is greater than a threshold, the template can be configured by adding lines or pixels, or the size of the template can be increased.

[0855] The threshold can be a value pre-configured in the encoder / decoder or a value sent from the encoder to the decoder by a signal.

[0856] As an example, it can be configured differently based on the motion information of neighboring blocks or their statistical values.

[0857] For example, when one of the statistical values ​​of the motion vectors of neighboring blocks is less than a threshold, the template can be configured by reducing lines or pixels, or the size of the template can be reduced.

[0858] For example, when one of the statistical values ​​of the motion vectors of neighboring blocks is less than a threshold, the template can be configured by adding lines or pixels, or the size of the template can be increased.

[0859] For example, when one of the statistical values ​​of the motion vectors of neighboring blocks is greater than a threshold, the template can be configured by reducing lines or pixels, or the size of the template can be reduced.

[0860] For example, when one of the statistical values ​​of the motion vectors of neighboring blocks is greater than a threshold, the template can be configured by adding lines or pixels, or the size of the template can be increased.

[0861] The threshold can be a value pre-configured in the encoder / decoder or a value sent from the encoder to the decoder by a signal.

[0862] Figure 21 illustrates a method for configuring a template by considering at least one of the prediction patterns of the current block or neighboring blocks.

[0863] As an example, it can be configured differently depending on the prediction mode of the current block or neighboring blocks. (Figure 21) For example, when the current block is performing inter-frame prediction, the template can be configured by selecting at least one of the blocks determined by inter-frame prediction.

[0864] For example, while the current block is performing inter-frame prediction, the template can be configured by selecting at least one of the blocks determined by intra-frame prediction.

[0865] For example, when intra-frame prediction is being performed on the current block (in an inter-frame frame), a template can be configured by selecting at least one of the blocks determined by inter-frame prediction.

[0866] For example, when intra-frame prediction is being performed on the current block (in an inter-frame frame), a template can be configured by selecting at least one of the blocks determined by intra-frame prediction.

[0867] In addition, the number of pixels, the number of lines, and the form of the template can be configured differently based on pixel position, distance between pixels, partitioned areas, etc.

[0868] As an example, the template format can be configured differently depending on the distance from the center of the search range.

[0869] For example, when the distance between the positions where a match is performed from the center of the search range is less than a threshold, the template can be configured by reducing the number of lines or pixels, or the size of the template can be reduced.

[0870] For example, when the distance between the positions where a match is performed from the center of the search range is less than a threshold, the template can be configured by adding lines or pixels, or the size of the template can be increased.

[0871] For example, when the distance between the positions where a match is performed from the center of the search range is greater than a threshold, the template can be configured by reducing the number of lines or pixels, or the size of the template can be reduced.

[0872] For example, when the distance between the positions where a match is performed from the center of the search range is greater than a threshold, the template can be configured by adding lines or pixels, or the size of the template can be increased.

[0873] The threshold can be a value pre-configured in the encoder / decoder or a value sent from the encoder to the decoder by a signal.

[0874] As an example, the template form can be configured differently based on the distance to the pixel that was first matched.

[0875] For example, when the distance to the pixel that first performs the match is less than a threshold, the template can be configured by reducing the number of lines or pixels, or the size of the template can be reduced.

[0876] For example, when the distance to the pixel that was first matched is less than a threshold, the template can be configured by adding lines or pixels, or the size of the template can be increased.

[0877] For example, when the distance to the pixel that first performs the match is greater than a threshold, the template can be configured by reducing the number of lines or pixels, or the size of the template can be reduced.

[0878] For example, when the distance to the first pixel to be matched is greater than a threshold, the template can be configured by adding lines or pixels, or the size of the template can be increased.

[0879] The threshold can be a value pre-configured in the encoder / decoder or a value sent from the encoder to the decoder by a signal.

[0880] As an example, when the search scope is divided into multiple partitioned regions, the template format can be configured differently for each partitioned region.

[0881] Additionally, the number of pixels, the number of lines, and the form of the template can be configured differently based on the steps involved in performing the matching.

[0882] For example, when trying to find neighboring blocks with the minimum error cost by using template matching, template matching can be performed only at specific locations within the style determined in the search range, similar to style matching methods, in order to reduce complexity.

[0883] In this case, the style matching method can be performed as follows.

[0884] In the first step, sparse matching can be performed by expanding the styles within the search range or by increasing the distance between the locations where a match will be performed.

[0885] In the next step, more precise matching can be performed by reducing the style or decreasing the distance between the positions to be matched, based on the position with the minimum error cost found in the previous step.

[0886] In this scenario, in the first step, the template can be configured by reducing lines or pixels, or by reducing the size of the template. Then, in the next step, the template can be configured by increasing lines or pixels, or by increasing the size of the template.

[0887] Additionally, in the first step, the template can be configured by adding lines or pixels, or the size of the template can be increased. Then, in the next step, the template can be configured by reducing lines or pixels, or the size of the template can be decreased.

[0888] Furthermore, matching based on matching techniques can be performed at at least one pixel location within the search range.

[0889] For example, it can be performed at at least one of the positions derived from the candidate list.

[0890] For example, it can be executed at at least one of the pixel positions at a specific distance from the current block.

[0891] For example, it can be executed at at least one of the pixel locations at a specific distance from the current block.

[0892] For example, when the search scope is partitioned, it can be executed at at least one location within a specific partition region.

[0893] For example, it can be performed at at least one of the pixel positions belonging to a specific style.

[0894] For example, it can be performed at at least one of the pixel positions derived from blocks (similar candidates) whose coding parameters are the same as or similar to those of the current block. Blocks whose coding parameters are the same as or similar to those of the current block may include blocks when 1) the statistics of the motion information of the corresponding block are the same as or similar to those of the motion information of the current block, 2) the prediction information (intra-frame, inter-frame) of the corresponding block is the same as or similar to the prediction information of the current block, 3) the prediction direction of the corresponding block is the same as or similar to that of the current block, and 4) the intra-frame prediction mode of the corresponding block is the same as or similar to that of the current block.

[0895] The same processing can be performed based on the statistical values ​​of the candidate encoding parameters.

[0896] The threshold can be a value pre-configured in the encoder / decoder or a value sent from the encoder to the decoder by a signal.

[0897] Figure 22 shows a flowchart of an image encoding / decoding method for generating chroma prediction signals using the inter-component prediction model of the present invention.

[0898] [E1 / D1] Steps for determining the chromaticity signal prediction method / model: When determining the chromaticity signal prediction method / model, the prediction signal (or prediction block) can be generated based on the correlation between chromaticity components.

[0899] In chromaticity signal prediction based on color relationships, the predicted signal can be derived based on the cross-component linear model (CCLM).

[0900] The chromaticity signal prediction method based on the cross-component linear model can generate a chromaticity prediction block for chromaticity prediction by using the sample values ​​of the reconstructed chromaticity samples adjacent to the current chromaticity block and the luminance samples reconstructed at their corresponding positions.

[0901] Here, a sample may include a pixel, a statistical value of the pixel value of at least one pixel, or encoded information derived from at least one pixel.

[0902] In this case, statistical values ​​(such as maximum and minimum values) or the correlation between neighboring pixels of the luminance block and neighboring pixels of the chrominance block can be used to derive the linear regression model in equation (1) below and calculate the model coefficients a and b. These two values, along with the pixels within the luminance block, can then be used to derive the chrominance block to be predicted.

[0903] Equation (1): C'(i,j) = a L'(i,j) + b or C'(i,j) = a In this case, L(x,y) + b, C(i,j) can refer to a chroma prediction block or a sample within the current chroma prediction block, and L'(i,j) can refer to the reconstructed luminance block / sample corresponding to the position of the chroma block to be predicted. However, L'(i,j) can refer to a luminance sample whose size has been adjusted to be the same as the size of the chroma block through subsampling, downsampling, etc., and L(x,y) can refer to a luminance sample whose size is different from the size of the chroma block. Therefore, the (x,y) of the luminance sample L(x,y) corresponding to the chroma sample C'(i,j) can be different from the position (i,j) of the chroma sample.

[0904] When deriving the model coefficients in equation (1), the model coefficients can be derived not only for each of the U and V signals, but also for both the U and V signals.

[0905] Whether to apply a linear model can be determined for each of U and V, which can be done by sending a signal.

[0906] To derive the values ​​of model coefficients a and b, information about at least one of the following can be used: the neighboring samples of the chroma block to be predicted, the neighboring samples of the luminance block corresponding to the position of the chroma block to be predicted, and the reference line index of the corresponding luminance block.

[0907] In this case, when determining the reference lines used to derive the model coefficients, at least N or more lines can be used as reference lines. Here, N is a positive number greater than or equal to 0, and can be a value pre-configured in the encoder / decoder or a value sent from the encoder to the decoder by a signal. It can be sent from the encoder.

[0908] In this case, the location of the neighboring blocks or sample points used can vary depending on the chroma format.

[0909] For example, in YUV4:4:4 or YUV:4:2:2 formats, the adjacent sample lines of the luma block, indicated by the reference line index of the luma block, and the adjacent neighboring sample lines of the chroma block at the same position can be used.

[0910] For example, in the YUV4:2:0 format, the adjacent sample line index (IntraChromaRefLineIdx) of the chroma block to be used can be derived as the floor (IntraLumaRefLineIdx / 2) value.

[0911] Furthermore, based on the subsampling method, the number of luminance samples can be matched to the number of chrominance samples using only an even or odd number of samples from the reconstructed luminance reference line. (The luminance sample corresponding to the chrominance sample can be determined.) Additionally, the luminance sample at a specific location can be determined as the sample corresponding to the chrominance sample. For example, when the luminance block size is 2M×2N and the chrominance block size is M×N, the sample at location (M+N) within the luminance block... X1 / 8, (M+N) X2 / 8, (M+N) X3 / 8 and (M+N) The sample at X4 / 8 can be identified as the luminance sample corresponding to the chrominance sample. In this case, Xi (i=1, 2, 3) is an arbitrary constant used to determine the sample position and can increase with the number of samples. For example, when using 4 samples as described above, the constants substituted from X1 to X4 can be arbitrarily determined as (1, 3, 5, 7), (2, 4, 6, 8), etc., and the values ​​can vary depending on the encoding information about neighboring blocks.

[0912] Alternatively, a luminance sample corresponding to one chrominance sample can be derived using N luminance samples. In this case, N can be any positive integer.

[0913] Figure 23 shows examples of downsampling filter coefficients and DNN-based downsampling.

[0914] For example, downsampling can be used to determine the luminance sample corresponding to the chrominance sample. In this case, downsampling can be performed not only based on the various filters shown in Figure 23, but also through neural network-based filtering.

[0915] In this case, at least one of the statistical values ​​(such as average, maximum, minimum, median, etc.) of at least one of the N luminance samples can be calculated to derive and use a representative luminance sample.

[0916] For example, when the luminance samples corresponding to the adjacent sample c(0,-1) of the current chroma block are L(-1,-2), L(0,-2), L(1,-2), L(-1,-1), L(0,-1), and L(1,-1), candidate values ​​can be derived using at least one of these samples. These candidate values ​​are used to derive representative values ​​through average, maximum, minimum, median, etc. Here, the candidate values ​​used to obtain representative values ​​can refer to statistical values. Here, the representative value can refer to the sample value of the luminance sample corresponding to the chroma sample used to derive the model coefficients. Specifically, the reference line of the chroma block can be derived using the following equation: c(0,-1)=(L(-1,-2)+2 L(0,-2)+L(1,-2)+L(-1,-1)+2 L(0,-1)+L(1,-1)+4)>>3. In this case, >> can refer to the right shift operator.

[0917] The luminance samples corresponding to the chrominance samples can be different, and the formulas used to derive the chrominance prediction samples can also be different. Information about the luminance samples corresponding to the neighboring samples of the current chrominance block, or information about the formulas used to derive the neighboring samples of the current chrominance block, can be sent via signals in the encoder.

[0918] The sample position can be a pre-configured value in the encoder / decoder or a value sent from the encoder to the decoder via a signal. It can be sent from the encoder. Alternatively, it can be derived from the encoding information of neighboring samples or blocks.

[0919] The values ​​of model coefficients a and b can be derived to generate the chromaticity block that will be predicted by equation (1). In this case, the values ​​of a and b can be derived by using at least one of N selected or derived representative values ​​(first representative value, second representative value, third representative value, and fourth representative value), as in equations (2) to (4). In this case, N can be a positive integer, or it can be at least 2, or it can be 4.

[0920] Equation (2): a = (CMAX - CMIN) / (LMAX - LMIN) Equation (3): b = CMIN - a In this case, the first representative value (LMAX) can represent the x-axis of the maximum value, the second representative value (CMAX) can represent the y-axis of the maximum value, the third representative value (LMIN) can represent the x-axis of the minimum value, and the fourth representative value (CMIN) can represent the y-axis of the minimum value.

[0921] For example, equation (2) may be difficult to implement for integer units because it uses division. Therefore, division can be replaced by multiplication and >> (shift operation).

[0922] As another example, equation (4) avoids the division operation in equation (2) and derives the value of a by subtraction using Log2.

[0923] Equation (4): a = Log2(CMAX-CMIN) - Log2(LMAX-LMIN)>>k In this case, since it is derived from Log2, the value of k can be used in >> (shift operation) to normalize the value of a.

[0924] In chromaticity signal prediction based on cross-component linear models, chromaticity signals / samples / blocks can be derived by using multiple models.

[0925] When deriving multiple models, they can be derived from brightness samples at different locations.

[0926] For example, one model can be derived from even-numbered brightness samples in a reference line, and another model can be derived from odd-numbered samples.

[0927] For example, the sample points at the top (above) position and the sample points at the left position of the brightness block can be divided, and each divided sample point can be used to derive different models.

[0928] Optionally, when deriving multiple models, the sample points used to derive model parameters can be divided based on a defined threshold, and different models can be derived using each divided sample point.

[0929] For example, by using a specific value as a threshold, samples with values ​​less than or equal to the threshold and samples with values ​​exceeding the threshold can be separated, and different models can be derived from each separated sample.

[0930] The threshold can be a value pre-configured in the encoder / decoder or a value sent from the encoder to the decoder by a signal.

[0931] Furthermore, when deriving multiple models, the models can be derived based on the statistical values ​​of brightness samples.

[0932] For example, by using the average value of the brightness samples as a threshold, samples with values ​​less than or equal to the threshold and samples with values ​​exceeding the threshold can be separated, and different models can be derived from each separated sample.

[0933] For example, arranging the luminance sample values ​​in an order close to the average value to select N samples from the top and S samples from the bottom can be used to derive different models. In this case, N and S can be pre-configured values ​​in the encoder / decoder or values ​​sent from the encoder to the decoder by signal.

[0934] For example, the range of sample values ​​between the maximum and minimum values ​​can be divided into N groups, and the samples in each group can be used to derive different models. In this case, N can be a pre-configured value in the encoder / decoder or a value sent from the encoder to the decoder by signal.

[0935] Figure 24 is an example of a sample point within a range of sample point values ​​between the minimum and maximum sample point values, divided into groups.

[0936] Figure 25 is an example of generating new groups by using the average of the samples calculated for each group.

[0937] For example, when dividing sample points into groups, different models can be derived using the statistical values ​​of the sample points in each group. In this case, the mean of the sample points in each group can be calculated, and new sample point groups can be generated based on the mean to derive different models. (Figure 24) For example, after calculating the mean of the sample points in each group, sample points with values ​​less than or equal to the mean of the sample point values ​​in each group and sample points with values ​​exceeding the mean can be grouped together to derive different models for each group. (Figure 24) For example, a new model can be derived by assigning the mean of the sample points calculated for each group to a new group. (Figure 25) When dividing sample points (reference area) for deriving multiple models, the sample points used for model derivation can be divided based on the centroid of the sample points.

[0938] For example, the centroids used to divide the sample points into two groups can be derived using the following method.

[0939] Step 1) Calculate the average value (AVG_OLD) of all pixel values.

[0940] Step 2) The samples can be divided into two groups based on the average value (AVG_OLD), and the average value (AVG_G1, AVG_G2) of the samples in each group can be recalculated by using the samples belonging to the group for each group.

[0941] Step 3) The median of the averages calculated in the two groups can be used as the new average (AVG_NEW=(AVG_G1+AVG_G2), and step 2 can be repeated.

[0942] Regarding this, when step 2 begins, AVG_OLD can be updated to AVG_NEW. In this case, the repetition can stop when AVG_NEW equals AVG_OLD.

[0943] Step 4) The AVG_NEW calculated in step 3 can be determined as the centroid.

[0944] Whether to use the method of dividing the sample points by using the centroid can be a value pre-configured in the encoder / decoder, or a value sent from the encoder to the decoder by a signal.

[0945] When determining a method for dividing reference samples (pixels) for generating multiple models, at least one of the proposed methods may be selected and used.

[0946] The proposed method can be used to partition reference regions and derive models for each method, and the optimal method can be determined based on the coding cost calculated in the predictive coding process using the method.

[0947] The model can be derived for each method using the proposed approach and applied to a template configured with neighboring samples of the current block to determine the optimal method based on the computed template matching cost.

[0948] Among the methods for dividing sample points, the best method can be either the values ​​pre-configured in the encoder / decoder or the values ​​of the signals sent from the encoder to the decoder.

[0949] When deriving multiple models, models derived from pre-encoded / decoded blocks can be used.

[0950] For example, the derived model can be stored for each block, and the stored model can be retrieved and used from the current block.

[0951] In this case, the model can be stored for each block, or it can be stored only when the chroma mode of the corresponding block is determined to be the chroma mode based on the inter-component model.

[0952] When determining the reference block for deriving the model (reference storage model coefficients) of the current block, at least one block can be selected and used, either a neighboring block or a non-neighboring block.

[0953] When determining the reference block (reference stored model coefficients) used to derive the model, at least one block from the reference frame can be selected and used.

[0954] When deriving multiple models, you can use models derived from blocks whose dimensions are different from the current block's dimensions.

[0955] For example, when the location of the reference region or reference points varies depending on the block size and therefore different models are generated, a model generated from a larger block can be used with a smaller block. Conversely, a model derived from a smaller block can be used with a larger block.

[0956] When deriving multiple models, models derived from blocks at different locations within a partitioned block can be used.

[0957] For example, when the location of a reference region or reference sample point differs depending on the block location within a partitioned block and thus different models are generated, a model derived from a block at a different location within the partitioned block can be used in the current block.

[0958] When deriving multiple models, a new model can be derived by modifying the coefficient values ​​of the previously derived (calculated / derived) model.

[0959] In this case, it can be derived by modifying at least one of the coefficients, biases, or nonlinear terms of the pre-derived (computed) model.

[0960] For example, in the model C'(i,j) = a L'(i,j) + b in the above equation (1), a new model can be generated by modifying the value of a or b, which are the model coefficients and the bias value, respectively.

[0961] Here, when the values obtained by modifying a and b are a' and b', a' and b' can be configured in the form of adding, subtracting, multiplying, or dividing a and b by a specific value v.

[0962] For example, a' = a + v, a' = a - v, a' = a v, or a' = a / v.

[0963] Here, a and b can be used to configure the formula for obtaining a' and b' in polynomial form.

[0964] For example, a'=a t + s.

[0965] Coefficients such as v, t, s, etc. used to derive a' and b' can be values pre-configured in the encoder / decoder or values signaled from the encoder to the decoder.

[0966] In this case, when the ranges of the values of a' and b' are W, X, Y, Z (W < a' < X, Y < b' < Z), W, X, Y, Z can be values pre-configured in the encoder / decoder or values signaled from the encoder to the decoder.

[0967] When deriving multiple models, a new model can be derived through the weighted sum of the coefficient values of pre-derived (calculated) models.

[0968] In this case, at least one of the coefficient values, bias, or non-linear terms for configuring the new model can be derived through the weighted sum of the coefficient values, bias, or non-linear terms of pre-derived (calculated) models.

[0969] For example, assume the pre-derived models are as follows.

[0970] Model 1: C'(i,j) = a L'(i,j) + b Model 2: C'(i,j) = c L'(i,j) + d In this case, assume the newly generated Model 3 using the above Model 1 and Model 2 is as follows.

[0971] Model 3: C'(i,j) = e L'(i,j) + f In this case, the coefficients e and bias f of Model 3 can be generated through the weighted sum of the model coefficients and bias of Model 1 and Model 2.

[0972] e = w a + (1-w) c, f = w b + (1-w) d(w: weight) In this case, the weight w can be a value pre-configured in the encoder / decoder or a value sent from the encoder to the decoder by signal.

[0973] The method used to derive models can be used not only to derive multiple models, but also to derive a single model for the current block.

[0974] For example, instead of calculating (deriving) a new model by using reference samples in the current block, a model pre-calculated in at least one of the blocks in an adjacent or neighboring block or a reference frame can be used as the model for the current block.

[0975] For example, instead of calculating (deriving) a new model by using neighboring samples in the current block, a model derived by a weighted sum of the coefficient values ​​of a pre-derived (calculated) model can be used as the model for the current block.

[0976] For example, instead of calculating (deriving) a new model by using neighboring samples in the current block, a model created by modifying the coefficient values ​​of a pre-derived (calculated) model can be used as the model for the current block.

[0977] In chromaticity signal prediction based on multiple models, the model information used to generate the chromaticity signal / block from the multiple models can be transmitted by signaling or derived using proximity coding information. In this case, the model information may include the equations required to derive the chromaticity signal, as well as the coefficients and constants included therein.

[0978] In this case, for example, we can assume the following derivation of four models.

[0979] Model 1: C'(I,j) = a Model 2: L'(I,j) + b: C''(I,j) = c L'(I,j) + d model 3: C'''(I,j) = e Model 4: L'(I,j) + f: C''''(I,j) = g L'(I,j) + h (a method for sending model information using signals) can be used to send model information used in generating the final chroma signal / block when generating chroma signals / blocks based on multiple models.

[0980] Figure 26 shows an example of a chroma prediction block and a luminance block that directly corresponds to the chroma block.

[0981] Figure 27 shows an example of a chromaticity prediction block.

[0982] Specifically, Figure 27 can 1) use the model to generate prediction samples (C', C'', C''', C'''') or chromaticity prediction blocks (BLK_C', BLK_C'', BLK_C'''', BLK_C'''') for the corresponding region, 2) select the chromaticity prediction blocks and models with the minimum error cost through chromaticity prediction, and 3) send the corresponding model information from the encoder to the decoder using signals.

[0983] In this case, after using multiple models to generate chromaticity prediction blocks and performing chromaticity prediction as shown in Figure 26 / Figure 27, the error cost can be calculated.

[0984] In this case, the final model can be determined based on the error cost, and the corresponding model information can be sent using signals.

[0985] For example, the model with the minimum error cost can be selected as the final prediction model.

[0986] In this case, the error costs of U and V can be calculated separately or together.

[0987] (Methods for deriving model information) When generating chroma signals / blocks based on multiple models, the final model for generating the final chroma block can be derived by performing chroma prediction through model derivation and model application in the reference area.

[0988] For example, to derive the final model information, the region in the reference region used for deriving the model and the region used for calculating the error cost by applying the derived model can be used separately.

[0989] Figure 28 shows an example of the region in the reference area used to derive the model and the region required to calculate the error cost between the chromaticity prediction signal and the chromaticity signal.

[0990] In this case, as shown in Figure 28, it can be assumed that the region in the reference region used for deriving the model is R1, and the reference region required to generate the chromaticity prediction signal by applying the derived model and to calculate the error cost of the chromaticity signal is R2.

[0991] In this scenario, at least one model can be derived using luminance samples and their corresponding chrominance samples in R1, and each derived model can be applied to luminance samples in region R2 to generate chrominance prediction samples (C', C'', C''', C'''') for each model. In this case, the chrominance prediction samples can also be combined and configured into a template (PRED_TMP_C', PRED_TMP_C'', PRED_TMP_C''', PRED_TMP_C'''). Figure 29 illustrates an example of calculating the error cost between the chrominance template and the chrominance prediction template generated for each model. In this scenario, as shown in Figure 29, the final prediction model can be derived by calculating the error cost between the chrominance prediction samples (or chrominance prediction template) generated for each model and the reconstructed chrominance samples (or chrominance template) in region R2.

[0992] For example, the model with the minimum error cost can be selected as the final prediction model.

[0993] In this case, the error costs of U and V can be calculated separately or together.

[0994] The size, location, and form of regions R1 and R2 can be variable. In this case, the corresponding information can be values ​​pre-configured in the encoder / decoder or values ​​sent from the encoder to the decoder via signals.

[0995] Regions R1 and R2 can be determined based on template matching. For example, as in TIMD, a reference region (R1, template) is determined by using a template to generate a prediction block (template), which can be compared with the adjacent templates (R2) of the corresponding block to calculate the error cost.

[0996] In chromaticity signal prediction based on multiple models, the number of bits required to send model information via signals can be reduced by using methods for deriving model information.

[0997] For example, when multiple models exist, a prediction pattern index (model index) can be defined for each model in the form of a table. In this case, the table could be called the "chromaticity prediction pattern table".

[0998] In this context, as in methods for deriving model information, chromaticity prediction samples (C', C'', C''', C'''') can be generated for each model around a chromaticity block, and the error cost between the generated chromaticity prediction samples and the reconstructed chromaticity samples (C) around the chromaticity block corresponding to the generated chromaticity prediction samples can be calculated.

[0999] In this case, the prediction pattern index can be arranged and used based on the magnitude of the calculated error cost.

[1000] Figure 30 shows an example of reconfiguring a model-based chromaticity prediction mode table.

[1001] In this case, for example, when models with small error costs are arranged to have high index values ​​on the "Model-Based Chromaticity Prediction Mode Table" in Figure 30, chromaticity prediction blocks can be generated only for the top N models, and the model information finally determined by chromaticity prediction can be signaled, saving signal transmission bits.

[1002] In this context, the decision to include a model in the chromaticity prediction block generation and model-based chromaticity prediction mode table can be based on the statistics of the error cost.

[1003] For example, only models with error costs less than or equal to the average error cost can be selected to generate chromaticity prediction pattern tables and chromaticity prediction blocks.

[1004] For example, an S-model with a small variance can be selected to generate a chromaticity prediction pattern table and chromaticity prediction blocks.

[1005] Here, N and S can be values ​​that are pre-configured in the encoder / decoder as being greater than or equal to 1 and less than or equal to 4, or values ​​that are sent from the encoder to the decoder by signal.

[1006] In chromaticity signal prediction based on the correlation between color coding information, the predicted signal can be derived based on the gradient linear model (GLM) representing the change of sample values ​​or statistical values ​​of sample values ​​in various coding information.

[1007] The method for predicting chroma signals using a gradient-value-based inter-color linear model generates chroma prediction blocks / signals by using gradient values ​​calculated using reconstructed chroma samples adjacent to the current chroma block and reconstructed luminance samples at positions corresponding to the reconstructed chroma samples.

[1008] Here, a sample may include a pixel, a statistical value of the pixel value of one or more pixels, or encoded information derived from one or more pixels.

[1009] In this case, the correlation or statistical values ​​(such as maximum and minimum values) between neighboring pixels of the luminance block and neighboring pixels of the chrominance block can be used to derive the linear regression model in equation (5) below and calculate the values ​​of model coefficients a and b. Furthermore, these two values ​​and the gradient values ​​of pixels within the luminance block can be used to derive the chrominance signal / block to be predicted.

[1010] Equation (5): C'(I,j) = a In this case, C(I,j) can refer to the chroma block to be predicted or the pixel value within the chroma block, and G(I,j) can refer to the gradient sample derived from the reconstructed luminance block corresponding to the position of the chroma block to be predicted.

[1011] When deriving the model coefficients in equation (5), the model coefficients can be derived not only for each of the U and V signals, but also for both the U and V signals.

[1012] Whether to apply a linear model can be determined for each of U and V, which can be done by sending a signal.

[1013] Figure 31 shows an example of a gradient detection filter or gradient detection style.

[1014] When detecting gradient values, as shown in Figure 31, multiple gradient detection filters can be used to derive the gradient samples corresponding to the chromaticity samples. In this case, the number of filters used can be at least one, and the number of filters can be a value pre-configured in the encoder / decoder or a value sent from the encoder to the decoder by a signal.

[1015] In this case, gradient values ​​can be detected separately for each filter that detects gradient values, so as to derive the corresponding gradient linear model.

[1016] For example, suppose the gradient values ​​detected by multiple gradient detection filters are used to derive multiple gradient linear models as follows. (This example derives four models using four gradient detection filters.) Model 1: C'(i,j) = a G(i,j) + b → Using filter 1, derive gradient value model 2: C'' (i,j) = c G(i,j) + d → Derive gradient value model 3 using filter 2: C''' (i,j) = e G(i,j) + f → Using filter 3 to derive gradient value model 4: C'''' (i,j) = g G(i,j) + h → Using filter 4 to derive gradient values. In chromaticity signal prediction using gradient linear models, chromaticity signals / samples / blocks can be derived by using multiple gradient-based models.

[1017] In deriving multiple models, gradient samples can be derived from brightness samples at different locations.

[1018] For example, gradient samples can be generated from even-numbered brightness samples on a reference line, and a model can be derived using these gradient samples. In the other case, gradient samples can be generated from odd-numbered brightness samples, and another model can be derived using these gradient samples.

[1019] For example, the sample points at the top (above) position of the brightness block and the sample points at the left position of the brightness block can be divided. Each divided sample point can be used to generate gradient sample points, and the generated gradient sample points can be used to derive different models.

[1020] In addition, when deriving multiple models, the models can be derived based on the statistical values ​​of gradient samples.

[1021] For example, by using the average value of the gradient samples as a threshold, samples with values ​​less than or equal to the threshold and samples with values ​​exceeding the threshold can be separated, and different models can be derived from each separated gradient sample.

[1022] For example, arranging gradient sample values ​​in order close to their average value to select N samples from the top and S samples from the bottom can be used to derive different models. In this case, N and S can be pre-configured values ​​in the encoder / decoder or values ​​sent from the encoder to the decoder by signal.

[1023] Alternatively, when deriving multiple models, the models can be derived based on defined thresholds.

[1024] For example, by using a specific value as a threshold, samples with values ​​less than or equal to the threshold and samples with values ​​exceeding the threshold can be separated, and different models can be derived from each separated sample.

[1025] The threshold can be a value pre-configured in the encoder / decoder or a value sent from the encoder to the decoder by a signal.

[1026] Multiple models can be derived using gradient samples derived through multiple gradient detection filters.

[1027] In this case, multiple models can be derived using gradient samples passed through a gradient detection filter.

[1028] When deriving multiple models, models derived from pre-encoded / decoded blocks can be used.

[1029] When deriving multiple models, you can use models derived from blocks whose dimensions are different from the current block's dimensions.

[1030] When deriving multiple models, models derived from blocks at different locations within a partitioned block can be used.

[1031] When deriving multiple models, a new model can be derived by modifying the coefficient values ​​of the pre-derived (calculated) model.

[1032] When deriving multiple models, a new model can be derived by weighting and summing the coefficient values ​​of the pre-derived (calculated) models.

[1033] The methods used to derive models can be used not only to derive multiple models, but also to derive a single model for the current block.

[1034] The methods for deriving multiple models and the methods for deriving a model for the current block can be used to derive the model by using at least one of the various methods for deriving the model used in "Cross-Component Linear Model (CCLM)".

[1035] Information about the filters among multiple gradient detection filters used to derive the model for generating the final chromaticity prediction block can be transmitted using signals or by using encoded information.

[1036] In this case, the number of filters used to derive the model that generates the final chromaticity prediction block can be at least one, and the number of filters can be a value pre-configured in the encoder / decoder or a value sent from the encoder to the decoder by a signal.

[1037] (A method for sending gradient detection filter information using signals) can be used to send information about the filters among multiple gradient detection filters used to derive the model for generating the final chromaticity prediction block.

[1038] Figure 32 shows the gradient block derived through four gradient detection filters.

[1039] In this case, multiple gradient detection filters can be used to derive the gradient linear model corresponding to each filter (as shown in Figure 32), and each model can be used to generate chromaticity prediction blocks.

[1040] In this case, after performing chromaticity prediction using the generated chromaticity prediction block, the error cost can be calculated, and based on the error cost, the gradient linear model used to generate the final chromaticity prediction block and the gradient detection filter used to derive the corresponding model can be derived.

[1041] In this case, for example, the gradient detection filter used to generate the chromaticity prediction block with the minimum error cost can be determined as the final gradient detection filter, and the corresponding filter information can be sent by signal.

[1042] In this case, the error costs of U and V can be calculated separately or together.

[1043] In this scenario, for the gradient detection filter, a different filter can be used for each of U and V, or the same filter can be used for both U and V. In this case, the filter information used can be transmitted via a signal.

[1044] (Method #1 for deriving gradient detection filter information) When searching for filter information among multiple gradient detection filters to derive the model required to generate the final chroma prediction block, the gradient detection filter used to derive the final model can be derived by performing chroma prediction through model derivation and model application in the reference area.

[1045] When deriving information about the gradient detection filter to derive the model used to generate the final chromaticity prediction block, a reference region (reference gradient sample) can be used alone.

[1046] In this case, when dividing the reference region, it is possible to divide the region into a region for deriving the gradient linear model corresponding to each filter through multiple gradient filters and a region for applying the derived model to generate the chromaticity prediction signal (or sample points) and calculate the error cost with respect to the chromaticity signal.

[1047] Figure 33 shows the reference region for deriving the gradient block derived through four gradient detection filters and the reference region required to calculate the error cost.

[1048] In this case, as shown in Figure 33, it can be assumed that the region used for deriving the model in the reference region is R1, and the reference region required to generate the chromaticity prediction signal by applying the derived model and calculating its error cost is R2.

[1049] In this scenario, multiple gradient detection filters can be used to generate gradient blocks. At least one model can be derived using gradient samples in region R1 of each gradient block and their corresponding chromaticity samples. Each derived model can be applied to gradient samples in region R2 to generate chromaticity prediction samples (C', C'', C''', C'''') for each model. In this case, the chromaticity prediction samples can also be combined and configured into a template. (PRED_TMP_C', PRED_TMP_C'', PRED_TMP_C''', PRED_TMP_C''') Figure 34 illustrates an example of calculating the error cost of a chromaticity prediction template derived from four gradient-based linear models and a chromaticity template.

[1050] Specifically, in Figure 34, [Step 1, Generating Gradient Samples] multiple gradient detection filters (styles) can be used to generate gradient samples (or gradient templates) for each filter. [Step 2, Deriving Model Coefficients] the model coefficients of the gradient linear model can be derived by using samples included in the reconstructed R1 region from the gradient samples directly corresponding to the chroma block and samples included in the R1 region from the reconstructed neighboring reference samples of the chroma block. In this case, when multiple gradient detection filters (styles) are used, a gradient linear model corresponding to each filter can be derived. In this case, multiple gradient linear models can be derived by dividing the samples used to derive the model coefficients (e.g., even encoding, odd encoding). [Step 3, Generating Chroma Prediction Samples (Template) Using the Model] the derived model can be used to generate chroma prediction samples (C', C'', C''', C'''') corresponding to the R2 region from the reconstructed neighboring reference samples of the chroma block. In this case, the chroma prediction samples can also be combined to configure a template. [Step 4, Calculate Error Cost] The error cost between the generated chromaticity prediction template (PRED_TMP_C', PRED_TMP_C'', PRED_TMP_C''', PRED_TMP_C'''') and the chromaticity template (TMP_C) generated using samples included in the R2 region from the reconstructed neighboring reference samples of its corresponding chromaticity block can be calculated. [Step 5, Select Model and Perform Chromaticity Prediction] A model can be selected based on the error cost, and a chromaticity prediction block can be generated to perform chromaticity prediction. In this case, the filter ultimately used among multiple gradient detection filters (styles) is the gradient detection filter used to derive the finally determined gradient linear model. In other words, it is the filter used to generate the gradient samples required to derive the finally determined gradient linear model.

[1051] In this case, as shown in Figure 34, the error cost between the chromaticity prediction samples (or chromaticity prediction templates) generated for each model and the reconstructed chromaticity samples (or chromaticity templates) in the R2 region can be calculated to determine the final gradient linear model.

[1052] For example, the model with the minimum error cost can be selected as the final prediction model.

[1053] In this case, the error costs of U and V can be calculated separately or together.

[1054] In this case, the gradient prediction filter used to derive the final model can be determined as the final gradient detection filter.

[1055] The size, location, and form of regions R1 and R2 can be variable. In this case, the corresponding information can be values ​​pre-configured in the encoder / decoder or values ​​sent from the encoder to the decoder via signals.

[1056] Regions R1 and R2 can be determined based on template matching. For example, as in TIMD, a reference region (R1, template) is determined by using a template to generate a prediction block (template), which can be compared with the adjacent templates (R2) of the corresponding block to calculate the error cost.

[1057] (Method #2 for deriving gradient detection filter information) When searching for information about the filter used to derive the model required to generate the final chroma prediction block among multiple gradient detection filters, it can be found by using intra-frame prediction mode information.

[1058] Figure 35 shows an example of a gradient detection filter (style).

[1059] For example, when a gradient detection filter (style) is designed to facilitate the detection of a specific gradient (gradient, direction) value (as shown in Figure 35), a gradient detection filter can be used to derive a gradient linear model. This gradient detection filter is designed to detect the gradient (gradient, direction) value that is most similar to the gradient (gradient, direction) value of the prediction pattern determined from the direction prediction of the neighboring block or the current block.

[1060] In this case, when the orientation prediction mode of the neighboring block or the current block is vertical, a gradient detection style designed to detect the vertical component can be selected from the given gradient prediction filter (style).

[1061] Similarly, a gradient detection filter can be used to derive a gradient linear model. This gradient detection filter is designed to detect the gradient (gradient, direction) component that is most similar to the gradient (gradient, direction) of the prediction pattern determined by DIMD in the neighboring block or the current block.

[1062] (Method #3 for deriving gradient detection filter information) When searching for filter information among multiple gradient detection filters to derive the model required to generate the final chromaticity prediction block, it can be found by using reconstructed luminance samples within the luminance block.

[1063] For example, samples within a luminance block can be used to obtain a histogram of the gradient values ​​of the corresponding samples (as in DIMD) to determine the prediction mode for the corresponding block, and based on this, a gradient detection filter can be used to derive a gradient linear model, which is designed to detect the gradient (gradient, direction) component that is most similar to the gradient (gradient, direction) of the corresponding prediction mode.

[1064] Furthermore, gradient values ​​or gradient samples can be derived through neural network-based filtering.

[1065] To derive the values ​​of model coefficients a and b, information about at least one of the following can be used: neighboring samples of the chroma block to be predicted; neighboring samples of the luminance block corresponding to the position of the chroma block to be predicted; gradient blocks or gradient samples generated by using samples within the luminance block and neighboring samples; the reference line index of the corresponding luminance block; and gradient values ​​derived from the luminance samples.

[1066] For example, N luminance samples can be used to derive the gradient sample of the luminance sample corresponding to a chrominance sample. In this case, N can be any positive integer.

[1067] For example, at least one of the statistical values ​​(such as average, maximum, minimum, median, etc.) of at least one of the gradient values ​​of N gradient samples can be calculated to derive and use a representative gradient value.

[1068] The luminance samples used to derive the gradient samples corresponding to the chrominance samples can vary, and furthermore, the formula used to derive the chrominance prediction signal can also vary. Information about the luminance samples corresponding to the adjacent samples of the current chrominance block or information about the formulas used to derive the adjacent samples of the current chrominance block can be sent via signals in the encoder.

[1069] The location of the brightness sample point used to derive the gradient sample point, or information about it, can be a pre-configured value in the encoder / decoder...

Claims

1. An image decoding method, the method comprising: Determine the inter-component prediction model for the current block; Chromaticity prediction signals are generated by using a defined inter-component prediction model.

2. The method according to claim 1, wherein, Determining the inter-component prediction model includes: generating a list of candidate mergers for the current block; and selecting at least one candidate from the list of candidate mergers.

3. The method according to claim 2, wherein: The merge candidate list includes candidates derived from neighboring blocks, which are either adjacent blocks that are adjacent to the current block or non-adjacent blocks that are not adjacent to the current block.

4. The method according to claim 3, wherein: The candidate model parameters derived from neighboring blocks are partly derived from the inter-component prediction patterns of neighboring blocks, and the remaining model parameters are derived using information from the current block.

5. The method according to claim 2, wherein: The merged candidate list includes candidates generated by modifying some or all of the model parameters of the candidates included in the merged candidate list based on the reference region of the current block.

6. The method according to claim 5, wherein: The reference region is either an adjacent reference region that is adjacent to the current block or a non-adjacent reference region that is not adjacent to the current block.

7. The method according to claim 2, wherein: The merged candidate list includes new candidates generated by weighting and summing the model parameters of the candidates included in the merged candidate list.

8. The method according to claim 7, wherein: The weights of the weighted sum are determined based on the index values ​​of the candidate list to be merged.

9. The method according to claim 2, wherein: The candidate list for merging includes candidates derived from inter-component prediction models of blocks at locations indicated by the block vectors of the current block's neighboring blocks.

10. The method according to claim 1, wherein: There are multiple prediction models between components.

11. The method of claim 10, wherein: The reference region used to derive the prediction model between multiple components is determined based on the average of samples from multiple groups used to divide all pixel values.

12. The method according to claim 10, wherein: The prediction model between multiple components is derived from different neighboring pre-decoded blocks.

13. The method according to claim 1, wherein: The inter-component prediction model for the current block is derived using neighboring or non-neighboring reference regions of the current block.

14. An image encoding method, the method comprising: Determine the inter-component prediction model for the current block; Chromaticity prediction signals are generated by using a defined inter-component prediction model.

15. A computer-readable recording medium storing a bitstream generated by an image encoding method, wherein, The image encoding method includes: determining the inter-component prediction model for the current block; and generating a chroma prediction signal by using the determined inter-component prediction model.