Method and apparatus for image encoding / decoding, and recording medium

By using subsampling technology to configure templates and cost functions in the decoder-side motion information export method, the problem of limited encoding efficiency in the decoder-side motion information export method is solved, and more efficient image encoding and decoding is achieved.

CN120345249APending Publication Date: 2025-07-18ELECTRONICS & TELECOMM RES INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202380084977.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-11
Filing Date
2023-10-11
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, the encoding efficiency of the decoder-side motion information derivation method is limited, and it is difficult to effectively use inter-frame prediction technology to perform efficient image encoding and decoding.

Method used

Subsampling technology is used to configure templates and cost functions in the decoder-side motion information export method to search and predict motion information to improve encoding and decoding efficiency.

Benefits of technology

The efficiency of image encoding and decoding is improved through subsampling technology, the prediction ability of target blocks is enhanced, and the encoding efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120345249A_ABST
    Figure CN120345249A_ABST
Patent Text Reader

Abstract

Disclosed herein are a method, apparatus, and storage medium for image encoding / decoding. In a typical image encoding / decoding method, a decoder-side motion information derivation method may be used in a limited manner. Thus, improvement in coding efficiency attributable to a decoder-side motion information derivation method may also be limited. In an embodiment, a motion information search method for use in prediction is disclosed. Sub-sampling is applied to various processes and targets used in a motion information search method, such as in a search area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to methods, devices, and storage media for image encoding / decoding.

[0002] This application claims the benefit of Korean Patent Application Nos. 10-2022-0129783, filed on October 11, 2022, and 10-2023-0135246, filed on October 11, 2023, which are hereby incorporated by reference in their entirety into this application. Background Art

[0003] With the continuous development of the information and communication industry, broadcast services supporting high-definition (HD) resolution have become widespread worldwide. Through this widespread adoption, a large number of users have become accustomed to high-resolution and high-definition images and / or videos.

[0004] To meet the users' demand for high definition, a large number of institutions have accelerated the development of next-generation imaging devices. In addition to high-definition TVs (HDTVs) and full high-definition (FHD) TVs, users' interest in ultra-high-definition (UHD) TVs has also increased, where the resolution of a UHD TV is more than four times that of an FHD TV. With this increasing interest, there is now a need for image encoding / decoding techniques for images with higher resolution and higher clarity.

[0005] As image compression techniques, there are various techniques such as inter-frame prediction techniques, intra-frame prediction techniques, transform, quantization techniques, and entropy encoding techniques.

[0006] The inter-frame prediction technique is a technique for predicting the values of pixels included in the current picture using the picture before the current picture and / or the picture after the current picture. The intra-frame prediction technique is a technique for predicting the values of pixels included in the current picture using information about the pixels in the current picture. The transform and quantization techniques can be techniques for compressing the energy of the residual signal. The entropy encoding technique is a technique for assigning short codewords to frequently occurring values and long codewords to less frequently occurring values.

[0007] By utilizing these image compression techniques, data about images can be effectively compressed, transmitted, and stored. Summary of the Invention

[0008] Technical Problem Embodiments are directed to providing a device, method, and storage medium for performing encoding / decoding on a target block using prediction.

[0009] Embodiments are directed to providing a device, method, and storage medium for performing encoding / decoding on a target block by utilizing subsampling for prediction.

[0010] Technical solution According to one aspect, there is provided an image encoding method, including: determining encoding information to be used for encoding a target block; generating encoded encoding information by performing encoding on the encoding information; and generating a bitstream including the encoded encoding information.

[0011] The image encoding method may further include performing inter prediction on the target block.

[0012] A decoder-side motion information derivation method for inter prediction is performed.

[0013] Subsampling may be used in the decoder-side motion information derivation method.

[0014] Subsampling may be used to configure a template in the decoder-side motion information derivation method.

[0015] Subsampling may be used to configure a cost function in the decoder-side motion information derivation method.

[0016] Subsampling may be used to perform a search in the decoder-side motion information derivation method.

[0017] The image encoding method may further include performing prediction on the target block.

[0018] Subsampling may be used for template matching in prediction.

[0019] Subsampling may be used to calculate a cost function between templates for template matching.

[0020] According to another aspect, there is provided an image decoding method, including: obtaining a bitstream including encoded encoding information; generating encoding information by performing decoding on the encoded encoding information; and using the encoding information to determine prediction information to be used for decoding a target block.

[0021] The image decoding method may further include performing inter prediction on the target block.

[0022] A decoder-side motion information derivation method for inter prediction may be performed.

[0023] Subsampling may be used in the decoder-side motion information derivation method.

[0024] Subsampling may be used to configure a template in the decoder-side motion information derivation method.

[0025] Subsampling may be used to configure a cost function in the decoder-side motion information derivation method.

[0026] Subsampling may be used to perform a search in the decoder-side motion information derivation method.

[0027] The image decoding method may further include performing prediction on the target block.

[0028] Subsampling may be used for template matching in the prediction.

[0029] Subsampling may be used to calculate the cost function between templates for template matching.

[0030] According to another aspect, there is provided a computer-readable storage medium for storing a bitstream for image decoding, wherein the bitstream may include encoded coding information.

[0031] The coding information may be generated by decoding the encoded coding information.

[0032] The prediction information to be used for decoding the target block may be determined using the coding information.

[0033] Inter-frame prediction of the target block may be performed.

[0034] A decoder-side motion information derivation method for inter-frame prediction may be performed.

[0035] Subsampling may be used in the decoder-side motion information derivation method.

[0036] Subsampling may be used to configure templates in the decoder-side motion information derivation method.

[0037] Subsampling may be used to configure the cost function in the decoder-side motion information derivation method.

[0038] Subsampling may be used to perform a search in the decoder-side motion information derivation method.

[0039] Prediction of the target block may be performed.

[0040] Subsampling may be used for template matching in inter-frame prediction.

[0041] Advantageous Effects There is provided an apparatus, method, and storage medium for performing encoding / decoding on a target block using prediction.

[0042] There is provided an apparatus, method, and storage medium for performing encoding / decoding on a target block by performing prediction using subsampling. Description of the Drawings

[0043] Figure 1 is a block diagram showing the configuration of an embodiment of an encoding device to which the present disclosure is applied; Figure 2 is a block diagram showing the configuration of an embodiment of a decoding device to which the present disclosure is applied; Figure 3is a diagram schematically showing a partition structure of an image when the image is encoded and decoded; Figure 4 is a diagram showing forms of prediction units (PUs) that a coding unit (CU) can include; Figure 5 is a diagram showing forms of transform units (TUs) that can be included in a CU; Figure 6 shows a division of blocks according to an example; Figure 7 is a diagram for explaining an example of an intra prediction process; Figure 8 is a diagram showing reference sample points used in the intra prediction process; Figure 9 is a diagram for explaining an example of an inter prediction process; Figure 10 shows spatial candidates according to an example; Figure 11 shows an order of adding motion information of spatial candidates to a merge list according to an example; Figure 12 shows transform and quantization processing according to an example; Figure 13 shows diagonal scanning according to an example; Figure 14 shows horizontal scanning according to an example; Figure 15 shows vertical scanning according to an example; Figure 16 is a configuration diagram of an encoding device according to an example; Figure 17 is a configuration diagram of a decoding device according to an example; Figure 18 is a flowchart showing a target block prediction method and a bitstream generation method according to an example; Figure 19 is a flowchart showing a target block prediction method using a bitstream according to an example; Figure 20 shows a motion vector difference in a symmetric motion vector difference mode according to an example; Figure 21 shows syntax elements signaled / encoded / decoded in a symmetric motion vector difference mode according to an example; Figures 22a to 22t shows a subsampling method in bilateral matching according to an example; Figure 23 shows bilateral matching according to an example; Figure 24 shows template matching according to an example; Figures 25a to 25n Shows a subsampling method for an area adjacent to a block according to an example; Figures 26a to 26n Shows a subsampling method for an area adjacent to a block according to an example; Figures 27a to 27j Shows subsampling using sub-block units for an area according to an example; Figures 28a to 28l Shows subsampling using sub-block units for an area adjacent to a block according to an example; Figures 29a to 29f Shows the relationship between a search pattern and a resolution according to an example; Figure 29a Shows a first relationship between a search pattern and a resolution according to an example; Figure 29b Shows a second relationship between a search pattern and a resolution according to an example; Figure 29c Shows a third relationship between a search pattern and a resolution according to an example; Figure 29d Shows a fourth relationship between a search pattern and a resolution according to an example; Figure 29e Shows a fifth relationship between a search pattern and a resolution according to an example; Figure 29f Shows a sixth relationship between a search pattern and a resolution according to an example; Figure 30 Shows a first template configuration method in an affine mode according to an example; Figure 31 Shows a second template configuration method in an affine mode according to an example; Figure 32 Is a first flowchart showing a method for selecting a scheme for configuring a candidate list of motion information according to an example; Figure 33 Is a second flowchart showing a method for selecting a scheme for configuring a candidate list of motion information according to an example; Figure 34 Is a third flowchart showing a method for selecting a scheme for configuring a candidate list of motion information according to an example; Figure 35 Is a fourth flowchart showing a method for selecting a scheme for configuring a candidate list of motion information according to an example; Figure 36 Shows the reconfiguration of a candidate list of motion information according to an embodiment; Figure 37 Shows the setting of a flag depending on the value of a motion information index according to an example; Figure 38 Shows a first setting of an index depending on the value of an index of motion information according to an example; Figure 39 Shows a second setting of an index depending on the value of an index of motion information according to an example; Figure 40 Shows a first setting of an index depending on the value of a reference image index according to an example; Figure 41 Shows a second setting of an index depending on the value of a reference image index according to an example; Figure 42 Shows a setting of a flag depending on the value of a reference image index according to an example; Figure 43 Shows a method for determining the positive or negative sign of each component of a motion vector difference according to an embodiment; Figure 44 Shows a first code for signaling / encoding / decoding a motion vector difference according to an example; Figure 45 Shows a second code for signaling / encoding / decoding a motion vector difference according to an example; Figure 46 Shows a third code for signaling / encoding / decoding a motion vector difference according to an example; Figure 47 Shows a fourth code for signaling / encoding / decoding a motion vector difference according to an example; Figure 48 Shows a fifth code for signaling / encoding / decoding a motion vector difference according to an example; Figure 49 Shows a method for signaling information for a motion information search mode according to an example; Figure 50 Shows a first code indicating encoded information signaled in association with a block partitioning structure according to an example; and Figure 51 Shows a second code indicating encoded information signaled in association with a block partitioning structure according to an example.

[0044] Figure 52 Shows the derivation of a template matching cost according to an example; Figure 53a Shows the prediction of the positive or negative sign of a BVD according to an example; Figure 53b Shows the prediction of the suffix binary bits of a BVD magnitude according to an example; Figure 54Shows the prediction of the MVD positive and negative signs and magnitude suffix binary bits according to an example; Figure 55 Shows the template in the reference picture and the reference sample points in the template according to an example; Figure 56 Shows the template for a block with sub-block motion and the sample points in the template according to an example, where the template uses multiple motion information of the sub-blocks of the current block; Figure 57 Shows the search area for intra-frame template matching according to an example; Figure 58 Shows the adjacent half-pixel positions in eight directions according to an example; Figure 59 Shows the integer bilateral matching for a sub-sampled sub-block according to an example; Figure 60 Shows the search for non-translation parameters according to an example; and Figure 61 Shows the independent bilateral matching search for CPMV according to an example. Detailed Description of the Invention

[0045] The present invention can be variously changed and can have various embodiments. Specific embodiments will be described in detail below with reference to the accompanying drawings. However, it should be understood that these embodiments are not intended to limit the present invention to a specific disclosed form, and they include all changes, equivalent forms, or modifications included within the spirit and scope of the present invention.

[0046] The following exemplary embodiments will be described in detail with reference to the accompanying drawings showing specific embodiments. These embodiments are described such that those of ordinary skill in the art to which the present disclosure pertains can easily implement these embodiments. It should be noted that the various embodiments are different from each other but do not need to be mutually exclusive. For example, the specific shapes, structures, and characteristics described herein can be implemented as other embodiments without departing from the spirit and scope of other embodiments related to one embodiment. In addition, it should be understood that the positions or arrangements of the respective components in each disclosed embodiment can be changed without departing from the spirit and scope of the embodiments. Therefore, the appended detailed description is not intended to limit the scope of the present disclosure, and the scope of the exemplary embodiments is only defined by the appended claims and their equivalents (as long as they are properly described).

[0047] In the drawings, like reference numerals are used to designate the same or similar functions in various aspects. The shapes, sizes, etc. of the components in the drawings may be exaggerated for clarity of description.

[0048] Terms such as "first" and "second" may be used to describe various components, but the components are not limited by the terms. The terms are only used to distinguish one component from another. For example, without departing from the scope of this specification, the first component may be referred to as the second component. Similarly, the second component may be referred to as the first component. The term "and / or" may include a combination of multiple related description items or any one of the multiple related description items.

[0049] It will be understood that when a component is referred to as "connected" or "coupled" to another component, the two components may be directly connected or coupled to each other, or there may be an intermediate component between the two components. On the other hand, it will be understood that when a component is referred to as "directly connected or coupled", there is no intermediate component between the two components.

[0050] The components described in the embodiments are independently shown to indicate different characteristic functions, but this does not mean that each component is formed by a single piece of hardware or software. That is, for convenience of description, multiple components are separately arranged and included. For example, at least two of the multiple components may be integrated into a single component. On the contrary, one component may be divided into multiple components. As long as it does not depart from the essence of this specification, embodiments in which multiple components are integrated or embodiments in which some components are separated are included in the scope of this specification.

[0051] The terms used in the embodiments are only used to describe specific embodiments and are not intended to limit the present invention. Singular expressions include plural expressions unless the context specifically indicates the contrary. In the embodiments, it should be understood that terms such as "including" or "having" are only intended to indicate the presence of features, numbers, steps, operations, components, parts, or combinations thereof, and are not intended to exclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof. That is, in the embodiments, the expression describing a component "including" a specific component means that other components may be included within the scope of the practice of the present invention or the technical spirit of the present invention, but does not exclude the presence of components other than the specific component.

[0052] In the embodiments, the term "at least one" may mean one of one or more quantities (such as 1, 2, 3, and 4). In the embodiments, the term "multiple" may mean one of two or more quantities (such as 2, 3, and 4).

[0053] Some components of the embodiments are not essential components for performing necessary functions, but may be optional components only for improving performance. The embodiments may be implemented by only using the essential components for realizing the essence of the embodiments. For example, a structure including only essential components (excluding optional components only for improving performance) is also included in the scope of the embodiments.

[0054] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings so that those of ordinary skill in the art to which the embodiments pertain can easily implement the embodiments. In the following description of the embodiments, a detailed description of well-known functions or configurations that are considered to obscure the gist of the present specification will be omitted. In addition, the same reference numerals are used throughout the drawings to designate the same components, and repeated descriptions of the same components will be omitted.

[0055] Hereinafter, an "image" may represent a single frame constituting a video, or may represent the video itself. For example, "encoding and / or decoding of an image" may represent "encoding and / or decoding of a video", and may also represent "encoding and / or decoding of any one of a plurality of images constituting a video".

[0056] Hereinafter, the terms "video" and "moving picture" may be used with the same meaning and may be used interchangeably with each other.

[0057] Hereinafter, a target image may be an encoding target image that is a target to be encoded and / or a decoding target image that is a target to be decoded. In addition, the target image may be an input image input to an encoding device or an input image input to a decoding device. Also, the target image may be a current image, that is, a target currently to be encoded and / or decoded. For example, the terms "target image" and "current image" may be used with the same meaning and may be used interchangeably with each other.

[0058] Hereinafter, the terms "image", "frame", "picture", and "screen" may be used with the same meaning and may be used interchangeably with each other.

[0059] Hereinafter, a target block may be an encoding target block (i.e., a target to be encoded) and / or a decoding target block (i.e., a target to be decoded). In addition, the target block may be a current block, that is, a target currently to be encoded and / or decoded. Here, the terms "target block" and "current block" may be used with the same meaning and may be used interchangeably with each other. The current block may represent an encoding target block that is an encoding target during encoding and / or a decoding target block that is a decoding target during decoding. In addition, the current block may be at least one of an encoding block, a prediction block, a residual block, and a transform block.

[0060] Hereinafter, the terms "block" and "unit" may be used with the same meaning and may be used interchangeably with each other. Alternatively, a "block" may represent a specific unit.

[0061] Hereinafter, the terms "region" and "segment" may be used interchangeably with each other.

[0062] In the following embodiments, specific information, data, flags, indices, elements, and attributes may have their respective values. The value "0" corresponding to each of the information, data, flags, indices, elements, and attributes may indicate false, logical false, or a first predefined value. In other words, the values "0", false, logical false, and the first predefined value may be used interchangeably with each other. The value "1" corresponding to each of the information, data, flags, indices, elements, and attributes may indicate true, logical true, or a second predefined value. In other words, the values "1", true, logical true, and the second predefined value may be used interchangeably with each other.

[0063] When variables such as i or j are used to indicate rows, columns, or indices, the value of i can be the integer 0 or an integer greater than 0, or can be the integer 1 or an integer greater than 1. In other words, in an embodiment, each of the rows, columns, and indices can be counted starting from 0, or can be counted starting from 1.

[0064] In an embodiment, the term "one or more" or the term "at least one" may represent the term "multiple". The term "one or more" or the term "at least one" may be used interchangeably with "multiple".

[0065] Next, terms to be used in the embodiments will be described.

[0066] Encoder: An encoder represents a device for performing encoding. That is, an encoder may represent an encoding device.

[0067] Decoder: A decoder represents a device for performing decoding. That is, a decoder may represent a decoding device.

[0068] Unit: A unit may represent a unit of image encoding and decoding. The terms "unit" and "block" may be used with the same meaning and may be used interchangeably with each other.

[0069] – A unit may be an array of samples of M×N. Each of M and N may be a positive integer. A unit generally may represent an array of samples in a two-dimensional form.

[0070] – During the encoding and decoding process of an image, a "unit" may be a region generated by partitioning an image. In other words, a "unit" may be a region specified in an image. A single image may be partitioned into multiple units. Alternatively, an image may be partitioned into sub-parts, and a unit may represent each of the partitioned sub-parts when encoding or decoding the partitioned sub-parts is performed.

[0071] – During the encoding and decoding process of an image, predefined processing may be performed on each unit according to the type of the unit.

[0072] – According to functions, unit types can be classified as macro units, coding units (CUs), prediction units (PUs), residual units, transform units (TUs), etc. Optionally, according to functions, a unit can represent a block, a macro block, a coding tree unit, a coding tree block, a coding unit, a coding block, a prediction unit, a prediction block, a residual unit, a residual block, a transform unit, a transform block, etc. For example, a target unit that is a target of encoding and / or decoding can be at least one of a CU, a PU, a residual unit, and a TU.

[0073] – The term "unit" can represent a block including a luma component block, a chroma component block corresponding to the luma component block, and information on syntax elements for each block, such that the unit is designated as being distinct from a block.

[0074] – The size and shape of a unit can be implemented differently. In addition, a unit can have any one of various sizes and shapes. Specifically, the shape of a unit can include not only a square but also geometric shapes (such as a rectangle, a trapezoid, a triangle, and a pentagon) that can be represented in two dimensions (2D).

[0075] In addition, unit information can include one or more of the type of the unit, the size of the unit, the depth of the unit, the encoding order of the unit, and the decoding order of the unit, etc. For example, the type of a unit can indicate one of a CU, a PU, a residual unit, and a TU.

[0076] – A unit can be partitioned into sub-units, each sub-unit having a size smaller than that of the associated unit.

[0077] Depth: The depth can represent the degree to which a unit is partitioned. In addition, the depth of a unit can indicate the level at which the corresponding unit exists when the unit is represented by a tree structure.

[0078] – Unit partition information can include a depth indicating the depth of the unit. The depth can indicate the number of times the unit is partitioned and / or the degree to which the unit is partitioned.

[0079] – In a tree structure, it can be considered that the depth of the root node is the smallest and the depth of the leaf node is the largest. The root node can be the highest (top) node. The leaf node can be the lowest node.

[0080] – A single unit can be hierarchically partitioned into multiple sub-units, while the single unit has depth information based on a tree structure. In other words, the unit and the sub-units generated by partitioning the unit can respectively correspond to a node and the sub-nodes of the node. Each partitioned sub-unit can have a unit depth. Since the depth indicates the number of times the unit is partitioned and / or the degree to which the unit is partitioned, the partition information of the sub-unit can include information on the size of the sub-unit.

[0081] In a tree structure, the top node may correspond to the initial node before partitioning. The top node may be referred to as the "root node". Additionally, the root node may have the minimum depth value. Here, the depth of the top node may be level "0".

[0082] – A node with a depth of level "1" may represent the unit generated when the initial unit is partitioned once. A node with a depth of level "2" may represent the unit generated when the initial unit is partitioned twice.

[0083] – A leaf node with a depth of level "n" may represent the unit generated when the initial unit is partitioned n times.

[0084] – A leaf node may be the bottom node that cannot be further partitioned. The depth of a leaf node may be the maximum level. For example, the predefined value for the maximum level may be 3.

[0085] – QT depth may represent the depth for quad - partitioning. BT depth may represent the depth for binary - partitioning. TT depth may represent the depth for ternary - partitioning.

[0086] – Sample: A sample may be the basic unit that constitutes a block. Samples can be represented by values from 0 to 2 Bd -1 according to the bit depth (Bd).

[0087] – A sample may be a pixel or a pixel value.

[0088] – Hereinafter, the terms "pixel" and "sample" may be used with the same meaning and may be used interchangeably with each other.

[0089] Coding Tree Unit (CTU): A CTU may be composed of a single - luminance - component (Y) coding - tree block and two chrominance - component (i.e., Cb, Cr) coding - tree blocks related to the luminance - component coding - tree block. Additionally, a CTU may represent the information including the above - mentioned blocks and the syntax elements for each block.

[0090] – Each Coding Tree Unit (CTU) may be partitioned using one or more partitioning methods such as quad - tree (QT), binary - tree (BT), and ternary - tree (TT) to configure sub - units such as coding units, prediction units, and transform units. A quad - tree may represent a quadtree. Additionally, each coding tree unit may be partitioned using one or more partitioning methods with a multi - type tree (MTT).

[0091] – "CTU" may be used as a term to specify a pixel block that serves as a processing unit in image decoding and encoding processes (such as in the case of partitioning an input image).

[0092] Coding Tree Block (CTB): "CTB" can be used as a term to designate any one of a Y coding tree block, a Cb coding tree block, and a Cr coding tree block.

[0093] Neighboring block: A neighboring block (or adjacent block) can represent a block adjacent to a target block. A neighboring block can represent a reconstructed neighboring block.

[0094] Hereinafter, the terms "neighboring block" and "adjacent block" can be used to have the same meaning and can be used interchangeably with each other.

[0095] A neighboring block can represent a reconstructed neighboring block.

[0096] Spatial neighboring block: A spatial neighboring block can be a block that is spatially adjacent to a target block. A neighboring block can include a spatial neighboring block.

[0097] – The target block and the spatial neighboring block can be included in the target picture.

[0098] – A spatial neighboring block can represent a block whose boundary touches the target block or a block located within a predetermined distance from the target block.

[0099] – A spatial neighboring block can represent a block adjacent to a vertex of the target block. Here, a block adjacent to a vertex of the target block can represent a block that is vertically adjacent to a neighboring block horizontally adjacent to the target block or a block that is horizontally adjacent to a neighboring block vertically adjacent to the target block.

[0100] Temporal neighboring block: A temporal neighboring block can be a block that is temporally adjacent to a target block. A neighboring block can include a temporal neighboring block.

[0101] – A temporal neighboring block can include a col block.

[0102] – A col block can be a block in a previously reconstructed col picture. The position of the col block in the col picture can correspond to the position of the target block in the target picture. Optionally, the position of the col block in the col picture can be equal to the position of the target block in the target picture. A col picture can be a picture included in a reference picture list.

[0103] – A temporal neighboring block can be a block that is temporally adjacent to a spatial neighboring block of the target block.

[0104] Prediction mode: A prediction mode can be information indicating a mode for intra prediction or a mode for inter prediction.

[0105] Prediction unit: A prediction unit can be a basic unit for prediction (such as inter prediction, intra prediction, inter compensation, intra compensation, and motion compensation).

[0106] – A single prediction unit can be divided into multiple partitions or sub - prediction units with smaller sizes. The multiple partitions can also be the basic units during prediction or compensation. The partitions generated by dividing the prediction unit can also be prediction units.

[0107] Prediction unit partition: A prediction unit partition can be the shape into which a prediction unit is divided.

[0108] Reconstructed neighboring unit: A reconstructed neighboring unit can be a unit that has been decoded and reconstructed and is adjacent to the target unit.

[0109] – A reconstructed neighboring unit can be a unit that is spatially adjacent or temporally adjacent to the target unit.

[0110] – A reconstructed spatial neighboring unit can be a unit that has been reconstructed through encoding and / or decoding and is included in the target picture.

[0111] – A reconstructed temporal neighboring unit can be a unit that has been reconstructed through encoding and / or decoding and is included in the reference picture. The position of the reconstructed temporal neighboring unit in the reference picture can be the same as the position of the target unit in the target picture, or can correspond to the position of the target unit in the target picture. In addition, a reconstructed temporal neighboring unit can be a block that is adjacent to the corresponding block in the reference picture. Here, the position of the corresponding block in the reference picture can correspond to the position of the target block in the target picture. Here, the fact that the positions of the blocks correspond to each other can mean that the positions of the blocks are the same, can mean that one block is included in another block, or can mean that one block occupies a specific position in another block.

[0112] Sub - picture: A picture can be divided into one or more sub - pictures. A sub - picture can be composed of one or more parallel block rows and one or more parallel block columns.

[0113] – A sub - picture can be an area in the picture with a square shape or a rectangular (i.e., non - square rectangle) shape. In addition, a sub - picture can include one or more CTUs.

[0114] – A sub - picture can be a rectangular area of one or more stripes in the picture.

[0115] – A single sub - picture can include one or more parallel blocks, one or more bricks, and / or one or more stripes.

[0116] Parallel block: A parallel block can be an area in the picture with a square shape or a rectangular (i.e., non - square rectangle) shape.

[0117] – A parallel block can include one or more CTUs.

[0118] – A parallel block can be partitioned into one or more chunks.

[0119] Chunk: A chunk can represent one or more CTU rows in a parallel block.

[0120] – A parallel block can be partitioned into one or more chunks. Each chunk can include one or more CTU rows.

[0121] – A parallel block that is not partitioned into two parts can also represent a chunk.

[0122] Strip: A strip can include one or more parallel blocks in a picture. Optionally, a strip can include one or more chunks in a parallel block.

[0123] – A sub - picture can include one or more strips that jointly cover a rectangular area of the picture. Thus, each sub - picture boundary is always a strip boundary, and each vertical sub - picture boundary is always a vertical parallel - block boundary.

[0124] Parameter set: A parameter set can correspond to header information in the internal structure of a bitstream.

[0125] A parameter set can include at least one of a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptive parameter set (APS), a decoding parameter set (DPS), etc.

[0126] – The information signaled by each parameter set can be applied to the pictures that reference the corresponding parameter set. For example, the information in the VPS can be applied to the pictures that reference the VPS. The information in the SPS can be applied to the pictures that reference the SPS. The information in the PPS can be applied to the pictures that reference the PPS.

[0127] – Each parameter set can reference a higher - level parameter set. For example, the PPS can reference the SPS. The SPS can reference the VPS.

[0128] – In addition, a parameter set can include a parallel - block group, strip - header information, and parallel - block header information. A parallel - block group can be a group that includes multiple parallel blocks. In addition, the meaning of “parallel - block group” can be the same as the meaning of “strip”.

[0129] Rate - distortion optimization: An encoding device can use rate - distortion optimization to provide high encoding efficiency by exploiting a combination of the following: the size of a coding unit (CU), a prediction mode, the size of a prediction unit (PU), motion information, and the size of a transform unit (TU).

[0130] – A rate - distortion optimization scheme can calculate the rate - distortion cost of each combination to select an optimal combination from these combinations. The equation “ ” to calculate the rate-distortion cost. Generally, the combination that minimizes the rate-distortion cost can be selected as the optimal combination under the rate-distortion optimization scheme.

[0131] – D can represent distortion. D can be the average of the squares of the differences between the original transform coefficients and the reconstructed transform coefficients in the transform unit (i.e., mean square error).

[0132] – R can represent the rate, which can use relevant context information to represent the bit rate.

[0133] – represents the Lagrange multiplier. R can include not only coding parameter information (such as prediction mode, motion information, and coding block flag), but also the bits generated due to coding the transform coefficients.

[0134] – The encoding device can perform processes such as inter prediction and / or intra prediction, transformation, quantization, entropy coding, inverse quantization (dequantization), and / or inverse transformation to calculate accurate D and R. These processes will greatly increase the complexity of the encoding device.

[0135] – Bitstream: The bitstream can represent a stream of bits including encoded image information.

[0136] Parsing: Parsing can be the determination of the value of a syntax element made by performing entropy decoding on the bitstream. Optionally, the term “parsing” can represent this entropy decoding itself.

[0137] Symbol: A symbol can be at least one of the syntax elements, coding parameters, and transform coefficients of the encoding target unit and / or the decoding target unit. In addition, a symbol can be the target of entropy coding or the result of entropy decoding.

[0138] Reference picture: A reference picture can be an image that is referenced by a unit to perform inter prediction or motion compensation. Optionally, a reference picture can be an image including reference units that are referenced by the target unit to perform inter prediction or motion compensation.

[0139] Hereinafter, the terms “reference picture” and “reference image” can be used with the same meaning and can be used interchangeably with each other.

[0140] Reference picture list: A reference picture list can be a list including one or more reference images used for inter prediction or motion compensation.

[0141] – The types of reference picture lists can include combination list (LC), list 0 (L0), list 1 (L1), list 2 (L2), list 3 (L3), etc.

[0142] – For inter prediction, one or more reference picture lists can be used.

[0143] Inter - frame prediction indicator: The inter - frame prediction indicator can indicate the inter - frame prediction direction for a target unit. The inter - frame prediction can be one of unidirectional prediction and bidirectional prediction. Optionally, the inter - frame prediction indicator can represent the number of reference pictures used to generate the prediction unit for the target unit. Optionally, the inter - frame prediction indicator can represent the number of prediction blocks used for inter - frame prediction or motion compensation of the target unit.

[0144] Prediction list utilization flag: The prediction list utilization flag can indicate whether at least one reference picture in a specific reference picture list is used to generate the prediction unit.

[0145] – The prediction list utilization flag can be used to derive the inter - frame prediction indicator. Conversely, the inter - frame prediction indicator can be used to derive the prediction list utilization flag. For example, when the prediction list utilization flag indicates "0" (as the first value), it can indicate that for the target unit, the reference pictures in the reference picture list are not used to generate the prediction block. When the prediction list utilization flag indicates "1" (as the second value), it can indicate that for the target unit, the reference picture list is used to generate the prediction unit.

[0146] Reference picture index: The reference picture index can be an index indicating a specific reference picture in the reference picture list.

[0147] Picture Order Count (POC): The POC value of a picture can represent the order of displaying the corresponding picture.

[0148] Motion Vector (MV): The motion vector can be a 2D vector used for inter - frame prediction or motion compensation. The motion vector can represent the offset between the target image and the reference image.

[0149] – For example, the MV can be represented in a form such as (mv x , mv y ). mv x can indicate the horizontal component, and mv y can indicate the vertical component.

[0150] – Search range: The search range can be a 2D area where the search for the MV is performed during inter - frame prediction. For example, the size of the search range can be M×N. M and N can be positive integers respectively.

[0151] Motion vector candidate: The motion vector candidate can be a block that is a prediction candidate when the motion vector is predicted or the motion vector of a block that is a prediction candidate.

[0152] – The motion vector candidate can be included in the motion vector candidate list.

[0153] Motion vector candidate list: The motion vector candidate list can be a list configured using one or more motion vector candidates.

[0154] Motion vector candidate index: The motion vector candidate index can be an indicator for indicating a motion vector candidate in a motion vector candidate list. Optionally, the motion vector candidate index can be an index of a motion vector predictor.

[0155] Motion information: The motion information can be information including at least one of a reference picture list, a reference image, a motion vector candidate, a motion vector candidate index, a merge candidate, and a merge index, as well as a motion vector, a reference picture index, and an inter prediction indicator.

[0156] Merge candidate list: The merge candidate list can be a list configured using one or more merge candidates.

[0157] Merge candidate: The merge candidate can be a spatial merge candidate, a temporal merge candidate, a combined merge candidate, a combined bi-prediction merge candidate, a history-based candidate, an average merge candidate based on the average of two candidates, a zero merge candidate, etc. The merge candidate can include an inter prediction indicator and can include motion information such as prediction type information, a reference picture index for each list, a motion vector, a prediction list utilization flag, and an inter prediction indicator.

[0158] Merge index: The merge index can be an indicator for indicating a merge candidate in a merge candidate list.

[0159] – The merge index can indicate a reconstructed unit for deriving a merge candidate among reconstructed units that are spatially adjacent to the target unit and reconstructed units that are temporally adjacent to the target unit.

[0160] – The merge index can indicate at least one of multiple pieces of motion information of a merge candidate.

[0161] Transform unit: The transform unit can be a basic unit for residual signal encoding and / or residual signal decoding (such as transform, inverse transform, quantization, dequantization, transform coefficient encoding, and transform coefficient decoding). A single transform unit can be partitioned into multiple sub-transform units with smaller sizes. Here, the transform can include one or more of a primary transform and a secondary transform, and the inverse transform can include one or more of a primary inverse transform and a secondary inverse transform.

[0162] Scaling: Scaling can represent a process of multiplying a factor by a transform coefficient level.

[0163] – As a result of scaling the transform coefficient level, transform coefficients can be generated. Scaling can also be referred to as "dequantization".

[0164] Quantization Parameter (QP): The quantization parameter can be a value used to generate transform coefficient levels for transform coefficients in quantization. Optionally, the quantization parameter can also be a value used to generate transform coefficients by scaling transform coefficient levels in dequantization. Optionally, the quantization parameter can be a value mapped to a quantization step size.

[0165] Delta Quantization Parameter: The delta quantization parameter can represent the difference between the quantization parameter of a target unit and the predicted quantization parameter.

[0166] Scanning: Scanning can represent a method for arranging the order of coefficients in a unit, block, or matrix. For example, a method for arranging a 2D array in the form of a one-dimensional (1D) array can be referred to as "scanning". Optionally, a method for arranging a 1D array in the form of a 2D array can also be referred to as "scanning" or "inverse scanning".

[0167] Transform Coefficient: The transform coefficient can be a coefficient value generated when a coding device performs a transform. Optionally, the transform coefficient can be a coefficient value generated when a decoding device performs at least one of entropy decoding and dequantization.

[0168] – Quantized levels or quantized transform coefficient levels generated by applying quantization to transform coefficients or residual signals can also be included in the meaning of the term "transform coefficient".

[0169] Quantized Level: The quantized level can be a value generated when a coding device performs quantization on transform coefficients or residual signals. Optionally, the quantized level can be a value targeted for dequantization when a decoding device performs dequantization.

[0170] – Quantized transform coefficient levels resulting from transform and quantization can also be included in the meaning of the quantized level.

[0171] Non-Zero Transform Coefficient: The non-zero transform coefficient can be a transform coefficient having a value other than 0, or can be a transform coefficient level having a value other than 0. Optionally, the non-zero transform coefficient can be a transform coefficient whose value magnitude is not 0, or can be a transform coefficient level whose value magnitude is not 0.

[0172] Quantization Matrix: The quantization matrix can be a matrix used in the quantization process or dequantization process to improve the subjective or objective image quality of an image. The quantization matrix can also be referred to as a "scaling list".

[0173] Quantization Matrix Coefficient: The quantization matrix coefficient can be each element in the quantization matrix. The quantization matrix coefficient can also be referred to as a "matrix coefficient".

[0174] Default matrix: The default matrix may be a quantization matrix predefined by an encoding device and a decoding device.

[0175] Non-default matrix: The non-default matrix may be a quantization matrix not predefined by an encoding device and a decoding device. The non-default matrix may represent a quantization matrix signaled by a user from the encoding device to the decoding device.

[0176] Most Probable Mode (MPM): The MPM may represent an intra prediction mode that is highly likely to be used for intra prediction of a target block.

[0177] The encoding device and the decoding device may determine one or more MPMs based on encoding parameters related to the target block and attributes of entities related to the target block.

[0178] The encoding device and the decoding device may determine one or more MPMs based on the intra prediction mode of a reference block. The reference block may include a plurality of reference blocks. The plurality of reference blocks may include spatially adjacent blocks adjacent to the left side of the target block and spatially adjacent blocks adjacent to the upper side of the target block. In other words, one or more different MPMs may be determined according to which intra prediction modes have been used for the reference block.

[0179] – One or more MPMs may be determined in the same way in both the encoding device and the decoding device. That is, the encoding device and the decoding device may share the same MPM list including one or more MPMs.

[0180] MPM list: The MPM list may be a list including one or more MPMs. The number of one or more MPMs in the MPM list may be predefined.

[0181] MPM indicator: The MPM indicator may indicate the MPM among one or more MPMs in the MPM list that will be used for intra prediction of a target block. For example, the MPM indicator may be an index for the MPM list.

[0182] – Since the MPM list is determined in the same way in both the encoding device and the decoding device, it may not be necessary to send the MPM list itself from the encoding device to the decoding device.

[0183] – The MPM indicator may be signaled from the encoding device to the decoding device. Since the MPM indicator is signaled, the decoding device may determine the MPM among the MPMs in the MPM list that will be used for intra prediction of a target block.

[0184] MPM usage indicator: The MPM usage indicator may indicate whether the MPM usage mode will be used for prediction of a target block. The MPM usage mode may be a mode of using the MPM list to determine the MPM that will be used for intra prediction of a target block.

[0185] – The MPM usage indicator can be signaled from an encoding device to a decoding device.

[0186] Signaling: "Signaling" can indicate that information is sent from an encoding device to a decoding device. Optionally, "signaling" can indicate that information is included by the encoding device in a bitstream or a recording medium. The information signaled by the encoding device can be used by the decoding device.

[0187] – The encoding device can generate encoded information by performing encoding on the information to be signaled. The encoded information can be sent from the encoding device to the decoding device. The decoding device can obtain the information by decoding the sent encoded information. Here, the encoding can be entropy encoding, and the decoding can be entropy decoding.

[0188] Selective signaling: Information can be selectively signaled. Selective signaling for information can mean that the encoding device selectively includes the information in the bitstream or the recording medium (according to specific conditions). Selective signaling for information can mean that the decoding device selectively extracts the information from the bitstream (according to specific conditions).

[0189] Omission of signaling: Signaling for information can be omitted. The omission of signaling for information regarding the information can mean that the encoding device does not include the information in the bitstream or the recording medium (according to specific conditions). The omission of signaling for information can mean that the decoding device does not extract the information from the bitstream (according to specific conditions).

[0190] Statistical value: A variable, an encoding parameter, a constant, etc. can have a computable value. A statistical value can be a value generated by performing a calculation (operation) on the value of a specified target. For example, a statistical value can indicate one or more of an average value, a weighted average value, a weighted sum, a minimum value, a maximum value, a mode, a median, and an interpolation of the value of a specific variable, a specific encoding parameter, a specific constant, etc.

[0191] Figure 1 is a block diagram showing the configuration of an embodiment of an encoding device to which the present disclosure is applied.

[0192] The encoding device 100 can be an encoder, a video encoding device, or an image encoding device. The video can include one or more images (frames). The encoding device 100 can sequentially encode one or more images of the video.

[0193] Refer to Figure 1, the encoding device 100 includes an inter-frame prediction unit 110, an intra-frame prediction unit 120, a switch 115, a subtractor 125, a transformation unit 130, a quantization unit 140, an entropy encoding unit 150, an inverse quantization unit 160, an inverse transformation unit 170, an adder 175, a filter unit 180, and a reference picture buffer 190.

[0194] The encoding device 100 can perform encoding on a target image using the intra-frame mode and / or the inter-frame mode. In other words, the prediction mode of the target block can be one of the intra-frame mode and the inter-frame mode.

[0195] Hereinafter, the terms "intra-frame mode", "intra-frame prediction mode", "intra-picture mode", and "intra-frame prediction mode" can be used with the same meaning and can be used interchangeably with each other.

[0196] Hereinafter, the terms "inter-frame mode", "inter-frame prediction mode", "inter-picture mode", and "inter-picture prediction mode" can be used with the same meaning and can be used interchangeably with each other.

[0197] Hereinafter, the term "image" can indicate only a partial image or can indicate a block. In addition, the processing of "image" can indicate the sequential processing of multiple blocks.

[0198] In addition, the encoding device 100 can generate a bitstream including encoded information by encoding the target image, and can output and store the generated bitstream. The generated bitstream can be stored in a computer-readable storage medium and can be streamed through wired and / or wireless transmission media.

[0199] When the intra-frame mode is used as the prediction mode, the switch 115 can switch to the intra-frame mode. When the inter-frame mode is used as the prediction mode, the switch 115 can switch to the inter-frame mode.

[0200] The encoding device 100 can generate a prediction block for the target block. In addition, after the prediction block has been generated, the encoding device 100 can encode the residual block for the target block using the residual between the target block and the prediction block.

[0201] When the prediction mode is the intra-frame mode, the intra-frame prediction unit 120 can use the pixels of the previously encoded / decoded neighboring blocks adjacent to the target block as reference samples. The intra-frame prediction unit 120 can perform spatial prediction on the target block using the reference samples and can generate prediction samples for the target block via spatial prediction. The prediction samples can represent the samples in the prediction block.

[0202] The inter-frame prediction unit 110 can include a motion prediction unit and a motion compensation unit.

[0203] When the prediction mode is an inter-frame mode, the motion prediction unit may search for the region in the reference image that best matches the target block during the motion prediction process, and may derive a motion vector for the target block and the found region based on the found region. Here, the motion prediction unit may use the search range as the target region for the search.

[0204] The reference image may be stored in the reference picture buffer 190. More specifically, when the encoding and / or decoding of the reference image has been processed, the encoded and / or decoded reference image may be stored in the reference picture buffer 190.

[0205] Since the decoded pictures are stored, the reference picture buffer 190 may be a decoded picture buffer (DPB).

[0206] The motion compensation unit may generate a prediction block for the target block by performing motion compensation using the motion vector. Here, the motion vector may be a two-dimensional (2D) vector for inter-frame prediction. In addition, the motion vector may indicate the offset between the target image and the reference image.

[0207] When the motion vector has a value other than an integer, the motion prediction unit and the motion compensation unit may generate a prediction block by applying an interpolation filter to a partial region of the reference image. To perform inter-frame prediction or motion compensation, it may be determined which of the skip mode, merge mode, advanced motion vector prediction (AMVP) mode, and current picture reference mode corresponds to the method for predicting and compensating for the motion of the PU included in the CU based on the CU, and inter-frame prediction or motion compensation may be performed according to the mode.

[0208] The subtractor 125 may generate a residual block, where the residual block is the difference between the target block and the prediction block. The residual block may also be referred to as a "residual signal".

[0209] The residual signal may be the difference between the original signal and the prediction signal. Alternatively, the residual signal may be a signal generated by transforming or quantizing the difference between the original signal and the prediction signal or a signal generated by transforming and quantizing the difference. The residual block may be the residual signal for a block unit.

[0210] The transform unit 130 may generate transform coefficients by transforming the residual block and may output the generated transform coefficients. Here, the transform coefficients may be the coefficient values generated by transforming the residual block.

[0211] The transform unit 130 may use one of a plurality of predefined transform methods when performing the transform.

[0212] The multiple predefined transformation methods may include, for example, Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Karhunen-Loeve Transform (KLT), etc.

[0213] The transformation method for transforming the residual block may be determined according to at least one of the coding parameters for the target block and / or neighboring blocks. For example, the transformation method may be determined based on at least one of the inter prediction mode for the PU, the intra prediction mode for the PU, the size of the TU, and the shape of the TU. Alternatively, the transformation information indicating the transformation method may be signaled from the coding device 100 to the decoding device 200.

[0214] When the transform skip mode is used, the transform unit 130 may omit the operation of transforming the residual block.

[0215] By performing quantization on the transform coefficients, quantized transform coefficient levels or quantized levels may be generated. Hereinafter, in the embodiments, each of the quantized transform coefficient levels and the quantized levels may also be referred to as "transform coefficients".

[0216] The quantization unit 140 may generate quantized transform coefficient levels (i.e., quantized levels or quantized coefficients) by quantizing the transform coefficients according to quantization parameters. The quantization unit 140 may output the generated quantized transform coefficient levels. In this case, the quantization unit 140 may use a quantization matrix to quantize the transform coefficients.

[0217] The entropy coding unit 150 may generate a bitstream by performing probability distribution-based entropy coding based on the value calculated by the quantization unit 140 and / or the coding parameter values calculated during the coding process. The entropy coding unit 150 may output the generated bitstream.

[0218] The entropy coding unit 150 may perform entropy coding on the information about the pixels of the image and the information required for decoding the image. For example, the information required for decoding the image may include syntax elements and the like.

[0219] When entropy coding is applied, fewer bits may be allocated to more frequently occurring symbols, and more bits may be allocated to less frequently occurring symbols. Since the symbols are represented by this allocation, the size of the bit string for the target symbols to be encoded may be reduced. Therefore, the compression performance of video coding may be improved by entropy coding.

[0220] In addition, for entropy coding, the entropy coding unit 150 may use coding methods such as exponential Golomb, context-adaptive variable length coding (CAVLC), or context-adaptive binary arithmetic coding (CABAC). For example, the entropy coding unit 150 may use a variable length coding / code (VLC) table to perform entropy coding. For example, the entropy coding unit 150 may derive a binarization method for a target symbol. In addition, the entropy coding unit 150 may derive a probability model for the target symbol / binary bit. The entropy coding unit 150 may perform arithmetic coding using the derived binarization method, probability model, and context model.

[0221] The entropy coding unit 150 may transform the coefficients in 2D block form into 1D vector form through a transform coefficient scanning method to encode the quantized transform coefficient levels.

[0222] Coding parameters may be information required for encoding and / or decoding. The coding parameters may include information encoded by the encoding device 100 and sent from the encoding device 100 to the decoding device, and may also include information that can be derived during the encoding or decoding process. For example, the information sent to the decoding device may include syntax elements.

[0223] Coding parameters may include not only information such as syntax elements (or flags or indices) encoded by an encoding device and signaled by the encoding device to a decoding device, but also information derived during the encoding or decoding process. Additionally, the coding parameters may include information required for encoding or decoding an image. For example, the coding parameters may include at least one value, a combination of the following items, or statistics of: the size of a unit / block, the shape / form of a unit / block, the depth of a unit / block, the partition information of a unit / block, the partition structure of a unit / block, information indicating whether a unit / block is partitioned in a quadtree structure, information indicating whether a unit / block is partitioned in a binary tree structure, the partition direction (horizontal or vertical direction) of a binary tree structure, the partition form (symmetric partition or asymmetric partition) of a binary tree structure, information indicating whether a unit / block is partitioned in a ternary tree structure, the partition direction (horizontal or vertical direction) of a ternary tree structure, the partition form (symmetric partition or asymmetric partition, etc.) of a ternary tree structure, information indicating whether a unit / block is partitioned in a multi-type tree structure, the combination and direction (horizontal or vertical direction, etc.) of partitions in a multi-type tree structure, the partition form (symmetric partition or asymmetric partition, etc.) of a multi-type tree structure, the partition tree (binary tree or ternary tree) of a multi-type tree form, the prediction type (intra prediction or inter prediction), the intra prediction mode / direction, the intra luminance prediction mode / direction, the intra chrominance prediction mode / direction, the intra partition information, the inter partition information, the coding block partition flag, the prediction block partition flag, the transform block partition flag, the reference sample filtering method, the reference sample filter taps, the reference sample filter coefficients, the prediction block filtering method, the prediction block filter taps, the prediction block filter coefficients, the prediction block boundary filtering method, the prediction block boundary filter taps, the prediction block boundary filter coefficients, the inter prediction mode, the motion information, the motion vector, the motion vector difference, the reference picture index, the inter prediction direction, the inter prediction indicator, the prediction list utilization flag, the reference picture list, the reference image, the POC, the motion vector prediction factor, the motion vector prediction index, the motion vector prediction candidate, the motion vector candidate list, information indicating whether the merge mode is used, the merge index, the merge candidate, the merge candidate list, information indicating whether the skip mode is used, the type of interpolation filter, the taps of the interpolation filter, the filter coefficients of the interpolation filter, the magnitude of the motion vector, the precision of the motion vector representation, the transform type, the transform size, information indicating whether the first transform is used, information indicating whether an additional (second) transform is used, the first transform selection information (or the first transform index), the second transform selection information (or the second transform index), information indicating the presence or absence of a residual signal, the coding block style, the coding block flag, the quantization parameter, the residual quantization parameter, the quantization matrix, information about the loop filter, information indicating whether the loop filter is applied, the coefficients of the loop filter, the taps of the loop filter, the shape / form of the loop filter, information indicating whether the deblocking filter is applied,Coefficients of the deblocking filter, taps of the deblocking filter, deblocking filter strength, shape / form of the deblocking filter, information indicating whether adaptive sample offset is applied, value of the adaptive sample offset, category of the adaptive sample offset, type of the adaptive sample offset, information indicating whether the adaptive loop filter is applied, coefficients of the adaptive loop filter, taps of the adaptive loop filter, shape / form of the adaptive loop filter, binarization / de-binarization method, context model, context model determination method, context model update method, information indicating whether the normal mode is executed, information indicating whether the bypass mode is executed, valid coefficient flag, last valid coefficient flag, coding flag of the coefficient group, position of the last valid coefficient, information indicating whether the value of the coefficient is greater than 1, information indicating whether the value of the coefficient is greater than 2, information indicating whether the value of the coefficient is greater than 3, remaining coefficient value information, positive / negative sign information, reconstructed luma sample, reconstructed chroma sample, context binary bit, bypass binary bit, residual luma sample, residual chroma sample, transform coefficient, luma transform coefficient, chroma transform coefficient, quantization level, luma quantization level, chroma quantization level, transform coefficient level, transform coefficient level scanning method, size of the motion vector search area on the decoding device side, shape / form of the motion vector search area on the decoding device side, number of times of motion vector search on the decoding device side, size of the CTU, minimum block size, maximum block size, maximum block depth, minimum block depth, image display / output order, slice identification information, slice type, slice partition information, parallel block group identification information, parallel block group type, parallel block group partition information, parallel block identification information, parallel block type, parallel block partition information, picture type, bit depth, input sample bit depth, reconstructed sample bit depth, residual sample bit depth, transform coefficient bit depth, quantization level bit depth, information about the luma signal, information about the chroma signal, color space of the target block and color space of the residual block. In addition, the above information related to the coding parameters may also be included in the coding parameters. Information for calculating and / or deriving the above coding parameters may also be included in the coding parameters. Information calculated or derived using the above coding parameters may also be included in the coding parameters.

[0224] The first transform selection information may indicate the first transform applied to the target block.

[0225] The second transform selection information may indicate the second transform applied to the target block.

[0226] The residual signal may represent the difference between the original signal and the predicted signal. Optionally, the residual signal may be a signal generated by transforming the difference between the original signal and the predicted signal. Optionally, the residual signal may be a signal generated by transforming and quantizing the difference between the original signal and the predicted signal. The residual block may be the residual signal for the block.

[0227] Here, sending information by a signal may indicate that the encoding device 100 includes the entropy-encoded information generated by performing entropy encoding on a flag or an index in a bitstream, and may indicate that the decoding device 200 obtains information by performing entropy decoding on the entropy-encoded information extracted from the bitstream. Here, the information may include a flag, an index, and the like.

[0228] A signal may mean the information to be sent by the signal. Hereinafter, information for an image and a block may be referred to as a "signal". In addition, hereinafter, the terms "information" and "signal" may be used to have the same meaning and may be used interchangeably with each other. For example, a specific signal may be a signal representing a specific block. An original signal may be a signal representing a target block. A prediction signal may be a signal representing a prediction block. A residual signal may be a signal representing a residual block.

[0229] A bitstream may include information based on a specific syntax. The encoding device 100 may generate a bitstream including information according to the specific syntax. The decoding device 200 may obtain information from the bitstream according to the specific syntax.

[0230] Since the encoding device 100 performs encoding via inter-frame prediction, the encoded target image may be used as a reference image for another image to be subsequently processed. Therefore, the encoding device 100 may reconstruct or decode the encoded target image and store the reconstructed or decoded image in the reference picture buffer 190 as a reference image. For decoding, inverse quantization and inverse transformation of the encoded target image may be performed.

[0231] The quantization level may be inverse-quantized by the inverse quantization unit 160 and may be inverse-transformed by the inverse transformation unit 170. The inverse quantization unit 160 may generate inverse-quantized coefficients by performing an inverse transformation on the quantization level. The inverse transformation unit 170 may generate coefficients that have been inverse-quantized and inverse-transformed by performing an inverse transformation on the inverse-quantized coefficients.

[0232] The coefficients that have been inverse-quantized and inverse-transformed may be added to the prediction block by the adder 175. Adding the coefficients that have been inverse-quantized and inverse-transformed and the prediction block may then generate a reconstructed block. Here, the coefficients that have been inverse-quantized and / or inverse-transformed may represent coefficients on which one or more of inverse quantization and inverse transformation have been performed, and may also represent a reconstructed residual block. Here, the reconstructed block may represent a restored block or a decoded block.

[0233] The reconstructed block can be filtered by the filter unit 180. The filter unit 180 can apply one or more of a deblocking filter, a sample adaptive offset (SAO) filter, an adaptive loop filter (ALF), and a non-local filter (NLF) to the reconstructed samples, the reconstructed block, or the reconstructed picture. The filter unit 180 may also be referred to as a "loop filter".

[0234] The deblocking filter can eliminate the block distortion that appears at the boundary between blocks in the reconstructed picture. To determine whether to apply the deblocking filter, the number of columns or rows that are included in the block and that include the pixels based on which it is determined whether to apply the deblocking filter to the target block can be determined.

[0235] When the deblocking filter is applied to the target block, the applied filter can be different according to the intensity of the required deblocking filtering. In other words, among different filters, the filter determined in consideration of the intensity of the deblocking filtering can be applied to the target block. When the deblocking filter is applied to the target block, one or more of a long-tap filter, a strong filter, a weak filter, and a Gaussian filter can be applied to the target block according to the required intensity of the deblocking filtering.

[0236] In addition, when performing vertical filtering and horizontal filtering on the target block, the horizontal filtering and the vertical filtering can be performed in parallel.

[0237] SAO can add an appropriate offset to the pixel value to compensate for the coding error. SAO can perform correction on the image to which deblocking has been applied based on pixels, where the correction uses the offset of the difference between the original image and the image to which deblocking has been applied. To perform the offset correction for the image, a method of dividing the pixels included in the image into a specific number of regions, determining the regions to which the offset will be applied among the divided regions, and applying the offset to the determined regions can be used, and a method of applying the offset in consideration of the edge information of each pixel can also be used.

[0238] ALF can perform filtering based on the value obtained by comparing the reconstructed image with the original image. After the pixels included in the image have been divided into a predetermined number of groups, the filter to be applied to each group can be determined, and filtering can be performed differently for each group. Information related to whether to apply the adaptive loop filter can be signaled for each CU. Such information can be signaled for the luminance signal. The shape and filter coefficients of the ALF to be applied to each block can be different for each block. Alternatively, an ALF having a fixed form can be applied to the block regardless of the characteristics of the block.

[0239] The non-local filter can perform filtering based on a reconstructed block similar to a target block. An area similar to the target block can be selected from the reconstructed picture, and the statistical attributes of the selected similar area can be used to perform filtering of the target block. Information on whether to apply the non-local filter can be signaled for a coding unit (CU). In addition, the shape and filter coefficients of the non-local filter applied to a block can vary according to the block.

[0240] The reconstructed block or reconstructed image filtered by the filter unit 180 can be stored in the reference picture buffer 190 as a reference picture. The reconstructed block filtered by the filter unit 180 can be part of the reference picture. In other words, the reference picture can be a reconstructed picture composed of the reconstructed blocks filtered by the filter unit 180. The stored reference picture can then be used for inter-frame prediction or motion compensation.

[0241] Figure 2 is a block diagram showing the configuration of an embodiment of a decoding device to which the present disclosure is applied.

[0242] The decoding device 200 can be a decoder, a video decoding device, or an image decoding device.

[0243] Referring to Figure 2 , the decoding device 200 may include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an intra prediction unit 240, an inter prediction unit 250, a switch 245, an adder 255, a filter unit 260, and a reference picture buffer 270.

[0244] The decoding device 200 can receive the bitstream output from the encoding device 100. The decoding device 200 can receive the bitstream stored in a computer-readable storage medium and can receive the bitstream streamed through a wired / wireless transmission medium.

[0245] The decoding device 200 can perform decoding on the bitstream in the intra mode and / or the inter mode. In addition, the decoding device 200 can generate a reconstructed image or a decoded image via decoding and can output the reconstructed image or the decoded image.

[0246] For example, the operation of switching to the intra mode or the inter mode based on the prediction mode for decoding can be performed by the switch 245. When the prediction mode for decoding is the intra mode, the switch 245 can be operated to switch to the intra mode. When the prediction mode for decoding is the inter mode, the switch 245 can be operated to switch to the inter mode.

[0247] The decoding device 200 can obtain the reconstructed residual block by decoding the input bitstream and can generate a prediction block. When the reconstructed residual block and the prediction block are obtained, the decoding device 200 can generate a reconstructed block as the target to be decoded by adding the reconstructed residual block and the prediction block.

[0248] The entropy decoding unit 210 can generate symbols by performing entropy decoding on the bitstream based on the probability distribution of the bitstream. The generated symbols can include symbols in the form of quantized transform coefficient levels (i.e., quantized levels or quantized coefficients). Here, the entropy decoding method can be similar to the entropy encoding method described above. That is, the entropy decoding method can be the inverse process of the entropy encoding method described above.

[0249] The entropy decoding unit 210 can change the coefficients in the form of a one-dimensional (1D) vector into a 2D block shape by a transform coefficient scanning method to decode the quantized transform coefficient levels.

[0250] For example, the coefficients of the block can be changed into a 2D block shape by scanning the block coefficients using a right upper diagonal scan. Alternatively, which one of the right upper diagonal scan, vertical scan, and horizontal scan will be used can be determined according to the size of the corresponding block and / or the intra prediction mode.

[0251] The quantized coefficients can be inverse quantized by the inverse quantization unit 220. The inverse quantization unit 220 can generate inverse quantized coefficients by performing inverse quantization on the quantized coefficients. In addition, the inverse quantized coefficients can be inverse transformed by the inverse transform unit 230. The inverse transform unit 230 can generate a reconstructed residual block by performing inverse transform on the inverse quantized coefficients. As a result of performing inverse quantization and inverse transform on the quantized coefficients, a reconstructed residual block can be generated. Here, when generating the reconstructed residual block, the inverse quantization unit 220 can apply a quantization matrix to the quantized coefficients.

[0252] When using the intra mode, the intra prediction unit 240 can generate a prediction block by performing spatial prediction on the target block, where the spatial prediction uses the pixel values of the previously decoded neighboring blocks adjacent to the target block.

[0253] The inter prediction unit 250 can include a motion compensation unit. Alternatively, the inter prediction unit 250 can be designated as the "motion compensation unit".

[0254] When using the inter mode, the motion compensation unit can generate a prediction block by performing motion compensation on the target block, where the motion compensation uses a motion vector and a reference image stored in the reference picture buffer 270.

[0255] The motion compensation unit may apply an interpolation filter to a partial region of a reference image when the motion vector has a value other than an integer, and may generate a prediction block using the reference image to which the interpolation filter has been applied. To perform motion compensation, the motion compensation unit may determine, based on the CU, which one of the skip mode, merge mode, advanced motion vector prediction (AMVP) mode, and current picture reference mode corresponds to the motion compensation method for the PUs included in the CU, and may perform motion compensation according to the determined mode.

[0256] The reconstructed residual block and the prediction block may be added to each other by an adder 255. The adder 255 may generate a reconstructed block by adding the reconstructed residual block and the prediction block.

[0257] The reconstructed block may be filtered by a filter unit 260. The filter unit 260 may apply at least one of a deblocking filter, SAO filter, ALF, and NLF to the reconstructed block or the reconstructed image. The reconstructed image may be a picture including the reconstructed block.

[0258] The filter unit may output the reconstructed image.

[0259] The reconstructed image and / or the reconstructed block filtered by the filter unit 260 may be stored as a reference picture in a reference picture buffer 270. The reconstructed block filtered by the filter unit 260 may be a part of the reference picture. In other words, the reference picture may be an image composed of the reconstructed blocks filtered by the filter unit 260. The stored reference picture may then be used for inter prediction or motion compensation.

[0260] Figure 3 is a diagram schematically showing the partitioning structure of an image when the image is encoded and decoded.

[0261] Figure 3 An example in which a single unit is partitioned into a plurality of sub-units may be schematically shown.

[0262] To partition an image effectively, a coding unit (CU) may be used in encoding and decoding. The term "unit" may be used to commonly specify 1) a block including image samples and 2) syntax elements. For example, "partitioning of a unit" may mean "partitioning of a block corresponding to the unit".

[0263] The CU may be used as a basic unit for image encoding / decoding. The CU may be used as a unit to which one mode selected from an intra mode and an inter mode is applied in image encoding / decoding. In other words, in image encoding / decoding, it may be determined which one of the intra mode and the inter mode will be applied to each CU.

[0264] In addition, a CU may be a basic unit for predicting, transforming, quantizing, inverse-transforming, dequantizing, and encoding / decoding transform coefficients.

[0265] Referring Figure 3 , image 300 may be sequentially partitioned into units corresponding to largest coding units (LCUs), and a partitioning structure may be determined for each LCU. Here, an LCU may be used to have the same meaning as a coding tree unit (CTU).

[0266] Partitioning a unit may represent partitioning a block corresponding to the unit. Block partitioning information may include depth information regarding the depth of the unit. The depth information may indicate the number of times the unit is partitioned and / or the degree to which the unit is partitioned. A single unit may be hierarchically partitioned into multiple sub-units while the single unit has depth information based on a tree structure.

[0267] Each partitioned sub-unit may have depth information. The depth information may be information indicating the size of a CU. The depth information may be stored for each CU.

[0268] Each CU may have depth information. When a CU is partitioned, the depth of the CU generated from the partition may increase by 1 from the depth of the partitioned CU.

[0269] The partitioning structure may represent the distribution of coding units (CUs) in LCU 310 for efficiently encoding an image. Such a distribution may be determined based on whether a single CU will be partitioned into multiple CUs. The number of CUs generated by partitioning may be a positive integer of 2 or greater, including 2, 3, 4, 8, 16, etc.

[0270] According to the number of CUs generated by partitioning, the horizontal size and vertical size of each CU generated by partitioning may be smaller than the horizontal size and vertical size of the CU before partitioning. For example, the horizontal size and vertical size of each CU generated by partitioning may be half of the horizontal size and vertical size of the CU before partitioning.

[0271] Each partitioned CU may be recursively partitioned into four CUs in the same manner. Compared with at least one of the horizontal size and vertical size of the CU before partitioning, at least one of the horizontal size and vertical size of each partitioned CU may be reduced via recursive partitioning.

[0272] Partitioning of a CU may be recursively performed until a predefined depth or a predefined size.

[0273] For example, the depth of a CU may have values ranging from 0 to 3. The size range of a CU may be from a size of 64×64 to a size of 8×8 depending on the depth of the CU.

[0274] For example, the depth of the LCU 310 can be 0, and the depth of the smallest coding unit (SCU) can be a predefined maximum depth. Here, as described above, the LCU can be a CU with the maximum coding unit size, and the SCU can be a CU with the smallest coding unit size.

[0275] Partitioning can start at the LCU 310, and whenever the horizontal size and / or vertical size of a CU is reduced by partitioning, the depth of the CU can be incremented by 1.

[0276] For example, for each depth, a non-partitioned CU can have a size of 2N×2N. Additionally, in the case where a CU is partitioned, a CU with a size of 2N×2N can be partitioned into four CUs each with a size of N×N. Whenever the depth is incremented by 1, the value of N can be halved.

[0277] Referring to Figure 3 , the LCU with a depth of 0 can have 64×64 pixels or a block of 64×64. 0 can be the minimum depth. The SCU with a depth of 3 can have 8×8 pixels or a block of 8×8. 3 can be the maximum depth. Here, the CU with a 64×64 block as the LCU can be represented by a depth of 0. The CU with a 32×32 block can be represented by a depth of 1. The CU with a 16×16 block can be represented by a depth of 2. The CU with an 8×8 block as the SCU can be represented by a depth of 3.

[0278] Information regarding whether a corresponding CU is partitioned can be represented by the partitioning information of the CU. The partitioning information can be 1-bit information. All CUs except the SCU can include the partitioning information. For example, the value of the partitioning information of a non-partitioned CU can be a first value. The value of the partitioning information of a partitioned CU can be a second value. When the partitioning information indicates whether a CU is partitioned, the first value can be "0" and the second value can be "1".

[0279] For example, when a single CU is partitioned into four CUs, the horizontal size and vertical size of each of the four CUs generated by partitioning can be half of the horizontal size and vertical size of the CU before partitioning. When a CU with a size of 32×32 is partitioned into four CUs, the size of each of the four partitioned CUs can be 16×16. When a single CU is partitioned into four CUs, it can be considered that the CU has been partitioned in a quadtree structure. In other words, it can be considered that quadtree partitioning has been applied to the CU.

[0280] For example, when a single CU is partitioned into two CUs, the horizontal or vertical size of each of the two CUs generated by the partitioning can be half of the horizontal or vertical size of the CU before partitioning. When a CU with a size of 32×32 is vertically partitioned into two CUs, the size of each of the two partitioned CUs can be 16×32. When a CU with a size of 32×32 is horizontally partitioned into two CUs, the size of each of the two partitioned CUs can be 32×16. When a single CU is partitioned into two CUs, it can be considered that the CU has been partitioned in a binary tree structure. In other words, it can be considered that binary tree partitioning has been applied to the CU.

[0281] For example, when a single CU is partitioned (or divided) into three CUs, the original CU before partitioning is partitioned such that its horizontal or vertical size is divided in a ratio of 1:2:1, thus enabling the generation of three sub-CUs. For example, when a CU with a size of 16×32 is horizontally partitioned into three sub-CUs, the three sub-CUs generated by the partitioning can have sizes of 16×8, 16×16, and 16×8 respectively in the top-to-bottom direction. For example, when a CU with a size of 32×32 is vertically partitioned into three sub-CUs, the three sub-CUs generated by the partitioning can have sizes of 8×32, 16×32, and 8×32 respectively in the left-to-right direction. When a single CU is partitioned into three CUs, it can be considered that the CU is partitioned in a ternary tree form. In other words, it can be considered that ternary tree partitioning has been applied to the CU.

[0282] Both quadtree partitioning and binary tree partitioning are applied to Figure 3 the LCU 310.

[0283] In the encoding device 100, a coding tree unit (CTU) with a size of 64×64 can be partitioned into multiple smaller CUs through a recursive quadtree structure. A single CU can be partitioned into four CUs with the same size. Each CU can be recursively partitioned and can have a quadtree structure.

[0284] Through the recursive partitioning of the CU, an optimal partitioning method that incurs the minimum rate-distortion cost can be selected.

[0285] Figure 3 The coding tree unit (CTU) 320 in

[0286] is an example of a CTU to which all of quadtree partitioning, binary tree partitioning, and ternary tree partitioning are applied. As described above, in order to partition a CTU, at least one of quadtree partitioning, binary tree partitioning, and ternary tree partitioning can be applied to the CTU. The partitioning can be applied based on a specific priority.

[0287] For example, quadtree partitioning can be preferentially applied to CTUs. CUs that cannot be further partitioned in the form of a quadtree can correspond to the leaf nodes of the quadtree. A CU corresponding to a leaf node of the quadtree can be the root node of a binary tree and / or a ternary tree. That is, a CU corresponding to a leaf node of the quadtree can be partitioned in the form of a binary tree or a ternary tree, or may not be further partitioned. In this case, it is prevented that each CU generated by applying binary tree partitioning or ternary tree partitioning to a CU corresponding to a leaf node of the quadtree is again quadtree partitioned, thereby effectively performing the operation of partitioning a block and / or signaling block partitioning information.

[0288] Quad-partitioning information can be used to signal the partitioning of the CU corresponding to each node of the quadtree. Quad-partitioning information having a first value (e.g., "1") can indicate that the corresponding CU is partitioned in the form of a quadtree. Quad-partitioning information having a second value (e.g., "0") can indicate that the corresponding CU is not partitioned in the form of a quadtree. The quad-partitioning information can be a flag having a specific length (e.g., 1 bit).

[0289] There may be no priority between binary tree partitioning and ternary tree partitioning. That is, a CU corresponding to a leaf node of the quadtree can be partitioned in the form of a binary tree or a ternary tree. In addition, a CU generated by binary tree partitioning or ternary tree partitioning can be further partitioned in the form of a binary tree or a ternary tree, or may not be further partitioned.

[0290] The partitioning performed when there is no priority between binary tree partitioning and ternary tree partitioning can be referred to as "multi-type tree partitioning". That is, a CU corresponding to a leaf node of the quadtree can be the root node of a multi-type tree. Information indicating whether the CU is partitioned according to a multi-type tree, partitioning direction information, and partitioning tree information can be used to signal the partitioning of the CU corresponding to each node of the multi-type tree. For the partitioning of the CU corresponding to each node of the multi-type tree, information indicating whether the partitioning according to the multi-type tree is performed, partitioning direction information, and partitioning tree information can be signaled sequentially.

[0291] For example, information indicating whether the CU is partitioned according to a multi-type tree and having a first value (e.g., "1") can indicate that the corresponding CU is partitioned in the form of a multi-type tree. Information indicating whether the CU is partitioned according to a multi-type tree and having a second value (e.g., "0") can indicate that the corresponding CU is not partitioned in the form of a multi-type tree.

[0292] When a CU corresponding to each node of the multi-type tree is partitioned in the form of a multi-type tree, the corresponding CU can further include partitioning direction information.

[0293] The partition direction information can indicate the partition direction of multiple types of tree partitions. The partition direction information having a first value (e.g., "1") can indicate that the corresponding CU is partitioned in the vertical direction. The partition direction information having a second value (e.g., "0") can indicate that the corresponding CU is partitioned in the horizontal direction.

[0294] When the CU corresponding to each node of the multiple-type tree is partitioned in the form of a multiple-type tree, the corresponding CU can further include partition tree information. The partition tree information can indicate the tree used for the multiple-type tree partition.

[0295] For example, the partition tree information having a first value (e.g., "1") can indicate that the corresponding CU is partitioned in the form of a binary tree. The partition tree information having a second value (e.g., "0") can indicate that the corresponding CU is partitioned in the form of a ternary tree.

[0296] Here, each of the above information indicating whether the partition of the multiple-type tree is performed, the partition tree information, and the partition direction information can be a flag having a specific length (e.g., 1 bit).

[0297] Entropy coding and / or entropy decoding can be performed on at least one of the above four-partition information, the information indicating whether the partition of the multiple-type tree is performed, the partition direction information, and the partition tree information. To perform the entropy coding / decoding of this information, the information of neighboring CUs adjacent to the target CU can be used.

[0298] For example, it can be considered that the probability that the partition form (i.e., partition / non-partition, partition tree, and / or partition direction) of the left CU and / or the upper CU is similar to that of the target CU is very high. Therefore, based on the information of the neighboring CUs, the context information for the entropy coding and / or entropy decoding of the information of the target CU can be derived. Here, the information of the neighboring CUs can include at least one of the following: 1) the four-partition information of the neighboring CUs, 2) the information indicating whether the neighboring CUs are partitioned in the form of a multiple-type tree, 3) the partition direction information of the neighboring CUs, and 4) the partition tree information of the neighboring CUs.

[0299] In another embodiment of the binary tree partition and the ternary tree partition, the binary tree partition can be preferentially performed. That is, the binary tree partition can be applied first, and then the CU corresponding to the leaf node of the binary tree can be set as the root node of the ternary tree. In this case, the four-tree partition or the binary tree partition is not performed on the CU corresponding to the node of the ternary tree.

[0300] A CU that is not further partitioned by quadtree partitioning, binary tree partitioning, and / or ternary tree partitioning can be a unit for encoding, prediction, and / or transformation. That is, the CU may not be further partitioned for prediction and / or transformation. Therefore, a partitioning structure for partitioning the CU into prediction units (PUs) / or transform units (TUs), its partitioning information, etc. may not exist in the bitstream.

[0301] However, when the size of the CU as a partitioning unit is larger than the size of the maximum transform block, the CU can be recursively partitioned until the size of the CU becomes less than or equal to the size of the maximum transform block. For example, when the size of the CU is 64×64 and the size of the maximum transform block is 32×32, the CU can be partitioned into four 32×32 blocks to perform transformation. For example, when the size of the CU is 32×64 and the size of the maximum transform block is 32×32, the CU can be partitioned into two 32×32 blocks.

[0302] In this case, information indicating whether the CU is partitioned for transformation may not be signaled separately. Without signaling, it can be determined whether the CU is partitioned via a comparison between the horizontal size (and / or vertical size) of the CU and the horizontal size (and / or vertical size) of the maximum transform block. For example, when the horizontal size of the CU is larger than the horizontal size of the maximum transform block, the CU can be vertically bisected. In addition, when the vertical size of the CU is larger than the vertical size of the maximum transform block, the CU can be horizontally bisected.

[0303] Information about the maximum size and / or minimum size of the CU and information about the maximum size and / or minimum size of the transform block can be signaled or determined at a level higher than the level of the CU. For example, the higher level can be the sequence level, picture level, parallel block level, parallel block group level, or slice level. For example, the minimum size of the CU can be set to 4×4. For example, the maximum size of the transform block can be set to 64×64. For example, the maximum size of the transform block can be set to 4×4.

[0304] Information about the minimum size of the CU corresponding to the leaf node of the quadtree (i.e., the minimum size of the quadtree) and / or information about the maximum depth of the path from the root node to the leaf node of the multi-type tree (i.e., the maximum depth of the multi-type tree) can be signaled or determined at a level higher than the level of the CU. For example, the higher level can be the sequence level, picture level, slice level, parallel block group level, or parallel block level. Information about the minimum size of the quadtree and / or information about the maximum depth of the multi-type tree can be signaled or determined separately at each of the intra-slice level and the inter-slice level.

[0305] Information about the difference between the size of a CTU and the maximum size of a transform block can be signaled or determined at a level higher than the level of the CU. For example, the higher level can be sequence level, picture level, slice level, parallel block group level, or parallel block level. Information about the maximum size (i.e., the maximum size of the binary tree) of the CU corresponding to each node of the binary tree can be determined based on the size of the CTU and the information about the difference. The maximum size (i.e., the maximum size of the ternary tree) of the CU corresponding to each node of the ternary tree can have different values according to the type of the slice. For example, the maximum size of the ternary tree at the intra-slice level can be 32×32. For example, the maximum size of the ternary tree at the inter-slice level can be 128×128. For example, the minimum size (i.e., the minimum size of the binary tree) of the CU corresponding to each node of the binary tree and / or the minimum size (i.e., the minimum size of the ternary tree) of the CU corresponding to each node of the ternary tree can be set to the minimum size of the CU.

[0306] In another example, the maximum size of the binary tree and / or the maximum size of the ternary tree can be signaled or determined at the slice level. In addition, the minimum size of the binary tree and / or the minimum size of the ternary tree can be signaled or determined at the slice level.

[0307] Based on the various block sizes and depths described above, the quad-partition information, the information indicating whether the partitioning by the multi-type tree is performed, the partition tree information, and / or the partition direction information may or may not be present in the bitstream.

[0308] For example, when the size of the CU is not greater than the minimum size of the quadtree, the CU may not include the quad-partition information, and the quad-partition information of the CU can be inferred as a second value.

[0309] For example, when the size (horizontal size and vertical size) of the CU corresponding to each node of the multi-type tree is greater than the maximum size (horizontal size and vertical size) of the binary tree and / or the maximum size (horizontal size and vertical size) of the ternary tree, the CU may not be partitioned in the form of the binary tree and / or the ternary tree. By this determination method, the information indicating whether the partitioning by the multi-type tree is performed may not be signaled, but can be inferred as a second value.

[0310] Optionally, when the size (horizontal size and vertical size) of the CU corresponding to each node of the multi-type tree is equal to the minimum size (horizontal size and vertical size) of the binary tree, or when the size (horizontal size and vertical size) of the CU is equal to twice the minimum size (horizontal size and vertical size) of the ternary tree, the CU may not be partitioned in the form of a binary tree and / or a ternary tree. By this determination, the information indicating whether the partitioning according to the multi-type tree is performed may not be signaled, but may be inferred as a second value. The reason is that when the CU is partitioned in the form of a binary tree and / or a ternary tree, CUs smaller than the minimum size of the binary tree and / or the minimum size of the ternary tree are generated.

[0311] Optionally, the binary tree partitioning or the ternary tree partitioning may be restricted based on the size of the virtual pipeline data unit (i.e., the size of the pipeline buffer). For example, when the CU is partitioned into sub-CUs that do not fit the size of the pipeline buffer by binary tree partitioning or ternary tree partitioning, the binary tree partitioning or the ternary tree partitioning may be restricted. The size of the pipeline buffer may be equal to the maximum size of the transform block (e.g., 64×64).

[0312] For example, when the size of the pipeline buffer is 64×64, the following partitioning may be restricted.

[0313] - Ternary tree partitioning for a CU of N×M (where N and / or M is 128) - Horizontal binary tree partitioning for a CU of 128×N (where N<=64) - Vertical binary tree partitioning for a CU of N×128 (where N<=64) Optionally, when the depth of the CU corresponding to each node of the multi-type tree is equal to the maximum depth of the multi-type tree, the CU may not be partitioned in the form of a binary tree and / or a ternary tree. By this determination, the information indicating whether the partitioning according to the multi-type tree is performed may not be signaled, but may be inferred as a second value.

[0314] Optionally, the information indicating whether the partitioning according to the multi-type tree is performed may be signaled only when at least one of the vertical binary tree partitioning, the horizontal binary tree partitioning, the vertical ternary tree partitioning, and the horizontal ternary tree partitioning is possible for the CU corresponding to each node of the multi-type tree. Otherwise, the CU may not be partitioned in the form of a binary tree and / or a ternary tree. By this determination, the information indicating whether the partitioning according to the multi-type tree is performed may not be signaled, but may be inferred as a second value.

[0315] Optionally, for a CU corresponding to each node of a multi-type tree, the partitioning direction information may be signaled only when both vertical binary tree partitioning and horizontal binary tree partitioning are feasible or only when both vertical ternary tree partitioning and horizontal ternary tree partitioning are feasible. Otherwise, the partitioning direction information may not be signaled, but may be inferred as a value indicating the direction in which the CU may be partitioned.

[0316] Optionally, for a CU corresponding to each node of a multi-type tree, the partitioning tree information may be signaled only when both vertical binary tree partitioning and vertical ternary tree partitioning are feasible or only when both horizontal binary tree partitioning and horizontal ternary tree partitioning are feasible. Otherwise, the partitioning tree information may not be signaled, but may be inferred as a value indicating the tree of the partitioning applicable to the CU.

[0317] Figure 4 is a diagram showing the forms of prediction units that a coding unit can include.

[0318] Among the CUs partitioned from an LCU, the CUs that are no longer partitioned may be divided into one or more prediction units (PUs). This division is also referred to as "partitioning".

[0319] A PU can be a basic unit for prediction. A PU can be encoded and decoded in any one of the skip mode, inter-frame mode, and intra-frame mode. A PU can be partitioned into various shapes according to each mode. For example, the target blocks described above with reference to Figure 1 and the target blocks described above with reference to Figure 2 can both be PUs.

[0320] A CU may not be divided into PUs. When a CU is not divided into PUs, the size of the CU and the size of the PU may be equal to each other.

[0321] In the skip mode, there may be no partitioning in the CU. In the skip mode, the 2N×2N mode 410 may be supported without partitioning, where in the 2N×2N mode 410, the size of the PU and the size of the CU are the same.

[0322] In the inter-frame mode, there may be 8 types of partitioning shapes in the CU. For example, in the inter-frame mode, the 2N×2N mode 410, 2N×N mode 415, N×2N mode 420, N×N mode 425, 2N×nU mode 430, 2N×nD mode 435, nL×2N mode 440, and nR×2N mode 445 may be supported.

[0323] In the intra-frame mode, the 2N×2N mode 410 and N×N mode 425 may be supported.

[0324] In the 2N×2N mode 410, a PU with a size of 2N×2N can be encoded. The PU with a size of 2N×2N can represent a PU having the same size as the CU. For example, the PU with a size of 2N×2N can have a size of 64×64, 32×32, 16×16, or 8×8.

[0325] In the N×N mode 425, a PU with a size of N×N can be encoded.

[0326] For example, in intra prediction, when the size of the PU is 8×8, the PUs in four partitions can be encoded. The size of each partitioned PU can be 4×4.

[0327] When encoding a PU in the intra mode, any one of multiple intra prediction modes can be used to encode the PU. For example, the HEVC technology can provide 35 intra prediction modes, and the PU can be encoded under any one of the 35 intra prediction modes.

[0328] It is possible to determine which of the 2N×2N mode 410 and the N×N mode 425 will be used to encode the PU based on the rate-distortion cost.

[0329] The encoding device 100 can perform an encoding operation on a PU with a size of 2N×2N. Here, the encoding operation can be an operation of encoding the PU under each of the multiple intra prediction modes that can be used by the encoding device 100. Through the encoding operation, the optimal intra prediction mode for the PU with a size of 2N×2N can be derived. The optimal intra prediction mode can be the intra prediction mode that incurs the minimum rate-distortion cost when encoding the PU with a size of 2N×2N among the multiple intra prediction modes that can be used by the encoding device 100.

[0330] In addition, the encoding device 100 can sequentially perform an encoding operation on each PU obtained by performing N×N partitioning. Here, the encoding operation can be an operation of encoding the PU under each of the multiple intra prediction modes that can be used by the encoding device 100. Through the encoding operation, the optimal intra prediction mode for the PU with a size of N×N can be derived. The optimal intra prediction mode can be the intra prediction mode that incurs the minimum rate-distortion cost when encoding the PU with a size of N×N among the multiple intra prediction modes that can be used by the encoding device 100.

[0331] The encoding device 100 can determine which of the PU with a size of 2N×2N and the PU with a size of N×N will be encoded based on a comparison between the rate-distortion cost of the PU with a size of 2N×2N and the rate-distortion cost of the PU with a size of N×N.

[0332] A single CU can be partitioned into one or more PUs, and a PU can be partitioned into multiple PUs.

[0333] For example, when a single PU is partitioned into four PUs, the horizontal size and vertical size of each of the four PUs generated by the partition can be half of the horizontal size and vertical size of the PU before partitioning. When a PU with a size of 32×32 is partitioned into four PUs, the size of each of the four partitioned PUs can be 16×16. When a single PU is partitioned into four PUs, it can be considered that the PU has been partitioned in a quadtree structure.

[0334] For example, when a single PU is partitioned into two PUs, the horizontal size or vertical size of each of the two PUs generated by the partition can be half of the horizontal size or vertical size of the PU before partitioning. When a PU with a size of 32×32 is vertically partitioned into two PUs, the size of each of the two partitioned PUs can be 16×32. When a PU with a size of 32×32 is horizontally partitioned into two PUs, the size of each of the two partitioned PUs can be 32×16. When a single PU is partitioned into two PUs, it can be considered that the PU has been partitioned in a binary tree structure.

[0335] Figure 5 is a diagram showing the form of transform units that can be included in a coding unit.

[0336] A transform unit (TU) can be the basic unit in a CU that is used for processes such as transformation, quantization, inverse transformation, dequantization, entropy coding, and entropy decoding.

[0337] A TU can have a square shape or a rectangular shape. The shape of the TU can be determined based on the size and / or shape of the CU.

[0338] In the CUs partitioned from an LCU, a CU that is no longer partitioned into CUs can be partitioned into one or more TUs. Here, the partition structure of the TUs can be a quadtree structure. For example, as Figure 5 shown, a single CU 510 can be partitioned one or more times according to the quadtree structure. Through this partition, a single CU 510 can be composed of TUs with various sizes.

[0339] It can be considered that when a single CU is divided two or more times, the CU is recursively divided. Through the division, a single CU can be composed of transform units (TUs) with various sizes.

[0340] Optionally, a single CU can be divided into one or more TUs based on the number of vertical lines and / or horizontal lines for partitioning the CU.

[0341] A CU can be divided into a symmetric TU or an asymmetric TU. To divide into an asymmetric TU, information about the size and / or shape of each TU can be signaled from the encoding device 100 to the decoding device 200. Optionally, the size and / or shape of each TU can be derived from the information about the size and / or shape of the CU.

[0342] A CU may not be divided into TUs. When a CU is not divided into TUs, the size of the CU and the size of the TU may be equal to each other.

[0343] A single CU can be partitioned into one or more TUs, and a TU can be partitioned into multiple TUs.

[0344] For example, when a single TU is partitioned into four TUs, the horizontal size and vertical size of each of the four TUs generated by the partitioning can be half of the horizontal size and vertical size of the TU before partitioning. When a TU with a size of 32×32 is partitioned into four TUs, the size of each of the four partitioned TUs can be 16×16. When a single TU is partitioned into four TUs, it can be considered that the TU has been partitioned in a quadtree structure.

[0345] For example, when a single TU is partitioned into two TUs, the horizontal size or vertical size of each of the two TUs generated by the partitioning can be half of the horizontal size or vertical size of the TU before partitioning. When a TU with a size of 32×32 is vertically partitioned into two TUs, the size of each of the two partitioned TUs can be 16×32. When a TU with a size of 32×32 is horizontally partitioned into two TUs, the size of each of the two partitioned TUs can be 32×16. When a single TU is partitioned into two TUs, it can be considered that the TU has been partitioned in a binary tree structure.

[0346] The CU can be divided in a way different from Figure 5 the way shown in

[0347] For example, a single CU can be divided into three CUs. The horizontal size or vertical size of the three CUs generated by the division can be 1 / 4, 1 / 2, and 1 / 4 of the horizontal size or vertical size of the original CU before division, respectively.

[0348] For example, when a CU with a size of 32×32 is vertically divided into three CUs, the sizes of the three CUs generated by the division can be 8×32, 16×32, and 8×32, respectively. In this way, when a single CU is divided into three CUs, it can be considered that the CU is divided in the form of a ternary tree.

[0349] One of the exemplary partitioning forms (i.e., quadtree partitioning, binary tree partitioning, and ternary tree partitioning) can be applied to the partitioning of the CU, and multiple partitioning schemes can be combined and used together for the partitioning of the CU. Here, the case of combining and using multiple partitioning schemes together can be referred to as "compound tree form partitioning".

[0350] Figure 6 Shows the partitioning of a block according to an example.

[0351] In video encoding and / or decoding processing, as Figure 6 shown, the target block can be partitioned. For example, the target block can be a CU.

[0352] For the partitioning of the target block, an indicator indicating the partitioning information can be signaled from the encoding device 100 to the decoding device 200. The partitioning information can be information indicating how the target block is partitioned.

[0353] The partitioning information can be one or more of a partitioning flag (hereinafter referred to as "split_flag"), a quadtree-binary flag (hereinafter referred to as "QB_flag"), a quadtree flag (hereinafter referred to as "quadtree_flag"), a binary tree flag (hereinafter referred to as "binarytree_flag"), and a binary type flag (hereinafter referred to as "Btype_flag").

[0354] "split_flag" can be a flag indicating whether the block is partitioned. For example, a split_flag value of 1 can indicate that the corresponding block is partitioned. A split_flag value of 0 can indicate that the corresponding block is not partitioned.

[0355] "QB_flag" can be a flag indicating which of the quadtree form and the binary tree form corresponds to the shape in which the block is partitioned. For example, a QB_flag value of 0 can indicate that the block is partitioned in the quadtree form. A QB_flag value of 1 can indicate that the block is partitioned in the binary tree form. Alternatively, a QB_flag value of 0 can indicate that the block is partitioned in the binary tree form. A QB_flag value of 1 can indicate that the block is partitioned in the quadtree form.

[0356] "quadtree_flag" can be a flag indicating whether the block is partitioned in the quadtree form. For example, a quadtree_flag value of 1 can indicate that the block is partitioned in the quadtree form. A quadtree_flag value of 0 can indicate that the block is not partitioned in the quadtree form.

[0357] "binarytree_flag" can be a flag indicating whether a block is divided in a binary tree form. For example, a binarytree_flag value of 1 can indicate that the block is divided in a binary tree form. A binarytree_flag value of 0 can indicate that the block is not divided in a binary tree form.

[0358] "Btype_flag" can be a flag indicating which one of the vertical division and the horizontal division corresponds to the division direction when the block is divided in a binary tree form. For example, a Btype_flag value of 0 can indicate that the block is divided in the horizontal direction. A Btype_flag value of 1 can indicate that the block is divided in the vertical direction. Optionally, a Btype_flag value of 0 can indicate that the block is divided in the vertical direction. A Btype_flag value of 1 can indicate that the block is divided in the horizontal direction.

[0359] For example, the division information of the block in Figure 6 can be derived by signaling at least one of quadtree_flag, binarytree_flag, and Btype_flag, as shown in Table 1 below.

[0360] Table 1

[0361] For example, the division information of the block in Figure 6 can be derived by signaling at least one of split_flag, QB_flag, and Btype_flag, as shown in Table 2 below.

[0362] Table 2

[0363] The division method may be limited to quadtree or binary tree according to the size and / or shape of the block. When this limitation is applied, split_flag can be a flag indicating whether the block is divided in a quadtree form or a flag indicating whether the block is divided in a binary tree form. The size and shape of the block can be derived from the depth information of the block, and the depth information can be signaled from the encoding device 100 to the decoding device 200.

[0364] When the size of the block falls within a specific range, it is only possible to divide in a quadtree form. For example, the specific range can be defined by at least one of the maximum block size and the minimum block size that can only be divided in a quadtree form.

[0365] Information indicating the maximum block size and the minimum block size that can be partitioned only in the form of a quadtree can be signaled from the encoding device 100 to the decoding device 200 via a bitstream. In addition, this information can be signaled for at least one of units such as video, sequence, picture, parameter, parallel block group, and slice (or segment).

[0366] Optionally, the maximum block size and / or the minimum block size can be fixed sizes predefined by the encoding device 100 and the decoding device 200. For example, when the size of a block is greater than 64×64 and less than 256×256, partitioning only in the form of a quadtree is possible. In this case, split_flag can be a flag indicating whether to perform partitioning in the form of a quadtree.

[0367] When the size of a block is greater than the maximum size of a transform block, partitioning only in the form of a quadtree is possible. Here, the sub-blocks generated by the partitioning can be at least one of a CU and a TU.

[0368] In this case, split_flag can be a flag indicating whether the CU is partitioned in the form of a quadtree.

[0369] When the size of a block falls within a specific range, partitioning only in the form of a binary tree or a ternary tree is possible. For example, the specific range can be defined by at least one of the maximum block size and the minimum block size that can be partitioned only in the form of a binary tree or a ternary tree.

[0370] Information indicating the maximum block size and / or the minimum block size that can be partitioned only in the form of a binary tree or in the form of a ternary tree can be signaled from the encoding device 100 to the decoding device 200 via a bitstream. In addition, this information can be signaled for at least one of units such as sequence, picture, and slice (or segment).

[0371] Optionally, the maximum block size and / or the minimum block size can be fixed sizes predefined by the encoding device 100 and the decoding device 200. For example, when the size of a block is greater than 8×8 and less than 16×16, partitioning only in the form of a binary tree is possible. In this case, split_flag can be a flag indicating whether to perform partitioning in the form of a binary tree or a ternary tree.

[0372] The above description of partitioning in the form of a quadtree can be equally applied to the form of a binary tree and / or the form of a ternary tree.

[0373] The partitioning of a block may be restricted by a previous partitioning. For example, when a block is partitioned in a specific binary tree form and multiple sub-blocks are generated from the partitioning, each sub-block may be further partitioned only in a specific tree form. Here, the specific tree form may be at least one of a binary tree form, a ternary tree form, and a quaternary tree form.

[0374] When the horizontal size or the vertical size of a partitioned block is a size that cannot be further divided, the above indicator may not be signaled.

[0375] Figure 7 is a diagram for explaining an embodiment of intra prediction processing.

[0376] From Figure 7 The arrow extending radially from the center of the diagram in indicates the prediction direction of the intra prediction mode. In addition, the numbers appearing near the arrow indicate examples of the mode values assigned to the intra prediction mode or the prediction direction of the intra prediction mode.

[0377] In Figure 7 , the number 0 may represent the planar mode as a non-directional intra prediction mode. The number 1 may represent the DC mode as a non-directional intra prediction mode.

[0378] Intra coding and / or decoding may be performed using reference samples of neighboring blocks of a target block. The neighboring blocks may be reconstructed neighboring blocks. The reference samples may represent neighboring samples.

[0379] For example, intra coding and / or decoding may be performed using the values of the reference samples included in the reconstructed neighboring blocks or the coding parameters of the reconstructed neighboring blocks.

[0380] The encoding device 100 and / or the decoding device 200 may generate a prediction block by performing intra prediction on a target block based on information about samples in the target image. When intra prediction is performed, the encoding device 100 and / or the decoding device 200 may generate a prediction block for the target block by performing intra prediction based on information about samples in the target image. When intra prediction is performed, the encoding device 100 and / or the decoding device 200 may perform directional prediction and / or non-directional prediction based on at least one reconstructed reference sample.

[0381] The prediction block may be a block generated as a result of performing intra prediction. The prediction block may correspond to at least one of a CU, a PU, and a TU.

[0382] The unit of the prediction block may have a size corresponding to at least one of a CU, a PU, and a TU. The prediction block may have a square shape with a size of 2N×2N or N×N. The size N×N may include sizes such as 4×4, 8×8, 16×16, 32×32, 64×64, etc.

[0383] Optionally, the prediction block may be a square block with a size of 2×2, 4×4, 8×8, 16×16, 32×32, 64×64, etc. or a rectangular block with a size of 2×8, 4×8, 2×16, 4×16, 8×16, etc.

[0384] Intra prediction may be performed considering the intra prediction mode for the target block. The number of intra prediction modes that the target block may have may be a predefined fixed value and may be a value determined differently according to the attributes of the prediction block. For example, the attributes of the prediction block may include the size of the prediction block, the type of the prediction block, etc. In addition, the attributes of the prediction block may indicate coding parameters for the prediction block.

[0385] For example, regardless of the size of the prediction block, the number of intra prediction modes may be fixed to N. Optionally, the number of intra prediction modes may be, for example, 3, 5, 9, 17, 34, 35, 36, 65, 67, or 95.

[0386] The intra prediction mode may be a non-directional mode or a directional mode.

[0387] For example, the intra prediction mode may include two non-directional modes corresponding to the numbers 0 to 66 shown in Figure 7 and 65 directional modes.

[0388] For example, in the case of using a specific intra prediction method, the intra prediction mode may include two non-directional modes corresponding to the numbers -14 to 80 shown in Figure 7 and 93 directional modes.

[0389] The two non-directional modes may include a DC mode and a planar mode.

[0390] The directional mode may be a prediction mode having a specific direction or a specific angle. The directional mode may also be referred to as an "angle mode".

[0391] The intra prediction mode may be represented by at least one of a mode number, a mode value, a mode angle, and a mode direction. In other words, the terms "(mode) number of the intra prediction mode", "(mode) value of the intra prediction mode", "(mode) angle of the intra prediction mode", and "(mode) direction of the intra prediction mode" may be used with the same meaning and may be used interchangeably with each other.

[0392] The number of intra prediction modes may be M. The value of M may be 1 or greater. In other words, the number of intra prediction modes may be M, where M includes the number of non-directional modes and the number of directional modes.

[0393] The number of intra prediction modes can be fixed to M, regardless of the size of the block and / or the color component. For example, the number of intra prediction modes can be fixed to either 35 or 67, regardless of the size of the block.

[0394] Optionally, the number of intra prediction modes can vary according to the shape, size, and / or type of the color component of the block.

[0395] For example, in Figure 7 the directional prediction mode shown by the dashed line can be applied only to the prediction for non-square blocks.

[0396] For example, the larger the size of the block, the more intra prediction modes there are. Optionally, the larger the size of the block, the fewer intra prediction modes there are. When the size of the block is 4×4 or 8×8, the number of intra prediction modes can be 67. When the size of the block is 16×16, the number of intra prediction modes can be 35. When the size of the block is 32×32, the number of intra prediction modes can be 19. When the size of the block is 64×64, the number of intra prediction modes can be 7.

[0397] For example, the number of intra prediction modes can vary according to whether the color component is a luminance signal or a chrominance signal. Optionally, the number of intra prediction modes corresponding to the luminance component block can be greater than the number of intra prediction modes corresponding to the chrominance component block.

[0398] For example, in the vertical mode with a mode value of 50, prediction can be performed in the vertical direction based on the pixel values of the reference samples. For example, in the horizontal mode with a mode value of 18, prediction can be performed in the horizontal direction based on the pixel values of the reference samples.

[0399] Even in the directional modes other than the above-mentioned modes, the encoding device 100 and the decoding device 200 can still perform intra prediction on the target unit using the reference samples according to the angle corresponding to the directional mode.

[0400] The intra prediction mode located to the right of the vertical mode can be referred to as the "vertical - right mode". The intra prediction mode located below the horizontal mode can be referred to as the "horizontal - below mode". For example, in Figure 7 the intra prediction mode with a mode value being one of 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, and 66 can be the vertical - right mode. The intra prediction mode with a mode value being one of 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, and 17 can be the horizontal - below mode.

[0401] The non-directional mode may include a DC mode and a planar mode. For example, the value of the DC mode may be 1. The value of the planar mode may be 0.

[0402] The directional mode may include an angular mode. Among the multiple intra-frame prediction modes, the remaining modes other than the DC mode and the planar mode may be directional modes.

[0403] When the intra-frame prediction mode is the DC mode, a prediction block may be generated based on the average value of the pixel values of multiple reference pixels. For example, the pixel value of the prediction block may be determined based on the average value of the pixel values of multiple reference pixels.

[0404] The number of the above-described intra-frame prediction modes and the mode values of each intra-frame prediction mode are merely exemplary. The number of the above-described intra-frame prediction modes and the mode values of each intra-frame prediction mode may be defined differently according to embodiments, implementations, and / or requirements.

[0405] To perform intra-frame prediction on a target block, a step of checking whether the samples included in the reconstructed neighboring blocks can be used as reference samples for the target block may be performed. When there are samples among the samples in the neighboring blocks that cannot be used as reference samples for the target block, the value generated by interpolation and / or copying of at least one of the sample values among the samples included in the reconstructed neighboring blocks may replace the sample value of the sample that cannot be used as a reference sample. When the value generated by copying and / or interpolation replaces the sample value of the existing sample, the sample may be used as a reference sample for the target block.

[0406] When intra-frame prediction is used, a filter may be applied to at least one of the reference samples and the prediction samples based on at least one of the size of the target block and the intra-frame prediction mode.

[0407] The type of the filter to be applied to at least one of the reference samples and the prediction samples may vary according to at least one of the intra-frame prediction mode of the target block, the size of the target block, and the shape of the target block. The type of the filter may be classified according to one or more of the length of the filter taps, the value of the filter coefficients, and the filter strength. The length of the filter taps may represent the number of the filter taps. In addition, the number of the filter taps may represent the length of the filter.

[0408] When the intra-frame prediction mode is the planar mode, when generating the prediction block of the target block, the sample value of the prediction target block may be generated using the weighted sum of the upper reference sample of the target block, the left reference sample of the target block, the upper right reference sample of the target block, and the lower left reference sample of the target block according to the position of the prediction target sample in the prediction block.

[0409] When the intra prediction mode is the DC mode, the average of the reference samples above the target block and the reference samples to the left of the target block can be used when generating the prediction block of the target block. In addition, filtering using the values of the reference samples can be performed on a specific row or a specific column in the target block. The specific row can be one or more upper rows adjacent to the reference samples. The specific column can be one or more left columns adjacent to the reference samples.

[0410] When the intra prediction mode is the direction mode, the upper reference sample, the left reference sample, the upper-right reference sample, and / or the lower-left reference sample of the target block can be used to generate the prediction block.

[0411] To generate the above prediction samples, interpolation based on real numbers can be performed.

[0412] The intra prediction mode of the target block can be predicted from the intra prediction modes of neighboring blocks adjacent to the target block, and the information used for the prediction can be entropy-coded / entropy-decoded.

[0413] For example, when the intra prediction modes of the target block and the neighboring block are the same, a predefined flag can be used to signal that the intra prediction modes of the target block and the neighboring block are the same.

[0414] For example, an indicator indicating the intra prediction mode that is the same as the intra prediction mode of the target block among the intra prediction modes of a plurality of neighboring blocks can be signaled.

[0415] When the intra prediction modes of the target block and the neighboring block are different from each other, the information about the intra prediction mode of the target block can be encoded and / or decoded using entropy coding and / or entropy decoding.

[0416] Figure 8 is a diagram showing the reference samples used in the intra prediction process.

[0417] The reconstructed reference samples used for intra prediction of the target block can include the lower-left reference sample, the left reference sample, the upper-left reference sample, the upper reference sample, and the upper-right reference sample.

[0418] For example, the left reference sample can represent a reconstructed reference pixel adjacent to the left side of the target block. The upper reference sample can represent a reconstructed reference pixel adjacent to the top of the target block. The upper-left reference sample can represent a reconstructed reference pixel located at the upper left corner of the target block. The lower-left reference sample can represent a reference sample located below the left sample line among the samples on the same line as the left sample line composed of the left reference samples. The upper-right reference sample can represent a reference sample located to the right of the upper sample line among the samples on the same line as the upper sample line composed of the upper reference samples.

[0419] When the size of the target block is N×N, the number of the lower-left reference sample points, left reference sample points, upper reference sample points, and upper-right reference sample points can all be N.

[0420] By performing intra prediction on the target block, a prediction block can be generated. The process of generating the prediction block can include determining the values of the pixels in the prediction block. The sizes of the target block and the prediction block can be the same.

[0421] The reference sample points for intra prediction of the target block can be changed according to the intra prediction mode of the target block. The direction of the intra prediction mode can represent the dependency relationship between the reference sample points and the pixels of the prediction block. For example, the value of the specified reference sample point can be used as the value of one or more specified pixels in the prediction block. In this case, the specified reference sample point and the one or more specified pixels in the prediction block can be the sample points and pixels located on a straight line along the direction of the intra prediction mode. In other words, the value of the specified reference sample point can be copied as the value of the pixels located in the direction opposite to the direction of the intra prediction mode. Alternatively, the value of the pixel in the prediction block can be the value of the reference sample point located in the direction of the intra prediction mode with respect to the position of the pixel.

[0422] In an example, when the intra prediction mode of the target block is the vertical mode, the upper reference sample points can be used for intra prediction. When the intra prediction mode is the vertical mode, the value of the pixel in the prediction block can be the value of the reference sample point vertically located above the position of the pixel. Therefore, the upper reference sample points adjacent to the top of the target block can be used for intra prediction. In addition, the values of the pixels in a row of the prediction block can be the same as the values of the pixels of the upper reference sample points.

[0423] In an example, when the intra prediction mode of the target block is the horizontal mode, the left reference sample points can be used for intra prediction. When the intra prediction mode is the horizontal mode, the value of the pixel in the prediction block can be the value of the reference sample point horizontally located to the left of the position of the pixel. Therefore, the left reference sample points adjacent to the left side of the target block can be used for intra prediction. In addition, the values of the pixels in a column of the prediction block can be the same as the values of the pixels of the left reference sample points.

[0424] In an example, when the mode value of the intra prediction mode of the current block is 34, at least some of the left reference sample points, the upper-left reference sample points, and at least some of the upper reference sample points can be used for intra prediction. When the mode value of the intra prediction mode is 34, the value of the pixel in the prediction block can be the value of the reference sample point diagonally located at the upper left corner of the pixel.

[0425] In addition, in the case of the intra prediction mode with the mode value in the range from 52 to 66, at least a part of the upper-right reference sample points can be used for intra prediction.

[0426] In addition, in the case of an intra prediction mode with a mode value in the range from 2 to 17, at least a part of the lower left reference samples can be used for intra prediction.

[0427] In addition, in the case of an intra prediction mode with a mode value in the range from 19 to 49, the upper left reference sample can be used for intra prediction.

[0428] The number of reference samples used to determine the pixel value of a pixel in the prediction block can be 1 or 2 or more.

[0429] As described above, the pixel value of a pixel in the prediction block can be determined according to the position of the pixel and the position of the reference sample indicated by the direction of the intra prediction mode. When the position of the pixel and the position of the reference sample indicated by the direction of the intra prediction mode are integer positions, the value of a reference sample indicated by the integer position can be used to determine the pixel value of the pixel in the prediction block.

[0430] When the position of the pixel and the position of the reference sample indicated by the direction of the intra prediction mode are not integer positions, an interpolated reference sample based on the two reference samples closest to the position of the reference sample can be generated. The value of the interpolated reference sample can be used to determine the pixel value of the pixel in the prediction block. In other words, when the position of the pixel in the prediction block and the position of the reference sample indicated by the direction of the intra prediction mode indicate a position between two reference samples, an interpolation based on the values of these two samples can be generated.

[0431] The prediction block generated via prediction can be different from the original target block. In other words, there may be a prediction error, which is the difference between the target block and the prediction block, and there may also be a prediction error between the pixels of the target block and the pixels of the prediction block.

[0432] Hereinafter, the terms "difference", "error" and "residual" can be used with the same meaning and can be used interchangeably with each other.

[0433] For example, in the case of directional intra prediction, the longer the distance between the pixels of the prediction block and the reference sample, the greater the possible prediction error. Such a prediction error can lead to discontinuity between the generated prediction block and adjacent blocks.

[0434] To reduce the prediction error, a filtering operation for the prediction block can be used. The filtering operation can be configured to adaptively apply a filter to a region in the prediction block that is considered to have a large prediction error. For example, the region considered to have a large prediction error can be the boundary of the prediction block. In addition, the region in the prediction block considered to have a large prediction error can be different according to the intra prediction mode, and the characteristics of the filter can also be different according to the intra prediction mode.

[0435] AsFigure 8 As shown, for intra prediction of a target block, at least one of reference lines 0 to 3 can be used.

[0436] In Figure 8 each of the reference lines can indicate a reference sample line including one or more reference samples. When the number of the reference line is small, it can indicate a reference sample line closer to the target block.

[0437] The samples in segment A and segment F can be obtained by padding instead of from the reconstructed neighboring blocks, where the padding uses the samples in segment B and segment E that are closest to the target block.

[0438] Index information indicating the reference sample line to be used for intra prediction of the target block can be signaled. The index information can indicate the reference sample line among the multiple reference sample lines that will be used for intra prediction of the target block. For example, the index information can have a value corresponding to any one of 0 to 3.

[0439] When the upper boundary of the target block is the boundary of the CTU, only reference line 0 can be available. Therefore, in this case, the index information may not be signaled. When additional reference sample lines other than reference line 0 are used, filtering of the prediction block described later may not be performed.

[0440] In the case of inter-color intra prediction, a prediction block of the target block of the second color component can be generated based on the corresponding reconstructed block of the first color component.

[0441] For example, the first color component can be a luminance component, and the second color component can be a chrominance component.

[0442] To perform inter-color intra prediction, parameters of a linear model between the first color component and the second color component can be derived based on a template.

[0443] The template can include reference samples (upper reference samples) above the target block and / or reference samples (left reference samples) to the left of the target block, and can include upper reference samples and / or left reference samples of the reconstructed blocks of the first color component corresponding to the reference samples.

[0444] For example, the following values can be used to derive the parameters of the linear model: 1) the value of the sample of the first color component having the maximum value among the samples in the template, 2) the value of the sample of the second color component corresponding to the sample of the first color component, 3) the value of the sample of the first color component having the minimum value among the samples in the template, and 4) the value of the sample of the second color component corresponding to the sample of the first color component.

[0445] When exporting the parameters of a linear model, a prediction block of a target block can be generated by applying a corresponding reconstruction block to the linear model.

[0446] According to the image format, subsampling can be performed on the samples adjacent to the reconstruction block of the first color component and the corresponding reconstruction block of the first color component. For example, when one sample of the second color component corresponds to four samples of the first color component, one corresponding sample can be calculated by performing subsampling on the four samples of the first color component. When performing subsampling, the derivation of the parameters of the linear model and the inter-color intra prediction can be performed based on the corresponding samples to be subsampled.

[0447] Information about whether to perform inter-color intra prediction and / or the range of the template can be signaled in the intra prediction mode.

[0448] The target block can be partitioned into two or four sub-blocks in the horizontal direction and / or the vertical direction.

[0449] The sub-blocks generated by partitioning can be reconstructed sequentially. That is, when performing intra prediction on each sub-block, a sub-prediction block of the sub-block can be generated. In addition, when performing inverse quantization and / or inverse transformation on each sub-block, a sub-residual block for the corresponding sub-block can be generated. The reconstructed sub-block can be generated by adding the sub-prediction block and the sub-residual block. The reconstructed sub-block can be used as a reference sample for the intra prediction of the sub-block with the next priority.

[0450] The sub-block can be a block including a specific number (e.g., 16) or more samples. For example, when the target block is an 8×4 block or a 4×8 block, the target block can be partitioned into two sub-blocks. In addition, when the target block is a 4×4 block, the target block cannot be partitioned into sub-blocks. When the target block has another size, the target block can be partitioned into four sub-blocks.

[0451] Information about whether to perform intra prediction based on these sub-blocks and / or information about the partitioning direction (horizontal direction or vertical direction) can be signaled.

[0452] This intra prediction based on sub-blocks can be restricted so that it is only performed when the reference sample line 0 is used. When performing intra prediction based on sub-blocks, the filtering of the prediction block described below may not be performed.

[0453] The final prediction block can be generated by performing filtering on the prediction block generated via intra prediction.

[0454] Filtering can be performed by applying specific weights to the filtering target samples, the left reference sample, the upper reference sample, and / or the upper left reference sample that are the targets to be filtered.

[0455] The weights for filtering and / or the reference samples (e.g., the range of reference samples, the positions of reference samples, etc.) can be determined based on at least one of the block size, the intra prediction mode, and the positions of the filtered target samples in the prediction block.

[0456] For example, filtering can be performed only in specific intra prediction modes (e.g., DC mode, planar mode, vertical mode, horizontal mode, diagonal mode, and / or adjacent diagonal mode).

[0457] The adjacent diagonal mode can be a mode having a number obtained by adding k to the number of the diagonal mode, and can be a mode having a number obtained by subtracting k from the number of the diagonal mode. In other words, the number of the adjacent diagonal mode can be the sum of the number of the diagonal mode and k, or can be the difference between the number of the diagonal mode and k. For example, k can be a positive integer of 8 or less.

[0458] The intra prediction mode of the target block can be derived using the intra prediction modes of the neighboring blocks appearing around the target block, and entropy coding and / or entropy decoding can be performed on such derived intra prediction mode.

[0459] For example, when the intra prediction mode of the target block is the same as the intra prediction mode of the neighboring block, specific flag information can be used to signal information indicating that the intra prediction mode of the target block is the same as the intra prediction mode of the neighboring block.

[0460] In addition, for example, indicator information of the neighboring blocks having the same intra prediction mode as the intra prediction mode of the target block among the intra prediction modes of multiple neighboring blocks can be signaled.

[0461] For example, when the intra prediction mode of the target block is different from the intra prediction mode of the neighboring block, entropy coding and / or entropy decoding can be performed on the information regarding the intra prediction mode of the target block by performing entropy coding and / or entropy decoding based on the intra prediction mode of the neighboring block.

[0462] Figure 9 is a diagram for explaining an embodiment of the inter prediction process.

[0463] Figure 9 The rectangle shown in can represent an image (or a picture). In addition, in Figure 9 The arrow can represent the prediction direction. The arrow pointing from the first picture to the second picture indicates that the second picture references the first picture. That is, each image can be encoded and / or decoded according to the prediction direction.

[0464] An image can be classified as an intra picture (I picture), a single predicted picture or a predicted coded picture (P picture), and a bi-predicted picture or a bi-predicted coded picture (B picture) according to the coding type. Each picture can be coded and / or decoded according to the coding type of each picture.

[0465] When the target image to be coded is an I picture, the target image can be coded using the data included in the image itself without performing inter prediction with reference to other images. For example, an I picture can be coded only via intra prediction.

[0466] When the target image is a P picture, the target image can be coded via inter prediction using a reference picture existing in one direction. Here, the one direction can be a forward direction or a backward direction.

[0467] When the target image is a B picture, the image can be coded via inter prediction using reference pictures existing in two directions, or the image can be coded via inter prediction using a reference picture existing in one of the forward direction and the backward direction. Here, the two directions can be the forward direction and the backward direction.

[0468] P pictures and B pictures coded and / or decoded using reference pictures can be regarded as images using inter prediction.

[0469] Hereinafter, inter prediction in an inter mode according to an embodiment will be described in detail.

[0470] Inter prediction or motion compensation can be performed using a reference image and motion information.

[0471] In the inter mode, the encoding device 100 can perform inter prediction and / or motion compensation on a target block. The decoding device 200 can perform inter prediction and / or motion compensation corresponding to the inter prediction and / or motion compensation performed by the encoding device 100 on the target block.

[0472] Motion information of a target block can be derived separately by the encoding device 100 and the decoding device 200 during inter prediction. Motion information can be derived using motion information of a reconstructed neighboring block, motion information of a col block, and / or motion information of a block adjacent to the col block.

[0473] For example, the encoding device 100 or the decoding device 200 can perform prediction and / or motion compensation by using motion information of a spatial candidate and / or a temporal candidate as the motion information of the target block. The target block can represent a PU and / or a PU partition.

[0474] The spatial candidate can be a reconstructed block spatially adjacent to the target block.

[0475] The temporal candidate may be a reconstructed block corresponding to the target block in a previously reconstructed co-located picture (col picture).

[0476] In inter prediction, the encoding device 100 and the decoding device 200 may improve the encoding efficiency and the decoding efficiency by using the motion information of spatial candidates and / or temporal candidates. The motion information of spatial candidates may be referred to as "spatial motion information". The motion information of temporal candidates may be referred to as "temporal motion information".

[0477] Hereinafter, the motion information of spatial candidates may be the motion information of PUs including spatial candidates. The motion information of temporal candidates may be the motion information of PUs including temporal candidates. The motion information of candidate blocks may be the motion information of PUs including candidate blocks.

[0478] Inter prediction may be performed using a reference picture.

[0479] The reference picture may be at least one of a picture before the target picture and a picture after the target picture. The reference picture may be an image for prediction of the target block.

[0480] In inter prediction, a reference picture index (or refIdx) for indicating a reference picture, a motion vector to be described later, etc. may be used to specify a region in the reference picture. Here, the region specified in the reference picture may indicate a reference block.

[0481] Inter prediction may select a reference picture, and may also select a reference block corresponding to the target block from the reference picture. In addition, inter prediction may use the selected reference block to generate a prediction block for the target block.

[0482] Motion information may be derived by each of the encoding device 100 and the decoding device 200 during inter prediction.

[0483] Spatial candidates may be blocks that 1) exist in the target picture, 2) have been previously reconstructed via encoding and / or decoding, and 3) are adjacent to or located at the corners of the target block. Here, a "block located at the corner of the target block" may be a block that is vertically adjacent to a neighboring block horizontally adjacent to the target block, or a block that is horizontally adjacent to a neighboring block vertically adjacent to the target block. In addition, a "block located at the corner of the target block" may have the same meaning as a "block adjacent to the corner of the target block". The meaning of a "block located at the corner of the target block" may be included in the meaning of a "block adjacent to the target block".

[0484] For example, spatial candidates may be a reconstructed block located to the left of the target block, a reconstructed block located above the target block, a reconstructed block located at the lower left corner of the target block, a reconstructed block located at the upper right corner of the target block, or a reconstructed block located at the upper left corner of the target block.

[0485] Each of the encoding device 100 and the decoding device 200 can identify a block at a position in the col picture that spatially corresponds to the target block. The position of the target block in the target picture and the position of the identified block in the col picture can correspond to each other.

[0486] Each of the encoding device 100 and the decoding device 200 can determine a col block at a predefined relative position for the identified block as a temporal candidate. The predefined relative position can be a position inside and / or outside the identified block.

[0487] For example, the col block can include a first col block and a second col block. When the coordinates of the identified block are (xP, yP) and the size of the identified block is represented by (nPSW, nPSH), the first col block can be the block located at the coordinates (xP + nPSW, yP + nPSH). The second col block can be the block located at the coordinates (xP + (nPSW >> 1), yP + (nPSH >> 1)). When the first col block is not available, the second col block can be selectively used.

[0488] The motion vector of the target block can be determined based on the motion vector of the col block. Each of the encoding device 100 and the decoding device 200 can scale the motion vector of the col block. The scaled motion vector of the col block can be used as the motion vector of the target block. In addition, the motion vector of the motion information of the temporal candidate stored in the list can be a scaled motion vector.

[0489] The ratio of the motion vector of the target block to the motion vector of the col block can be the same as the ratio of the first temporal distance to the second temporal distance. The first temporal distance can be the distance between the reference picture and the target picture of the target block. The second temporal distance can be the distance between the reference picture and the col picture of the col block.

[0490] The scheme for deriving motion information can be changed according to the inter-frame prediction mode of the target block. For example, as the inter-frame prediction mode applied to inter-frame prediction, there can be an Advanced Motion Vector Prediction (AMVP) mode, a merge mode, a skip mode, a merge mode with a motion vector difference, a sub-block merge mode, a triangular partitioning mode, an inter-frame / intra-frame combined prediction mode, an affine inter-frame mode, a current picture reference mode, etc. The merge mode can also be referred to as the "motion merge mode". Each mode will be described in detail below.

[0491] 1) AMVP mode When using the AMVP mode, the encoding device 100 may search for similar blocks in the neighboring area of the target block. The encoding device 100 may perform prediction on the target block by using the motion information of the found similar blocks to obtain a predicted block. The encoding device 100 may encode the residual block that is the difference between the target block and the predicted block.

[0492] 1-1) Create a list of predicted motion vector candidates When the AMVP mode is used as a prediction mode, each of the encoding device 100 and the decoding device 200 may use a spatial candidate motion vector, a temporal candidate motion vector, and a zero vector to create a list of predicted motion vector candidates. The list of predicted motion vector candidates may include one or more predicted motion vector candidates. At least one of the spatial candidate motion vector, the temporal candidate motion vector, and the zero vector may be determined and used as a predicted motion vector candidate.

[0493] Hereinafter, the terms "predicted motion vector (candidate)" and "motion vector (candidate)" may be used to have the same meaning and may be used interchangeably with each other.

[0494] Hereinafter, the terms "predicted motion vector candidate" and "AMVP candidate" may be used to have the same meaning and may be used interchangeably with each other.

[0495] Hereinafter, the terms "list of predicted motion vector candidates" and "list of AMVP candidates" may be used to have the same meaning and may be used interchangeably with each other.

[0496] Spatial candidates may include reconstructed spatially neighboring blocks. In other words, the motion vector of the reconstructed neighboring block may be referred to as a "spatial predicted motion vector candidate".

[0497] Temporal candidates may include col blocks and blocks adjacent to the col blocks. In other words, the motion vector of the col block or the motion vector of the block adjacent to the col block may be referred to as a "temporal predicted motion vector candidate".

[0498] The zero vector may be a (0,0) motion vector.

[0499] A predicted motion vector candidate may be a motion vector predictor for predicting a motion vector. In addition, in the encoding device 100, each predicted motion vector candidate may be an initial search position for the motion vector.

[0500] 1-2) Search for the motion vector using the list of predicted motion vector candidates The encoding device 100 can determine a motion vector to be used for encoding a target block within a search range by using a list of predicted motion vector candidates. In addition, the encoding device 100 can determine, among the predicted motion vector candidates present in the predicted motion vector candidate list, a predicted motion vector candidate to be used as the predicted motion vector of the target block.

[0501] The motion vector to be used for encoding the target block can be a motion vector that can be encoded at the minimum cost.

[0502] In addition, the encoding device 100 can determine whether to use the AMVP mode to encode the target block.

[0503] 1-3) Transmission of inter-frame prediction information The encoding device 100 can generate a bitstream including inter-frame prediction information required for inter-frame prediction. The decoding device 200 can perform inter-frame prediction on the target block by using the inter-frame prediction information of the bitstream.

[0504] The inter-frame prediction information can include 1) mode information indicating whether the AMVP mode is used, 2) a predicted motion vector index, 3) a motion vector difference (MVD), 4) a reference direction, and 5) a reference picture index.

[0505] Hereinafter, the terms "predicted motion vector index" and "AMVP index" can be used with the same meaning and can be used interchangeably with each other.

[0506] In addition, the inter-frame prediction information can include a residual signal.

[0507] When the mode information indicates that the AMVP mode is used, the decoding device 200 can obtain the predicted motion vector index, MVD, reference direction, and reference picture index from the bitstream through entropy decoding.

[0508] The predicted motion vector index can indicate a predicted motion vector candidate to be used for predicting the target block among the predicted motion vector candidates included in the predicted motion vector candidate list.

[0509] 1-4) Inter-frame prediction in AMVP mode using inter-frame prediction information The decoding device 200 can use the predicted motion vector candidate list to derive predicted motion vector candidates, and can determine the motion information of the target block based on the derived predicted motion vector candidates.

[0510] The decoding device 200 can use the predicted motion vector index to determine a motion vector candidate for the target block among the predicted motion vector candidates included in the predicted motion vector candidate list. The decoding device 200 can select, as the predicted motion vector of the target block, the predicted motion vector candidate indicated by the predicted motion vector index from among the predicted motion vector candidates included in the predicted motion vector candidate list.

[0511] The encoding device 100 can generate an entropy-coded predicted motion vector index by applying entropy coding to the predicted motion vector index, and can generate a bitstream including the entropy-coded predicted motion vector index. The entropy-coded predicted motion vector index can be signaled from the encoding device 100 to the decoding device 200 via the bitstream. The decoding device 200 can extract the entropy-coded predicted motion vector index from the bitstream, and can obtain the predicted motion vector index by applying entropy decoding to the entropy-coded predicted motion vector index.

[0512] The motion vector that will actually be used for inter prediction of the target block may not match the predicted motion vector. To indicate the difference between the motion vector that will actually be used for inter prediction of the target block and the predicted motion vector, an MVD can be used. The encoding device 100 can derive a predicted motion vector similar to the motion vector that will actually be used for inter prediction of the target block so as to use the smallest possible MVD.

[0513] The motion vector difference (MVD) can be the difference between the motion vector of the target block and the predicted motion vector. The encoding device 100 can calculate the MVD, and can generate an entropy-coded MVD by applying entropy coding to the MVD. The encoding device 100 can generate a bitstream including the entropy-coded MVD.

[0514] The MVD can be sent from the encoding device 100 to the decoding device 200 via the bitstream. The decoding device 200 can extract the entropy-coded MVD from the bitstream, and can obtain the MVD by applying entropy decoding to the entropy-coded MVD.

[0515] The decoding device 200 can derive the motion vector of the target block by summing the MVD and the predicted motion vector. In other words, the motion vector of the target block derived by the decoding device 200 can be the sum of the MVD and the motion vector candidate.

[0516] In addition, the encoding device 100 can generate an entropy-coded MVD resolution information by applying entropy coding to the calculated MVD resolution information, and can generate a bitstream including the entropy-coded MVD resolution information. The decoding device 200 can extract the entropy-coded MVD resolution information from the bitstream, and can obtain the MVD resolution information by applying entropy decoding to the entropy-coded MVD resolution information. The decoding device 200 can use the MVD resolution information to adjust the resolution of the MVD.

[0517] Additionally, the encoding device 100 can calculate the MVD based on an affine model. The decoding device 200 can derive the affine control motion vector of the target block by the sum of the MVD and the affine control motion vector candidate, and can use the affine control motion vector to derive the motion vector of the sub-block.

[0518] The reference direction may indicate a list of reference pictures to be used for predicting a target block. For example, the reference direction may indicate one of reference picture list L0 and reference picture list L1.

[0519] The reference direction only indicates the list of reference pictures to be used for predicting a target block, and does not necessarily mean that the direction of the reference picture is limited to the forward direction or the backward direction. In other words, each of reference picture list L0 and reference picture list L1 may include pictures in the forward direction and / or the backward direction.

[0520] That the reference direction is unidirectional may mean using a single reference picture list. That the reference direction is bidirectional may mean using two reference picture lists. In other words, the reference direction may indicate one of the following cases: the case of using only reference picture list L0, the case of using only reference picture list L1, and the case of using two reference picture lists.

[0521] The reference picture index may indicate the reference picture among the reference pictures existing in the reference picture list for predicting the target block. The encoding device 100 may generate an entropy-encoded reference picture index by applying entropy encoding to the reference picture index, and may generate a bitstream including the entropy-encoded reference picture index. The entropy-encoded reference picture index may be signaled from the encoding device 100 to the decoding device 200 through the bitstream. The decoding device 200 may extract the entropy-encoded reference picture index from the bitstream, and may obtain the reference picture index by applying entropy decoding to the entropy-encoded reference picture index.

[0522] When two reference picture lists are used for predicting a target block, a single reference picture index and a single motion vector may be used for each of the reference picture lists. In addition, when two reference picture lists are used for predicting a target block, it may be that the target block designates two prediction blocks. For example, the average value or weighted sum of two prediction blocks for the target block may be used to generate the (final) prediction block of the target block.

[0523] The motion vector of the target block may be derived by the predicted motion vector index, MVD, reference direction, and reference picture index.

[0524] The decoding device 200 may generate a prediction block for the target block based on the derived motion vector and reference picture index. For example, the prediction block may be a reference block indicated by the derived motion vector in the reference picture indicated by the reference picture index.

[0525] Since the predicted motion vector index and MVD are encoded while the motion vector of the target block itself is not encoded, the number of bits sent from the encoding device 100 to the decoding device 200 may be reduced, and the encoding efficiency may be improved.

[0526] For a target block, the motion information of reconstructed neighboring blocks can be used. In a specific inter prediction mode, the encoding device 100 may not encode the actual motion information of the target block separately. Instead of encoding the motion information of the target block, additional information may be encoded, where the additional information enables the motion information of the target block to be derived using the motion information of reconstructed neighboring blocks. Since the additional information is encoded, the number of bits sent to the decoding device 200 can be reduced, and the encoding efficiency can be improved.

[0527] For example, as inter prediction modes in which the motion information of the target block is not directly encoded, there may be a skip mode and / or a merge mode. Here, each of the encoding device 100 and the decoding device 200 may use an identifier and / or an index of a unit indicating that the motion information of the unit among the reconstructed neighboring units will be used as the motion information of the target unit.

[0528] 2) Merge mode As a scheme for deriving the motion information of the target block, there is merge. The term "merge" may mean merging the motions of multiple blocks. "Merge" may mean that the motion information of one block is also applied to other blocks. In other words, the merge mode may be a mode of deriving the motion information of the target block from the motion information of neighboring blocks.

[0529] When using the merge mode, the encoding device 100 may use the motion information of spatial candidates and / or the motion information of temporal candidates to predict the motion information of the target block. Spatial candidates may include reconstructed spatial neighboring blocks that are spatially adjacent to the target block. Spatial neighboring blocks may include a left neighboring block and an upper neighboring block. Temporal candidates may include col blocks.

[0530] The terms "spatial candidate" and "spatial merge candidate" may be used with the same meaning and may be used interchangeably with each other. The terms "temporal candidate" and "temporal merge candidate" may be used with the same meaning and may be used interchangeably with each other.

[0531] The encoding device 100 may obtain a predicted block through prediction. The encoding device 100 may encode a residual block that is the difference between the target block and the predicted block.

[0532] 2-1) Create a merge candidate list When using the merge mode, each of the encoding device 100 and the decoding device 200 may use the motion information of spatial candidates and / or the motion information of temporal candidates to create a merge candidate list. The motion information may include 1) a motion vector, 2) a reference picture index, and 3) a reference direction. The reference direction may be unidirectional or bidirectional. The reference direction may represent an inter prediction indicator.

[0533] The merge candidate list may include merge candidates. A merge candidate may be motion information. In other words, the merge candidate list may be a list storing multiple pieces of motion information.

[0534] A merge candidate may be motion information of multiple temporal candidates and / or spatial candidates. In other words, the merge candidate list may include motion information of temporal candidates and / or spatial candidates, etc.

[0535] In addition, the merge candidate list may include new merge candidates generated by combining merge candidates already existing in the merge candidate list. In other words, the merge candidate list may include new motion information generated by combining multiple pieces of motion information previously existing in the merge candidate list.

[0536] In addition, the merge candidate list may include history-based merge candidates. A history-based merge candidate may be motion information of a block that was encoded and / or decoded before the target block.

[0537] In addition, the merge candidate list may include merge candidates based on the average value of two merge candidates.

[0538] A merge candidate may be a specific mode for deriving inter-frame prediction information. A merge candidate may be information indicating a specific mode for deriving inter-frame prediction information. The inter-frame prediction information of the target block may be derived according to the specific mode indicated by the merge candidate. In addition, the specific mode may include a process for deriving a series of inter-frame prediction information. Such a specific mode may be an inter-frame prediction information derivation mode or a motion information derivation mode.

[0539] The inter-frame prediction information of the target block may be derived according to the mode indicated by the merge candidate selected from the merge candidates in the merge candidate list through a merge index.

[0540] For example, the motion information derivation mode in the merge candidate list may be at least one of the following modes: 1) a motion information derivation mode for a sub-block unit and 2) an affine motion information derivation mode.

[0541] In addition, the merge candidate list may include motion information of a zero vector. The zero vector may also be referred to as a "zero merge candidate".

[0542] In other words, multiple pieces of motion information in the merge candidate list may be at least one of the following information: 1) motion information of a spatial candidate, 2) motion information of a temporal candidate, 3) motion information generated by combining multiple pieces of motion information previously existing in the merge candidate list, and 4) a zero vector.

[0543] The motion information may include 1) a motion vector, 2) a reference picture index, and 3) a reference direction. The reference direction may also be referred to as an "inter-frame prediction indicator". The reference direction may be unidirectional or bidirectional. The unidirectional reference direction may indicate L0 prediction or L1 prediction.

[0544] A merge candidate list may be created before performing prediction in the merge mode.

[0545] The number of merge candidates in the merge candidate list may be predefined. Each of the encoding device 100 and the decoding device 200 may add merge candidates to the merge candidate list according to a predefined scheme and a predefined priority so that the merge candidate list has a predefined number of merge candidates. The merge candidate list of the encoding device 100 and the merge candidate list of the decoding device 200 may be made the same as each other using the predefined scheme and the predefined priority.

[0546] Merge may be applied based on a CU or a PU. When performing merge based on a CU or a PU, the encoding device 100 may send a bitstream including predefined information to the decoding device 200. For example, the predefined information may include 1) information indicating whether merge is performed for each block partition, and 2) information about the block to be merged among the blocks that are spatial candidates and / or temporal candidates for the target block.

[0547] 2-2) Search for the motion vector using the merge candidate list The encoding device 100 may determine a merge candidate to be used for encoding a target block. For example, the encoding device 100 may perform prediction on the target block using a merge candidate in the merge candidate list and may generate a residual block for the merge candidate. The encoding device 100 may encode the target block using the merge candidate that generates the minimum cost in the prediction and the encoding of the residual block.

[0548] In addition, the encoding device 100 may determine whether to encode the target block using the merge mode.

[0549] 2-3) Transmission of inter-frame prediction information The encoding device 100 may generate a bitstream including inter-frame prediction information required for inter-frame prediction. The encoding device 100 may generate entropy-coded inter-frame prediction information by performing entropy coding on the inter-frame prediction information and may send the bitstream including the entropy-coded inter-frame prediction information to the decoding device 200. The entropy-coded inter-frame prediction information may be signaled by the encoding device 100 to the decoding device 200 through the bitstream. The decoding device 200 may extract the entropy-coded inter-frame prediction information from the bitstream and may obtain the inter-frame prediction information by applying entropy decoding to the entropy-coded inter-frame prediction information.

[0550] The decoding device 200 may perform inter-frame prediction on the target block using the inter-frame prediction information of the bitstream.

[0551] Inter-frame prediction information may include 1) mode information indicating whether the merge mode is used, 2) a merge index, and 3) correction information.

[0552] In addition, the inter-frame prediction information may include a residual signal.

[0553] The decoding device 200 may obtain the merge index from the bitstream only when the mode information indicates that the merge mode is used.

[0554] The mode information may be a merge flag. The unit of the mode information may be a block. Information about the block may include the mode information, and the mode information may indicate whether the merge mode is applied to the block.

[0555] The merge index may indicate a merge candidate among the merge candidates included in the merge candidate list that will be used to predict the target block. Alternatively, the merge index may indicate a block among neighboring blocks that are spatially or temporally adjacent to the target block and that will be merged with the target block.

[0556] The encoding device 100 may select a merge candidate having the highest encoding performance among the merge candidates included in the merge candidate list, and may set the value of the merge index to indicate the selected merge candidate.

[0557] The correction information may be information for correcting a motion vector. The encoding device 100 may generate the correction information. The decoding device 200 may correct the motion vector of the merge candidate selected by the merge index based on the correction information.

[0558] The correction information may include at least one of information indicating whether correction will be performed, correction direction information, and correction size information. A prediction mode in which the motion vector is corrected based on the signaled correction information may be referred to as a "merge mode with motion vector difference".

[0559] 2-4) Inter-frame prediction in merge mode using inter-frame prediction information The decoding device 200 may perform prediction on the target block using the merge candidate indicated by the merge index among the merge candidates included in the merge candidate list.

[0560] The motion vector of the target block may be specified by the motion vector, reference picture index, and reference direction of the merge candidate indicated by the merge index.

[0561] 3) Skip mode The skip mode may be a mode in which the motion information of a spatial candidate or the motion information of a temporal candidate is applied to the target block without change. In addition, the skip mode may be a mode in which the residual signal is not used. In other words, when the skip mode is used, the reconstructed block may be the same as the predicted block.

[0562] The difference between the merge mode and the skip mode lies in whether to send or use the residual signal. That is, the skip mode can be similar to the merge mode except that the residual signal is not sent or used.

[0563] When using the skip mode, the encoding device 100 can send information related to the block whose motion information among the blocks as spatial candidates or temporal candidates will be used as the motion information of the target block to the decoding device 200 through the bitstream. The encoding device 100 can generate entropy-coded information by performing entropy coding on this information, and can signal the entropy-coded information to the decoding device 200 through the bitstream. The decoding device 200 can extract the entropy-coded information from the bitstream, and can obtain the information by applying entropy decoding to the entropy-coded information.

[0564] In addition, when using the skip mode, the encoding device 100 may not send other syntax information (such as MVD) to the decoding device 200. For example, when using the skip mode, the encoding device 100 may not signal the syntax elements related to at least one of MVD, coded block flag, and transform coefficient level to the decoding device 200.

[0565] 3-1) Create a merge candidate list The skip mode can also use the merge candidate list. In other words, the merge candidate list can be used in both the merge mode and the skip mode. In this regard, the merge candidate list can also be referred to as the "skip candidate list" or the "merge / skip candidate list".

[0566] Optionally, the skip mode can use an additional candidate list different from the candidate list of the merge mode. In this case, in the following description, the merge candidate list and the merge candidate can be replaced by the skip candidate list and the skip candidate, respectively.

[0567] The merge candidate list can be created before performing prediction in the skip mode.

[0568] 3-2) Search for the motion vector using the merge candidate list The encoding device 100 can determine the merge candidate to be used for encoding the target block. For example, the encoding device 100 can perform prediction on the target block using the merge candidate in the merge candidate list. The encoding device 100 can encode the target block using the merge candidate that generates the minimum cost in the prediction.

[0569] In addition, the encoding device 100 can determine whether to use the skip mode to encode the target block.

[0570] 3-3) Transmission of inter-frame prediction information The encoding device 100 may generate a bitstream including inter-prediction information required for inter-prediction. The decoding device 200 may perform inter-prediction on a target block using the inter-prediction information of the bitstream.

[0571] The inter-prediction information may include 1) mode information indicating whether the skip mode is used and 2) a skip index.

[0572] The skip index may be the same as the merge index described above.

[0573] When the skip mode is used, the target block may be encoded without using a residual signal. The inter-prediction information may not include a residual signal. Optionally, the bitstream may not include a residual signal.

[0574] The decoding device 200 may obtain the skip index from the bitstream only when the mode information indicates that the skip mode is used. As described above, the merge index and the skip index may be the same as each other. The decoding device 200 may obtain the skip index from the bitstream only when the mode information indicates that the merge mode or the skip mode is used.

[0575] The skip index may indicate a merge candidate among the merge candidates included in the merge candidate list that will be used to predict the target block.

[0576] 3-4) Inter-frame prediction in skip mode using inter-frame prediction information The decoding device 200 may perform prediction on the target block using the merge candidate indicated by the skip index among the merge candidates included in the merge candidate list.

[0577] The motion vector of the target block may be specified by the motion vector, reference picture index, and reference direction of the merge candidate indicated by the skip index.

[0578] 4) Current picture reference mode The current picture reference mode may represent a prediction mode that uses a previously reconstructed region in the target picture to which the target block belongs.

[0579] A motion vector for specifying the previously reconstructed region may be used. The reference picture index of the target block may be used to determine whether the target block has been encoded in the current picture reference mode.

[0580] A flag or index indicating whether the target block is a block encoded in the current picture reference mode may be signaled from the encoding device 100 to the decoding device 200. Optionally, it may be inferred whether the target block is a block encoded in the current picture reference mode based on the reference picture index of the target block.

[0581] When the target block is encoded in the current picture reference mode, the current picture may be present at a fixed position or an arbitrary position in the reference picture list for the target block.

[0582] For example, the fixed position may be a position where the value of the reference picture index is 0 or the last position.

[0583] When the target picture exists at any position in the reference picture list, an additional reference picture index indicating such an arbitrary position may be signaled by the encoding device 100 to the decoding device 200.

[0584] 5) Sub-block merge mode The sub-block merge mode may be a mode of deriving motion information from sub-blocks of a CU.

[0585] When the sub-block merge mode is applied, motion information of a co-located sub-block (col-sub-block) of a target sub-block in a reference picture (i.e., a time-based merge candidate based on the sub-block) and / or an affine control point motion vector merge candidate may be used to generate a sub-block merge candidate list.

[0586] 6) Triangular partitioning mode In the triangular partitioning mode, a target block may be partitioned in a diagonal direction, and sub-target blocks generated by the partitioning may be generated. For each sub-target block, motion information corresponding to the sub-target block may be derived, and the derived motion information may be used to derive predicted samples of each sub-target block. Predicted samples of the target block may be derived by a weighted sum of the predicted samples of the sub-target blocks generated via the partitioning.

[0587] 7) Combined inter-frame - intra-frame prediction mode The combined inter-intra prediction mode may be a mode of deriving predicted samples of a target block by using a weighted sum of predicted samples generated via inter prediction and predicted samples generated via intra prediction.

[0588] In the above modes, the decoding device 200 may autonomously correct the derived motion information. For example, the decoding device 200 may search for motion information having a minimum sum of absolute differences (SAD) in a specific region based on a reference block indicated by the derived motion information, and may derive the found motion information as the corrected motion information.

[0589] In the above modes, the decoding device 200 may use optical flow to compensate for predicted samples derived via inter prediction.

[0590] In the AMVP mode, merge mode, skip mode, etc. described above, index information of a list may be used to specify motion information among multiple pieces of motion information in the list that will be used to predict a target block.

[0591] To improve the coding efficiency, the coding device 100 may signal only the index of the element among the elements in the signal sending list that generates the minimum cost in the inter prediction of the target block. The coding device 100 may encode the index and signal the encoded index.

[0592] Therefore, it must be possible for the coding device 100 and the decoding device 200 to derive the above-described lists (i.e., the predicted motion vector candidate list and the merge candidate list) based on the same data using the same scheme. Here, the same data may include the reconstructed picture and the reconstructed block. In addition, in order to specify an element using an index, the order of the elements in the list must be fixed.

[0593] Figure 10 Shows spatial candidates according to an embodiment.

[0594] In Figure 10 the positions of the spatial candidates are shown.

[0595] The large block at the center of the figure may represent the target block. The five small blocks may represent spatial candidates.

[0596] The coordinates of the target block may be (xP, yP), and the size of the target block may be represented by (nPSW, nPSH).

[0597] Spatial candidate A0 may be a block adjacent to the lower left corner of the target block. A0 may be a block occupying the pixels located at coordinates (xP - 1, yP + nPSH).

[0598] Spatial candidate A1 may be a block adjacent to the left side of the target block. A1 may be the lowermost block among the blocks adjacent to the left side of the target block. Alternatively, A1 may be a block adjacent to the top of A0. A1 may be a block occupying the pixels located at coordinates (xP - 1, yP + nPSH - 1).

[0599] Spatial candidate B0 may be a block adjacent to the upper right corner of the target block. B0 may be a block occupying the pixels located at coordinates (xP + nPSW, yP - 1).

[0600] Spatial candidate B1 may be a block adjacent to the top of the target block. B1 may be the rightmost block among the blocks adjacent to the top of the target block. Alternatively, B1 may be a block adjacent to the left of B0. B1 may be a block occupying the pixels located at coordinates (xP + nPSW - 1, yP - 1).

[0601] Spatial candidate B2 may be a block adjacent to the upper left corner of the target block. B2 may be a block occupying the pixels located at coordinates (xP - 1, yP - 1).

[0602] Determination of the availability of spatial and temporal candidates In order to include the motion information of a spatial candidate or the motion information of a temporal candidate in a list, it is necessary to determine whether the motion information of the spatial candidate or the motion information of the temporal candidate is available.

[0603] Hereinafter, a candidate block may include a spatial candidate and a temporal candidate.

[0604] For example, the determination may be performed by sequentially applying the following steps 1) to 4).

[0605] Step 1) When the PU including the candidate block is outside the boundary of the picture, the availability of the candidate block may be set to "false". The expression "the availability is set to false" may have the same meaning as "set to unavailable".

[0606] Step 2) When the PU including the candidate block is outside the boundary of the slice, the availability of the candidate block may be set to "false". When the target block and the candidate block are in different slices, the availability of the candidate block may be set to "false".

[0607] Step 3) When the PU including the candidate block is outside the boundary of the parallel block, the availability of the candidate block may be set to "false". When the target block and the candidate block are in different parallel blocks, the availability of the candidate block may be set to "false".

[0608] Step 4) When the prediction mode of the PU including the candidate block is the intra prediction mode, the availability of the candidate block may be set to "false". When the PU including the candidate block does not use inter prediction, the availability of the candidate block may be set to "false".

[0609] Figure 11 Shows the order of adding the motion information of a spatial candidate to the merge list according to an embodiment.

[0610] As Figure 11 shown, when multiple motion information of spatial candidates is added to the merge list, the order of A1, B1, B0, A0, and B2 may be used. That is, multiple motion information of available spatial candidates may be added to the merge list in the order of A1, B1, B0, A0, and B2.

[0611] Method for deriving a merge list in merge mode and skip mode As described above, the maximum number of merge candidates in the merge list may be set. The set maximum number may be indicated by "N". The set number may be sent from the encoding device 100 to the decoding device 200. The slice header of the slice may include N. In other words, the maximum number of merge candidates in the merge list for the target block of the slice may be set by the slice header. For example, the value of N may be substantially 5.

[0612] Multiple pieces of motion information (i.e., merge candidates) can be added to the merge list in the order of steps 1) to 4) below.

[0613] Step 1) Among the spatial candidates, available spatial candidates can be added to the merge list. Multiple pieces of motion information of the available spatial candidates can be added to the merge list in the order shown in Figure 11 Here, when the motion information of the available spatial candidate overlaps with other motion information already existing in the merge list, the motion information of the available spatial candidate may not be added to the merge list. The operation of checking whether the corresponding motion information overlaps with other motion information existing in the list can be abbreviated as "overlap check".

[0614] The maximum number of pieces of motion information to be added can be N.

[0615] Step 2) When the number of pieces of motion information in the merge list is less than N and time candidates are available, the motion information of the time candidates can be added to the merge list. Here, when the motion information of the available time candidate overlaps with other motion information already existing in the merge list, the motion information of the available time candidate may not be added to the merge list.

[0616] Step 3) When the number of pieces of motion information in the merge list is less than N and the type of the target strip is "B", the combined motion information generated by combining bidirectional prediction (bi-prediction) can be added to the merge list.

[0617] The target strip can be a strip including the target block.

[0618] The combined motion information can be a combination of L0 motion information and L1 motion information. The L0 motion information can be motion information that only refers to the L0 reference picture list. The L1 motion information can be motion information that only refers to the L1 reference picture list.

[0619] In the merge list, there can be one or more pieces of L0 motion information. In addition, in the merge list, there can be one or more pieces of L1 motion information.

[0620] The combined motion information can include one or more pieces of combined motion information. When generating the combined motion information, the L0 motion information and the L1 motion information among the one or more pieces of L0 motion information and the one or more pieces of L1 motion information that will be used for the steps of generating the combined motion information can be predefined. One or more pieces of combined motion information can be generated in a predefined order through combined bidirectional prediction using a combination of a pair of different motion information in the merge list. One piece of motion information in the pair of different motion information can be L0 motion information, and the other piece of motion information in the pair of different motion information can be L1 motion information.

[0621] For example, the combined motion information with the highest priority added thereto may be a combination of L0 motion information with a merge index of 0 and L1 motion information with a merge index of 1. When the motion information with a merge index of 0 is not L0 motion information or when the motion information with a merge index of 1 is not L1 motion information, the combined motion information may neither be generated nor added. Next, the combined motion information with the next highest priority added thereto may be a combination of L0 motion information with a merge index of 1 and L1 motion information with a merge index of 0. Subsequent detailed combinations may conform to other combinations in the field of video encoding / decoding.

[0622] Here, when the combined motion information overlaps with other motion information already present in the merge list, the combined motion information may not be added to the merge list.

[0623] Step 4) When the number of motion information in the merge list is less than N, motion information of a zero vector may be added to the merge list.

[0624] The motion information of a zero vector may be motion information whose motion vector is a zero vector.

[0625] The number of motion information of a zero vector may be one or more. The reference picture indices of one or more pieces of motion information of a zero vector may be different from each other. For example, the value of the reference picture index of the first motion information of a zero vector may be 0. The value of the reference picture index of the second motion information of a zero vector may be 1.

[0626] The number of motion information of a zero vector may be the same as the number of reference pictures in the reference picture list.

[0627] The reference direction of the motion information of a zero vector may be bidirectional. Both motion vectors may be zero vectors. The number of motion information of a zero vector may be the smaller of the number of reference pictures in reference picture list L0 and the number of reference pictures in reference picture list L1. Optionally, when the number of reference pictures in reference picture list L0 and the number of reference pictures in reference picture list L1 are different from each other, a unidirectional reference direction may be used for the reference picture index that can be applied only to a single reference picture list.

[0628] The encoding device 100 and / or the decoding device 200 may subsequently add the motion information of a zero vector to the merge list while changing the reference picture index.

[0629] When the motion information of a zero vector overlaps with other motion information already present in the merge list, the motion information of a zero vector may not be added to the merge list.

[0630] The order of the above steps 1) to 4) is merely exemplary and can be changed. In addition, some of the above steps can be omitted according to predefined conditions.

[0631] Method for deriving a list of predicted motion vector candidates in AMVP mode The maximum number of predicted motion vector candidates in the predicted motion vector candidate list can be predefined. The predefined maximum number can be indicated by N. For example, the predefined maximum number can be 2.

[0632] Multiple pieces of motion information (i.e., predicted motion vector candidates) can be added to the predicted motion vector candidate list in the order of the following steps 1) to 3).

[0633] Step 1) Available spatial candidates among the spatial candidates can be added to the predicted motion vector candidate list. The spatial candidates can include a first spatial candidate and a second spatial candidate.

[0634] The first spatial candidate can be one of A0, A1, scaled A0, and scaled A1. The second spatial candidate can be one of B0, B1, B2, scaled B0, scaled B1, and scaled B2.

[0635] Multiple pieces of motion information of the available spatial candidates can be added to the predicted motion vector candidate list in the order of the first spatial candidate and the second spatial candidate. In this case, when the motion information of the available spatial candidate overlaps with other motion information already existing in the predicted motion vector candidate list, the motion information of the available spatial candidate may not be added to the predicted motion vector candidate list. In other words, when the value of N is 2, if the motion information of the second spatial candidate is the same as that of the first spatial candidate, the motion information of the second spatial candidate may not be added to the predicted motion vector candidate list.

[0636] The maximum number of added motion information can be N.

[0637] Step 2) When the number of pieces of motion information in the predicted motion vector candidate list is less than N and the temporal candidate is available, the motion information of the temporal candidate can be added to the predicted motion vector candidate list. In this case, when the motion information of the available temporal candidate overlaps with other motion information already existing in the predicted motion vector candidate list, the motion information of the available temporal candidate may not be added to the predicted motion vector candidate list.

[0638] Step 3) When the number of pieces of motion information in the predicted motion vector candidate list is less than N, zero vector motion information can be added to the predicted motion vector candidate list.

[0639] The zero vector motion information may include one or more pieces of zero vector motion information. The reference picture indices of the one or more pieces of zero vector motion information may be different from each other.

[0640] The encoding device 100 and / or the decoding device 200 may sequentially add multiple pieces of zero vector motion information to the prediction motion vector candidate list while changing the reference picture index.

[0641] When the zero vector motion information overlaps with other motion information already existing in the prediction motion vector candidate list, the zero vector motion information may not be added to the prediction motion vector candidate list.

[0642] The description of the zero vector motion information made above in combination with the merge list may also be applied to the zero vector motion information. The repetitive description thereof will be omitted.

[0643] The order of steps 1) to 3) described above is merely exemplary and may be changed. In addition, some of the steps may be omitted according to predefined conditions.

[0644] Figure 12 Shows the transform and quantization processing according to an example.

[0645] As Figure 12 shown, the quantization levels may be generated by performing transform and / or quantization processing on the residual signal.

[0646] The residual signal may be generated as the difference between the original block and the prediction block. Here, the prediction block may be a block generated via intra prediction or inter prediction.

[0647] The residual signal may be transformed into a signal in the frequency domain through a transform process that is part of the quantization process.

[0648] The transform kernels for transform may include various DCT kernels, such as discrete cosine transform (DCT) type 2 (DCT-II) and discrete sine transform (DST) kernels.

[0649] These transform kernels may perform separable transform or two-dimensional (2D) non-separable transform on the residual signal. The separable transform may be a transform indicating that a one-dimensional (1D) transform is performed on the residual signal in each of the horizontal and vertical directions.

[0650] The DCT types and DST types adaptively used for 1D transform may include DCT-V, DCT-VIII, DST-I, and DST-VII in addition to DCT-II, as shown in each of Table 3 and Table 4 below.

[0651] Table 3

[0652] Table 4

[0653] As shown in Table 3 and Table 4, a transform set can be used when exporting the DCT type or DST type to be used for transformation. Each transform set can include multiple transform candidates. Each transform candidate can be a DCT type or a DST type.

[0654] The following Table 5 shows an example of the transform set to be applied to the horizontal direction and the transform set to be applied to the vertical direction according to the intra prediction mode.

[0655] Table 5

[0656] In Table 5, the numbers of the vertical transform set and the horizontal transform set to be applied to the horizontal direction of the residual signal according to the intra prediction mode of the target block are shown.

[0657] As Figure 4 and Figure 5 illustrated, the transform sets to be applied to the horizontal direction and the vertical direction can be predefined according to the intra prediction mode of the target block. The encoding device 100 can use the transforms included in the transform set corresponding to the intra prediction mode of the target block to perform transformation and inverse transformation on the residual signal. In addition, the decoding device 200 can use the transforms included in the transform set corresponding to the intra prediction mode of the target block to perform inverse transformation on the residual signal.

[0658] In the transformation and inverse transformation, as illustrated in Table 3, Table 4, and Table 5, the transform set to be applied to the residual signal can be determined and may not be signaled. The transform indication information can be signaled from the encoding device 100 to the decoding device 200. The transform indication information can be information indicating which one of the multiple transform candidates included in the transform set to be applied to the residual signal is used.

[0659] For example, when the size of the target block is 64×64, transform sets each having three transforms can be configured according to the intra prediction mode. The optimal transform method can be selected from a total of nine multi-transform methods generated by the combination of three transforms in the horizontal direction and three transforms in the vertical direction. Through such an optimal transform method, the residual signal can be encoded and / or decoded, and thus the encoding efficiency can be improved.

[0660] Here, the information indicating which one of the multiple transforms belonging to each transform set has been used for at least one of the vertical transformation and the horizontal transformation can be entropy encoded and / or entropy decoded. Here, truncated unary binarization can be used to encode and / or decode such information.

[0661] As described above, various transform methods can be applied to a residual signal generated via intra prediction or inter prediction.

[0662] The transform may include at least one of a first transform and a secondary transform. The transform coefficients can be generated by performing the first transform on the residual signal, and the secondary transform coefficients can be generated by performing the secondary transform on the transform coefficients.

[0663] The first transform may be referred to as a "primary transform". In addition, the first transform may also be referred to as an "adaptive multi-transform (AMT) scheme". As described above, AMT may mean applying different transforms to each 1D direction (i.e., the vertical direction and the horizontal direction).

[0664] The secondary transform can be a transform for improving the energy concentration of the transform coefficients generated by the first transform. Similar to the first transform, the secondary transform can be a separable transform or a non-separable transform. Such a non-separable transform can be a non-separable secondary transform (NSST).

[0665] At least one of a plurality of predefined transform methods can be used to perform the first transform. For example, the plurality of predefined transform methods may include a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loeve transform (KLT), etc.

[0666] In addition, according to the kernel functions defining the discrete cosine transform (DCT) or the discrete sine transform (DST), the first transform can be a transform of various types.

[0667] For example, the transform type can be determined based on at least one of the following items: 1) the prediction mode of the target block (e.g., one of intra prediction and inter prediction), 2) the size of the target block, 3) the shape of the target block, 4) the intra prediction mode of the target block, 5) the component of the target block (e.g., one of the luminance component and the chrominance component), and 6) the partition type applied to the target block (e.g., one of a quadtree, a binary tree, and a ternary tree).

[0668] For example, according to the transform kernels presented in Table 6 below, the first transform may include transforms such as DCT-2, DCT-5, DCT-7, DST-7, DST-1, DST-8, and DCT-8. In Table 6 below, various transform types and transform kernel functions for multi-transform selection (MTS) are illustrated.

[0669] MTS may refer to the selection of a combination of one or more DCT and / or DST kernels for transforming the residual signal in the horizontal and / or vertical directions.

[0670] Table 6

[0671] In Table 6, i and j can be integer values equal to or greater than 0 and less than or equal to N - 1.

[0672] A secondary transformation can be performed on the transform coefficients generated by performing a first transformation.

[0673] As in the first transformation, a set of transformations can also be defined in the secondary transformation. The method for deriving and / or determining the above set of transformations can be applied not only to the first transformation but also to the secondary transformation.

[0674] The first transformation and the secondary transformation can be determined for a specific target.

[0675] For example, the first transformation and the secondary transformation can be applied to signal components corresponding to one or more of a luma component and a chroma component. Whether to apply the first transformation and / or the secondary transformation can be determined based on at least one of the coding parameters for a target block and / or neighboring blocks. For example, whether to apply the first transformation and / or the secondary transformation can be determined based on the size and / or shape of the target block.

[0676] In the encoding device 100 and the decoding device 200, transform information indicating the transformation method to be used for a target can be derived by using specified information.

[0677] For example, the transform information can include a transform index to be used for a primary transformation and / or a secondary transformation. Alternatively, the transform information can indicate that the primary transformation and / or the secondary transformation is not used.

[0678] For example, when the target of the primary transformation and the secondary transformation is a target block, the transformation method indicated by the transform information to be applied to the primary transformation and / or the secondary transformation can be determined based on at least one of the coding parameters for the target block and / or a block adjacent to the target block.

[0679] Alternatively, the transform information indicating the transformation method for a specific target can be signaled from the encoding device 100 to the decoding device 200.

[0680] For example, for a single CU, the decoding device 200 can derive as transform information whether to use the primary transformation, the index indicating the primary transformation, whether to use the secondary transformation, and the index indicating the secondary transformation. Alternatively, for a single CU, the transform information indicating the following items can be signaled: whether to use the primary transformation, the index indicating the primary transformation, whether to use the secondary transformation, and the index indicating the secondary transformation.

[0681] Quantized transform coefficients (i.e., quantized levels) can be generated by performing quantization on the result generated by performing the first transformation and / or the secondary transformation or on the residual signal.

[0682] Figure 13 Shows a diagonal scan according to an example.

[0683] Figure 14 Shows a horizontal scan according to an example.

[0684] Figure 15 Shows a vertical scan according to an example.

[0685] The quantized transform coefficients can be scanned via at least one of (upper right) diagonal scan, vertical scan, and horizontal scan according to at least one of the intra prediction mode, block size, and block shape. The block can be a transform unit (TU).

[0686] Each scan can be initiated at a specific start point and can be terminated at a specific end point.

[0687] For example, by using Figure 13 the diagonal scan of to scan the coefficients of a block to change the quantized transform coefficients into a 1D vector form. Optionally, the horizontal scan of Figure 14 or Figure 15 the vertical scan of can be used according to the block size and / or intra prediction mode without using the diagonal scan.

[0688] The vertical scan can be an operation of scanning 2D block type coefficients in the column direction. The horizontal scan can be an operation of scanning 2D block type coefficients in the row direction.

[0689] In other words, which one of the diagonal scan, vertical scan, and horizontal scan will be used can be determined according to the block size and / or inter prediction mode.

[0690] As Figure 13 , Figure 14 and Figure 15 shown, the quantized transform coefficients can be scanned along the diagonal direction, horizontal direction, or vertical direction.

[0691] The quantized transform coefficients can be represented by the block shape. Each block can include a plurality of sub-blocks. Each sub-block can be defined according to the minimum block size or minimum block shape.

[0692] In the scan, the scan order according to the type or direction of the scan can be first applied to the sub-blocks. In addition, the scan order according to the direction of the scan can be applied to the quantized transform coefficients in each sub-block.

[0693] For example, as Figure 13 , Figure 14 and Figure 15As shown, when the size of the target block is 8×8, the quantized transform coefficients can be generated by first transformation, secondary transformation and quantization of the residual signal of the target block. Therefore, one of three types of scanning orders can be applied to four 4×4 sub-blocks, and the quantized transform coefficients can also be scanned for each 4×4 sub-block according to the scanning order.

[0694] The encoding device 100 can generate entropy-coded quantized transform coefficients by performing entropy encoding on the scanned quantized transform coefficients, and can generate a bitstream including the entropy-coded quantized transform coefficients.

[0695] The decoding device 200 can extract the entropy-coded quantized transform coefficients from the bitstream, and can generate quantized transform coefficients by performing entropy decoding on the entropy-coded quantized transform coefficients. The quantized transform coefficients can be arranged in the form of a 2D block via inverse scanning. Here, as a method of inverse scanning, at least one of a top-right diagonal scan, a vertical scan, and a horizontal scan can be performed.

[0696] In the decoding device 200, inverse quantization can be performed on the quantized transform coefficients. Depending on whether secondary inverse transformation is performed, the result generated by performing inverse quantization can be subjected to secondary inverse transformation. In addition, depending on whether first inverse transformation will be performed, the result generated by performing secondary inverse transformation can be subjected to first inverse transformation. The reconstructed residual signal can be generated by performing first inverse transformation on the result generated by performing secondary inverse transformation.

[0697] For the luminance component reconstructed by intra prediction or inter prediction, inverse mapping with a dynamic range can be performed before loop filtering.

[0698] The dynamic range can be divided into 16 equal segments, and the mapping function of the corresponding segment can be signaled. Such a mapping function can be signaled at the slice level or the parallel block group level.

[0699] The inverse mapping function for performing inverse mapping can be derived based on the mapping function.

[0700] Loop filtering, storage of reference pictures, and motion compensation can be performed in the inverse mapping region.

[0701] The prediction block generated by inter prediction can be transformed into the mapping region by mapping using the mapping function, and the transformed prediction block can be used to generate the reconstructed block. However, since intra prediction is performed in the mapping region, the prediction block generated by intra prediction can be used to generate the reconstructed block without the need for mapping and / or inverse mapping.

[0702] For example, when the target block is the residual block of the chrominance component, the residual block can be transformed into the inverse mapping region by scaling the chrominance component of the mapping region.

[0703] Scaling availability can be signaled at the stripe level or at the parallel block group level.

[0704] For example, scaling can be applied only to cases where the mapping is available for the luminance component and the partitions of the luminance component and the partitions of the chrominance component follow the same tree structure.

[0705] Scaling can be performed based on the average value of the samples in the luminance prediction block corresponding to the chrominance prediction block. Here, when the target block uses inter prediction, the luminance prediction block can represent the mapped luminance prediction block.

[0706] The values required for scaling can be derived by referring to a lookup table using the index of the segment to which the average value of the sample values of the luminance prediction block belongs.

[0707] The residual block can be transformed into the inverse mapped region by scaling the residual block using the finally derived values. Thereafter, for the blocks of the chrominance component, reconstruction, intra prediction, inter prediction, loop filtering, and storage of reference pictures can be performed in the inverse mapped region.

[0708] For example, information indicating whether mapping and / or inverse mapping of the luminance component and the chrominance component is available can be signaled by the sequence parameter set.

[0709] A prediction block of the target block can be generated based on a block vector. The block vector can indicate the displacement between the target block and the reference block. The reference block can be a block in the target image.

[0710] In this way, the prediction mode of generating a prediction block by referring to the target image can be referred to as the "intra block copy (IBC) mode".

[0711] The IBC mode can be applied to CUs with a specific size. For example, the IBC mode can be applied to an M×N CU. Here, M and N can be less than or equal to 64.

[0712] The IBC mode can include a skip mode, a merge mode, an AMVP mode, etc. In the case of the skip mode or the merge mode, a merge candidate list can be configured, and a merge index can be signaled, and thus a single merge candidate can be specified among the merge candidates existing in the merge candidate list. The block vector of the specified merge candidate can be used as the block vector of the target block.

[0713] In the case of the AMVP mode, a differential block vector can be signaled. In addition, a prediction block vector can be derived from the left neighboring block and the upper neighboring block of the target block. In addition, an index indicating which neighboring block will be used can be signaled.

[0714] Prediction blocks in the IBC mode can be included in the target CTU or the left CTU, and can be limited to blocks within the previously reconstructed region. For example, the value of the block vector can be restricted such that the prediction block of the target block is located in a specific region. The specific region can be a region defined by three 64×64 blocks that were encoded and / or decoded before the 64×64 block including the target block. Restricting the value of the block vector in this way can thus reduce the memory consumption and device complexity caused by the implementation of the IBC mode.

[0715] Figure 16 is a configuration diagram of an encoding device according to an embodiment.

[0716] The encoding device 1600 may correspond to the encoding device 100 described above.

[0717] The encoding device 1600 may include a processing unit 1610, a memory 1630, a user interface (UI) input device 1650, a UI output device 1660, and a memory 1640 that communicate with each other via a bus 1690. The encoding device 1600 may also include a communication unit 1620 connected to a network 1699.

[0718] The processing unit 1610 may be a central processing unit (CPU) or a semiconductor device for running processing instructions stored in the memory 1630 or the memory 1640. The processing unit 1610 may be at least one hardware processor.

[0719] The processing unit 1610 may generate and process signals, data, or information that is input to, output from, or used in the encoding device 1600, and may perform checks, comparisons, determinations, etc. related to the signals, data, or information. In other words, in an embodiment, the generation and processing of data or information and the checks, comparisons, and determinations related to the data or information may be performed by the processing unit 1610.

[0720] The processing unit 1610 may include an inter-frame prediction unit 110, an intra-frame prediction unit 120, a switch 115, a subtractor 125, a transform unit 130, a quantization unit 140, an entropy encoding unit 150, an inverse quantization unit 160, an inverse transform unit 170, an adder 175, a filter unit 180, and a reference picture buffer 190.

[0721] At least some of the inter-frame prediction unit 110, intra-frame prediction unit 120, switch 115, subtractor 125, transform unit 130, quantization unit 140, entropy encoding unit 150, dequantization unit 160, inverse transform unit 170, adder 175, filter unit 180, and reference picture buffer 190 may be program modules and may communicate with an external device or system. The program modules may be included in the encoding device 1600 in the form of an operating system, application program module, or other program module.

[0722] The program modules may be physically stored in various types of well-known storage devices. In addition, at least some of the program modules may also be stored in a remote storage device capable of communicating with the encoding device 1600.

[0723] The program modules may include, but are not limited to, routines, subroutines, programs, objects, components, and data structures for performing functions or operations according to an embodiment or for implementing abstract data types according to an embodiment.

[0724] The program modules may be implemented using instructions or code run by at least one processor of the encoding device 1600.

[0725] The processing unit 1610 may run instructions or code in the inter-frame prediction unit 110, intra-frame prediction unit 120, switch 115, subtractor 125, transform unit 130, quantization unit 140, entropy encoding unit 150, dequantization unit 160, inverse transform unit 170, adder 175, filter unit 180, and reference picture buffer 190.

[0726] The storage unit may represent the memory 1630 and / or the storage 1640. Each of the memory 1630 and the storage 1640 may be any of various types of volatile or non-volatile storage media. For example, the memory 1630 may include at least one of a read-only memory (ROM) 1631 and a random access memory (RAM) 1632.

[0727] The storage unit may store data or information for the operation of the encoding device 1600. In an embodiment, the data or information of the encoding device 1600 may be stored in the storage unit.

[0728] For example, the storage unit may store pictures, blocks, lists, motion information, inter-frame prediction information, bitstreams, etc.

[0729] The encoding device 1600 may be implemented in a computer system including a computer-readable storage medium.

[0730] The storage medium may store at least one module required for the operation of the encoding device 1600. The memory 1630 may store at least one module and may be configured such that the at least one module is run by the processing unit 1610.

[0731] Functions related to the communication of data or information with the encoding device 1600 may be performed by the communication unit 1620.

[0732] For example, the communication unit 1620 may send a bitstream to the decoding device 1700 which will be described later.

[0733] Figure 17 is a configuration diagram of a decoding device according to an embodiment.

[0734] The decoding device 1700 may correspond to the decoding device 200 described above.

[0735] The decoding device 1700 may include a processing unit 1710, a memory 1730, a user interface (UI) input device 1750, a UI output device 1760, and a memory 1740 that communicate with each other via a bus 1790. The decoding device 1700 may further include a communication unit 1720 connected to a network 1799.

[0736] The processing unit 1710 may be a central processing unit (CPU) or a semiconductor device for running processing instructions stored in the memory 1730 or the memory 1740. The processing unit 1710 may be at least one hardware processor.

[0737] The processing unit 1710 may generate and process signals, data, or information that is input to, output from, or used in the decoding device 1700, and may perform checks, comparisons, determinations, etc. related to the signals, data, or information. In other words, in an embodiment, the generation and processing of data or information and the checks, comparisons, and determinations related to the data or information may be performed by the processing unit 1710.

[0738] The processing unit 1710 may include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an intra prediction unit 240, an inter prediction unit 250, a switcher 245, an adder 255, a filter unit 260, and a reference picture buffer 270.

[0739] At least some of the entropy decoding unit 210, dequantization unit 220, inverse transformation unit 230, intra prediction unit 240, inter prediction unit 250, adder 255, switch 245, filter unit 260, and reference picture buffer 270 of the decoding device 200 may be program modules and may communicate with an external device or system. The program modules may be included in the decoding device 1700 in the form of an operating system, application program module, or other program module.

[0740] The program modules may be physically stored in various types of well-known storage devices. In addition, at least some of the program modules may also be stored in a remote storage device capable of communicating with the decoding device 1700.

[0741] The program modules may include, but are not limited to, routines, subroutines, programs, objects, components, and data structures for performing functions or operations according to an embodiment or for implementing abstract data types according to an embodiment.

[0742] The program modules may be implemented using instructions or code run by at least one processor of the decoding device 1700.

[0743] The processing unit 1710 may run instructions or code in the entropy decoding unit 210, dequantization unit 220, inverse transformation unit 230, intra prediction unit 240, inter prediction unit 250, switch 245, adder 255, filter unit 260, and reference picture buffer 270.

[0744] The storage unit may represent the memory 1730 and / or the storage 1740. Each of the memory 1730 and the storage 1740 may be any of various types of volatile or non-volatile storage media. For example, the memory 1730 may include at least one of a ROM 1731 and a RAM 1732.

[0745] The storage unit may store data or information for the operation of the decoding device 1700. In an embodiment, the data or information of the decoding device 1700 may be stored in the storage unit.

[0746] For example, the storage unit may store pictures, blocks, lists, motion information, inter prediction information, bitstreams, etc.

[0747] The decoding device 1700 may be implemented in a computer system including a computer-readable storage medium.

[0748] The storage medium may store at least one module required for the operation of the decoding device 1700. The memory 1730 may store at least one module and may be configured such that the at least one module is run by the processing unit 1710.

[0749] Functions related to communication of data or information with the decoding device 1700 may be performed by the communication unit 1720.

[0750] For example, the communication unit 1720 may receive a bitstream from the encoding device 1700.

[0751] Hereinafter, the processing unit may represent the processing unit 1610 of the encoding device 1600 and / or the processing unit 1710 of the decoding device 1700. For example, regarding functions related to prediction, the processing unit may represent the switch 115 and / or the switch 245. Regarding functions related to inter-frame prediction, the processing unit may represent the inter-frame prediction unit 110, the subtractor 125, and the adder 175, and may represent the inter-frame prediction unit 250 and the adder 255. Regarding functions related to intra-frame prediction, the processing unit may represent the intra-frame prediction unit 120, the subtractor 125, and the adder 175, and may represent the intra-frame prediction unit 240 and the adder 255. Regarding functions related to transformation, the processing unit may represent the transformation unit 130 and the inverse transformation unit 170, and may represent the inverse transformation unit 230. Regarding functions related to quantization, the processing unit may represent the quantization unit 140 and the inverse quantization unit 160, and may indicate the inverse quantization unit 220. Regarding functions related to entropy encoding and / or entropy decoding, the processing unit may represent the entropy encoding unit 150 and / or the entropy decoding unit 210. Regarding functions related to filtering, the processing unit may represent the filter unit 180 and / or the filter unit 260. Regarding functions related to reference pictures, the processing unit may indicate the reference picture buffer 190 and / or the reference picture buffer 270.

[0752] In a typical image encoding / decoding method, the decoder-side motion information derivation method may be used limitedly. Therefore, the improvement in encoding efficiency attributable to the decoder-side motion information derivation method may also be limited.

[0753] In an embodiment, in order to improve the encoding efficiency in inter-frame prediction, an encoding / decoding method, device, and storage medium using a motion information search method may be provided.

[0754] The processing unit may perform inter-frame prediction on a target block. Here, when performing prediction on the target block by the motion information search method, the processing unit may derive the motion information of the target block from available information.

[0755] By means of the derivation of motion information, the encoding efficiency may be improved by minimizing the number of bits of the encoding information used to signal the information required for decoding the target block.

[0756] In an embodiment, the encoding information may include the information required for decoding the target block (sent from the encoding device 1600 to the decoding device 1700) described in the embodiment. For example, the encoding information may include encoding parameters.

[0757] In an embodiment, information described as being signaled via a bitstream may be included in the coded information. Further, the coded information may be signaled via the bitstream.

[0758] In an embodiment, the available information may refer to information that was previously decoded (or reconstructed) before performing prediction on a target block.

[0759] In an embodiment, the available information may include at least one of coded parameters, motion information, predicted sample information, reconstructed sample information, and loop filter decoded sample information at a specific sample position in a target picture.

[0760] In an embodiment, the available information may include at least one of coded parameters, motion information, predicted sample information, reconstructed sample information, and loop filter decoded sample information for a specific block in a target picture that includes a specific sample position.

[0761] In an embodiment, the specific sample position may refer to the position of neighboring samples adjacent to the left, above, and / or left - upper of the target block.

[0762] In an embodiment, the available information may include at least one of coded parameters, motion information, predicted sample information, reconstructed sample information, and loop filter reconstructed sample information at a specific sample position in a reference picture for the target block.

[0763] In an embodiment, the available information may include at least one of the following: coded parameters, motion information, predicted sample information, reconstructed sample information, and loop filter reconstructed sample information for a block that includes a specific sample position in a reference picture for the target block.

[0764] For example, the specific sample position may be indicated by the motion information of the target block.

[0765] In an embodiment, the available information may include at least one of coded parameters, motion information, predicted sample information, reconstructed sample information, and loop filter reconstructed sample information at a specific sample position in a col picture for the target block.

[0766] In an embodiment, the available information may include at least one of the following: coded parameters, motion information, predicted sample information, reconstructed sample information, and loop filter reconstructed sample information for a block that includes a specific sample position in a col picture for the target block.

[0767] For example, the specific sample position may be indicated by the motion information of the target block.

[0768] For example, the specific sample position may refer to the position of neighboring samples adjacent to the left, above, and / or left - upper of the co1 block.

[0769] In the following, embodiments in which a motion information search method is implemented will be described.

[0770] Figure 18 is a flowchart showing a target block prediction method and a bitstream generation method according to an embodiment.

[0771] The target block prediction method and the bitstream generation method according to the embodiment may be executed by the encoding device 1600. This embodiment may be part of a target block encoding method or a video encoding method.

[0772] The prediction may be one of the prediction methods described above in the embodiments. For example, the prediction may be inter prediction or intra prediction.

[0773] In step 1810, the processing unit 1610 may determine prediction information to be used for encoding the target block.

[0774] The prediction information may include the information for prediction described in the embodiments.

[0775] For example, the prediction information may include inter prediction information. For example, the prediction information may include intra prediction information.

[0776] For example, the prediction information may include a motion information candidate list and final motion information to be described later. The determination of the prediction information may include the configuration of the motion information candidate list and the determination of the final motion information.

[0777] In step 1820, the encoded encoding information may be generated by performing encoding on the encoding information.

[0778] The encoding information may refer to the information signaled / encoded / decoded in the embodiments. In other words, the encoding information may be information for allowing the decoding device 1700 to perform a prediction corresponding to the prediction performed by the encoding device 1600.

[0779] In step 1830, the processing unit 1610 may generate a bitstream.

[0780] The bitstream may include information about the target block. In addition, the bitstream may include the information described above in the embodiments.

[0781] For example, the bitstream may include the encoded encoding information or the encoding information.

[0782] For example, the bitstream may include encoding parameters related to the target block and / or the attributes of the target block.

[0783] The information included in the bitstream may be generated in step 1820, or may be generated at least in part in steps 1810 and 1820.

[0784] The processing unit 1610 may store the generated bitstream in the memory 1640. Optionally, the communication unit 1620 may send the bitstream to the decoding device 1700.

[0785] The bitstream may include coding information about the target block. The processing unit 1610 may generate the coding information about the target block by performing entropy coding on the information about the target block.

[0786] In step 1840, the processing unit 1610 may perform prediction on the target block using the information about the target block and the prediction information.

[0787] The processing unit 1610 may use the coding information when performing prediction on the target block. Optionally, the coding information may be generated to correspond to the information used for the prediction of the target block.

[0788] The prediction block may be generated by predicting the target block. The residual block may be generated as the difference between the target block and the prediction block. The information about the target block may be generated by applying transform and quantization to the residual block.

[0789] The information about the target block may include the transform and quantization coefficients for the target block. The reconstructed residual block may be generated by applying inverse quantization and inverse transform to the transform and quantization coefficients of the target block. The reconstructed block may be generated as the sum of the prediction block and the reconstructed residual block.

[0790] Figure 19 is a flowchart showing a method for predicting a target block using a bitstream according to an embodiment.

[0791] The method for predicting a target block using a bitstream according to an embodiment may be executed by the decoding device 1700. This embodiment may be part of a method for decoding a target block or a video decoding method.

[0792] The prediction may be one of the prediction methods described above in the embodiments. For example, the second prediction may be inter prediction or intra prediction.

[0793] In step 1910, the communication unit 1720 may acquire the bitstream. The communication unit 1720 may receive the bitstream from the encoding device 1600. The processing unit 1710 may store the acquired bitstream in the memory 1740.

[0794] The processing unit 1710 may read the bitstream from the memory 1740.

[0795] The bitstream may include information about the target block.

[0796] The information about the target block may include the transform and quantization coefficients for the target block.

[0797] In addition, the bitstream may include the information described above in the embodiments.

[0798] For example, the bitstream may include encoded coded information or coded information.

[0799] For example, the bitstream may include coding parameters related to the target block and / or attributes of the target block.

[0800] The computer-readable storage medium may include the bitstream, and may use the information about the target block included in the bitstream to perform prediction and decoding of the target block.

[0801] The computer-readable storage medium may be a non-transitory computer-readable storage medium.

[0802] The bitstream may include coding information about the target block. The processing unit 1710 may generate information about the target block by performing entropy decoding on the coding information about the target block.

[0803] In step 1920, the processing unit 1710 may obtain the coding information from the bitstream.

[0804] The processing unit 1710 may generate the coding information by performing decoding on the encoded coding information of the bitstream.

[0805] The coding information may refer to the information signaled / encoded / decoded as described in the embodiments. In other words, the coding information may be information for allowing the decoding device 1700 to perform a prediction corresponding to the prediction performed by the encoding device 1600.

[0806] In step 1930, the processing unit 1710 may determine the prediction information to be used for decoding the target block.

[0807] The prediction information may include the information for prediction as described in the embodiments.

[0808] The prediction information in the decoding device 1700 may be the same as the prediction information in the encoding device 1600. In other words, the processing unit 1710 may generate the same prediction information as the prediction information used in step 1840 in order to perform the same prediction as the prediction performed in step 1840.

[0809] For example, the prediction information may include inter-frame prediction information. For example, the prediction information may include intra-frame prediction information.

[0810] For example, the prediction information may include a motion information candidate list and final motion information to be described later. The determination of the prediction information may include the configuration of the motion information candidate list and the determination of the final motion information.

[0811] The processing unit 1710 may use the method used in the embodiments to determine the prediction information.

[0812] The processing unit 1710 may determine prediction information for a target block based on prediction method-related information obtained from a bitstream.

[0813] The prediction information may include inter prediction information. The prediction information may include intra prediction information.

[0814] In step 1940, the processing unit 1710 may perform prediction on the target block using information about the target block and the prediction information.

[0815] The processing unit 1710 may use coding information when performing prediction on the target block. A prediction block may be generated through the prediction of the target block.

[0816] The information about the target block may include transform and quantization coefficients for the target block. A reconstructed residual block may be generated by applying inverse quantization and inverse transform to the transform and quantization coefficients for the target block. A reconstructed block may be generated as the sum of the prediction block and the reconstructed residual block.

[0817] Intra-block copy (IBC) The intra block copy (IBC) mode may refer to a mode in which a region indicated by a block vector of a target block is used as a prediction block of the target block.

[0818] The target block may be encoded / decoded in one of an intra prediction mode, an inter prediction mode, and an intra block copy mode.

[0819] 1) When the luminance component and the chrominance component have independent block partition structures (i.e., when using a dual-tree structure), and 2) when the luminance component and the chrominance component have the same block partition structure (i.e., when using a single-tree structure), an encoding / decoding method based on prediction using intra block copy may be used.

[0820] The intra block copy mode may be a method for deriving a block (e.g., a reference block or a prediction block) from a previously encoded / decoded region in a target image using an exported block vector (BV).

[0821] The target image may be an image including the target block. Here, since the block is derived from the image including the target block, intra block copy may correspond to intra prediction.

[0822] The block vector may refer to an intra block vector.

[0823] The previously encoded / decoded region may be a region in a reconstructed image or a decoded image for a target picture (target image). Here, the region in the reconstructed image may refer to a reconstructed region. The region in the decoded image may refer to a decoded region.

[0824] The previously encoded / decoded region in the target image may be a reconstructed region to which at least one loop filter scheme has not been applied.

[0825] In an embodiment, loop filtering may include 1) chroma scaling and luminance mapping, and 2) deblocking filtering, adaptive sample offset (ASO), and adaptive loop filtering.

[0826] In an embodiment, a previously encoded / decoded region in a target image may be a reconstructed / decoded region on which at least one loop filtering scheme is performed.

[0827] Adaptive motion vector resolution (AMVR) In an embodiment, the term "resolution" may refer to the term "motion vector resolution".

[0828] In adaptive motion vector resolution, the resolution of a motion vector difference may be adjusted on a block-by-block basis.

[0829] Adaptive motion vector resolution information may indicate the resolution of a motion vector difference. The resolution of a motion vector difference of a target block may be determined by signaling / encoding / decoding the adaptive motion vector resolution information.

[0830] The motion vector resolutions applicable to blocks may be the same as or different from each other.

[0831] For example, the resolution of a motion vector applicable to a target block may be determined based on at least one of the encoding parameters, motion information, and mode information of the target block.

[0832] Adaptive motion vector resolution may improve the encoding efficiency by adjusting the resolution of a motion vector difference.

[0833] For example, the adjusted resolution may be one of 16 pixels (px), 8 pixels, 4 pixels, full pixel, half pixel, and quarter pixel, and is not limited to the pixel values listed above.

[0834] When the value of a component of a motion vector difference is changed by 1 in the case where the adjusted resolution is n pixels, the position indicated by the motion vector difference may be changed by n pixels. In other words, when the adjusted resolution of a target block is n pixels, each component of the motion vector difference may indicate a reference block in units of n pixels.

[0835] For example, when the motion vector difference actually applied to a target block is (a, b) and the adjusted resolution is p pixels, (a / p, b / p) other than (a, b) may be encoded. That is, the motion vector difference signaled / encoded by the encoding device 1600 may be (a / p, b / p). The decoding device 1700 may derive the original motion vector difference (a, b) by multiplying the signaled motion vector difference (a / p, b / p) by p.

[0836] Decoder-side motion vector derivation In an embodiment, the decoder-side motion information derivation method may derive the motion information of a target block by 1) deriving the initial motion information of the target block and 2) performing refinement on the initial motion information using a predefined operation.

[0837] In an embodiment, the refinement of specific information may involve modifying, correcting, or updating the specific information. In an embodiment, the terms "refinement", "correction", and "revision" may be used interchangeably with each other. Refined information may be generated by performing refinement on specific information.

[0838] For example, the decoder-side motion information derivation method 1) may derive the motion information offset and / or the motion vector difference in a second direction by applying a predefined operation to the motion information offset and / or the motion vector difference in a first direction, and 2) may refine the motion information by adding the derived motion information offset and / or the derived motion vector difference to the motion information in the second direction.

[0839] For example, the predefined operation may include at least one of mirroring, scaling, and copying, and is not limited to the operations listed above.

[0840] For example, when mirroring is applied to a specific motion vector MV, the result of mirroring may be -MV.

[0841] For example, when scaling is applied to a specific motion vector MV, the scaled motion vector may be derived by changing the size of the specific motion vector MV based on 1) the picture order count (POC) interval between the image including the specific motion vector MV and the reference image indicated by MV and 2) the POC interval between the image including the target block and the reference image for the target block. The direction of the specific motion vector MV and the direction of the scaled motion vector generated by applying scaling to MV may be the same as each other.

[0842] For example, when copying is applied to a specific motion vector MV, the result of copying may be MV.

[0843] In an embodiment, the decoder-side motion information derivation method may refine the motion information by searching the position indicated by the initial motion information.

[0844] In an embodiment, the decoder-side motion information derivation method may refine the motion information candidate list by searching an initial motion information candidate list. The initial motion information candidate list may include multiple (initial) motion information. Through refinement, at least some of the multiple motion information in the motion information candidate list may be refined.

[0845] For example, each piece of motion information in the refined motion information candidate list can be one of 1) the initial motion information or 2) the refined motion information generated by performing refinement on the initial motion information using the decoder-side motion information derivation method. In addition, each piece of motion information is not limited to 1) the initial motion information and 2) the refined motion information. The initial motion information can refer to one of multiple pieces of motion information included in the initial motion information candidate list.

[0846] In an embodiment, the decoder-side motion information derivation method 1) can compare the matching costs of multiple candidates in the (initial) motion information candidate list, and 2) can adjust the order of multiple candidates in the motion information candidate list depending on the matching costs of the multiple candidates. In other words, reordering of multiple candidates can be performed based on the matching costs of multiple candidates in the motion information candidate list. The multiple candidates can be multiple pieces of motion information.

[0847] For example, multiple candidates in the (initial) motion information candidate list can be sorted in ascending order of the matching costs of the multiple candidates.

[0848] In an embodiment, the initial motion information can be 1) the motion information to which the decoder-side motion information derivation method is applied and / or 2) the motion information of each search step to which the decoder-side motion information derivation method is applied.

[0849] For example, the initial motion information can be at least one of a motion vector predictor (MVP), a motion information candidate, and the motion information of a neighboring block.

[0850] At least one of the following processes can be applied to the initial motion information. The motion information to which the following processes are applied can be used as the initial motion information of the decoder-side motion information derivation method.

[0851] [Process 1]: The decoder-side motion information derivation method can be used to refine the initial motion information.

[0852] [Process 2]: Motion information offset can be used to refine the initial motion information.

[0853] For example, a motion vector difference can be added to the initial motion information.

[0854] For example, a motion information offset can be added to the initial motion information.

[0855] For example, the initial motion information or a part of the initial motion information can be changed to the same value as the motion information offset.

[0856] For example, when the advanced motion vector prediction (AMVP) mode is used for a target block, the initial motion information can be a motion vector predictor.

[0857] For example, when the AMVP mode is used for a target block, the initial motion information may be the sum of a motion vector predictor (MVP) and a motion vector difference.

[0858] For example, when performing a decoder-side motion information derivation method, at least one motion information offset may be added to at least one of a plurality of refined motion information generated at a search step.

[0859] For example, when the decoder-side motion information derivation method consists of N search steps, first refined motion information may be generated at a first search step. A first motion information offset may be added to the first refined motion information. A second search step may be performed on the first refined motion information. Thereafter, a second motion information offset may be added to the (N-1)th refined motion information generated at the (N-1)th search step. An Nth search step may be performed on the (N-1)th refined motion information.

[0860] For example, information regarding a motion information offset may be signaled / encoded / decoded.

[0861] The motion information offset may be determined by a rate-distortion optimization process. In this case, the coding efficiency of the target block may be improved.

[0862] For example, the motion information offset may be a predefined value. The predefined value may be 0. The predefined value may be (0, 0).

[0863] For example, the motion information offset may refer to a motion vector difference.

[0864] For example, when performing bi-directional inter prediction on a target block, at least one of the motion information offsets may be added to each of the motion information in the L0 direction and the motion information in the L1 direction.

[0865] For example, at least one of the motion information offsets may be added only to the motion information in the LX direction.

[0866] For example, when performing bi-directional inter prediction for a target block, only the corresponding motion information offset may be added to the motion information in the LX direction.

[0867] X may be 0, 1, or a positive integer.

[0868] X may be a predefined value.

[0869] For example, the predefined value may be 0.

[0870] For example, the predefined value may be a value indicating a direction having a lower matching cost between the L0 direction and the L1 direction. Here, the matching cost may be the matching cost of the motion information.

[0871] For example, the predefined value may be a value indicating the direction with a higher matching cost among the L0 direction and the L1 direction. Here, the matching cost may be the matching cost of motion information.

[0872] When using the predefined value X, the decoding device 1700 can determine the motion information or direction (to which the motion information offset will be added) without signaling / encoding / decoding the predefined value. By this determination, the number of bits required for signaling can be reduced.

[0873] The information indicating the predefined value X can be signaled / encoded / decoded. When signaling / encoding / decoding the information indicating the predefined value X, the predefined value X used in each block can be determined by rate-distortion optimization. By this determination, the prediction performance of the target block can be improved.

[0874] For example, the motion information offset may be a motion vector.

[0875] For example, a combination of an angle in an angle list and a distance offset in a distance offset list can be used to determine the motion information offset. The angle list may include multiple angles. The distance offset list may include multiple distance offsets.

[0876] In an embodiment, the angle list may be a list of angles with respect to the X-axis / Y-axis.

[0877] In an embodiment, "width axis / height axis" or "horizontal axis / vertical axis" may be used instead of "X-axis / Y-axis".

[0878] The angle list may be configured to include angles having a value of "i_ANGLE×π / NUM_ANGLE". i_ANGLE may be an integer from 0, 1, ..., or (2×NUM_ANGLE - 1).

[0879] The number of angles in the angle list may be 2×NUM_ANGLE.

[0880] NUM_ANGLE may be a predefined value. NUM_ANGLE may be 4, 8, or 16. NUM_ANGLE may be a positive integer.

[0881] NUM_ANGLE and / or the angles constituting the angle list may be determined based on the encoding parameters of the target block. For example, the encoding parameters may include motion information.

[0882] For example, when the affine mode is used for the target block, the value of NUM_ANGLE may be 8 (or a first value). When the affine mode is not used for the target block, the value of NUM_ANGLE may be 16 (or a second value).

[0883] For example, when the affine mode is used for the current prediction block, the value of NUM_ANGLE can be 8 (or the third value). When the affine mode is not used for the current prediction block, the value of NUM_ANGLE can be 4 (or the fourth value).

[0884] The values of NUM_ANGLE mentioned above can be merely examples. Each of the first value, the second value, the third value, and the fourth value can be a specific integer of 1 or greater.

[0885] For example, the distance offset list can be a list indicating the distance from the position indicated by the current motion vector. Alternatively, the distance offset list can be a list indicating the position relative to the position indicated by the current motion vector.

[0886] For example, the current motion vector can refer to the motion vector of the target block.

[0887] For example, when the motion vector (of the target block) indicates a first position, the second position can be specified by the angles in the angle list and the distance offsets in the distance offset list. Here, the angles in the angle list can indicate the angle between a first line and a second line. The first line can be the x-axis or the y-axis. The second line can be the straight line passing through the first position and the second position. The distance offsets in the distance offset list can refer to the distance between the first position and the second position.

[0888] For example, the distance offset list can be configured to include at least one of 0, 1, 2, 4, and 8. In other words, the distance offsets in the distance offset list can be some or all of 0, 1, 2, 4, and 8.

[0889] For example, the distance offset list can include values corresponding to multi offset -pixels as the distance offsets. Each distance offset in the distance offset list can represent a value of multi offset -pixels.

[0890] For example, multi offset can be 1 / 8, 1 / 4, 1 / 2, 1, 2, 4, 8, 16, 32, or a positive integer.

[0891] For example, the distance offsets constituting the distance offset list and the size of the distance offset list can be determined based on the coding parameters of the target block or the motion information of the target block.

[0892] In an embodiment, the size of the list can represent the number of elements (or candidates) in the corresponding list.

[0893] For example, when the affine mode is not used for the target block, the distance offset list can be {4, 8, 16, 32, 64, 128}. When the affine mode is used for the target block, the distance offset list can be {1, 2, 4, 8, 16}.

[0894] The values of the distance offsets in the foregoing distance offset list can be merely examples. Each value of the distance offset can be a specific integer that is 1 or greater.

[0895] For example, a position determined using a combination of a specific angle θ and a specific distance offset ρ can indicate a position at a distance ρ in a direction that forms an angle θ with the X-axis / Y-axis.

[0896] For example, a motion vector determined using a combination of a specific angle θ and a specific distance offset ρ can refer to a motion vector that indicates a position at a distance ρ in a direction that forms an angle θ with the X-axis / Y-axis.

[0897] For example, a motion information offset can be added to a reference picture index.

[0898] For example, when the value of the reference picture index of a target block is a first value and the value of the motion information offset added to the reference picture index is a second value, the value of the reference picture index of the target block can be changed to the first value + the second value. The first value can be 0, 1, or a positive integer. The second value can be -2, -1, 0, 1, 2, or an integer.

[0899] For example, when the value of the reference picture index of a target block is a first value and the value of the motion information offset added to the reference picture index is a second value, the value of the reference picture index of the target block can be changed to the second value. The first value can be 0, 1, or a positive integer. The second value can be 0, 1, or a positive integer.

[0900] For example, when performing a decoder-side motion information derivation method, at least one of the multiple pieces of refined motion information generated in the corresponding search step can be changed to have a value equal to the value of at least one motion information offset.

[0901] For example, when the decoder-side motion information derivation method consists of N search steps, the motion vector of the refined motion information generated in the first search step can be changed to a first motion information offset. A second search step can be performed on the motion information with the changed motion vector. Thereafter, the reference picture index of the refined motion information generated in the N-1th search step can be changed to a second motion information offset. An Nth search step can be performed on the motion information with the changed reference picture index.

[0902] For example, the motion information offset can refer to an encoding parameter. Optionally, the motion information offset can refer to motion information.

[0903] For example, motion information offset may refer to a motion vector or a reference picture index.

[0904] For example, when bidirectional inter prediction is used for a target block, the motion information in the L0 direction and the motion information in the L1 direction may be changed to values equal to the motion information offsets in the respective directions.

[0905] For example, only the motion information in the LX direction among multiple pieces of motion information in multiple directions may be changed to a value equal to the motion information offset.

[0906] For example, when bidirectional inter prediction is used for a target block, only the motion information in the LX direction among multiple pieces of motion information in multiple directions may be changed to a value equal to the motion information offset.

[0907] X may be 0, 1, or a positive integer.

[0908] X may be a predefined value.

[0909] For example, the predefined value may be 0.

[0910] For example, the predefined value may be a value corresponding to the direction with a lower matching cost of motion information among the L0 direction and the L1 direction.

[0911] For example, the predefined value may be a value corresponding to the direction with a higher matching cost of motion information among the L0 direction and the L1 direction.

[0912] When using the predefined value X, the decoder side may determine X without signaling / encoding / decoding information indicating the predefined value. Through this determination, the number of bits required for signaling can be reduced.

[0913] Information indicating X may be signaled / encoded / decoded. When information indicating X is signaled / encoded / decoded, X used in each block may be determined through rate-distortion optimization. Through this determination, the prediction performance of the target block can be improved.

[0914] For example, the search in the embodiments may be performed using the calculation of a cost function for determining the similarity between NUM_TEMPLATE_COMPARE templates.

[0915] The search in the embodiments may include a process of determining motion information that satisfies a specific condition within a predefined search range. Based on at least one piece of motion information determined through the search, the motion information of the target block may be determined and / or changed.

[0916] For example, the motion information that satisfies a specific condition may refer to the motion information with the lowest matching cost among multiple pieces of motion information within the search range.

[0917] In an embodiment, the search may include a process of determining at least one block that satisfies a specific condition within a predefined search range. The motion information indicating the block determined by the search may be used as the motion information of the target block.

[0918] For example, the block that satisfies the specific condition may be one of the reference blocks within the search range.

[0919] In an embodiment, NUM_TEMPLATE_COMPARE may be 0, 1, 2, or a positive integer.

[0920] Refined motion information may refer to motion information that satisfies a specific condition.

[0921] Refined motion information may refer to at least one of the following: 1) motion information determined by a decoder-side motion information derivation method, and 2) motion information determined at each search step of the decoder-side motion information derivation method.

[0922] In an embodiment, the search range may be a specific range centered at the position indicated by the initial motion information. In other words, the center of the search range may be the position indicated by the initial motion information. The specific range may be a range with a predefined area.

[0923] Optionally, the search range may be a specific range where the position indicated by the initial motion information is the upper left position. In other words, the upper left position of the search range may be the position indicated by the initial motion information.

[0924] Optionally, the search range may be a specific range around the position indicated by the initial motion information. In other words, the center position of the search range may be the position indicated by the initial motion information. Optionally, the search range may include at least one of the samples in the lower left area, left area, upper left area, upper area, and upper right area located around the target block.

[0925] Optionally, the search range may include at least one of the positions of the samples in the lower left area, left area, upper left area, upper area, and upper right area located around the target block.

[0926] At least one of the size and shape of the search range for the target block may be predefined by the encoding device 1600 and the decoding device 1700.

[0927] Optionally, at least one of the size and shape of the search range for the target block may be determined based on at least one of the size of the target block, the encoding parameters of the target block, the motion information of the target block, and the prediction mode of the target block.

[0928] Optionally, information indicating one of the size and shape of the search range for the target block may be encoded / decoded / signaled.

[0929] The search range may have a rectangular shape with a horizontal length of SR_X and a vertical length of SR_Y. Optionally, the search range may have a diamond shape with a horizontal length of SR_X and a vertical length of SR_Y. However, the shape and size of the search range are not limited to the above forms.

[0930] Each of SR_X and SR_Y can be a positive integer. Each of SR_X and SR_Y can be a predefined value or a value determined based on the information sent / coded / decoded by the signal.

[0931] The initial motion information can be determined based on at least one of the motion information of the target block, the coding parameters of the target block, the motion vector of the target block, the reference image of the target block, the block vector of the target block, the motion vector predictor of the target block, the block vector predictor of the target block, the motion information of at least one neighboring block of the target block, the merge candidate of the target block, the motion vector difference of the target block, and the block vector difference of the target block.

[0932] Each template can be a subset of the pixels included in a specific area.

[0933] For example, the template of the target block may include one or more of the following: 1) a subset of the pixels included in the TMSIZE_LEFT line adjacent to the left side of the target block and 2) a subset of the pixels included in the TMSIZE_ABOVE line adjacent to the upper (top) side of the target block. However, the relationship between the position of the template and the position of the target block or the template configuration method is not limited to the above embodiments.

[0934] For example, the template of the reference block may include one or more of the following: 1) a subset of the pixels included in the TMSIZE_LEFT line adjacent to the left side of the reference block, and 2) a subset of the pixels included in the TMSIZE_ABOVE line adjacent to the top of the reference block. However, the relationship between the position of the template and the position of the reference block or the template configuration method is not limited to the above embodiments.

[0935] Each of TMSIZE_LEFT and TMSIZE_ABOVE can be a predefined integer of 0 or greater. TMSIZE_LEFT and TMSIZE_ABOVE can be the same as or different from each other.

[0936] For example, the cost function of the template of the target block and the cost function of the template of the reference block in the L0 direction and / or the L1 direction can be calculated.

[0937] For example, the cost function of the template of the reference block in the L0 direction and the cost function of the template of the reference block in the L1 direction can be calculated.

[0938] In an embodiment, the cost function may be one or more of the sum of absolute differences (SAD), the sum of absolute transform differences (SATD), the sum of mean removed absolute differences (MR - SAD), the mean squared error (MSE), and the sum of squared errors (SSE). However, the cost function is not limited to the items listed above.

[0939] When performing the decoder - side motion information derivation method on a target block, the cost function used in the decoder - side motion information derivation method may be predefined.

[0940] When performing the decoder - side motion information derivation method on a target block, information about the cost function used in the decoder - side motion information derivation method may be signaled / encoded / decoded.

[0941] A template may be all or part of the block information of one or more blocks included in a specific region.

[0942] For example, the template of a target block may include one or more of the following: 1) a subset of the block information of the blocks included in the TMSIZE_LEFT line adjacent to the left side of the target block and 2) a subset of the block information of the blocks included in the TMSIZE_ABOVE line adjacent to the top of the target block. However, the relationship between the template and the position of the target block or the template configuration method is not limited to the above - mentioned embodiment.

[0943] For example, the template of a reference block may include one or more of the following: 1) a subset of the block information of the blocks included in the TMSIZE_LEFT line adjacent to the left side of the reference block and 2) a subset of the block information of the blocks included in the TMSIZE_ABOVE line adjacent to the top of the reference block. However, the relationship between the template and the position of the reference block or the template configuration method is not limited to the above - mentioned embodiment.

[0944] Each of TMSIZE_LEFT and TMSIZE_ABOVE may be a predefined integer of 0 or greater.

[0945] In an embodiment, the block information may refer to at least one of information about the target block, information about neighboring blocks, information about reference blocks, and information about col blocks.

[0946] In addition, the block information may include at least one of the encoding parameters.

[0947] In addition, the block information may include at least one of multiple pieces of information used in inter - frame prediction, intra - frame prediction, transformation, inverse transformation, quantization, de - quantization, entropy encoding / decoding, and loop filtering.

[0948] That is to say, the block information may include at least one value, combined value or statistical value among the following: block size, block depth, block partition information, block shape (square or non-square), information indicating whether the block is partitioned in a quadtree form, information indicating whether the block is partitioned in a binary tree form, partition direction (horizontal or vertical) in the binary tree form, partition shape (symmetric or asymmetric) in the binary tree form, prediction mode (intra prediction or inter prediction), intra luminance prediction mode / direction, intra chrominance prediction mode / direction, intra partition information, inter partition information, coding block partition flag, block partition flag, transform block partition flag, reference sample filter taps, reference sample filter coefficients, block filter taps, block filter coefficients, block boundary filter taps, block boundary filter coefficients, motion vector (motion vector of at least one of L0, L1, L2, and L3), motion vector difference (motion vector difference of at least one of L0, L1, L2, and L3), inter prediction direction (inter prediction direction of at least one of uni-directional prediction and bi-directional prediction), reference image (picture) index (reference image index of at least one of L0, L1, L2, and L3), inter prediction indicator, prediction list utilization flag, reference image list, motion vector prediction index, motion vector prediction candidate, motion information candidate list, information indicating whether the merge mode is used, merge index, merge candidate, merge candidate list, information indicating whether the skip mode is used, interpolation filter type, interpolation filter taps, interpolation filter coefficients, motion vector magnitude, precision of motion vector representation (unit or resolution representing the motion vector,Such as integer samples, 1 / 2 samples, 1 / 4 samples, 1 / 8 samples, 1 / 16 samples, and 1 / 32 samples), transform type, transform size, information indicating whether to use a primary transform, information indicating whether to use a secondary transform, primary transform index, secondary transform index, information indicating whether there is a residual signal, coding block style, coding block flag, quantization parameter, residual quantization parameter, quantization matrix, information indicating whether a loop filter is applied, loop filter coefficient, loop filter tap, loop filter shape / type, information indicating whether a deblocking filter is applied, deblocking filter coefficient, deblocking filter tap, deblocking filter strength, deblocking filter shape / type, information indicating whether an adaptive sample offset is applied, adaptive sample offset value, adaptive sample offset category, adaptive sample offset type, information indicating whether an adaptive loop filter is applied, adaptive loop filter coefficient, adaptive loop filter tap, adaptive loop filter shape / type, binarization / inverse binarization method, context model determination method, context model update method, information indicating whether to execute a normal mode, information indicating whether to execute a bypass mode, context binary bits, bypass binary bits, valid coefficient flag, last valid coefficient flag, flag indicating whether to encode / decode a unit of a coefficient group, position of the last valid coefficient, flag indicating whether a coefficient value is greater than 1, flag indicating whether a coefficient value is greater than 2, flag indicating whether a coefficient value is greater than 3, remaining coefficient value information, positive / negative sign information, reconstructed luminance samples, reconstructed chrominance samples, residual luminance samples, residual chrominance samples, luminance transform coefficients, chrominance transform coefficients, luminance quantization level, chrominance quantization level, transform coefficient horizontal scan method, size of the motion information search range on the decoder side, shape of the motion information search range on the decoder side, number of motion information search iterations on the decoder side, CTU size information, minimum block size information, maximum block size information, maximum block depth information, minimum block depth information, slice identification information, slice partition information, parallel block identification information, parallel block type, parallel block partition information, input sample bit depth, reconstructed sample bit depth, residual sample bit depth, transform coefficient bit depth, quantization level bit depth, indicator for the overlapping block motion compensation (OBMC) mode, indicator for the local illumination compensation (LIC) mode, and motion information offset.,

[0949] For example, the cost function can be a function that determines the similarity between multiple pieces of motion information and / or coding parameters of a template.

[0950] In an embodiment, at least one of the following can be used to determine the similarity between a first value and a second value: 1) the difference between the two values, 2) the ratio of the two values, and 3) an operation of comparing the difference between the two values with a specific value.

[0951] For example, the fact that multiple pieces of motion information (or coding parameters) of a template are similar to each other may mean a case where the difference between the multiple pieces of motion information (or coding parameters) of the template is less than or equal to a predefined positive number THRES_PARAMETER. However, the criteria based on which similarity is determined are not limited to the above comparison.

[0952] For example, the fact that multiple pieces of motion information (or coding parameters) of a template are similar to each other may mean a case where the sum of the differences between the multiple pieces of motion information (or coding parameters) of the template is less than or equal to a predefined positive number THRES_PARAMETER. However, the criteria based on which similarity is determined are not limited to the above comparison.

[0953] For example, the fact that multiple pieces of motion information (or coding parameters) of a template are similar to each other may mean a case where the ratio of the multiple pieces of motion information (or coding parameters) of the template is less than or equal to a predefined positive number WEIGHT_THRES_PARAMETER. However, the criteria based on which similarity is determined are not limited to the above comparison.

[0954] For example, the fact that multiple pieces of motion information (or coding parameters) of a template are similar to each other may mean a case where the ratio of the multiple pieces of motion information (or coding parameters) of the template is equal to or greater than a predefined positive number WEIGHT_THRES_PARAMETER. However, the criteria based on which similarity is determined are not limited to the above comparison.

[0955] For example, the fact that multiple pieces of motion information (or coding parameters) of a template are similar to each other may mean a case where the number of multiple identical pieces of motion information (or coding parameters) in the template is equal to or greater than a predefined positive number THRES_PARAMETER, but the criteria based on which similarity is determined are not limited to the above comparison.

[0956] Each of THRES_PARAMETER and WEIGHT_THRES_PARAMETER can be a predefined positive integer.

[0957] Each of THRES_PARAMETER and WEIGHT_THRES_PARAMETER can be set based on the motion information (or coding parameters). Each of THRES_PARAMETER and WEIGHT_THRES_PARAMETER can vary according to the motion information (or coding parameters). Each of THRES_PARAMETER and WEIGHT_THRES_PARAMETER can be set based on the type of the motion information (or coding parameters). Each of THRES_PARAMETER and WEIGHT_THRES_PARAMETER can vary according to the type of the motion information (or coding parameters).

[0958] For example, the refined motion information candidate list may include DMVG_NUM pieces of motion information.

[0959] DMVG_NUM can be a predefined positive integer of 1 or greater.

[0960] The information indicating DMVG_NUM may be signaled / encoded / decoded.

[0961] DMVG_NUM may be equal to the size of the initial motion information candidate list.

[0962] DMVG_NUM may be different from the size of the initial motion information candidate list.

[0963] For example, DMVD_NUM may be less than the size of the initial motion information candidate list.

[0964] Each candidate in the refined motion information candidate list may be the initial motion information or the refined motion information generated based on the initial motion information.

[0965] For example, the refined motion information candidate list may be configured to include DMVG_NUM pieces of initial motion information having the lowest matching cost among the candidates in the initial motion information candidate list.

[0966] In an embodiment, "multiple pieces of information having the lowest matching cost" may refer to "multiple pieces of information selected in ascending order of the matching cost".

[0967] For example, multiple pieces of refined motion information may be generated by applying refinement to the candidates in the initial motion information candidate list. The refined motion information candidate list may be configured to include DMVG_NUM pieces of refined motion information having the lowest matching cost among the multiple pieces of refined motion information.

[0968] For example, in a case where the motion information in the LX direction has been determined as MVP_LX when refining the initial motion information candidate list for the L(1 - X) direction, each candidate in the initial motion information candidate list for the L(1 - X) direction may be refined to minimize the bilateral matching cost with MVP_LX.

[0969] For example, in a case where the motion information in the LX direction has been determined as MVP_LX when refining the initial motion information candidate list for the L(1 - X) direction, the order of the candidates in the initial motion information candidate list for the L(1 - X) direction may be re - sorted in ascending order of the bilateral matching cost with MVP_LX.

[0970] X can be 0, 1, or a positive integer.

[0971] X can be a predefined value. The predefined value can be 0.

[0972] For example, the predefined value may refer to a value indicating the direction with a lower matching cost among the L0 direction and the L1 direction. Here, the matching cost of a specific direction may be the matching cost of the motion information in the specific direction.

[0973] For example, the predefined value may refer to a value indicating the direction with a higher matching cost among the L0 direction and the L1 direction. Here, the matching cost of a specific direction may be the matching cost of the motion information in the specific direction.

[0974] Information indicating X can be signaled / encoded / decoded.

[0975] The information indicating X can be determined by rate-distortion optimization processing. Through this determination, the encoding efficiency of the target block can be improved.

[0976] For example, when performing a search in a motion information derivation method on the decoder side, the position-based weight w can be used to match the cost of the motion information in each search process.

[0977] For example, the matching cost of the specific motion information in each search process can be refined by an operation using the weight w. This operation can be at least one of the four arithmetic operations.

[0978] Whether to use the weight w and / or the value of the weight w can be determined based on the position indicated by the initial motion information.

[0979] The initial motion information may refer to the initial motion information of the motion information derivation method on the decoder side. Optionally, the initial motion information may refer to the motion information that is the target to be refined in each search process of the motion information derivation method on the decoder side.

[0980] In an embodiment, the weight w may be a value that multiplies the matching cost of the specific motion information in each search process.

[0981] For example, the weight w can be determined based on the distance between the position indicated by the specific motion information and the position indicated by the initial motion information.

[0982] For example, based on the distance between the position indicated by the specific motion information and the position indicated by the initial motion information, it can be determined whether the weight w has a value other than 1 in the search process.

[0983] For example, based on the distance between the position indicated by the specific motion information and the position indicated by the initial motion information, it can be determined whether the weight w multiplies the matching cost of the specific motion information in the search process.

[0984] The weight w can be predefined.

[0985] For example, the predefined weight W may be multiplied by the matching cost at the position indicated by the initial motion information. Here, W may be a value less than or equal to 1, or a value equal to or greater than 1.

[0986] For example, when the AMVP mode is used for the target block, W at the position indicated by the initial motion information may be a value equal to or greater than 1.

[0987] For example, when the merge mode is used for the target block, W at the position indicated by the initial motion information may be a value less than or equal to 1.

[0988] For example, when the AMVP mode is used for the target block, since a specific position is closer to the position indicated by the initial motion information, the weight W may have a larger value.

[0989] For example, when the merge mode is used for the target block, since a specific position is closer to the position indicated by the initial motion information, the weight W may have a smaller value.

[0990] For example, in a case where the value obtained by multiplying the matching cost of the specific motion information by the weight W when performing a specific search step in the decoder-side motion information derivation method is less than or equal to 1, the possibility that the initial motion information of the current search step will be refined into the specific motion information may increase.

[0991] For example, in a case where the value obtained by multiplying the matching cost of the specific motion information by the weight W when performing a specific search step in the decoder-side motion information derivation method is equal to or greater than 1, the possibility that the initial motion information of the current search step will be refined into the specific motion information may decrease.

[0992] The decoder-side motion information derivation method may include one or more search steps.

[0993] In each search step, the refined motion information may be determined by searching from the initial motion information in the search step.

[0994] Two different search steps may differ from each other in at least one of the initial motion information, the template configuration method, the type of cost function for determining similarity, the unit that performs the refinement of the motion information, the search range, the search pattern, the search resolution, and the search method.

[0995] For example, each search step of the decoder-side motion information derivation method may be regarded as an independent decoder-side motion information derivation method implemented using one search step.

[0996] For example, a first decoder-side motion information derivation method consisting of L search steps and a second decoder-side motion information derivation method consisting of M search steps can be considered as a decoder-side motion information derivation method consisting of N search steps. N can be the sum of L and M.

[0997] The types of search methods may differ from each other based on one or more of the following: 1) search pattern, 2) search resolution, 3) search range, and 4) unit for deriving motion information. However, the criteria for classifying the types of search methods are not limited to such conditions.

[0998] The search resolution can be one of 4 pixels, full pixel, half pixel, and quarter pixel. However, the search resolution is not limited to the above pixel values.

[0999] The search pattern can be one of the diamond pattern, cross pattern, and full search pattern. However, the search pattern is not limited to the patterns listed above.

[1000] Search using the diamond pattern may refer to the operation of searching one or more of the positions (0, 2×RR), (RR, RR), (2×RR, 0), (RR, -RR), (0, -RR), (-RR, -RR), (-RR, 0), (-RR, RR), and (0, 0) when (0, 0) represents the position indicated by the initial motion information. RR may refer to the search resolution and can be a predefined positive number.

[1001] Search using the cross pattern can be the operation of searching one or more of the positions (0, RR), (RR, 0), (0, -RR), (-RR, 0), and (0, 0) when (0, 0) represents the position indicated by the initial motion information.

[1002] Search using the full search pattern can be the operation of searching all positions within a predefined search range.

[1003] For example, assuming that FS_i has values ranging from -FS_X to FS_X and FS_j has values from -FS_Y to FS_Y, search using the full search pattern may refer to the operation of searching the position corresponding to (FS_i×RR, FX_j×RR). Here, (0,0) can be the position indicated by the initial motion information. However, the search range is not limited to the above positions. Each of FS_X and FS_Y can be a predefined positive number.

[1004] For example, the decoder-side motion information derivation method may be configured in the following order: 1) a step of deriving motion information of an entire block, and 2) a step of deriving motion information of a sub-block of the block. However, the method of deriving motion information performed at the corresponding steps and the order of the corresponding steps are not limited to the methods and order described in the embodiments above.

[1005] For example, the decoder-side motion information derivation method may consist of N search steps. In the first search step, motion information of a first block may be derived. In a second search step following the first search step, motion information of a second block may be derived. Here, the unit of the second block may be smaller than the unit of the first block. In other words, the second block may be a sub-block of the first block. Alternatively, the unit of the first block and the unit of the second block may be the same as each other. In other words, the first block and the second block may be the same as each other.

[1006] For example, assuming that the derivation of motion information of an entire block is performed in the first search step, motion information of the entire block or motion information of a sub-block of the block may be derived in the second search step.

[1007] The search method for each step may be predefined based on the coding parameters or motion information of the target block.

[1008] When the decoder-side motion information derivation method consists of two or more search steps, the initial motion information input at a specific search step after the first search step may be the motion information refined at a search step before the specific search step.

[1009] For example, the decoder-side motion information derivation method may consist of three search steps. 1) In the first search step, motion information may be refined using bilateral matching for an entire block. 2) In the second search step, motion information may be refined by bilateral matching on a per-sub-block basis, where each sub-block has a size of 16×16. 3) In the third search step, motion information may be refined using optical flow-based motion information prediction on a per-sub-block basis, where each sub-block has a size of 8×8.

[1010] The motion information refined by the decoder-side motion information derivation method according to the embodiment may replace the initial motion information. For example, the refined motion information may be used to deter...

Claims

1. An image encoding method, comprising: Determining encoding information to be used for encoding a target block; Generating encoded encoding information by performing encoding on the encoding information; And Generating a bitstream including the encoded encoding information.

2. The image encoding method according to claim 1, further comprising: Performing inter-frame prediction on the target block, Wherein, a decoder-side motion information derivation method for inter-frame prediction is performed, and Wherein, subsampling is used in the decoder-side motion information derivation method.

3. The image encoding method according to claim 2, wherein, The subsampling is used to configure a template in the decoder-side motion information derivation method.

4. The image encoding method according to claim 2, wherein, The subsampling is used to configure a cost function in the decoder-side motion information derivation method.

5. The image encoding method according to claim 2, wherein, The subsampling is used to perform a search in the decoder-side motion information derivation method.

6. The image encoding method according to claim 1, further comprising: Performing prediction on the target block, Wherein, subsampling is used for template matching of the prediction.

7. The image encoding method according to claim 6, wherein, The subsampling is used to calculate a cost function between templates for the template matching.

8. An image decoding method, comprising: Obtaining a bitstream including encoded encoding information; Generating encoding information by performing decoding on the encoded encoding information; And Using the encoding information to determine prediction information to be used for decoding a target block.

9. The image decoding method according to claim 8, further comprising: Performing inter-frame prediction on the target block, Wherein, a decoder-side motion information derivation method for inter-frame prediction is performed, and Wherein, subsampling is used in the decoder-side motion information derivation method.

10. The image decoding method according to claim 8, wherein, The subsampling is used to configure a template in the decoder-side motion information derivation method.

11. The image decoding method according to claim 8, wherein, The subsampling is used to configure a cost function in the decoder-side motion information derivation method.

12. The image decoding method according to claim 8, wherein, The subsampling is used to perform a search in the decoder-side motion information derivation method.

13. The image decoding method according to claim 1, further comprising: Performing prediction on the target block, Wherein, subsampling is used for template matching of the prediction.

14. The image decoding method according to claim 13, wherein, The subsampling is used to calculate a cost function between templates for the template matching.

15. A computer-readable storage medium for storing a bitstream for image decoding, wherein: The bitstream includes encoded encoding information, The encoding information is generated by performing decoding on the encoded encoding information, and Prediction information to be used for decoding a target block is determined using the encoding information.

16. The computer-readable storage medium according to claim 15, wherein: Inter-frame prediction of the target block is performed, A decoder-side motion information derivation method for inter-frame prediction is performed, and Subsampling is used in the decoder-side motion information derivation method.

17. The computer-readable storage medium according to claim 16, wherein, The subsampling is used to configure a template in the decoder-side motion information derivation method.

18. The computer-readable storage medium according to claim 16, wherein, The subsampling is used to configure a cost function in the decoder-side motion information derivation method.

19. The computer-readable storage medium according to claim 16, wherein, The subsampling is used to perform a search in the decoder-side motion information derivation method.

20. The computer-readable storage medium according to claim 15, wherein: Prediction of the target block is performed; Subsampling is used for template matching in the inter-frame prediction.

Citation Information

Patent Citations

  • Formation method of lithium secondary battery

    KR1020220129783A

  • Battery pack structure including heater plate

    KR1020230135246A