Method for image encoding / decoding and recording medium storing bitstream

By using a BM-based image encoding/decoding method, the reconstruction of image prediction blocks is optimized through BM search and weighted summation operations, which solves the problem of insufficient prediction accuracy in high-resolution image encoding/decoding and achieves more efficient image data compression.

CN120883609APending Publication Date: 2025-10-31ELECTRONICS & TELECOMM RES INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480019068.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-03-15
Filing Date
2024-03-14
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing image encoding/decoding technologies lack sufficient prediction accuracy in high-resolution and high-definition image processing, making it difficult to effectively compress and decode high-resolution image data.

Method used

A bilateral matching (BM)-based image encoding/decoding method is adopted. By deriving the initial motion vector of the target block, BM search and weighted summation are performed. Combined with intra-frame prediction and inter-frame prediction, the reconstruction process of the image prediction block is optimized.

Benefits of technology

It improves the prediction accuracy of image encoding/decoding, reduces the amount of residual data, and enhances image compression efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120883609A_ABST
    Figure CN120883609A_ABST
Patent Text Reader

Abstract

The invention provides a bilateral matching (BM)-based image encoding / decoding method and a recording medium. As an example, an image decoding method according to the present disclosure may comprise the steps of: deriving a first initial motion vector and a second initial motion vector of a target block; performing a BM search on the first BM reference image and the second BM reference image by using the first initial motion vector and the second initial motion vector; and reconstructing the target block by using the first BM reference block and the second BM reference block acquired by the BM search.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to methods, apparatus, and recording media for image encoding / decoding. Background Technology

[0002] With the continued development of the information and communication industry, broadcast services supporting high-definition (HD) resolution have become widespread throughout the world. Through this widespread adoption, a large number of users have become accustomed to high-resolution and high-definition images and / or videos.

[0003] To meet user demand for high definition, numerous organizations have accelerated the development of next-generation imaging devices. In addition to High Definition TV (HDTV) and Full High Definition (FHD) TV, user interest in UHD TV has also increased, with UHD TV offering more than four times the resolution of FHD TV. With this growing interest, there is now a need for image encoding / decoding technologies specifically designed for images with higher resolution and greater definition.

[0004] As an image compression technique, there are various techniques (such as inter-frame prediction, intra-frame prediction, transform, quantization, filtering, and entropy coding).

[0005] Inter-frame prediction techniques are used to predict the values ​​of pixels included in the current frame using frames preceding and / or following the current frame. Intra-frame prediction techniques are used to predict the values ​​of pixels included in the current frame using information about the pixels in the current frame. Transform and quantization techniques can be used to compress the energy of residual signals. Entropy coding techniques are used to assign short codewords to frequently occurring values ​​and long codewords to less frequent values.

[0006] By utilizing these image compression techniques, data about images can be effectively compressed, sent, and stored. Summary of the Invention

[0007] Technical issues

[0008] This disclosure aims to provide a BM-based image encoding / decoding apparatus, method, and recording medium.

[0009] Specifically, this disclosure aims to provide a method for predicting target blocks based on BM when encoding / decoding an image.

[0010] Specifically, this disclosure aims to provide a method for predicting the residual signal of a target block based on BM when encoding / decoding an image.

[0011] Specifically, this disclosure aims to provide a method for refining the motion vectors of a target block based on BM when encoding / decoding an image.

[0012] Technical solution

[0013] The image decoding method according to embodiments of the present disclosure may include: deriving a first initial motion vector and a second initial motion vector of a target block; performing a BM search on a first bilateral matching (BM) reference image and a second BM reference image using the first initial motion vector and the second initial motion vector; and reconstructing the target block using the first BM reference block and the second BM reference block obtained by the BM search.

[0014] In the image decoding method according to embodiments of the present disclosure, the optimal BM block can be derived based on the average operation or weighted summation operation of the first BM reference block and the second BM reference block.

[0015] In the image decoding method according to embodiments of the present disclosure, the best block of the BM can be configured as the prediction block of the target block.

[0016] In the image decoding method according to embodiments of the present disclosure, the final prediction block of the target block can be obtained by performing a weighted summation operation between the BM block and the intra-predicted block obtained by performing intra-prediction on the target block.

[0017] In the image decoding method according to embodiments of the present disclosure, intra-prediction blocks can be obtained by using an intra-prediction mode predefined in the decoder.

[0018] In the image decoding method according to embodiments of the present disclosure, an intra-prediction mode for intra-prediction can be derived from a first BM reference block or a second BM reference block encoded by intra-prediction.

[0019] In the image decoding method according to embodiments of the present disclosure, a first weight applied to an intra-prediction block and a second weight applied to an optimal BM block can be determined by comparing a threshold with a matching criterion between a first BM reference block and a second BM reference block.

[0020] In the image decoding method according to embodiments of the present disclosure, the final prediction block of the target block can be obtained by performing a weighted summation operation between the BM block and the inter-prediction block obtained by performing inter-frame prediction on the target block.

[0021] In the image decoding method according to embodiments of the present disclosure, a first BM reference image or a second BM reference image may be configured as a reference image for inter-frame prediction.

[0022] In the image decoding method according to embodiments of the present disclosure, the residual block of the target block can be derived as the sum of the prediction residual block and the BM residual block, and the BM residual block can be obtained based on the residual coefficients decoded from the bitstream, and the residual block of the BM optimal block can be configured as the prediction residual block.

[0023] In the image decoding method according to embodiments of the present disclosure, the residual block of the BM optimal block can be derived by subtracting the prediction block of the target block from the BM optimal block.

[0024] In the image decoding method according to embodiments of the present disclosure, residual samples in the residual block can be reconstructed by using only the predicted residual samples in the predicted residual block in a portion of the target block.

[0025] In the image decoding method according to embodiments of the present disclosure, a first initial motion vector may be derived based on advanced motion vector prediction (AMVP) coding information, and a second initial motion vector may be derived based on merged coding information.

[0026] In the image decoding method according to an embodiment of the present disclosure, a first initial motion vector may be configured to be the same as the motion vector of a merge candidate included in a merge candidate list, and a vector with the same magnitude but opposite direction to the first initial motion vector may be configured as a second initial motion vector.

[0027] In the image decoding method according to embodiments of the present disclosure, BM search may be performed only in the BM search region within the first BM reference image and the second BM reference image.

[0028] In the image decoding method according to an embodiment of the present disclosure, the shape of the BM search region is determined based on index information indicating one of the shape candidates, and among the shape candidates, the first shape candidate may be a rectangular shape and the second shape candidate may be a rhombus shape.

[0029] In the image decoding method according to embodiments of the present disclosure, based on the position indicated by the initial motion vector, the BM search region is configured to be N times the size of the basic unit (N is a natural number greater than or equal to 1), and the basic unit may be at least one of the target unit, the coding tree unit, or the virtual pipeline data unit (VPDU).

[0030] In the image decoding method according to embodiments of the present disclosure, when at least a portion of the region to be configured as the BM search region extends beyond the image boundary, the residual region other than the region extending beyond the image boundary can be configured as the BM search region.

[0031] The image encoding method according to embodiments of the present disclosure may include: deriving a first initial motion vector and a second initial motion vector of a target block; performing a BM search on a first bilateral matching (BM) reference image and a second BM reference image using the first initial motion vector and the second initial motion vector; and reconstructing the target block using the first BM reference block and the second BM reference block obtained by the BM search.

[0032] In this disclosure, a recording medium for recording a bitstream generated by an image encoding method may be provided.

[0033] Technical effect

[0034] According to this disclosure, a BM-based image encoding / decoding apparatus, method, and recording medium are provided.

[0035] According to this disclosure, when encoding / decoding an image, it has the effect of improving prediction accuracy by predicting target blocks based on BM.

[0036] According to this disclosure, when encoding / decoding an image, it has the effect of reducing the amount of residual data to be encoded / decoded by predicting the residual signal based on BM.

[0037] According to this disclosure, when encoding / decoding an image, it has the effect of improving prediction accuracy by refining the motion vector of the target block based on BM. Attached Figure Description

[0038] Figure 1 This is a block diagram illustrating the configuration of an embodiment of an encoding device to which this disclosure is applied;

[0039] Figure 2 This is a block diagram illustrating the configuration of an embodiment of a decoding device to which this disclosure is applied;

[0040] Figure 3 It is a schematic diagram illustrating the partitioning structure of an image as it is encoded and decoded;

[0041] Figure 4 This is a diagram showing the form of prediction units (PUs) that a coding unit (CU) can include;

[0042] Figure 5 This is a diagram showing the form of a transformation unit (TU) that can be included in a CU;

[0043] Figure 6 This shows the block division based on the example;

[0044] Figure 7 This is a diagram illustrating an embodiment used to explain the intra-frame prediction process;

[0045] Figure 8 This is a diagram showing the reference samples used in the intra-frame prediction process;

[0046] Figure 9 This is a diagram illustrating an embodiment used to explain the inter-frame prediction process;

[0047] Figure 10 Spatial candidates are shown according to an embodiment;

[0048] Figure 11 The order in which motion information of spatial candidates is added to the merging list is shown according to an embodiment;

[0049] Figure 12 The transformation and quantization processes are shown based on the example;

[0050] Figure 13 The diagonal scan is shown according to the example;

[0051] Figure 14 Showing a horizontal scan based on an example;

[0052] Figure 15 The vertical scan is shown according to the example;

[0053] Figure 16 This is a configuration diagram of the encoding device according to an embodiment;

[0054] Figure 17 This is a configuration diagram of the decoding device according to an embodiment;

[0055] Figure 18 This is a flowchart of a method for obtaining a prediction block based on the BM method according to an embodiment of the present disclosure.

[0056] Figure 19 and Figure 20 An example is shown where the BM reference block is determined using a BM reference image.

[0057] Figure 21 and Figure 22 The BM motion vector derived through BM search is shown.

[0058] Figures 23 to 34 Examples related to the BM search area are shown.

[0059] Figure 35 This shows an example where the location of the target block is configured as the start location of the BM search.

[0060] Figure 36 An example is shown of deriving the initial motion vector from a block adjacent to the target block.

[0061] Figure 37 An example is shown where matching criteria are calculated using pixels at subsampling locations.

[0062] Figure 38 An example of deriving the BM residual signal based on the BM optimal block is shown.

[0063] Figures 39 to 42 The region predicted by the residual signal is shown in the target block.

[0064] Figure 43This is a diagram used to illustrate an example of correcting the optimal block for BM.

[0065] Figure 44 An example is shown where linear model coefficients are derived using only some location pixels within the template.

[0066] Figure 45 An example of applying a downsampling filter to a template is shown.

[0067] Figure 46 This is a diagram used to describe the application aspects of intra-frame and BM hybrid prediction.

[0068] Figure 47 This is a diagram used to describe the application aspects of inter-frame and BM hybrid prediction. Detailed Implementation

[0069] This invention can be modified in various ways and can have various embodiments, which will be described in detail below with reference to the accompanying drawings. However, it should be understood that these embodiments are not intended to limit the invention to the specific forms disclosed, and they include all variations, equivalents, or modifications included within the spirit and scope of the invention.

[0070] The following exemplary embodiments will be described in detail with reference to the accompanying drawings, which illustrate specific embodiments. These embodiments are described to enable those skilled in the art to readily implement them. It should be noted that the various embodiments differ from one another but are not necessarily mutually exclusive. For example, specific shapes, structures, and characteristics described herein may be implemented as other embodiments without departing from the spirit and scope of other embodiments associated with one embodiment. Furthermore, it should be understood that the position or arrangement of various components in each disclosed embodiment can be changed without departing from the spirit and scope of the embodiments. Therefore, the appended detailed description is not intended to limit the scope of this disclosure, and the scope of the exemplary embodiments is defined only by the appended claims and their equivalents (provided they are properly described).

[0071] In the accompanying drawings, similar reference numerals are used to designate the same or similar functions in various respects. The shape, size, etc., of the components in the drawings may be exaggerated for clarity of description.

[0072] Terms such as “first” and “second” may be used to describe various components, but components are not limited by these terms. These terms are used only to distinguish one component from another. For example, without departing from the scope of this specification, a first component may be referred to as a second component. Similarly, a second component may be referred to as a first component. The term “and / or” may include a combination of multiple related descriptive terms or any one of multiple related descriptive terms.

[0073] It will be understood that when a component is referred to as "connected" or "combined" to another component, the two components may be directly connected or combined with each other, or there may be an intermediate component between the two components. On the other hand, it will be understood that when a component is referred to as "directly connected or combined," there is no intermediate component between the two components.

[0074] Furthermore, the components described in the embodiments are shown independently to indicate different functional characteristics, but this does not mean that each component is formed by a single piece of hardware or software. That is, for ease of description, multiple components are arranged and included separately. For example, at least two of the multiple components may be integrated into a single component. Conversely, a component may be divided into multiple components. As long as it does not depart from the spirit of this specification, embodiments in which multiple components are integrated or embodiments in which some components are separated are included within the scope of this specification.

[0075] The terminology used in the embodiments is for describing particular embodiments only and is not intended to limit the invention. Singular expressions include plural expressions unless the context specifically indicates the contrary. In the embodiments, it should be understood that terms such as “comprising” or “having” are intended only to indicate the presence of features, numbers, steps, operations, components, parts, or combinations thereof, and are not intended to exclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof. That is, in the embodiments, the expression describing a component as “comprising” a specific component means that additional components may be included within the practice or technical spirit of the invention, but do not exclude the presence of components other than the specific component stated therein.

[0076] In embodiments, the term "at least one" may mean one of one or more quantities (such as 1, 2, 3, and 4). In embodiments, the term "a plurality of" may mean one of two or more quantities (such as 2, 3, and 4).

[0077] Some components of the embodiments are not essential components for performing the necessary functions, but may be optional components used only to improve performance. Embodiments may be implemented using only the essential components necessary to achieve the essence of the embodiments. For example, a structure that includes only the essential components (excluding optional components used only to improve performance) is also included within the scope of the embodiments.

[0078] The embodiments will now be described in detail with reference to the accompanying drawings, enabling those skilled in the art to readily implement the embodiments. In the following description of the embodiments, detailed descriptions of well-known functions or configurations that are considered to obscure key points of this specification will be omitted. Furthermore, the same reference numerals are used throughout the drawings to designate the same components, and repeated descriptions of the same components will be omitted.

[0079] In the following text, "image" may refer to a single frame that constitutes a video, or it may refer to the video itself. For example, "encoding and / or decoding of an image" may mean "encoding and / or decoding of a video," and may also mean "encoding and / or decoding of any one of the multiple images that constitute a video."

[0080] In the following text, the terms “video” and “moving footage” may be used to have the same meaning and may be used interchangeably.

[0081] In the following text, the target image can be an encoded target image that is a target to be encoded and / or a decoded target image that is a target to be decoded. Furthermore, the target image can be an input image input to an encoding device or an input image input to a decoding device. Also, the target image can be the current image, i.e., the target currently to be encoded and / or decoded. For example, the terms "target image" and "current image" can be used to have the same meaning and can be used interchangeably.

[0082] In the following text, the terms “image,” “picture,” “frame,” and “screen” may be used to have the same meaning and may be used interchangeably with each other.

[0083] In the following text, a target block can be an encoding target block (i.e., the target to be encoded) and / or a decoding target block (i.e., the target to be decoded). Furthermore, a target block can be the current block, i.e., the target currently to be encoded and / or decoded. Here, the terms "target block" and "current block" can be used to have the same meaning and are interchangeable. A current block can represent an encoding target block that is the encoding target during encoding and / or a decoding target block that is the decoding target during decoding. Furthermore, a current block can be at least one of an encoding block, a prediction block, a residual block, and a transform block.

[0084] In the following text, the terms “block” and “unit” may be used to have the same meaning and may be used interchangeably. Optionally, “block” may refer to a specific unit.

[0085] In the following text, the terms “region” and “fragment” are used interchangeably.

[0086] In the following embodiments, specific information, data, flags, indexes, elements, and attributes may have their own values. The value "0" corresponding to each of the information, data, flags, indexes, elements, and attributes may indicate false, logical false, or a first predefined value. In other words, the values ​​"0," false, logical false, and the first predefined value are interchangeable. The value "1" corresponding to each of the information, data, flags, indexes, elements, and attributes may indicate true, logical true, or a second predefined value. In other words, the values ​​"1," true, logical true, and the second predefined value are interchangeable.

[0087] When variables such as i or j are used to indicate rows, columns, or indices, the value i can be an integer 0 or a greater than 0, or an integer 1 or a greater than 1. In other words, in the embodiments, each of the rows, columns, and indices can be counted starting from 0 or starting from 1.

[0088] In the embodiments, the term "one or more" or the term "at least one" may mean the term "multiple". The terms "one or more" or the term "at least one" may be used interchangeably with "multiple".

[0089] The terminology used in the embodiments will be described below.

[0090] Encoder: An encoder represents a device used to perform encoding. In other words, an encoder can represent an encoding device.

[0091] Decoder: A decoder represents a device used to perform decoding. In other words, a decoder can represent a decoding device.

[0092] Unit: A unit can represent a component of image encoding and decoding. The terms "unit" and "block" can be used interchangeably and have the same meaning.

[0093] – A cell can be an M×N sample array. Each of M and N can be a positive integer. Cells can typically represent sample arrays in two-dimensional form.

[0094] In the process of image encoding and decoding, a "unit" can be a region generated by partitioning an image. In other words, a "unit" can be a specified region within an image. A single image can be partitioned into multiple units. Optionally, an image can be partitioned into sub-parts, and a unit can represent each sub-part created when encoding or decoding is performed on the partitioned sub-parts.

[0095] – During the encoding and decoding of an image, predefined processing can be performed on each unit according to its type.

[0096] Based on function, unit types can be classified as macrounits, coding units (CUs), prediction units (PUs), residual units, transform units (TUs), etc. Optionally, based on function, units can represent blocks, macroblocks, coding tree units, coding tree blocks, coding units, coding blocks, prediction units, prediction blocks, residual units, residual blocks, transform units, transform blocks, etc. For example, the target unit that serves as the object of encoding and / or decoding can be at least one of CUs, PUs, residual units, and TUs.

[0097] – The term “unit” can refer to a block of luma components, a block of chroma components corresponding to the luma components, and information about the syntax elements for each block, such that the unit is specified to be distinct from the block.

[0098] The size and shape of the unit can be implemented differently. In addition, the unit can have any of a variety of sizes and shapes. Specifically, the shape of the unit can include not only squares, but also geometric shapes that can be represented in two dimensions (2D) (such as rectangles, trapezoids, triangles and pentagons).

[0099] In addition, cell information may include one or more of the following: cell type, cell size, cell depth, cell encoding order, and cell decoding order. For example, the cell type may indicate one of CU, PU, ​​residual cell, and TU.

[0100] – A cell can be divided into sub-cells, each sub-cell having a smaller size than the related cell.

[0101] Depth: Depth represents the degree to which a cell is partitioned. Furthermore, cell depth indicates the level at which a cell exists when represented by a tree structure.

[0102] – Cell partitioning information may include depth, which indicates the depth of the cell. Depth may indicate the number of times the cell is partitioned and / or the degree to which the cell is partitioned.

[0103] In a tree structure, the root node can be considered to have the smallest depth and the leaf nodes have the largest depth. The root node can be the highest (top) node. The leaf nodes can be the lowest nodes.

[0104] A single unit can be hierarchically partitioned into multiple sub-units, with each unit possessing depth information based on a tree structure. In other words, a unit and the sub-units generated by partitioning that unit can correspond to a node and its child nodes, respectively. Each partitioned sub-unit can have a unit depth. Since depth indicates the number of times a unit is partitioned and / or the degree to which a unit is partitioned, the partitioning information of a sub-unit can include information about the size of that sub-unit.

[0105] In a tree structure, the top node corresponds to the initial node before partitioning. The top node can be called the "root node". Furthermore, the root node can have a minimum depth value. Here, the depth of the top node can be level "0".

[0106] A node with a depth of level "1" can represent a cell generated when the initial cell is partitioned once. A node with a depth of level "2" can represent a cell generated when the initial cell is partitioned twice.

[0107] – A leaf node of depth “n” can represent a cell generated when the initial cell is partitioned n times.

[0108] – Leaf nodes can be bottom nodes that cannot be further partitioned. The depth of a leaf node can be the maximum level. For example, a predefined value for the maximum level could be 3.

[0109] –QT depth can represent the depth for a four-partition drive. BT depth can represent the depth for a two-partition drive. TT depth can represent the depth for a three-partition drive.

[0110] – Sample: A sample can be the basic unit that makes up a block. It can be 0 to 2 based on the bit depth (Bd). Bd The value of -1 is used to represent the sample point.

[0111] – A sample point can be a pixel or a pixel value.

[0112] – In the following text, the terms “pixel” and “sample” may be used to have the same meaning and may be used interchangeably.

[0113] Code Tree Unit (CTU): A CTU can consist of a single luma component (Y) code tree block and two chroma component (i.e., Cb, Cr) code tree blocks associated with the luma component code tree block. Furthermore, a CTU can represent information including the aforementioned blocks and the syntax elements used for each block.

[0114] – Each coding tree unit (CTU) can be partitioned using one or more partitioning methods (such as quadtree (QT), binary tree (BT), and ternary tree (TT)) to configure sub-units, such as coding units, prediction units, and transform units. A quadtree can represent a quaternion tree. Alternatively, one or more partitioning methods can be used to partition each coding tree unit using multi-type tree (MTT).

[0115] – “CTU” can be used as a term to specify a pixel block as a processing unit in image decoding and encoding processes (such as in the case of partitioning an input image).

[0116] Coding Tree Block (CTB): "CTB" can be used as a term to specify any one of the Y coding tree block, Cb coding tree block, and Cr coding tree block.

[0117] Neighboring Blocks: Neighboring blocks (or adjacent blocks) can represent blocks that are adjacent to the target block. Neighboring blocks can also represent reconstructed neighboring blocks.

[0118] In the following text, the terms “nearby block” and “adjacent block” may be used to have the same meaning and may be used interchangeably.

[0119] Neighboring blocks can represent reconstructed neighboring blocks.

[0120] Spatial neighbor block: A spatial neighbor block can be a block that is spatially adjacent to the target block. Neighbor blocks can include spatial neighbor blocks.

[0121] – Target blocks and spatially adjacent blocks can be included in the target frame.

[0122] – A spatially adjacent block can represent a block whose boundary is in contact with the target block or a block located within a predetermined distance from the target block.

[0123] – A spatial neighbor block can represent a block that is adjacent to the vertex of the target block. Here, a block that is adjacent to the vertex of the target block can be a block that is vertically adjacent to a neighbor block that is horizontally adjacent to the target block, or a block that is horizontally adjacent to a neighbor block that is vertically adjacent to the target block.

[0124] Temporally neighboring blocks: Temporally neighboring blocks can be blocks that are temporally adjacent to the target block. Neighboring blocks can include temporally neighboring blocks.

[0125] –Time-proximity blocks can include col blocks.

[0126] The col block can be a block in a previously reconstructed co-location frame (col frame). The position of the col block in the col frame can correspond to the position of the target block in the target frame. Optionally, the position of the col block in the col frame can be equal to the position of the target block in the target frame. The col frame can be a frame included in the list of reference frames.

[0127] - A temporally neighboring block can be a block that is spatially adjacent to the target block in time.

[0128] Prediction mode: Prediction mode can be information indicating the mode used for intra-frame prediction or the mode used for inter-frame prediction.

[0129] Prediction Unit: A prediction unit can be a basic unit used for prediction (such as inter-frame prediction, intra-frame prediction, inter-frame compensation, intra-frame compensation, and motion compensation).

[0130] A single prediction unit can be divided into multiple partitions or sub-prediction units with smaller sizes. These partitions can also be the basic units used in performing prediction or compensation. Partitions generated by dividing the prediction unit can also be prediction units.

[0131] Prediction cell partitioning: Prediction cell partitioning can be the shape in which prediction cells are divided.

[0132] Reconstructed neighboring cells: The reconstructed neighboring cells can be cells that are adjacent to the target cell and have already been decoded and reconstructed.

[0133] - The reconstructed neighboring units can be those that are spatially adjacent to the target unit or temporally adjacent to the target unit.

[0134] – The reconstructed spatial neighboring unit can be a unit included in the target image that has already been reconstructed through encoding and / or decoding.

[0135] – The reconstructed temporal neighbor unit can be a unit included in the reference image that has already been reconstructed through encoding and / or decoding. The position of the reconstructed temporal neighbor unit in the reference image can be the same as the position of the target unit in the target image, or it can correspond to the position of the target unit in the target image. Furthermore, the reconstructed temporal neighbor unit can be a block adjacent to a corresponding block in the reference image. Here, the position of the corresponding block in the reference image can correspond to the position of the target block in the target image. The fact that the positions of the blocks correspond to each other can indicate that the positions of the blocks are the same, that one block is included in another block, or that one block occupies a specific position in another block.

[0136] Sub-screen: A screen can be divided into one or more sub-screens. A sub-screen can consist of one or more parallel block rows and one or more parallel block columns.

[0137] – A sub-screen can be an area within a screen that has a square or rectangular shape (i.e., a non-square rectangle). Furthermore, a sub-screen can include one or more CTUs.

[0138] – A sub-picture can be a rectangular area of ​​one or more strips in the picture.

[0139] A sub-picture may include one or more parallel blocks, one or more bricks, and / or one or more stripes.

[0140] Parallel blocks: Parallel blocks can be areas in the image that have a square or rectangular shape (i.e., a non-square rectangle).

[0141] – A parallel block may include one or more CTUs.

[0142] – Parallel blocks can be partitioned into one or more blocks.

[0143] Blocking: Blocks can represent one or more CTU lines in a parallel block.

[0144] – A parallel block can be partitioned into one or more blocks. Each block may contain one or more CTU rows.

[0145] – Parallel blocks that are not partitioned into two parts can also be represented as blocks.

[0146] Strip: A strip may include one or more parallel blocks of a frame. Optionally, a strip may include one or more sub-blocks of a parallel block.

[0147] A sub-picture may contain one or more stripes that share a rectangular area covering the picture. Therefore, each sub-picture boundary is always a stripe boundary, and each vertical sub-picture boundary is always a vertical parallel block boundary.

[0148] Parameter set: The parameter set corresponds to the header information in the internal structure of the bitstream.

[0149] The parameter set may include at least one of the following: Video Parameter Set (VPS), Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Adaptive Parameter Set (APS), Decoding Parameter Set (DPS).

[0150] Information sent via signals for each parameter set can be applied to the screen referencing the corresponding parameter set. For example, information in a VPS can be applied to the screen referencing the VPS. Information in an SPS can be applied to the screen referencing the SPS. Information in a PPS can be applied to the screen referencing the PPS.

[0151] - Each parameter set can reference a higher-level parameter set. For example, PPS can reference SPS. SPS can reference VPS.

[0152] Additionally, the parameter set may include parallel block groups, stripe header information, and parallel block header information. A parallel block group can be a group comprising multiple parallel blocks. Furthermore, the meaning of "parallel block group" can be the same as that of "strip".

[0153] Rate-distortion optimization: The coding device can use rate-distortion optimization to provide high coding efficiency by utilizing a combination of the following: the size of the coding unit (CU), the prediction mode, the size of the prediction unit (PU), motion information, and the size of the transform unit (TU).

[0154] Rate distortion optimization schemes calculate the rate distortion cost of each combination to select the optimal combination. The rate distortion cost can be calculated using the equation "D + λ * R". Typically, the combination that minimizes the rate distortion cost is selected as the optimal combination under the rate distortion optimization scheme.

[0155] –D can represent distortion. D can be the average of the squares of the differences between the original transform coefficients and the reconstructed transform coefficients in the transform unit (i.e., mean square error).

[0156] –R can represent the rate, which can use relevant context information to represent the bit rate.

[0157] –λ represents the Lagrange multiplier. R can include not only coding parameter information (such as prediction mode, motion information, and coding block flags), but also bits generated from encoding the transform coefficients.

[0158] The coding device can perform processes such as inter-frame prediction and / or intra-frame prediction, transform, quantization, entropy coding, inverse quantization (dequantization), and / or inverse transform to compute accurate D and R. These processes greatly increase the complexity of the coding device.

[0159] – Bitstream: A bitstream can represent a stream of bits that encode image information.

[0160] Parsing: Parsing can be the determination of the value of a syntax element by performing entropy decoding on a bitstream. Alternatively, the term "parsing" can refer to this entropy decoding itself.

[0161] Symbols: A symbol can be at least one of the syntax elements, encoding parameters, and transform coefficients of the encoding target unit and / or the decoding target unit. Furthermore, a symbol can be the target of entropy encoding or the result of entropy decoding.

[0162] Reference frame: The reference frame can be an image referenced by the cell to perform inter-frame prediction or motion compensation. Optionally, the reference frame can be an image that includes reference cells referenced by the target cell to perform inter-frame prediction or motion compensation.

[0163] In the following text, the terms “reference screen” and “reference image” may be used to have the same meaning and may be used interchangeably.

[0164] Reference frame list: The reference frame list can be a list of one or more reference images that are used for inter-frame prediction or motion compensation.

[0165] – The types of reference screen lists may include combined list (LC), list 0 (L0), list 1 (L1), list 2 (L2), list 3 (L3), etc.

[0166] – For inter-frame prediction, one or more reference frame lists can be used.

[0167] Inter-frame prediction indicator: The inter-frame prediction indicator indicates the direction of inter-frame prediction for the target cell. Inter-frame prediction can be either unidirectional or bidirectional. Optionally, the inter-frame prediction indicator can represent the number of reference frames used to generate the prediction cell for the target cell. Optionally, the inter-frame prediction indicator can represent the number of prediction blocks used for inter-frame prediction or motion compensation of the target cell.

[0168] Prediction list utilization flags: Prediction list utilization flags indicate whether to use at least one reference screen from a specific reference screen list to generate prediction cells.

[0169] – Inter-frame prediction indicators can be derived using prediction list utilization flags. Conversely, prediction list utilization flags can be derived using inter-frame prediction indicators. For example, a prediction list utilization flag of "0" (as a first value) indicates that reference frames from the reference frame list are not used to generate prediction blocks for the target cell. A prediction list utilization flag of "1" (as a second value) indicates that the reference frame list is used to generate prediction cells for the target cell.

[0170] Reference screen index: The reference screen index can be an index that indicates a specific reference screen in the list of reference screens.

[0171] Screen Order Count (POC): The POC value of a screen indicates the order in which the corresponding screens are displayed.

[0172] Motion Vector (MV): A motion vector can be a 2D vector used for inter-frame prediction or motion compensation. A motion vector can represent the offset between a target image and a reference image.

[0173] – For example, it can be in the form of (mv x ,mv y MV is represented in the form of ) x It can indicate the horizontal component, mv y It can indicate the vertical component.

[0174] – Search Range: The search range can be a 2D region where a search for the MV is performed during inter-frame prediction. For example, the size of the search range can be M×N. M and N can both be positive integers.

[0175] Motion vector candidates: Motion vector candidates can be blocks that are used as prediction candidates when motion vectors are predicted, or motion vectors that are used as prediction candidates.

[0176] – Motion vector candidates can be included in the motion vector candidate list.

[0177] Motion vector candidate list: The motion vector candidate list can be a list using one or more motion vector candidate configurations.

[0178] Motion vector candidate index: The motion vector candidate index can be an indicator used to indicate motion vector candidates in the motion vector candidate list. Optionally, the motion vector candidate index can be an index of motion vector predictors.

[0179] Motion information: Motion information may include at least one of the following: a list of reference frames, a reference image, motion vector candidates, a motion vector candidate index, a merge candidate and a merge index, as well as information on motion vectors, reference frame indexes and inter-frame prediction indicators.

[0180] Merge candidate list: The merge candidate list can be a list that uses one or more merge candidate configurations.

[0181] Merging Candidates: Merging candidates can be spatial merging candidates, temporal merging candidates, combined merging candidates, combined bidirectional prediction merging candidates, history-based candidates, candidates based on the average of two candidates, zero merging candidates, etc. Merging candidates may include inter-frame prediction indicators and may include motion information, such as prediction type information, reference frame index for each list, motion vectors, prediction list utilization flags, and inter-frame prediction indicators.

[0182] Merge index: A merge index can be an indicator used to indicate merge candidates in a merge candidate list.

[0183] - The merge index can indicate which reconstruction cell is used to derive the merge candidate among the reconstruction cells that are spatially adjacent to the target cell and temporally adjacent to the target cell.

[0184] – The merge index can indicate at least one of the multiple motion information candidates to be merged.

[0185] Transform Unit: A transform unit can be the basic unit for residual signal encoding and / or residual signal decoding (such as transform, inverse transform, quantization, dequantization, transform coefficient encoding, and transform coefficient decoding). A single transform unit can be partitioned into multiple sub-transform units with smaller sizes. Here, the transform may include one or more primary transforms and secondary transforms, and the inverse transform may include one or more primary inverse transforms and secondary inverse transforms.

[0186] Scaling: Scaling can represent the process of multiplying a factor by a transformation coefficient.

[0187] – As a result of scaling the levels of the transform coefficients, transform coefficients can be generated. Scaling can also be referred to as “inverse quantization”.

[0188] Quantization parameter (QP): The quantization parameter can be a value used to generate a transform coefficient level for the transform coefficients during quantization. Optionally, the quantization parameter can also be a value used to generate transform coefficient values ​​by scaling the transform coefficient level during dequantization. Optionally, the quantization parameter can be a value mapped to the quantization step size.

[0189] Delta quantization parameter: The Delta quantization parameter represents the difference between the quantization parameter of the target cell and the predicted quantization parameter.

[0190] Scan: A scan can refer to a method of arranging the coefficients in a cell, block, or matrix in order. For example, a method for arranging a 2D array in the form of a one-dimensional (1D) array can be called a "scan". Alternatively, a method for arranging a 1D array in the form of a 2D array can also be called a "scan" or "inverse scan".

[0191] Transform coefficients: Transform coefficients can be coefficient values ​​generated when the encoding device performs a transform. Optionally, transform coefficients can be coefficient values ​​generated when the decoding device performs at least one of entropy decoding and dequantization.

[0192] – The level of quantization or the level of quantized transform coefficients generated by applying quantization to transform coefficients or residual signals can also be included in the meaning of the term “transform coefficients”.

[0193] Quantization level: The quantization level can be a value generated when the encoding device performs quantization on the transform coefficients or residual signal. Optionally, the quantization level can be a value that serves as the target for dequantization when the decoding device performs dequantization.

[0194] – The transformation coefficient levels of quantization, as a result of transformation and quantization, can also be included in the meaning of the quantization levels.

[0195] Non-zero transform coefficients: Non-zero transform coefficients can be transform coefficients with values ​​other than 0, or transform coefficient levels with values ​​other than 0. Optionally, non-zero transform coefficients can be transform coefficients with a value amplitude that is not zero, or transform coefficient levels with a value amplitude that is not zero.

[0196] Quantization matrix: A quantization matrix is ​​a matrix used during the quantization or dequantization process to improve the subjective or objective image quality of an image. A quantization matrix can also be referred to as a "scaling list".

[0197] Quantization matrix coefficients: Quantization matrix coefficients can be each element in the quantization matrix. Quantization matrix coefficients are also referred to as "matrix coefficients".

[0198] Default matrix: The default matrix can be a quantization matrix predefined by the encoding and decoding devices.

[0199] Non-default matrix: A non-default matrix can be a quantization matrix that is not predefined by the encoding and decoding devices. A non-default matrix can represent a quantization matrix that is sent by the user from the encoding device to the decoding device.

[0200] Most Probable Mode (MPM): MPM can represent an intra-prediction mode that is highly likely to be used for intra-prediction of the target block.

[0201] The encoding and decoding devices can determine one or more MPMs based on encoding parameters associated with the target block and attributes of entities associated with the target block.

[0202] The encoding and decoding apparatus can determine one or more MPMs based on the intra-prediction modes of the reference blocks. A reference block may include multiple reference blocks. These multiple reference blocks may include spatially adjacent blocks to the left of the target block and spatially adjacent blocks above the target block. In other words, one or more distinct MPMs can be determined based on which intra-prediction modes have been used for the reference blocks.

[0203] – One or more MPMs can be determined in the same way in both the encoding and decoding devices. That is, the encoding and decoding devices can share the same list of MPMs, which includes one or more MPMs.

[0204] MPM List: The MPM list can be a list that includes one or more MPMs. The number of one or more MPMs in the MPM list can be predefined.

[0205] MPM Indicator: The MPM indicator can indicate one or more MPMs in the MPM list that will be used for intra-prediction against the target block. For example, the MPM indicator can be an index used for the MPM list.

[0206] Since the MPM list is determined in the same way in both the encoding and decoding devices, it is not necessary to send the MPM list itself from the encoding device to the decoding device.

[0207] – An MPM indicator can be signaled from the encoding device to the decoding device. Because the MPM indicator is signaled, the decoding device can determine which MPM in the MPM list will be used for intra-frame prediction of the target block.

[0208] MPM Usage Indicator: The MPM usage indicator indicates whether an MPM usage mode will be used for prediction of the target block. The MPM usage mode can be determined using a list of MPMs to identify the MPMs that will be used for intra-frame prediction of the target block.

[0209] –MPM uses indicators that can be sent from the encoding device to the decoding device via signals.

[0210] Signaling: "Signaling" can indicate that information is sent from the encoding device to the decoding device. Optionally, "signaling" can indicate that information is included by the encoding device in a bitstream or recording medium. Information sent by the encoding device using signals can be used by the decoding device.

[0211] – An encoding device generates encoded information by encoding the information to be transmitted as a signal. The encoded information can be sent from the encoding device to a decoding device. The decoding device obtains the information by decoding the transmitted encoded information. Here, the encoding can be entropy encoding, and the decoding can be entropy decoding.

[0212] Selective signaling: Information can be selectively transmitted using signals. Selective signaling for information can mean that an encoding device (based on specific conditions) selectively includes information in a bitstream or recording medium. Selective signaling for information can mean that a decoding device (based on specific conditions) selectively extracts information from a bitstream.

[0213] Omission of signaling: Signaling used for information can be omitted. Omission of signaling for information may mean that the encoding device (under certain conditions) does not include information in the bitstream or recording medium. Omission of signaling for information may mean that the decoding device (under certain conditions) does not extract information from the bitstream.

[0214] Statistical values: Variables, coding parameters, constants, etc., can have computable values. Statistical values ​​can be values ​​generated by performing calculations (operations) on the values ​​of a specified target. For example, a statistical value can indicate one or more of the following: the mean, weighted average, weighted sum, minimum, maximum, mode, median, and interpolation of the values ​​of a specific variable, a specific coding parameter, a specific constant, etc.

[0215] Figure 1 This is a block diagram illustrating the configuration of an embodiment of the encoding apparatus to which this disclosure is applied.

[0216] The encoding device 100 may be an encoder, a video encoding device, or an image encoding device. The video may include one or more images. The encoding device 100 may encode one or more images of the video sequentially.

[0217] Reference Figure 1 The encoding device 100 includes an inter-frame prediction unit 110, an intra-frame prediction unit 120, a switcher 115, a subtractor 125, a transform unit 130, a quantization unit 140, an entropy coding unit 150, an inverse quantization unit 160, an inverse transform unit 170, an adder 175, a filter unit 180, and a reference frame buffer 190.

[0218] The encoding device 100 can perform encoding on the target image using intra-frame mode and / or inter-frame mode. In other words, the prediction mode of the target block can be one of intra-frame mode and inter-frame mode.

[0219] In the following text, the terms "intra-frame mode", "intra-frame prediction mode", "in-frame mode" and "in-frame prediction mode" may be used to have the same meaning and may be used interchangeably.

[0220] In the following text, the terms "inter-frame mode", "inter-frame prediction mode", "inter-picture mode" and "inter-picture prediction mode" may be used to have the same meaning and may be used interchangeably.

[0221] In the following text, the term "image" may refer to only a portion of an image, or it may refer to a block. Furthermore, the processing of an "image" may refer to the sequential processing of multiple blocks.

[0222] Furthermore, the encoding device 100 can generate a bitstream including encoded information by encoding the target image, and can output and store the generated bitstream. The generated bitstream can be stored in a computer-readable storage medium and can be streamed via wired and / or wireless transmission media.

[0223] When intra-frame mode is used as prediction mode, switcher 115 can switch to intra-frame mode. When inter-frame mode is used as prediction mode, switcher 115 can switch to inter-frame mode.

[0224] The encoding device 100 can generate a prediction block for the target block. Furthermore, after the prediction block has been generated, the encoding device 100 can use the residual between the target block and the prediction block to encode the residual block for the target block.

[0225] When the prediction mode is intra-frame mode, the intra-frame prediction unit 120 can use pixels from previously encoded / decoded neighboring blocks adjacent to the target block as reference samples. The intra-frame prediction unit 120 can use the reference samples to perform spatial prediction on the target block and can generate prediction samples for the target block via spatial prediction. The prediction samples can represent samples in the prediction block.

[0226] The inter-frame prediction unit 110 may include a motion prediction unit and a motion compensation unit.

[0227] When the prediction mode is inter-frame mode, the motion prediction unit can search for the region in the reference image that best matches the target block during motion prediction, and can derive motion vectors for both the target block and the found region based on the found region. Here, the motion prediction unit can use the search range as the target region for the search.

[0228] A reference image may be stored in a reference frame buffer 190. More specifically, when the encoding and / or decoding of a reference image has been processed, the encoded and / or decoded reference image may be stored in the reference frame buffer 190.

[0229] Since it stores the decoded screen, the reference screen buffer 190 can be a decoded screen buffer (DPB).

[0230] The motion compensation unit can generate a predicted block for the target block by performing motion compensation using motion vectors. Here, the motion vector can be a two-dimensional (2D) vector used for inter-frame prediction. Furthermore, the motion vector can indicate the offset between the target image and the reference image.

[0231] When the motion vector has values ​​other than integers, the motion prediction unit and motion compensation unit can generate prediction blocks by applying interpolation filters to a portion of the reference image. To perform inter-frame prediction or motion compensation, it can be determined which of the following modes—skip mode, merge mode, advanced motion vector prediction (AMVP) mode, and current frame reference mode—corresponds to the method for predicting and compensating for the motion of the PU included in the CU based on the CU, and inter-frame prediction or motion compensation can be performed according to that mode.

[0232] Subtractor 125 generates a residual block, which is the difference between the target block and the prediction block. The residual block can also be referred to as the "residual signal".

[0233] The residual signal can be the difference between the original signal and the predicted signal. Optionally, the residual signal can be a signal generated by transforming or quantizing the difference between the original signal and the predicted signal, or a signal generated by transforming and quantizing the difference. The residual block can be the residual signal for a block unit.

[0234] The transformation unit 130 can generate transformation coefficients by transforming the residual block, and can output the generated transformation coefficients. Here, the transformation coefficients can be coefficient values ​​generated by transforming the residual block.

[0235] Transformation unit 130 may use one of a number of predefined transformation methods when performing a transformation.

[0236] The predefined transformation methods may include Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Karhunen-Loeve Transform (KLT), etc.

[0237] The transformation method for transforming the residual block can be determined based on at least one of the coding parameters for the target block and / or neighboring blocks. For example, the transformation method can be determined based on at least one of the inter-frame prediction mode for the PU, the intra-frame prediction mode for the PU, the size of the TU, and the shape of the TU. Optionally, transformation information indicating the transformation method can be transmitted from the encoding device 100 to the decoding device 200 by a signal.

[0238] When using the transform skip mode, the transform unit 130 can omit the operation of transforming the residual block.

[0239] By quantizing the transform coefficients, a quantized transform coefficient level or a quantized level can be generated. In the following examples, each of the quantized transform coefficient level and the quantized level may also be referred to as a "transform coefficient".

[0240] The quantization unit 140 can generate a quantized transform coefficient level (i.e., a quantized level or quantized coefficient) by quantizing the transform coefficients according to quantization parameters. The quantization unit 140 can output the generated quantized transform coefficient level. In this case, the quantization unit 140 can use a quantization matrix to quantize the transform coefficients.

[0241] Entropy coding unit 150 can generate a bitstream by performing probability distribution-based entropy coding based on values ​​calculated by quantization unit 140 and / or coding parameter values ​​calculated during the encoding process. Entropy coding unit 150 can output the generated bitstream.

[0242] The entropy coding unit 150 can perform entropy coding on information about the pixels of the image and information required to decode the image. For example, the information required to decode the image may include syntax elements, etc.

[0243] When applying entropy coding, fewer bits can be allocated to more frequently occurring symbols, and more bits can be allocated to less frequently occurring symbols. Because symbols are represented through this allocation, the size of the bit string used to encode the target symbol can be reduced. Therefore, entropy coding can improve the compression performance of video coding.

[0244] Furthermore, for entropy coding, the entropy coding unit 150 can use coding methods such as Exponential Golomb, Context Adaptive Variable Length Coding (CAVLC), or Context Adaptive Binary Arithmetic Coding (CABAC). For example, the entropy coding unit 150 can use a variable-length code / code (VLC) table to perform entropy coding. For example, the entropy coding unit 150 can derive a binarization method for the target symbol. Furthermore, the entropy coding unit 150 can derive a probabilistic model for the target symbol / bit. The entropy coding unit 150 can use the derived binarization method, probabilistic model, and context model to perform arithmetic coding.

[0245] The entropy coding unit 150 can transform the coefficients in 2D block form into 1D vector form through the transform coefficient scanning method, so as to encode the quantized transform coefficient levels.

[0246] Encoding parameters can be information required for encoding and / or decoding. Encoding parameters may include information encoded by encoding device 100 and transmitted from encoding device 100 to decoding device, and may also include information that can be derived during encoding or decoding. For example, information transmitted to decoding device may include syntax elements.

[0247] Encoding parameters can include not only information such as syntax elements (or flags or indexes) encoded by the encoding device and transmitted by the encoding device to the decoding device via signals, but also information derived during the encoding or decoding process. Furthermore, encoding parameters can include information required for encoding or decoding the image. For example, encoding parameters can include at least one value of the following, a combination of the following, or statistics of the following: unit / block size, unit / block shape / form, unit / block depth, unit / block partitioning information, unit / block partitioning structure, information indicating whether the unit / block is partitioned in a quadtree structure, information indicating whether the unit / block is partitioned in a binary tree structure, partitioning direction of the binary tree structure (horizontal or vertical), partitioning form of the binary tree structure (symmetric or asymmetric partitioning), information indicating whether the unit / block is partitioned in a ternary tree structure, partitioning direction of the ternary tree structure (horizontal or vertical), partitioning form of the ternary tree structure (symmetric or asymmetric partitioning), and partitioning direction of the ternary tree structure (horizontal or vertical). Information including: symmetric partitioning, whether the indicator unit / block is partitioned in a multi-type tree structure, the combination and direction of partitions in the multi-type tree structure (horizontal or vertical, etc.), partition form of the multi-type tree structure (symmetric or asymmetric partitioning, etc.), partition tree form of the multi-type tree (binary or ternary tree), prediction type (intra-frame prediction or inter-frame prediction), intra-frame prediction mode / direction, intra-frame luma prediction mode / direction, intra-frame chroma prediction mode / direction, intra-frame partition information, inter-frame partition information, coded block partition flag, prediction block partition flag, transform block partition flag, reference sample filtering method, reference sample filter taps, reference sample filter coefficients, and prediction block filtering method. Prediction block filter taps, prediction block filter coefficients, prediction block boundary filtering method, prediction block boundary filter taps, prediction block boundary filter coefficients, inter-frame prediction mode, motion information, motion vector, motion vector difference, reference frame index, inter-frame prediction direction, inter-frame prediction indicator, prediction list utilization flag, reference frame list, reference image, POC, motion vector prediction factor, motion vector prediction index, motion vector prediction candidate, motion vector candidate list, information indicating whether merging mode is used, merging index, merging candidate, merging candidate list, information indicating whether skip mode is used, interpolation filter type, interpolation filter taps, interpolation filter... The filter coefficients, magnitude of the motion vector, accuracy of the motion vector representation, transform type, transform magnitude, information indicating whether the first transform is used, information indicating whether the additional (second) transform is used, first transform selection information (or first transform index), second transform selection information (or second transform index), information indicating the presence or absence of the residual signal, code block pattern, code block flag, quantization parameters, residual quantization parameters, quantization matrix, information about the loop filter, information indicating whether the loop filter is applied, loop filter coefficients, loop filter taps, loop filter shape / form, information indicating whether the deblocking filter is applied.Deblocking filter coefficients, deblocking filter taps, deblocking filter strength, deblocking filter shape / form, information indicating whether adaptive sample offset is applied, adaptive sample offset value, adaptive sample offset category, adaptive sample offset type, information indicating whether adaptive loop filter is applied, adaptive loop filter coefficients, adaptive loop filter taps, adaptive loop filter shape / form, binarization / debinarization method, context model, context model determination method, context model update method, information indicating whether normal mode is executed, information indicating whether bypass mode is executed, valid coefficient flag, last valid coefficient flag, coefficient group encoding flag, last valid coefficient position, information indicating whether the coefficient value is greater than 1, information indicating whether the coefficient value is greater than 2, information indicating whether the coefficient value is greater than 3, remaining coefficient value information, positive / negative sign information, reconstructed luminance sample, reconstructed chrominance sample, context binary bits. The following parameters are included: bypass binary bits, residual luminance samples, residual chrominance samples, transform coefficients, luminance transform coefficients, chrominance transform coefficients, quantization level, luminance quantization level, chrominance quantization level, transform coefficient level, transform coefficient level scanning method, size of the motion vector search area on the decoding device side, shape / form of the motion vector search area on the decoding device side, number of motion vector searches on the decoding device side, CTU size, minimum block size, maximum block size, maximum block depth, minimum block depth, image display / output order, stripe identification information, stripe type, stripe partition information, parallel block group identification information, parallel block group type, parallel block group partition information, parallel block identification information, parallel block type, parallel block partition information, screen type, bit depth, input sample bit depth, reconstructed sample bit depth, residual sample bit depth, transform coefficient bit depth, quantization level bit depth, information about the luminance signal, information about the chrominance signal, color space of the target block, and color space of the residual block. Furthermore, information related to the above encoding parameters may also be included in the encoding parameters. Information used to calculate and / or derive the above encoding parameters may also be included in the encoding parameters. Information calculated or derived using the above encoding parameters can also be included in the encoding parameters.

[0248] The first transformation selection information can indicate the first transformation to be applied to the target block.

[0249] The second transformation selection information can indicate the second transformation to be applied to the target block.

[0250] The residual signal can represent the difference between the original signal and the predicted signal. Optionally, the residual signal can be a signal generated by transforming the difference between the original signal and the predicted signal. Optionally, the residual signal can be a signal generated by transforming and quantizing the difference between the original signal and the predicted signal. The residual block can be the residual signal for a block.

[0251] Here, sending information via a signal can indicate that the encoding device 100 includes entropy-encoded information generated by performing entropy encoding on flags or indices in the bitstream, and can also indicate that the decoding device 200 obtains information by performing entropy decoding on the entropy-encoded information extracted from the bitstream. Here, the information may include flags, indices, etc.

[0252] A signal can refer to information that will be transmitted using a signal. In the following text, information used for images and blocks may be referred to as a "signal". Furthermore, in the following text, the terms "information" and "signal" may be used to have the same meaning and may be used interchangeably. For example, a specific signal may be a signal representing a specific block. A raw signal may be a signal representing a target block. A prediction signal may be a signal representing a predicted block. A residual signal may be a signal representing a residual block.

[0253] The bitstream may include information based on a specific syntax. Encoding device 100 may generate a bitstream that includes information according to the specific syntax. Decoding device 200 may obtain information from the bitstream according to the specific syntax.

[0254] Since the encoding device 100 performs encoding via inter-frame prediction, the encoded target image can be used as a reference image for other images to be processed subsequently. Therefore, the encoding device 100 can reconstruct or decode the encoded target image and store the reconstructed or decoded image as a reference image in the reference frame buffer 190. For decoding, inverse quantization and inverse transform of the encoded target image can be performed.

[0255] The quantization levels can be dequantized by the dequantization unit 160 and inversely transformed by the inverse transform unit 170. The dequantization unit 160 can generate dequantized coefficients by performing an inverse transform on the quantization levels. The inverse transform unit 170 can generate coefficients that have undergone both dequantization and inverse transform by performing an inverse transform on the dequantized coefficients.

[0256] The coefficients that have undergone dequantization and inverse transform can be added to the prediction block by adder 175. Adding the coefficients that have undergone dequantization and inverse transform to the prediction block generates a reconstructed block. Here, the coefficients that have undergone dequantization and / or inverse transform can represent one or more coefficients that have undergone dequantization and inverse transform, and can also represent the reconstructed residual block. Here, the reconstructed block can represent either the recovered block or the decoded block.

[0257] The reconstructed blocks can be filtered by filter unit 180. Filter unit 180 can apply one or more filters, including deblocking filter, sample adaptive offset (SAO) filter, adaptive loop filter (ALF), and nonlocal filter (NLF), to the reconstructed samples, reconstructed blocks, or reconstructed images. Filter unit 180 may also be referred to as a "loop filter".

[0258] Deblocking filters eliminate block distortion that occurs at the boundaries between blocks in a reconstructed image. To determine whether to apply a deblocking filter, the number of columns or rows of pixels included in the block and on which the determination of whether to apply the deblocking filter to the target block is based can be determined.

[0259] When a deblocking filter is applied to a target block, the applied filter can vary depending on the required deblocking strength. In other words, among different filters, one that takes into account the strength of the deblocking filter can be applied to the target block. When a deblocking filter is applied to a target block, one or more filters, such as long-tap filters, strong filters, weak filters, and Gaussian filters, can be applied to the target block according to the required deblocking strength.

[0260] Furthermore, when performing vertical and horizontal filtering on the target block, horizontal and vertical filtering can be performed in parallel.

[0261] SAO can add an appropriate offset to the pixel value to compensate for coding errors. SAO can perform correction on the image to which deblocking is applied based on pixels, where the correction uses an offset of the difference between the original image and the image to which deblocking is applied. To perform offset correction on an image, methods can be used to divide the pixels included in the image into a specific number of regions, determine the region to be offset within the divided regions, and apply the offset to the determined region, or methods can be used to apply the offset taking into account the edge information of each pixel.

[0262] The ALF can perform filtering based on values ​​obtained by comparing the reconstructed image with the original image. After the pixels included in the image have been divided into a predetermined number of groups, the filter to be applied to each group can be determined, and filtering can be performed differently for each group. Information related to whether an adaptive loop filter is applied can be sent to each CU via a signal. This information can be sent via a signal for the luminance signal. The shape and filter coefficients of the ALF to be applied to each block can be different for each block. Alternatively, an ALF with a fixed form can be applied to the block regardless of its characteristics.

[0263] Nonlocal filters can perform filtering based on reconstructed blocks similar to the target block. Regions similar to the target block can be selected from the reconstructed image, and the statistical properties of the selected similar regions can be used to perform filtering on the target block. Information regarding whether a nonlocal filter is applied can be sent to the coding unit (CU) via a signal. Furthermore, the shape and filter coefficients of the nonlocal filter applied to the block can vary depending on the block.

[0264] The reconstructed blocks or reconstructed image filtered by filter unit 180 can be stored as a reference frame in reference frame buffer 190. The reconstructed blocks filtered by filter unit 180 can be part of the reference frame. In other words, the reference frame can be a reconstructed frame composed of reconstructed blocks filtered by filter unit 180. The stored reference frame can then be used for inter-frame prediction or motion compensation.

[0265] Figure 2 This is a block diagram illustrating the configuration of an embodiment of the decoding apparatus to which this disclosure is applied.

[0266] The decoding device 200 can be a decoder, a video decoding device, or an image decoding device.

[0267] Reference Figure 2 The decoding device 200 may include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an intra-frame prediction unit 240, an inter-frame prediction unit 250, a switcher 245, an adder 255, a filter unit 260, and a reference frame buffer 270.

[0268] The decoding device 200 can receive bit streams output from the encoding device 100. The decoding device 200 can receive bit streams stored in a computer-readable storage medium and can also receive bit streams transmitted via wired / wireless transmission media.

[0269] The decoding device 200 can perform decoding on the bitstream in intra-frame mode and / or inter-frame mode. Furthermore, the decoding device 200 can generate a reconstructed image or a decoded image via decoding, and can output the reconstructed image or the decoded image.

[0270] For example, switcher 245 can be used to switch between intra-frame mode and inter-frame mode based on the prediction mode used for decoding. When the prediction mode used for decoding is intra-frame mode, switcher 245 can be operated to switch to intra-frame mode. When the prediction mode used for decoding is inter-frame mode, switcher 245 can be operated to switch to inter-frame mode.

[0271] The decoding device 200 can obtain the reconstructed residual block by decoding the input bitstream and can generate the prediction block. When the reconstructed residual block and the prediction block are obtained, the decoding device 200 can generate the reconstructed block, which is the target to be decoded, by adding the reconstructed residual block and the prediction block.

[0272] The entropy decoding unit 210 can generate symbols by performing entropy decoding on the bitstream based on the probability distribution of the bitstream. The generated symbols may include symbols in the form of quantized transform coefficient levels (i.e., quantized levels or quantized coefficients). Here, the entropy decoding method may be similar to the entropy coding method described above. That is, the entropy decoding method may be the inverse process of the entropy coding method described above.

[0273] The entropy decoding unit 210 can transform coefficients in one-dimensional (1D) vector form into 2D block shapes by a transform coefficient scanning method in order to decode the quantized transform coefficient levels.

[0274] For example, the coefficients of a block can be transformed into a 2D block shape by scanning the block coefficients using a top-right diagonal scan. Optionally, which of the top-right diagonal scan, vertical scan, and horizontal scan will be used can be determined based on the size of the corresponding block and / or the intra-frame prediction mode.

[0275] The quantized coefficients can be dequantized by the dequantization unit 220. The dequantization unit 220 generates dequantized coefficients by performing dequantization on the quantized coefficients. Furthermore, the dequantized coefficients can be inversely transformed by the inverse transform unit 230. The inverse transform unit 230 generates a reconstructed residual block by performing an inverse transform on the dequantized coefficients. As a result of performing dequantization and inverse transform on the quantized coefficients, a reconstructed residual block can be generated. Here, when generating the reconstructed residual block, the dequantization unit 220 can apply the quantization matrix to the quantized coefficients.

[0276] When using intra-frame mode, intra-frame prediction unit 240 can generate a prediction block by performing spatial prediction on the target block, wherein the spatial prediction uses the pixel values ​​of previously decoded neighboring blocks adjacent to the target block.

[0277] The inter-frame prediction unit 250 may include a motion compensation unit. Optionally, the inter-frame prediction unit 250 may be designated as a "motion compensation unit".

[0278] When using inter-frame mode, the motion compensation unit can generate a prediction block by performing motion compensation on the target block, wherein the motion compensation uses motion vectors and a reference image stored in the reference frame buffer 270.

[0279] The motion compensation unit can apply an interpolation filter to a portion of the reference image when the motion vector has values ​​other than integers, and can use the reference image with the interpolation filter applied to generate prediction blocks. To perform motion compensation, the motion compensation unit can determine, based on the CU, which of the following modes—skip mode, merge mode, advanced motion vector prediction (AMVP) mode, and current frame reference mode—corresponds to the motion compensation method used by the PU included in the CU, and can perform motion compensation according to the determined mode.

[0280] The reconstructed residual block and the predicted block can be added to each other by adder 255. Adder 255 generates a reconstructed block by adding the reconstructed residual block and the predicted block.

[0281] The reconstructed blocks can be filtered by filter unit 260. Filter unit 260 can apply at least one of a deblocking filter, a SAO filter, an ALF filter, and an NLF filter to the reconstructed blocks or the reconstructed image. The reconstructed image can be a picture that includes the reconstructed blocks.

[0282] The filter unit can output a reconstructed image.

[0283] The reconstructed image and / or reconstructed blocks filtered by filter unit 260 can be stored as a reference image in reference image buffer 270. The reconstructed blocks filtered by filter unit 260 can be part of the reference image. In other words, the reference image can be an image composed of reconstructed blocks filtered by filter unit 260. The stored reference image can then be used for inter-frame prediction or motion compensation.

[0284] Figure 3 It is a schematic diagram illustrating the partitioning structure of an image when it is encoded and decoded.

[0285] Figure 3 An example can be illustrated by showing a single cell divided into multiple sub-cells.

[0286] To effectively partition an image, coding units (CUs) can be used in encoding and decoding. The term "unit" can be used to collectively specify 1) a block comprising image samples and 2) a syntax element. For example, "partition of a unit" can mean "partition of a block corresponding to a unit".

[0287] A CU can be used as the basic unit for image encoding / decoding. A CU can be used as the unit to which one of the intra-frame and inter-frame modes is applied during image encoding / decoding. In other words, during image encoding / decoding, it can be determined which of the intra-frame and inter-frame modes will be applied to each CU.

[0288] Furthermore, the CU can be the basic unit for predicting, transforming, quantizing, inverse transforming, dequantizing, and encoding / decoding transform coefficients.

[0289] Reference Figure 3 Image 300 can be sequentially partitioned into units corresponding to the largest coding unit (LCU), and the partitioning structure can be determined for each LCU. Here, LCU can be used to have the same meaning as coding tree unit (CTU).

[0290] Partitioning a cell can represent partitioning the block corresponding to the cell. Block partitioning information may include depth information about the depth of the cell. The depth information may indicate the number of times the cell is partitioned and / or the degree to which the cell is partitioned. A single cell may be hierarchically partitioned into multiple sub-cells, and the single cell may have depth information based on a tree structure.

[0291] Each partitioned sub-unit can have depth information. The depth information can be information indicating the size of the CU. Depth information can be stored for each CU.

[0292] Each CU can have depth information. When a CU is partitioned, the depth of the CU generated from the partition can be increased by 1 from the depth of the partitioned CU.

[0293] The partitioning structure represents the distribution of coding units (CUs) in the LCU 310 used for effective encoding of an image. This distribution can be determined by whether a single CU will be partitioned into multiple CUs. The number of CUs generated by partitioning can be a positive integer of 2 or greater, including 2, 3, 4, 8, 16, etc.

[0294] Depending on the number of CUs generated through partitioning, the horizontal and vertical dimensions of each CU generated through partitioning can be smaller than the horizontal and vertical dimensions of the CU before partitioning. For example, the horizontal and vertical dimensions of each CU generated through partitioning can be half the horizontal and vertical dimensions of the CU before partitioning.

[0295] Each partitioned CU can be recursively partitioned into four CUs in the same manner. Through recursive partitioning, at least one of the horizontal and vertical dimensions of each partitioned CU can be reduced compared to at least one of the horizontal and vertical dimensions of the CU before partitioning.

[0296] The partitioning of a CU can be performed recursively until a predefined depth or predefined size is reached.

[0297] For example, the depth of a CU can range from 0 to 3. The size of a CU can range from 64×64 to 8×8, depending on its depth.

[0298] For example, the depth of LCU 310 can be 0, and the depth of the minimum coding unit (SCU) can be a predefined maximum depth. Here, as mentioned above, the LCU can be a CU with the maximum coding unit size, and the SCU can be a CU with the minimum coding unit size.

[0299] Partitioning can begin at LCU 310, and the depth of the CU can be increased by 1 whenever the horizontal and / or vertical dimensions of the CU are reduced by partitioning.

[0300] For example, for each depth, an unpartitioned CU can have a size of 2N×2N. Furthermore, when CUs are partitioned, a CU of size 2N×2N can be partitioned into four CUs, each with a size of N×N. The value of N is halved each time the depth increases by 1.

[0301] Reference Figure 3 An LCU with a depth of 0 can have 64×64 pixels or 64×64 blocks. 0 can be the minimum depth. An SCU with a depth of 3 can have 8×8 pixels or 8×8 blocks. 3 can be the maximum depth. Here, a CU with 64×64 blocks as an LCU can be represented by depth 0. A CU with 32×32 blocks can be represented by depth 1. A CU with 16×16 blocks can be represented by depth 2. A CU with 8×8 blocks as an SCU can be represented by depth 3.

[0302] Information about whether a corresponding CU is partitioned can be represented by the CU's partition information. Partition information can be 1 bit. All CUs except the SCU can include partition information. For example, the partition information value for an unpartitioned CU can be a first value. The partition information value for a partitioned CU can be a second value. When the partition information indicates whether a CU is partitioned, the first value can be "0" and the second value can be "1".

[0303] For example, when a single CU is partitioned into four CUs, the horizontal and vertical dimensions of each of the four CUs created through partitioning can be half the horizontal and vertical dimensions of the CU before partitioning. When a 32×32 CU is partitioned into four CUs, the size of each of the four partitioned CUs can be 16×16. When a single CU is partitioned into four CUs, the CU can be considered to have been partitioned using a quadtree structure. In other words, quadtree partitioning can be considered to have been applied to the CU.

[0304] For example, when a single CU is partitioned into two CUs, the horizontal or vertical dimension of each of the two resulting CUs can be half the horizontal or vertical dimension of the CU before partitioning. When a 32×32 CU is vertically partitioned into two CUs, the size of each of the two resulting CUs can be 16×32. When a 32×32 CU is horizontally partitioned into two CUs, the size of each of the two resulting CUs can be 32×16. When a single CU is partitioned into two CUs, the CU can be considered to have been partitioned using a binary tree structure. In other words, binary tree partitioning can be considered to have been applied to the CU.

[0305] For example, when a single CU is partitioned (or divided) into three CUs, the original CU before partitioning is partitioned such that its horizontal or vertical dimensions are divided in a 1:2:1 ratio, thus enabling the generation of three sub-CUs. For example, when a 16×32 CU is horizontally partitioned into three sub-CUs, the three sub-CUs generated by the partitioning can have dimensions of 16×8, 16×16, and 16×8 respectively from top to bottom. For example, when a 32×32 CU is vertically partitioned into three sub-CUs, the three sub-CUs generated by the partitioning can have dimensions of 8×32, 16×32, and 8×32 respectively from left to right. When a single CU is partitioned into three CUs, the CU can be considered to be partitioned in the form of a ternary tree. In other words, ternary tree partitioning can be considered to have been applied to the CU.

[0306] Quadtree partitioning and binary tree partitioning are both used... Figure 3 LCU 310.

[0307] In the encoding device 100, a 64×64 coding tree unit (CTU) can be partitioned into multiple smaller CUs using a recursive quadtree structure. A single CU can be partitioned into four CUs of the same size. Each CU can be recursively partitioned and can have a quadtree structure.

[0308] By using recursive partitioning of the CU, the optimal partitioning method that causes the minimum rate distortion cost can be selected.

[0309] Figure 3 The Coding Tree Unit (CTU) 320 in the example is a CTU in which quadtree partitioning, binary tree partitioning and ternary tree partitioning are all applied.

[0310] As described above, to partition a CTU, at least one of quadtree partitioning, binary tree partitioning, and ternary tree partitioning can be applied to the CTU. Partitioning can be applied based on specific priorities.

[0311] For example, quadtree partitioning can be preferentially applied to CTUs. CUs that cannot be further partitioned in quadtree form can correspond to the leaf nodes of a quadtree. CUs corresponding to the leaf nodes of a quadtree can be the root nodes of a binary tree and / or a ternary tree. That is, CUs corresponding to the leaf nodes of a quadtree can be partitioned in binary or ternary tree form, or may not be further partitioned. In this case, to prevent each CU generated by applying binary or ternary tree partitioning to the CUs corresponding to the leaf nodes of a quadtree from being quadtree partitioned again, the operations of block partitioning and / or signaling block partitioning information are effectively performed.

[0312] Four-partition information can be used to signal the partitions of a CU corresponding to each node of a quadtree. A four-partition message with a first value (e.g., "1") indicates that the corresponding CU is partitioned in quadtree form. A four-partition message with a second value (e.g., "0") indicates that the corresponding CU is not partitioned in quadtree form. The four-partition message can be a flag with a specific length (e.g., 1 bit).

[0313] There may be no priority between binary tree partitioning and ternary tree partitioning. That is, the CU corresponding to the leaf node of a quadtree can be partitioned in either binary or ternary tree form. Furthermore, CUs generated by binary or ternary tree partitioning can be further partitioned in either binary or ternary tree form, or they may not be further partitioned.

[0314] A partition executed when there is no priority between a binary tree partition and a ternary tree partition can be called a "multi-type tree partition". That is, the CU corresponding to a leaf node of a quadtree can be the root node of a multi-type tree. The partitioning of the CU corresponding to each node of the multi-type tree can be signaled using at least one of the following: information indicating whether the CU is partitioned according to the multi-type tree, partitioning direction information, and partitioning tree information. For the partitioning of the CU corresponding to each node of the multi-type tree, the information indicating whether the multi-type tree partitioning is executed, the partitioning direction information, and the partitioning tree information can be signaled sequentially.

[0315] For example, information indicating whether a CU is partitioned in a multi-type tree and having a first value (e.g., "1") can indicate that the corresponding CU is partitioned in a multi-type tree format. Information indicating whether a CU is partitioned in a multi-type tree and having a second value (e.g., "0") can indicate that the corresponding CU is not partitioned in a multi-type tree format.

[0316] When the CU corresponding to each node of the multi-type tree is partitioned in the form of a multi-type tree, the corresponding CU may further include partitioning direction information.

[0317] Partition direction information indicates the partitioning direction of a multi-type tree partition. Partition direction information with a first value (e.g., "1") indicates that the corresponding CU is partitioned in the vertical direction. Partition direction information with a second value (e.g., "0") indicates that the corresponding CU is partitioned in the horizontal direction.

[0318] When a CU corresponding to each node of a multi-type tree is partitioned in the form of a multi-type tree, the corresponding CU may further include partition tree information. The partition tree information can indicate the tree used for multi-type tree partitioning.

[0319] For example, partition tree information with a first value (e.g., "1") can indicate that the corresponding CU is partitioned in a binary tree format. Partition tree information with a second value (e.g., "0") can indicate that the corresponding CU is partitioned in a ternary tree format.

[0320] Here, each of the above-mentioned information indicating whether partitioning by multi-type tree is performed, partition tree information, and partition direction information can be a flag with a specific length (e.g., 1 bit).

[0321] Entropy encoding and / or entropy decoding can be performed on at least one of the above four partition information, information indicating whether partitioning according to multiple tree types has been executed, partition direction information, and partition tree information. To perform entropy encoding / entropy decoding of this information, information from neighboring CUs adjacent to the target CU can be used.

[0322] For example, it can be assumed that the partitioning patterns (i.e., partitioned / non-partitioned, partitioned tree, and / or partitioned direction) of the left and / or upper CUs are highly similar to the partitioning patterns of the target CU. Therefore, based on the information of neighboring CUs, contextual information for entropy encoding and / or entropy decoding of the information for the target CU can be derived. Here, the information of neighboring CUs may include at least one of the following: 1) the four-partition information of the neighboring CUs, 2) information indicating whether the neighboring CUs are partitioned according to multiple types of trees, 3) the partitioned direction information of the neighboring CUs, and 4) the partitioned tree information of the neighboring CUs.

[0323] In another embodiment of binary tree partitioning and ternary tree partitioning, binary tree partitioning can be performed first. That is, binary tree partitioning can be applied first, and then the CUs corresponding to the leaf nodes of the binary tree can be set as the root nodes of the ternary tree. In this case, quadtree partitioning or binary tree partitioning may not be performed on the CUs corresponding to the nodes of the ternary tree.

[0324] A CU that is not further partitioned by quadtree partitioning, binary tree partitioning, and / or ternary tree partitioning can be a unit for encoding, prediction, and / or transformation. That is, a CU may not be further partitioned for use in prediction and / or transformation. Therefore, the partitioning structure used to partition CUs into prediction units (PUs) and / or transformation units (TUs), its partitioning information, etc., may not exist in the bitstream.

[0325] However, when the size of a CU (Computer Unit) used as a partitioning unit is larger than the size of the largest transform block, the CU can be recursively partitioned until the size of the CU becomes smaller than or equal to the size of the largest transform block. For example, when the size of the CU is 64×64 and the size of the largest transform block is 32×32, the CU can be partitioned into four 32×32 blocks to perform the transform. Similarly, when the size of the CU is 32×64 and the size of the largest transform block is 32×32, the CU can be partitioned into two 32×32 blocks.

[0326] In this case, it is not necessary to separately send a signal indicating whether the CU has been partitioned for transformation. Without signal transmission, partitioning of the CU can be determined by comparing its horizontal (and / or vertical) dimensions with the horizontal (and / or vertical) dimensions of the largest transform block. For example, when the horizontal dimension of the CU is greater than the horizontal dimension of the largest transform block, the CU can be vertically bisected. Furthermore, when the vertical dimension of the CU is greater than the vertical dimension of the largest transform block, the CU can be horizontally bisected.

[0327] Information regarding the maximum and / or minimum size of the CU and the maximum and / or minimum size of the transform block can be signaled or determined at a level higher than the CU level. For example, higher levels could be sequence level, frame level, parallel block level, parallel block group level, or stripe level. For example, the minimum size of the CU could be set to 4×4. For example, the maximum size of the transform block could be set to 64×64. For example, the maximum size of the transform block could be set to 4×4.

[0328] Information regarding the minimum size of the CU corresponding to the leaf node of the quadtree (i.e., the minimum size of the quadtree) and / or the maximum depth of the path from the root node to the leaf node of the multi-type tree (i.e., the maximum depth of the multi-type tree) can be signaled or determined at a level higher than the CU level. For example, higher levels could be sequence level, frame level, stripe level, parallel block group level, or parallel block level. Information regarding the minimum size of the quadtree and / or the maximum depth of the multi-type tree can be signaled or determined individually at each of the intra-strip and inter-strip levels.

[0329] Information regarding the difference between the size of the CTU and the maximum size of the transform block can be signaled or determined at a level higher than the CU level. For example, higher levels could be sequence level, frame level, stripe level, parallel block group level, or parallel block level. Information regarding the maximum size of the CU corresponding to each node of the binary tree (i.e., the maximum size of the binary tree) can be determined based on the size of the CTU and the aforementioned difference. The maximum size of the CU corresponding to each node of the ternary tree (i.e., the maximum size of the ternary tree) can have different values ​​depending on the stripe type. For example, the maximum size of the ternary tree in an intra-strip level could be 32×32. For example, the maximum size of the ternary tree in an inter-strip level could be 128×128. For example, the minimum size of the CU corresponding to each node of the binary tree (i.e., the minimum size of the binary tree) and / or the minimum size of the CU corresponding to each node of the ternary tree (i.e., the minimum size of the ternary tree) can be set to the minimum size of the CU.

[0330] In another example, the maximum size of a binary tree and / or the maximum size of a ternary tree can be signaled or determined at the stripe level. Furthermore, the minimum size of a binary tree and / or the minimum size of a ternary tree can be signaled or determined at the stripe level.

[0331] Based on the various block sizes and depths described above, the four partition information, information indicating whether partitioning by multiple tree types has been performed, partition tree information, and / or partition direction information may or may not exist in the bitstream.

[0332] For example, when the size of the CU is not greater than the minimum size of the quadtree, the CU may not include the four-partition information, and the four-partition information of the CU can be inferred as the second value.

[0333] For example, when the size (horizontal and vertical dimensions) of the CU corresponding to each node of the multi-type tree is greater than the maximum size (horizontal and vertical dimensions) of the binary tree and / or the maximum size (horizontal and vertical dimensions) of the ternary tree, the CU may not be partitioned in binary and / or ternary tree form. In this determination method, information indicating whether partitioning by multi-type tree is performed may not be sent by signal, but can be inferred as a second value.

[0334] Optionally, when the size (horizontal and vertical dimensions) of the CU corresponding to each node of the multi-type tree is equal to the minimum size (horizontal and vertical dimensions) of the binary tree, or when the size (horizontal and vertical dimensions) of the CU is equal to twice the minimum size (horizontal and vertical dimensions) of the ternary tree, the CU may not be partitioned in binary and / or ternary tree form. With this determination method, information indicating whether partitioning by multi-type tree is performed can be sent without signaling, but can be inferred as a second value. This is because when the CU is partitioned in binary and / or ternary tree form, it generates CUs smaller than the minimum size of the binary tree and / or the minimum size of the ternary tree.

[0335] Optionally, binary or ternary partitioning can be limited based on the size of the virtual pipeline data unit (i.e., the size of the pipeline buffer). For example, binary or ternary partitioning can be limited when a CU is partitioned into sub-CUs that do not fit the size of the pipeline buffer. The size of the pipeline buffer can be equal to the maximum size of the transform block (e.g., 64×64).

[0336] For example, when the size of the pipeline buffer is 64×64, the following partitions can be restricted.

[0337] - A ternary tree partition for an N×M CU (where N and / or M are 128).

[0338] - Horizontal binary tree partitioning for a 128×N CU (where N<=64)

[0339] - Vertical binary tree partitioning for N×128 CUs (where N<=64)

[0340] Optionally, when the depth of the CU corresponding to each node of the multi-type tree is equal to the maximum depth of the multi-type tree, the CU may not be partitioned in binary and / or ternary tree form. With this determination method, information indicating whether partitioning by the multi-type tree is performed can be sent without signaling, but can be inferred as a second value.

[0341] Optionally, information indicating whether partitioning by the multi-type tree has been performed may be signaled only if at least one of vertical binary tree partitioning, horizontal binary tree partitioning, vertical ternary tree partitioning, and horizontal ternary tree partitioning is possible for each CU corresponding to each node of the multi-type tree. Otherwise, the CU may not be partitioned in binary and / or ternary tree form. With this determination method, the information indicating whether partitioning by the multi-type tree has been performed may not be signaled, but rather inferred as a second value.

[0342] Optionally, for each CU corresponding to a node in a multi-type tree, partitioning direction information may be signaled only when both vertical binary tree partitioning and horizontal binary tree partitioning are feasible, or only when both vertical ternary tree partitioning and horizontal ternary tree partitioning are feasible. Otherwise, partitioning direction information may not be signaled, but may be inferred as a value indicating the direction in which the CU can be partitioned.

[0343] Optionally, for each CU corresponding to a node of a multi-type tree, partition tree information may be signaled only when both vertical binary tree partitioning and vertical ternary tree partitioning are feasible, or only when both horizontal binary tree partitioning and horizontal ternary tree partitioning are feasible. Otherwise, partition tree information may not be signaled, but may be inferred as the value of the tree indicating the partitions that can be applied to the CU.

[0344] Figure 4 This is a diagram showing the form of prediction units that a coding unit can include.

[0345] In the CUs partitioned from the LCU, the CUs that are no longer partitioned can be divided into one or more prediction units (PUs).

[0346] A PU (Program Unit) can be the basic unit used for prediction. A PU can be encoded and decoded in any of the following modes: skip mode, inter-frame mode, and intra-frame mode. A PU can be partitioned into various shapes according to each mode. For example, refer to the above... Figure 1 The target block described and the above references Figure 2 The target blocks described can all be PUs.

[0347] A CU may not be classified as a PU. When a CU is not classified as a PU, the dimensions of the CU and the PU can be equal.

[0348] In skip mode, partitioning may not be present in the CU. Skip mode also supports a 2N×2N mode 410 without partitioning, where the PU and CU have the same size.

[0349] In inter-frame mode, eight types of partition shapes can exist in the CU. For example, in inter-frame mode, 2N×2N mode 410, 2N×N mode 415, N×2N mode 420, N×N mode 425, 2N×nU mode 430, 2N×nD mode 435, nL×2N mode 440 and nR×2N mode 445 are supported.

[0350] In intra-frame mode, 2N×2N mode 410 and N×N mode 425 are supported.

[0351] In 2N×2N mode 410, a PU with a size of 2N×2N can be encoded. A PU with a size of 2N×2N can represent a PU with the same size as the CU. For example, a PU with a size of 2N×2N can have a size of 64×64, 32×32, 16×16, or 8×8.

[0352] In N×N mode 425, PUs with an N×N size can be encoded.

[0353] For example, in intra-frame prediction, when the PU size is 8×8, the PUs from four partitions can be encoded. The size of the PU from each partition can be 4×4.

[0354] When encoding a PU in intra-frame mode, the PU can be encoded using any of several intra-frame prediction modes. For example, HEVC technology provides 35 intra-frame prediction modes, and the PU can be encoded in any of these 35 intra-frame prediction modes.

[0355] The rate-distortion cost can be used to determine which of the 2N×2N modes 410 and N×N modes 425 will be used to encode the PU.

[0356] The encoding device 100 can perform encoding operations on a PU of size 2N×2N. Here, the encoding operation can be an operation of encoding the PU in each of a plurality of intra-prediction modes that can be used by the encoding device 100. Through the encoding operation, an optimal intra-prediction mode for the PU of size 2N×2N can be derived. This optimal intra-prediction mode can be the intra-prediction mode that incurs the minimum rate-distortion cost when encoding a PU of size 2N×2N, among the plurality of intra-prediction modes that can be used by the encoding device 100.

[0357] Furthermore, the encoding device 100 can sequentially perform encoding operations on each PU obtained by performing N×N partitioning. Here, the encoding operation can be an operation of encoding the PU in each of a plurality of intra-prediction modes that can be used by the encoding device 100. Through the encoding operation, the optimal intra-prediction mode for a PU of size N×N can be derived. This optimal intra-prediction mode can be the intra-prediction mode that results in the minimum rate-distortion cost when encoding a PU of size N×N among the plurality of intra-prediction modes that can be used by the encoding device 100.

[0358] The encoding device 100 can determine which of the PUs, one of size 2N×2N and one of size N×N, will be encoded based on a comparison between the rate-distortion cost of the PU with size 2N×2N and the rate-distortion cost of the PU with size N×N.

[0359] A single CU can be partitioned into one or more PUs, and a PU can be partitioned into multiple PUs.

[0360] For example, when a single PU is partitioned into four PUs, the horizontal and vertical dimensions of each of the four PUs created through partitioning can be half the horizontal and vertical dimensions of the original PU. When a 32×32 PU is partitioned into four PUs, the size of each of the four partitioned PUs can be 16×16. When a single PU is partitioned into four PUs, the PU can be considered to have been partitioned in a quadtree structure.

[0361] For example, when a single PU is partitioned into two PUs, the horizontal or vertical dimension of each of the two resulting PUs can be half the horizontal or vertical dimension of the original PU. When a 32×32 PU is vertically partitioned into two PUs, the size of each of the two partitioned PUs can be 16×32. When a 32×32 PU is horizontally partitioned into two PUs, the size of each of the two partitioned PUs can be 32×16. When a single PU is partitioned into two PUs, the PU can be considered to have been partitioned in a binary tree structure.

[0362] Figure 5 This is a diagram showing the form of a transformation unit that can be included in an encoding unit.

[0363] A transform unit (TU) can be a basic unit in a CU used for processes such as transform, quantization, inverse transform, dequantization, entropy coding, and entropy decoding.

[0364] The TU can be square or rectangular. The shape of the TU can be determined based on the size and / or shape of the CU.

[0365] Within a CU partitioned from an LCU, CUs that are no longer designated as CUs can be partitioned into one or more TUs. Here, the partitioning structure of a TU can be a quadtree structure. For example, ... Figure 5 As shown, a single CU 510 can be partitioned once or more according to a quadtree structure. Through this partitioning, a single CU 510 can be composed of TUs of various sizes.

[0366] The CU can be considered to be recursively partitioned when a single CU is partitioned two or more times. Through partitioning, a single CU can be composed of transformation units (TUs) of various sizes.

[0367] Optionally, a single CU can be divided into one or more TUs based on the number of vertical and / or horizontal lines dividing the CU.

[0368] The CU can be divided into symmetrical TUs or asymmetrical TUs. To divide into asymmetrical TUs, information about the size and / or shape of each TU can be transmitted from the encoding device 100 to the decoding device 200 via signals. Optionally, the size and / or shape of each TU can be derived from the information about the size and / or shape of the CU.

[0369] A CU may not be classified as a TU. When a CU is not classified as a TU, the dimensions of the CU and the TU may be equal.

[0370] A single CU can be partitioned into one or more TUs, and a TU can be partitioned into multiple TUs.

[0371] For example, when a single TU is partitioned into four TUs, the horizontal and vertical dimensions of each of the four TUs generated by the partitioning can be half the horizontal and vertical dimensions of the original TU. When a TU of size 32×32 is partitioned into four TUs, the size of each of the four partitioned TUs can be 16×16. When a single TU is partitioned into four TUs, the TU can be considered to have been partitioned in a quadtree structure.

[0372] For example, when a single TU is partitioned into two TUs, the horizontal or vertical dimension of each of the two resulting TUs can be half the horizontal or vertical dimension of the original TU. When a TU of size 32×32 is vertically partitioned into two TUs, the size of each of the two partitioned TUs can be 16×32. When a TU of size 32×32 is horizontally partitioned into two TUs, the size of each of the two partitioned TUs can be 32×16. When a single TU is partitioned into two TUs, the TU can be considered to have been partitioned in a binary tree structure.

[0373] Can be with Figure 5 The different methods shown illustrate how CUs are divided.

[0374] For example, a single CU can be divided into three CUs. The horizontal or vertical dimensions of the three CUs generated by the division can be 1 / 4, 1 / 2, and 1 / 4 of the horizontal or vertical dimensions of the original CU before the division, respectively.

[0375] For example, when a 32×32 CU is vertically divided into three CUs, the resulting three CUs can have sizes of 8×32, 16×32, and 8×32, respectively. In this way, when a single CU is divided into three CUs, the CU can be considered to be divided in the form of a ternary tree.

[0376] One of the exemplary partitioning forms (i.e., quadtree partitioning, binary tree partitioning, and ternary tree partitioning) can be applied to the partitioning of the CU, and multiple partitioning schemes can be combined together for the partitioning of the CU. Here, the combination of multiple partitioning schemes is referred to as "composite tree partitioning".

[0377] Figure 6 This shows the block division based on the example.

[0378] In video encoding and / or decoding processes, such as Figure 6 As shown, the target block can be divided. For example, the target block can be a CU.

[0379] For the partitioning of the target block, an indicator indicating the partitioning information can be sent from the encoding device 100 to the decoding device 200 by a signal. The partitioning information can be information indicating how the target block is partitioned.

[0380] The partitioning information can be one or more of the following: a partitioning flag (hereinafter referred to as "split_flag"), a quad-binary flag (hereinafter referred to as "QB_flag"), a quadtree flag (hereinafter referred to as "quadtree_flag"), a binary tree flag (hereinafter referred to as "binarytree_flag"), and a binary type flag (hereinafter referred to as "Btype_flag").

[0381] The "split_flag" can be a flag indicating whether a block has been split. For example, a split_flag value of 1 indicates that the corresponding block has been split, while a split_flag value of 0 indicates that the corresponding block has not been split.

[0382] The `QB_flag` can be a flag indicating whether the block is partitioned into a quadtree or a binary tree. For example, a `QB_flag` value of 0 indicates that the block is partitioned into a quadtree, while a `QB_flag` value of 1 indicates that the block is partitioned into a binary tree. Alternatively, a `QB_flag` value of 0 indicates that the block is partitioned into a binary tree, while a `QB_flag` value of 1 indicates that the block is partitioned into a quadtree.

[0383] The "quadtree_flag" can be a flag indicating whether the block is partitioned as a quadtree. For example, a quadtree_flag value of 1 indicates that the block is partitioned as a quadtree, while a quadtree_flag value of 0 indicates that the block is not partitioned as a quadtree.

[0384] The "binarytree_flag" can be a flag indicating whether the block is partitioned as a binary tree. For example, a binarytree_flag value of 1 indicates that the block is partitioned as a binary tree, while a binarytree_flag value of 0 indicates that the block is not partitioned as a binary tree.

[0385] The `Btype_flag` can be a flag indicating which of the vertical or horizontal partitions corresponds to the partitioning direction when a block is partitioned in a binary tree format. For example, a `Btype_flag` value of 0 indicates that the block was partitioned horizontally, and a `Btype_flag` value of 1 indicates that the block was partitioned vertically. Alternatively, a `Btype_flag` value of 0 indicates that the block was partitioned vertically, and a `Btype_flag` value of 1 indicates that the block was partitioned horizontally.

[0386] For example, it can be derived by sending at least one of quadtree_flag, binarytree_flag, and Btype_flag using signals. Figure 6 The block division information is shown in Table 1 below.

[0387] Table 1

[0388]

[0389] For example, it can be derived by sending at least one of split_flag, QB_flag, and Btype_flag using signals. Figure 6 The block division information is shown in Table 2 below.

[0390] Table 2

[0391]

[0392] The partitioning method may be limited to quadtrees or binary trees depending on the size and / or shape of the blocks. When this restriction is applied, the split_flag may be a flag indicating whether the blocks are partitioned in a quadtree or a binary tree format. The size and shape of the blocks can be derived from the block depth information, and the depth information can be transmitted from the encoding device 100 to the decoding device 200 by a signal.

[0393] When the block size falls within a certain range, it is possible to partition in quadtree form only. For example, the certain range can be defined by at least one of the maximum block size and the minimum block size that can be partitioned in quadtree form only.

[0394] Information indicating the maximum and minimum block sizes that can be partitioned solely in quadtree form can be transmitted from the encoding device 100 to the decoding device 200 via a bitstream. Furthermore, this information can be transmitted via a signal for at least one of the units such as video, sequences, frames, parameters, parallel block groups, and stripes (or segments).

[0395] Optionally, the maximum block size and / or minimum block size can be fixed sizes predefined by the encoding device 100 and the decoding device 200. For example, when the block size is greater than 64×64 and less than 256×256, it is possible to partition in quadtree form only. In this case, split_flag can be a flag indicating whether to perform partitioning in quadtree form.

[0396] When the size of a block is larger than the maximum size of a transform block, it is possible to partition it only in the form of a quadtree. Here, the sub-blocks generated by partitioning can be at least one of CU and TU.

[0397] In this case, split_flag can be a flag indicating whether the CU is partitioned in the form of a quadtree.

[0398] When the size of a block falls within a certain range, it is possible to partition it using only a binary tree or a ternary tree. For example, the specific range can be defined by at least one of the maximum block size and the minimum block size that allows partitioning using only a binary tree or a ternary tree.

[0399] Information indicating the maximum and / or minimum block size, which can be partitioned in a binary tree or a ternary tree manner, can be transmitted from the encoding device 100 to the decoding device 200 via a bitstream. Furthermore, this information can be transmitted via a signal for at least one of the units such as sequences, frames, and stripes (or segments).

[0400] Optionally, the maximum block size and / or minimum block size can be fixed sizes predefined by the encoding device 100 and the decoding device 200. For example, when the block size is greater than 8×8 and less than 16×16, it is possible to partition in binary tree form only. In this case, split_flag can be a flag indicating whether to perform partitioning in binary tree or ternary tree form.

[0401] The above description of partitioning in the form of a quadtree can be applied equally to binary and / or ternary tree forms.

[0402] The partitioning of a block may be limited by previous partitioning. For example, when a block is partitioned in a specific binary tree form and multiple sub-blocks are generated from said partition, each sub-block may be further partitioned only in that specific tree form. Here, the specific tree form can be at least one of a binary tree form, a ternary tree form, and a quadtree form.

[0403] When the horizontal or vertical dimensions of a partition block are dimensions that cannot be further subdivided, the aforementioned indicator does not need to be sent.

[0404] Figure 7 This is a diagram illustrating an embodiment of intra-frame prediction processing.

[0405] from Figure 7 The radially extending arrow from the center of the diagram indicates the prediction direction of the intra-prediction mode. Furthermore, the numbers appearing near the arrows indicate examples of mode values ​​assigned to the intra-prediction mode or its prediction direction.

[0406] exist Figure 7 In the diagram, number 0 can represent the planar mode as a non-directional intra-prediction mode. Number 1 can represent the DC mode as a non-directional intra-prediction mode.

[0407] Intra-frame coding and / or decoding can be performed using reference samples from neighboring blocks of the target block. A neighboring block can be a reconstructed neighboring block. Reference samples can represent neighboring samples.

[0408] For example, intra-frame encoding and / or decoding can be performed using the values ​​of reference samples included in the reconstructed neighboring blocks or the encoding parameters of the reconstructed neighboring blocks.

[0409] The encoding device 100 and / or the decoding device 200 can generate a prediction block by performing intra-frame prediction on the target block based on information about samples in the target image. When intra-frame prediction is performed, the encoding device 100 and / or the decoding device 200 can generate a prediction block for the target block by performing intra-frame prediction based on information about samples in the target image. When intra-frame prediction is performed, the encoding device 100 and / or the decoding device 200 can perform directional prediction and / or non-directional prediction based on at least one reconstructed reference sample.

[0410] A prediction block can be a block generated as a result of performing intra-frame prediction. A prediction block can correspond to at least one of CU, PU, ​​and TU.

[0411] The cells of the prediction block may have a size corresponding to at least one of CU, PU, ​​and TU. The prediction block may have a square shape with a size of 2N×2N or N×N. The size N×N may include sizes such as 4×4, 8×8, 16×16, 32×32, 64×64, etc.

[0412] Optionally, the prediction block can be a square block with a size of 2×2, 4×4, 8×8, 16×16, 32×32, 64×64, etc., or a rectangular block with a size of 2×8, 4×8, 2×16, 4×16, 8×16, etc.

[0413] Intra-prediction can be performed using intra-prediction modes for the target block. The number of intra-prediction modes that a target block can have can be a predefined fixed value, or it can be a value determined differently based on the attributes of the prediction block. For example, the attributes of the prediction block can include the size of the prediction block, the type of the prediction block, etc. In addition, the attributes of the prediction block can indicate the coding parameters used for the prediction block.

[0414] For example, the number of intra-prediction modes can be fixed at N, regardless of the size of the prediction block. Alternatively, the number of intra-prediction modes can be, for example, 3, 5, 9, 17, 34, 35, 36, 65, 67, or 95.

[0415] Intra-frame prediction mode can be either non-directional or directional.

[0416] For example, intra-frame prediction modes may include... Figure 7 The numbers 0 to 66 shown correspond to two non-directional modes and 65 directional modes.

[0417] For example, using a specific intra-prediction method, the intra-prediction mode may include... Figure 7 The numbers -14 to 80 shown correspond to the two non-directional patterns and 93 directional patterns.

[0418] The two non-directional modes may include DC mode and planar mode.

[0419] Directional patterns can be prediction patterns with a specific direction or angle. Directional patterns can also be called "angle patterns".

[0420] An intra-prediction mode can be represented by at least one of a mode number, a mode value, a mode angle, and a mode direction. In other words, the terms "(mode) number of intra-prediction mode", "(mode) value of intra-prediction mode", "(mode) angle of intra-prediction mode", and "(mode) direction of intra-prediction mode" can be used to have the same meaning and can be used interchangeably with each other.

[0421] The number of intra-prediction modes can be M. The value of M can be 1 or greater. In other words, the number of intra-prediction modes can be M, where M includes the number of non-directional modes and the number of directional modes.

[0422] The number of intra-prediction modes can be fixed at M, regardless of the block size and / or color components. For example, the number of intra-prediction modes can be fixed at either 35 or 67, regardless of the block size.

[0423] Optionally, the number of intra-frame prediction modes may vary depending on the shape, size, and / or type of color components of the block.

[0424] For example, in Figure 7 In the diagram, the direction prediction pattern shown by the dashed line can only be applied to the prediction of non-square blocks.

[0425] For example, the larger the block size, the more intra-prediction modes there are. Alternatively, the larger the block size, the fewer intra-prediction modes there are. When the block size is 4×4 or 8×8, the number of intra-prediction modes can be 67. When the block size is 16×16, the number of intra-prediction modes can be 35. When the block size is 32×32, the number of intra-prediction modes can be 19. When the block size is 64×64, the number of intra-prediction modes can be 7.

[0426] For example, the number of intra-prediction modes can vary depending on whether the color component is a luma signal or a chrominance signal. Optionally, the number of intra-prediction modes corresponding to the luma component block can be greater than the number of intra-prediction modes corresponding to the chrominance component block.

[0427] For example, in the vertical mode with a mode value of 50, prediction can be performed along the vertical direction based on the pixel values ​​of the reference sample. Similarly, in the horizontal mode with a mode value of 18, prediction can be performed along the horizontal direction based on the pixel values ​​of the reference sample.

[0428] Even in directional modes other than those described above, the encoding device 100 and the decoding device 200 can still perform intra-frame prediction on the target unit using reference samples based on the angle corresponding to the directional mode.

[0429] Intra-prediction modes located to the right of the vertical mode can be called "vertical-right mode". Intra-prediction modes located below the horizontal mode can be called "horizontal-bottom mode". For example, in Figure 7 In the frame prediction mode, the mode value being one of 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, and 66 can be a vertical-right mode. The mode value being one of 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, and 17 can be a horizontal-downward mode.

[0430] Non-directional modes can include DC mode and planar mode. For example, the value for DC mode can be 1, and the value for planar mode can be 0.

[0431] Orientation modes can include angle modes. Among the various intra-frame prediction modes, all modes except DC mode and planar mode can be orientation modes.

[0432] When the intra-frame prediction mode is DC mode, a prediction block can be generated based on the average pixel values ​​of multiple reference pixels. For example, the pixel values ​​of the prediction block can be determined based on the average pixel values ​​of multiple reference pixels.

[0433] The number of intra-prediction modes and the mode values ​​of each intra-prediction mode described above are merely exemplary. The number of intra-prediction modes and the mode values ​​of each intra-prediction mode described above may be defined differently depending on the embodiment, implementation, and / or requirements.

[0434] To perform intra-frame prediction on a target block, a step can be performed to check whether samples included in the reconstructed neighboring blocks can be used as reference samples for the target block. When there are samples among the samples in the neighboring blocks that cannot be used as reference samples for the target block, a value generated by interpolation and / or duplication using at least one sample value from the samples included in the reconstructed neighboring blocks can replace the sample value of the sample that cannot be used as a reference sample. When a value generated by duplication and / or interpolation replaces the sample value of an existing sample, that sample can be used as a reference sample for the target block.

[0435] When using intra-frame prediction, filters can be applied to at least one of the reference samples and the prediction samples based on at least one of the target block size and the intra-frame prediction mode.

[0436] The type of filter to be applied to at least one of the reference sample and the prediction sample can vary based on at least one of the intra-prediction mode of the target block, the size of the target block, and the shape of the target block. The filter type can be classified based on one or more of the filter tap length, the value of the filter coefficients, and the filter strength. The filter tap length can represent the number of filter taps. Furthermore, the number of filter taps can represent the length of the filter.

[0437] When the intra-frame prediction mode is planar mode, the sample value of the predicted target block can be generated by weighting the top reference sample, left reference sample, upper right reference sample, and lower left reference sample of the target block according to the position of the predicted target sample in the prediction block.

[0438] When the intra-frame prediction mode is DC mode, the average of reference samples above and to the left of the target block can be used when generating the prediction block for the target block. Furthermore, filtering using the values ​​of the reference samples can be performed on specific rows or columns within the target block. The specific row can be one or more upper rows adjacent to the reference samples. The specific column can be one or more left columns adjacent to the reference samples.

[0439] When the intra-frame prediction mode is directional mode, the top reference sample, left reference sample, top right reference sample, and / or bottom left reference sample of the target block can be used to generate the prediction block.

[0440] To generate the above predicted samples, real-number-based interpolation can be performed.

[0441] The intra-prediction mode of the target block can be predicted from the intra-prediction modes of neighboring blocks adjacent to the target block, and the information used for prediction can be entropy encoded / entropy decoded.

[0442] For example, when the intra-prediction modes of the target block and neighboring blocks are the same, a predefined flag can be used to signal that the intra-prediction modes of the target block and neighboring blocks are the same.

[0443] For example, an indicator can be sent to indicate an intra-prediction mode that is the same as the intra-prediction mode of the target block among the intra-prediction modes of multiple neighboring blocks.

[0444] When the intra-prediction modes of the target block and neighboring blocks are different from each other, entropy coding and / or entropy decoding can be used to encode and / or decode information about the intra-prediction mode of the target block.

[0445] Figure 8 This is a diagram showing the reference samples used in the intra-frame prediction process.

[0446] The reconstruction reference points used for intra-frame prediction of the target block may include the lower left reference point, the left reference point, the upper left corner reference point, the upper reference point, and the upper right reference point.

[0447] For example, a left reference sample may represent a reconstructed reference pixel adjacent to the left side of the target block. A top reference sample may represent a reconstructed reference pixel adjacent to the top of the target block. A top-left reference sample may represent a reconstructed reference pixel located at the top-left corner of the target block. A bottom-left reference sample may represent a reference sample located below the left-side sample line, which is the same line as the left-side sample line formed by the left reference samples. A top-right reference sample may represent a reference sample located to the right of the upper sample line, which is the same line as the upper sample line formed by the upper reference samples.

[0448] When the size of the target block is N×N, the number of the lower left reference point, the left reference point, the upper reference point, and the upper right reference point can all be N.

[0449] A prediction block can be generated by performing intra-frame prediction on the target block. The process of generating a prediction block may include determining the values ​​of the pixels in the prediction block. The target block and the prediction block can have the same size.

[0450] The reference sample used for intra-prediction of the target block can be changed according to the intra-prediction mode of the target block. The direction of the intra-prediction mode can represent the dependency between the reference sample and the pixels of the prediction block. For example, the value of a specified reference sample can be used as the value of one or more specified pixels in the prediction block. In this case, the specified reference sample and the one or more specified pixels in the prediction block can be samples and pixels located on a straight line along the direction of the intra-prediction mode. In other words, the value of the specified reference sample can be copied as the value of a pixel located in the opposite direction to the direction of the intra-prediction mode. Optionally, the value of a pixel in the prediction block can be the value of a reference sample located in the direction of the intra-prediction mode relative to the pixel's position.

[0451] In the example, when the intra-prediction mode of the target block is vertical, the upper reference sample can be used for intra-prediction. When the intra-prediction mode is vertical, the value of a pixel in the prediction block can be the value of a reference sample vertically above that pixel. Therefore, the upper reference sample adjacent to the top of the target block can be used for intra-prediction. Furthermore, the value of a pixel in a row of the prediction block can be the same as the value of a pixel at the upper reference sample.

[0452] In the example, when the intra-prediction mode of the target block is horizontal, the left reference sample can be used for intra-prediction. When the intra-prediction mode is horizontal, the value of a pixel in the prediction block can be the value of a reference sample horizontally to the left of that pixel. Therefore, the left reference sample adjacent to the left side of the target block can be used for intra-prediction. Furthermore, the value of a pixel in a column of the prediction block can be the same as the value of the pixel in the left reference sample.

[0453] In the example, when the mode value of the intra-prediction mode for the current block is 34, at least some of the left reference samples, the top-left reference sample, and the top reference sample can be used for intra-prediction. When the mode value of the intra-prediction mode is 34, the value of a pixel in the prediction block can be the value of a reference sample located diagonally at the top-left corner of that pixel.

[0454] Furthermore, in the case of an intra-prediction mode with mode values ​​ranging from 52 to 66, at least a portion of the upper right reference samples can be used for intra-prediction.

[0455] Furthermore, in the case of intra-prediction modes with mode values ​​ranging from 2 to 17, at least a portion of the lower left reference samples can be used for intra-prediction.

[0456] Furthermore, in the case of intra-prediction modes with mode values ​​ranging from 19 to 49, the upper left reference sample can be used for intra-prediction.

[0457] The number of reference samples used to determine the pixel value of a pixel in the prediction block can be 1, 2, or more.

[0458] As described above, the pixel value of a pixel in a prediction block can be determined based on the pixel's position and the position of a reference sample indicated by the direction of the intra-prediction mode. When both the pixel's position and the position of the reference sample indicated by the direction of the intra-prediction mode are integer positions, the value of a reference sample indicated by the integer position can be used to determine the pixel value of the pixel in the prediction block.

[0459] When the pixel position and the position of the reference sample indicated by the direction of the intra-prediction mode are not integer positions, an interpolated reference sample can be generated based on the two reference samples closest to that reference sample position. The value of the interpolated reference sample can be used to determine the pixel value of the pixel in the prediction block. In other words, when the pixel position in the prediction block and the position of the reference sample indicated by the direction of the intra-prediction mode indicate the position between two reference samples, an interpolation based on the values ​​of those two samples can be generated.

[0460] The predicted block generated by prediction may differ from the original target block. In other words, there may be prediction errors, which are the differences between the target block and the predicted block, and there may also be prediction errors between pixels in the target block and pixels in the predicted block.

[0461] In the following text, the terms “difference,” “error,” and “residual” may be used to have the same meaning and may be used interchangeably with each other.

[0462] For example, in the case of intra-frame prediction, the greater the distance between the pixels of the predicted block and the reference sample, the greater the potential prediction error. This prediction error can lead to discontinuities between the generated predicted block and its neighboring blocks.

[0463] To reduce prediction error, filtering operations can be used for prediction blocks. These filtering operations can be configured to adaptively apply filters to regions within the prediction block that are considered to have large prediction errors. For example, regions considered to have large prediction errors could be the boundaries of the prediction block. Furthermore, the regions within the prediction block considered to have large prediction errors can vary depending on the intra-prediction mode, and the characteristics of the filters can also vary depending on the intra-prediction mode.

[0464] like Figure 8 As shown, for intra-frame prediction of the target block, at least one of reference lines 0 to 3 can be used.

[0465] exist Figure 8 Each reference line in the code can indicate a reference point line that includes one or more reference points. When the reference line number is smaller, it can indicate a reference point line that is closer to the target block.

[0466] Samples in fragments A and F can be obtained by padding instead of from reconstructed neighboring blocks, wherein the padding uses the samples from fragments B and E that are closest to the target block.

[0467] An index information indicating the reference sample lines to be used for intra-frame prediction of the target block can be transmitted using a signal. The index information can indicate which of a plurality of reference sample lines will be used for intra-frame prediction of the target block. For example, the index information can have a value corresponding to any one of 0 to 3.

[0468] When the upper boundary of the target block is the boundary of the CTU, only reference sample line 0 can be available. Therefore, in this case, index information does not need to be sent. When additional reference sample lines besides reference sample line 0 are used, filtering of the prediction block, which will be described later, is not required.

[0469] In the case of intra-frame prediction between colors, the predicted block of the target block of the second color component can be generated based on the corresponding reconstructed block of the first color component.

[0470] For example, the first color component can be the luminance component, and the second color component can be the chromaticity component.

[0471] To perform inter-color intra-frame prediction, the parameters of a linear model between the first and second color components can be derived based on a template.

[0472] The template may include a reference sample point above the target block (upper reference sample point) and / or a reference sample point to the left of the target block (left reference sample point), and may include the upper reference sample point and / or the left reference sample point of the reconstructed block of the first color component corresponding to the reference sample point.

[0473] For example, the following values ​​can be used to derive the parameters of a linear model: 1) the value of the sample point of the first color component with the maximum value among the samples in the template, 2) the value of the sample point of the second color component corresponding to the sample point of the first color component, 3) the value of the sample point of the first color component with the minimum value among the samples in the template, and 4) the value of the sample point of the second color component corresponding to the sample point of the first color component.

[0474] When exporting the parameters of a linear model, the predicted block of the target block can be generated by applying the corresponding reconstructed block to the linear model.

[0475] Depending on the image format, subsampling can be performed on samples adjacent to the reconstructed block of the first color component and on the corresponding reconstructed block of the first color component. For example, when one sample of the second color component corresponds to four samples of the first color component, a corresponding sample can be calculated by subsampling the four samples of the first color component. When subsampling is performed, the parameters of the linear model and inter-color intra-frame prediction can be performed based on the subsampled corresponding sample.

[0476] In intra-frame prediction mode, information about whether to perform inter-color intra-frame prediction and / or the range of templates can be sent by signaling.

[0477] The target block can be divided into two or four sub-blocks in the horizontal and / or vertical directions.

[0478] Sub-blocks generated by partitioning can be reconstructed sequentially. That is, when intra-prediction is performed on each sub-block, a sub-prediction block for that sub-block can be generated. Furthermore, when inverse quantization and / or inverse transform is performed on each sub-block, a sub-residual block for the corresponding sub-block can be generated. Reconstructed sub-blocks can be generated by adding the sub-prediction blocks to the sub-residual blocks. The reconstructed sub-blocks can be used as reference samples for intra-prediction of sub-blocks with the next higher priority.

[0479] A sub-block can be a block containing a specific number (e.g., 16) or more samples. For example, when the target block is an 8×4 block or a 4×8 block, the target block can be divided into two sub-blocks. Furthermore, when the target block is a 4×4 block, it cannot be divided into sub-blocks. When the target block has another size, it can be divided into four sub-blocks.

[0480] Signals can be used to send information about whether to perform intra-frame prediction based on these sub-blocks and / or about the partitioning direction (horizontal or vertical).

[0481] This sub-block-based intra-prediction can be restricted so that it is performed only when reference sample line 0 is used. When performing sub-block-based intra-prediction, filtering of the prediction block, which will be described below, may not be performed.

[0482] The final prediction block can be generated by filtering the prediction block generated via intra-frame prediction.

[0483] Filtering can be performed by applying specific weights to the target sample, left reference sample, top reference sample, and / or top-left reference sample, which are the targets to be filtered.

[0484] The weights and / or reference samples (e.g., the range of reference samples, the location of reference samples, etc.) used for filtering can be determined based on at least one of the block size, intra-frame prediction mode, and the location of the target filter sample in the prediction block.

[0485] For example, filtering can be performed only in specific intra-frame prediction modes (e.g., DC mode, planar mode, vertical mode, horizontal mode, diagonal mode, and / or adjacent diagonal mode).

[0486] Adjacent diagonal patterns can be patterns with numbers obtained by adding k to the diagonal pattern's number, or patterns with numbers obtained by subtracting k from the diagonal pattern's number. In other words, the number of an adjacent diagonal pattern can be the sum of the diagonal pattern's number and k, or the difference between the diagonal pattern's number and k. For example, k can be a positive integer of 8 or less.

[0487] The intra prediction mode of the target block can be derived using the intra prediction modes of neighboring blocks that exist near the target block, and this derived intra prediction mode can be entropy encoded and / or entropy decoded.

[0488] For example, when the intra prediction mode of the target block is the same as that of the neighboring blocks, specific flag information can be used to signal information indicating that the intra prediction mode of the target block is the same as that of the neighboring blocks.

[0489] Furthermore, for example, an indicator information of neighboring blocks whose intra-prediction modes are the same as those of the target block can be transmitted using signals.

[0490] For example, when the intra prediction mode of the target block is different from that of the neighboring blocks, entropy coding and / or entropy decoding can be performed on the information about the intra prediction mode of the target block by performing entropy coding and / or entropy decoding based on the intra prediction modes of the neighboring blocks.

[0491] Figure 9 This is a diagram illustrating an embodiment used to explain the inter-frame prediction process.

[0492] Figure 9 The rectangles shown can represent images (or screens). Furthermore, in... Figure 9 In the image, arrows indicate the prediction direction. An arrow pointing from the first frame to the second frame indicates that the second frame references the first frame. In other words, each image can be encoded and / or decoded based on the prediction direction.

[0493] Images can be classified according to their encoding type into intra-frame frames (I-frames), one-way predictive frames or predictive-coded frames (P-frames), and two-way predictive frames or two-way predictive-coded frames (B-frames). Each frame can be encoded and / or decoded according to its encoding type.

[0494] When the target image to be encoded is an I-frame, the target image can be encoded using the data contained in the image itself without inter-frame prediction referencing other images. For example, an I-frame can be encoded solely via intra-frame prediction.

[0495] When the target image is a P-frame, it can be encoded using inter-frame prediction with reference frames existing in one direction. Here, the one direction can be a forward direction or a backward direction.

[0496] When the target image is a B-frame, the image can be encoded via inter-frame prediction using reference frames present in both directions, or via inter-frame prediction using reference frames present in one of the forward and backward directions. Here, the two directions can be the forward and backward directions.

[0497] P-frames and B-frames that are encoded and / or decoded using reference frames can be considered as images using inter-frame prediction.

[0498] The following will describe in detail the inter-frame prediction in inter-frame mode according to the embodiments.

[0499] Reference images and motion information can be used to perform inter-frame prediction or motion compensation.

[0500] In inter-frame mode, the encoding device 100 may perform inter-frame prediction and / or motion compensation on the target block. The decoding device 200 may perform inter-frame prediction and / or motion compensation on the target block corresponding to the inter-frame prediction and / or motion compensation performed by the encoding device 100.

[0501] Motion information of the target block can be derived independently by the encoding unit 100 and the decoding unit 200 during inter-frame prediction. Motion information can be derived using the motion information of reconstructed neighboring blocks, the motion information of the col block, and / or the motion information of blocks adjacent to the col block.

[0502] For example, encoding device 100 or decoding device 200 can perform prediction and / or motion compensation by using motion information of spatial candidates and / or temporal candidates as motion information of a target block. The target block may represent a PU and / or a PU partition.

[0503] Spatial candidates can be reconstructed blocks that are spatially adjacent to the target block.

[0504] The time candidate can be a reconstructed block that corresponds to the target block in a previously reconstructed co-location frame (col frame).

[0505] In inter-frame prediction, the coding unit 100 and the decoding unit 200 can improve coding efficiency and decoding efficiency by utilizing motion information from spatial candidates and / or temporal candidates. The motion information from spatial candidates can be referred to as "spatial motion information." The motion information from temporal candidates can be referred to as "temporal motion information."

[0506] Below, the motion information of spatial candidates can be the motion information of PUs including spatial candidates. The motion information of temporal candidates can be the motion information of PUs including temporal candidates. The motion information of candidate blocks can be the motion information of PUs including candidate blocks.

[0507] Inter-frame prediction can be performed using a reference frame.

[0508] The reference image can be at least one of an image preceding or following the target image. The reference image can be an image used for prediction of the target block.

[0509] In inter-frame prediction, regions within a reference frame can be specified using a reference frame index (or refIdx) used to indicate the reference frame, motion vectors that will be described subsequently, and so on. Here, the region specified in the reference frame can indicate a reference block.

[0510] Inter-frame prediction can select a reference frame, and can also select a reference block corresponding to the target block from the reference frame. In addition, inter-frame prediction can use the selected reference block to generate a prediction block for the target block.

[0511] Motion information can be derived by each of the encoding device 100 and the decoding device 200 during inter-frame prediction.

[0512] Spatial candidates can be 1) blocks existing in the target frame, 2) blocks that have been previously reconstructed via encoding and / or decoding, and 3) blocks adjacent to or located at the corner of the target block. Here, a "block located at the corner of the target block" can be a block that is vertically adjacent to a horizontally adjacent neighboring block, or a block that is horizontally adjacent to a vertically adjacent neighboring block. Furthermore, "block located at the corner of the target block" can have the same meaning as "block adjacent to the corner of the target block." The meaning of "block located at the corner of the target block" can be included within the meaning of "block adjacent to the target block."

[0513] For example, a spatial candidate can be a reconstruction block located to the left of the target block, a reconstruction block located above the target block, a reconstruction block located at the lower left corner of the target block, a reconstruction block located at the upper right corner of the target block, or a reconstruction block located at the upper left corner of the target block.

[0514] Each of the encoding device 100 and the decoding device 200 can identify a block existing in the col frame at a spatial position corresponding to the target block. The position of the target block in the target frame and the position of the identified block in the col frame can correspond to each other.

[0515] Each of the encoding device 100 and the decoding device 200 can identify a col block existing at a predefined relevant location for the identified block as a time candidate. The predefined relevant location can be a location existing inside and / or outside the identified block.

[0516] For example, a col block can include a first col block and a second col block. When the coordinates of the identified block are (xP, yP) and the size of the identified block is represented by (nPSW, nPSH), the first col block can be the block located at coordinates (xP+nPSW, yP+nPSH). The second col block can be the block located at coordinates (xP+(nPSW>>1), yP+(nPSH>>1)). When the first col block is unavailable, the second col block can be used selectively.

[0517] The motion vector of the target block can be determined based on the motion vector of the col block. Each of the encoding device 100 and the decoding device 200 can scale the motion vector of the col block. The scaled motion vector of the col block can be used as the motion vector of the target block. Furthermore, the motion vectors of motion information for time candidates stored in a list can also be scaled motion vectors.

[0518] The ratio of the motion vector of the target block to the motion vector of the col block can be the same as the ratio of the first time distance to the second time distance. The first time distance can be the distance between the reference frame and the target frame of the target block. The second time distance can be the distance between the reference frame and the col frame of the col block.

[0519] The scheme used to derive motion information can be varied depending on the inter-frame prediction mode of the target block. For example, inter-frame prediction modes applied to inter-frame prediction may include Advanced Motion Vector Prediction Factor (AMVP) mode, merge mode, skip mode, merge mode with motion vector difference, sub-block merge mode, triangular partitioning mode, inter-frame / intra-frame combined prediction mode, affine inter-frame mode, and current frame reference mode. The merge mode can also be called "motion merge mode." Each mode will be described in detail below.

[0520] 1) AMVP mode

[0521] When using AMVP mode, the encoding device 100 can search for similar blocks in the neighborhood of the target block. The encoding device 100 can obtain a predicted block by performing a prediction on the target block using the motion information of the found similar blocks. The encoding device 100 can encode the residual block, which is the difference between the target block and the predicted block.

[0522] 1-1) Create a list of candidate motion vectors for prediction.

[0523] When the AMVP mode is used as the prediction mode, each of the encoding device 100 and the decoding device 200 can create a list of prediction motion vector candidates using spatial candidate motion vectors, temporal candidate motion vectors, and zero vectors. The prediction motion vector candidate list may include one or more prediction motion vector candidates. At least one of the spatial candidate motion vectors, temporal candidate motion vectors, and zero vectors can be determined and used as a prediction motion vector candidate.

[0524] In the following text, the terms “predicted motion vector (candidate)” and “motion vector (candidate)” can be used to have the same meaning and can be used interchangeably.

[0525] In the following text, the terms “predicted motion vector candidate” and “AMVP candidate” can be used to have the same meaning and can be used interchangeably.

[0526] In the following text, the terms “predicted motion vector candidate list” and “AMVP candidate list” can be used to have the same meaning and can be used interchangeably.

[0527] Spatial candidates can include reconstructed spatial neighbor blocks. In other words, the motion vectors of the reconstructed neighbor blocks can be referred to as "spatial prediction motion vector candidates".

[0528] A temporal candidate can include the col block and the blocks adjacent to the col block. In other words, the motion vector of the col block or the motion vector of the blocks adjacent to the col block can be called a "temporal prediction motion vector candidate".

[0529] The zero vector can be a (0,0) motion vector.

[0530] The predicted motion vector candidate can be a motion vector prediction factor used to predict the motion vector. Furthermore, in the encoding device 100, each predicted motion vector candidate can be an initial search position for the motion vector.

[0531] 1-2) Search for motion vectors using the list of predicted motion vector candidates.

[0532] The encoding device 100 can use a list of predicted motion vector candidates to determine the motion vector to be used for encoding the target block within a search range. Furthermore, the encoding device 100 can determine a predicted motion vector candidate that will be used as the target block from among the predicted motion vector candidates present in the list of predicted motion vector candidates.

[0533] The motion vector used to encode the target block can be a motion vector that can be encoded at the minimum cost.

[0534] In addition, the encoding device 100 can determine whether to use the AMVP mode to encode the target block.

[0535] 1-3) Transmission of inter-frame prediction information

[0536] The encoding device 100 can generate a bitstream that includes inter-frame prediction information required for inter-frame prediction. The decoding device 200 can use the inter-frame prediction information of the bitstream to perform inter-frame prediction on the target block.

[0537] Inter-frame prediction information may include 1) mode information indicating whether AMVP mode is used, 2) predicted motion vector index, 3) motion vector difference (MVD), 4) reference direction and 5) reference frame index.

[0538] In the following text, the terms “predicted motion vector index” and “AMVP index” can be used interchangeably and have the same meaning.

[0539] In addition, inter-frame prediction information may include residual signals.

[0540] When the mode information indicates that the AMVP mode is used, the decoding device 200 can obtain the predicted motion vector index, MVD, reference direction and reference frame index from the bitstream through entropy decoding.

[0541] The predicted motion vector index indicates which of the predicted motion vector candidates included in the predicted motion vector candidate list will be used to predict the target block.

[0542] 1-4) Inter-frame prediction in AMVP mode using inter-frame prediction information

[0543] The decoding device 200 can use the list of predicted motion vector candidates to derive predicted motion vector candidates, and can determine the motion information of the target block based on the derived predicted motion vector candidates.

[0544] The decoding device 200 can use the predicted motion vector index to determine a motion vector candidate for the target block from among the predicted motion vector candidates included in the predicted motion vector candidate list. The decoding device 200 can select the predicted motion vector candidate indicated by the predicted motion vector index from among the predicted motion vector candidates included in the predicted motion vector candidate list as the predicted motion vector for the target block.

[0545] Encoding device 100 can generate an entropy-coded predicted motion vector index by applying entropy coding to the predicted motion vector index, and can generate a bitstream including the entropy-coded predicted motion vector index. The entropy-coded predicted motion vector index can be transmitted from encoding device 100 to decoding device 200 via the bitstream. Decoding device 200 can extract the entropy-coded predicted motion vector index from the bitstream, and can obtain the predicted motion vector index by applying entropy decoding to the entropy-coded predicted motion vector index.

[0546] The motion vector actually used for inter-frame prediction of the target block may not match the predicted motion vector. MVD (Motion Vector Difference) can be used to indicate the difference between the actual motion vector used for inter-frame prediction of the target block and the predicted motion vector. The encoding device 100 can derive a predicted motion vector similar to the actual motion vector used for inter-frame prediction of the target block in order to use the smallest possible MVD.

[0547] Motion Vector Difference (MVD) can be the difference between the motion vector of the target block and the predicted motion vector. Encoding device 100 can calculate the MVD and generate an entropy-coded MVD by applying entropy coding to the MVD. Encoding device 100 can generate a bitstream including the entropy-coded MVD.

[0548] The MVD can be sent from the encoding device 100 to the decoding device 200 via a bitstream. The decoding device 200 can extract the entropy-encoded MVD from the bitstream and obtain the MVD by applying entropy decoding to the entropy-encoded MVD.

[0549] The decoding device 200 can derive the motion vector of the target block by summing the MVD and the predicted motion vector. In other words, the motion vector of the target block derived by the decoding device 200 can be the sum of the MVD and the motion vector candidate.

[0550] Furthermore, the encoding device 100 can generate entropy-coded MVD resolution information by applying entropy coding to the calculated MVD resolution information, and can generate a bitstream including the entropy-coded MVD resolution information. The decoding device 200 can extract the entropy-coded MVD resolution information from the bitstream, and can obtain the MVD resolution information by applying entropy decoding to the entropy-coded MVD resolution information. The decoding device 200 can use the MVD resolution information to adjust the MVD resolution.

[0551] Furthermore, the encoding device 100 can calculate the MVD based on an affine model. The decoding device 200 can derive the affine control motion vector of the target block from the sum of the MVD and the affine control motion vector candidates, and can use the affine control motion vector to derive the motion vector of the sub-block.

[0552] The reference direction can indicate a list of reference frames that will be used to predict the target block. For example, the reference direction can indicate one of reference frame list L0 and reference frame list L1.

[0553] The reference direction only indicates the list of reference frames that will be used to predict the target block, and does not necessarily mean that the direction of the reference frames is limited to the forward or backward direction. In other words, each of the reference frame lists L0 and L1 can include frames in the forward and / or backward directions.

[0554] A unidirectional reference direction can mean using a single reference screen list. A bidirectional reference direction can mean using two reference screen lists. In other words, the reference direction can indicate one of the following: using only reference screen list L0, using only reference screen list L1, or using both reference screen lists.

[0555] The reference screen index indicates a reference screen in the reference screen list for predicting the target block. Encoding device 100 can generate an entropy-coded reference screen index by applying entropy coding to the reference screen index, and can generate a bitstream including the entropy-coded reference screen index. The entropy-coded reference screen index can be signaled from encoding device 100 to decoding device 200 via the bitstream. Decoding device 200 can extract the entropy-coded reference screen index from the bitstream, and can obtain the reference screen index by applying entropy decoding to the entropy-coded reference screen index.

[0556] When two reference frame lists are used to predict a target block, a single reference frame index and a single motion vector can be used for each of the reference frame lists. Furthermore, when two reference frame lists are used to predict a target block, two prediction blocks can be specified for the target block. For example, the (final) prediction block for the target block can be generated using the average or weighted sum of the two prediction blocks for the target block.

[0557] The motion vector of a target block can be derived by predicting the motion vector index, MVD, reference direction, and reference screen index.

[0558] The decoding device 200 can generate a prediction block for a target block based on the derived motion vectors and a reference frame index. For example, the prediction block can be a reference block indicated by the derived motion vectors in a reference frame indicated by the reference frame index.

[0559] Since the predicted motion vector index and MVD are encoded, while the motion vector of the target block itself is not encoded, the number of bits sent from the encoding device 100 to the decoding device 200 can be reduced, and the encoding efficiency can be improved.

[0560] For the target block, motion information from reconstructed neighboring blocks can be used. In certain inter-frame prediction modes, the encoding device 100 may not encode the actual motion information of the target block separately. Instead of encoding the target block's motion information, additional information can be encoded, which enables the derivation of the target block's motion information using the reconstructed motion information from neighboring blocks. Because this additional information is encoded, the number of bits sent to the decoding device 200 can be reduced, and encoding efficiency can be improved.

[0561] For example, in inter-frame prediction modes where motion information of the target block is not directly encoded, skipping modes and / or merging modes may exist. Here, the motion information of each of the encoding device 100 and the decoding device 200 among the neighboring units that indicate reconstruction will be used as the identifier and / or index of the unit's motion information for the target unit.

[0562] 2) Merge Mode

[0563] Merging is a scheme used to derive motion information for a target block. The term "merging" can mean merging the motion of multiple blocks. Merging can also mean that the motion information of one block is applied to other blocks. In other words, a merging pattern can be a mode for deriving the motion information of a target block from the motion information of neighboring blocks.

[0564] When using the merging mode, the encoding device 100 can use motion information from spatial candidates and / or temporal candidates to predict motion information of the target block. Spatial candidates may include reconstructed spatially adjacent blocks that are spatially adjacent to the target block. Spatially adjacent blocks may include left-side adjacent blocks and top-side adjacent blocks. Temporal candidates may include col blocks. The terms "spatial candidate" and "spatial merging candidate" are used interchangeably and have the same meaning. The terms "temporal candidate" and "temporal merging candidate" are used interchangeably and have the same meaning.

[0565] The encoding device 100 can obtain a prediction block through prediction. The encoding device 100 can encode a residual block, which is the difference between the target block and the prediction block.

[0566] 2-1) Create a list of candidate mergers

[0567] When using the merging mode, each of the encoding device 100 and the decoding device 200 can create a merging candidate list using motion information of spatial candidates and / or motion information of temporal candidates. The motion information may include 1) a motion vector, 2) a reference frame index, and 3) a reference direction. The reference direction can be unidirectional or bidirectional. The reference direction may represent an inter-frame prediction indicator.

[0568] The merge candidate list can include merge candidates. Merge candidates can be motion information. In other words, the merge candidate list can be a list that stores multiple pieces of motion information.

[0569] The merged candidate can be the motion information of multiple temporal and / or spatial candidates. In other words, the merged candidate list can include the motion information of temporal and / or spatial candidates, etc.

[0570] Furthermore, the merge candidate list may include new merge candidates generated by combining merge candidates that already exist in the merge candidate list. In other words, the merge candidate list may include new motion information generated by combining multiple motion information items that previously existed in the merge candidate list.

[0571] In addition, the merge candidate list may include history-based merge candidates. History-based merge candidates may be motion information of blocks that were encoded and / or decoded before the target block.

[0572] In addition, the list of merge candidates may include merge candidates based on the average of two merge candidates.

[0573] Merging candidates can be specific patterns for deriving inter-frame prediction information. Merging candidates can be information indicating specific patterns for deriving inter-frame prediction information. Inter-frame prediction information for a target block can be derived based on the specific patterns indicated by the merging candidates. Furthermore, the specific patterns can include the processing of deriving a series of inter-frame prediction information. Such specific patterns can be inter-frame prediction information deriving patterns or motion information deriving patterns.

[0574] Inter-frame prediction information for the target block can be derived based on the pattern indicated by the merge candidate selected in the merge candidate list via the merge index.

[0575] For example, the motion information export mode in the candidate list can be at least one of the following modes: 1) motion information export mode for sub-block units and 2) affine motion information export mode.

[0576] In addition, the list of merged candidates may include motion information for the zero vector. The zero vector may also be referred to as a "zero merged candidate".

[0577] In other words, the multiple motion information in the merged candidate list can be at least one of the following: 1) motion information of spatial candidates, 2) motion information of temporal candidates, 3) motion information generated by combining multiple motion information that previously existed in the merged candidate list, and 4) zero vector.

[0578] Motion information may include 1) motion vectors, 2) reference frame indexes, and 3) reference directions. The reference direction can also be referred to as an "inter-frame prediction indicator." The reference direction can be unidirectional or bidirectional. A unidirectional reference direction can indicate L0 prediction or L1 prediction.

[0579] A list of merge candidates can be created before performing predictions in merge mode.

[0580] The number of merging candidates in the merging candidate list can be predefined. Each of the encoding device 100 and the decoding device 200 can add merging candidates to the merging candidate list according to a predefined scheme and a predefined priority, such that the merging candidate list has a predefined number of merging candidates. The merging candidate list of the encoding device 100 and the merging candidate list of the decoding device 200 can be made identical to each other using a predefined scheme and a predefined priority.

[0581] Merging can be applied based on CU or PU. When merging is performed based on CU or PU, the encoding device 100 can send a bit stream including predefined information to the decoding device 200. For example, the predefined information may include 1) information indicating whether merging is performed for each block partition, and 2) information about the blocks to be merged among the blocks that are spatial candidates and / or temporal candidates for the target block.

[0582] 2-2) Search for motion vectors using a merged candidate list

[0583] The encoding device 100 can determine merge candidates to be used for encoding the target block. For example, the encoding device 100 can perform prediction on the target block using merge candidates from the merge candidate list, and can generate residual blocks for the merge candidates. The encoding device 100 can encode the target block using the merge candidate that generates the minimum cost in both the prediction and the encoding of the residual blocks.

[0584] In addition, the encoding device 100 can determine whether to use a merging mode to encode the target block.

[0585] 2-3) Transmission of inter-frame prediction information

[0586] Encoding device 100 can generate a bitstream including inter-frame prediction information required for inter-frame prediction. Encoding device 100 can generate entropy-coded inter-frame prediction information by performing entropy coding on the inter-frame prediction information, and can send the bitstream including the entropy-coded inter-frame prediction information to decoding device 200. The entropy-coded inter-frame prediction information can be sent by encoding device 100 to decoding device 200 via a bitstream signal. Decoding device 200 can extract the entropy-coded inter-frame prediction information from the bitstream, and can obtain inter-frame prediction information by applying entropy decoding to the entropy-coded inter-frame prediction information.

[0587] The decoding device 200 can use the inter-frame prediction information of the bitstream to perform inter-frame prediction on the target block.

[0588] Information indicating whether the merge mode is used, 2) merge index, and 3) correction information.

[0589] In addition, inter-frame prediction information may include residual signals.

[0590] The decoding device 200 can obtain the merge index from the bitstream only when the mode information indicates that the merge mode is used.

[0591] Pattern information can be merge flags. The unit of pattern information can be a block. Information about a block can include pattern information, and the pattern information can indicate whether a merge pattern is applied to the block.

[0592] The merge index can indicate which merge candidate from the merge candidate list will be used to predict the target block. Optionally, the merge index can indicate which block from the spatially or temporally adjacent neighboring blocks will be merged with the target block.

[0593] The encoding device 100 can select the merge candidate with the highest encoding performance from the merge candidate list, and can set the value of the merge index to indicate the selected merge candidate.

[0594] The correction information can be information used to correct motion vectors. Encoding device 100 can generate the correction information. Decoding device 200 can correct the motion vectors of the merging candidates selected by the merging index based on the correction information.

[0595] The correction information may include at least one of information indicating whether correction will be performed, correction direction information, and correction size information. A prediction mode that corrects motion vectors based on correction information transmitted by a signal may be referred to as a "merging mode with motion vector difference".

[0596] 2-4) Inter-frame prediction using merging mode with inter-frame prediction information

[0597] The decoding device 200 can perform prediction on the target block using a merge candidate indicated by a merge index from among the merge candidates included in the merge candidate list.

[0598] The motion vector of the target block can be specified by the motion vector of the merge candidate indicated by the merge index, the reference screen index, and the reference direction.

[0599] 3) Skip Mode

[0600] Skip mode can be a mode that applies spatial or temporal motion information to the target block without alteration. Furthermore, skip mode can be a mode that does not use the residual signal. In other words, when using skip mode, the reconstructed block can be identical to the predicted block.

[0601] The difference between merge mode and skip mode lies in whether or not residual signals are sent or used. In other words, skip mode is similar to merge mode except that residual signals are not sent or used.

[0602] When using the skip mode, the encoding device 100 can transmit information related to blocks whose motion information will be used as motion information for target blocks via a bitstream to the decoding device 200. The encoding device 100 can generate entropy-encoded information by performing entropy encoding on this information, and can transmit the entropy-encoded information as a signal to the decoding device 200 via the bitstream. The decoding device 200 can extract the entropy-encoded information from the bitstream and obtain information by applying entropy decoding to the entropy-encoded information.

[0603] Furthermore, when using the skip mode, the encoding device 100 may not send other syntax information (such as MVD) to the decoding device 200. For example, when using the skip mode, the encoding device 100 may not send the syntax elements associated with at least one of MVD, code block flag, and transform coefficient level to the decoding device 200.

[0604] 3-1) Create a list of candidate mergers

[0605] The merge candidate list can also be used in skip mode. In other words, the merge candidate list can be used in both merge mode and skip mode. In this respect, the merge candidate list can also be referred to as the "skip candidate list" or the "merge / skip candidate list".

[0606] Optionally, the skip mode may use an additional candidate list that differs from the candidate list used in the merge mode. In this case, in the following description, the merge candidate list and merge candidate may be replaced by the skip candidate list and the skip candidate, respectively.

[0607] A list of merged candidates can be created before performing predictions in skip mode.

[0608] 3-2) Search for motion vectors using a merged candidate list

[0609] Encoding device 100 can determine merge candidates to be used for encoding the target block. For example, encoding device 100 can use merge candidates from the merge candidate list to perform prediction on the target block. Encoding device 100 can use the merge candidate that generates the minimum cost in the prediction to encode the target block.

[0610] In addition, the encoding device 100 can determine whether to use a skip mode to encode the target block.

[0611] 3-3) Transmission of inter-frame prediction information

[0612] The encoding device 100 can generate a bitstream that includes inter-frame prediction information required for inter-frame prediction. The decoding device 200 can use the inter-frame prediction information of the bitstream to perform inter-frame prediction on the target block.

[0613] Inter-frame prediction information may include 1) mode information indicating whether a skip mode is used and 2) a skip index.

[0614] Skipping indexes is the same as merging indexes as described above.

[0615] When using skip mode, the target block can be encoded without using the residual signal. Inter-frame prediction information may not include the residual signal. Optionally, the bitstream may not include the residual signal.

[0616] The decoding device 200 may obtain the skip index from the bitstream only when the mode information indicates that a skip mode is used. As mentioned above, the merge index and the skip index may be the same. The decoding device 200 may obtain the skip index from the bitstream only when the mode information indicates that either a merge mode or a skip mode is used.

[0617] Skip index indicates which of the merge candidates included in the merge candidate list will be used to predict the target block.

[0618] 3-4) Inter-frame prediction in skip mode using inter-frame prediction information

[0619] The decoding device 200 can perform prediction on the target block using a merge candidate indicated by a skip index from among the merge candidates included in the merge candidate list.

[0620] The motion vector of the target block can be specified by the motion vector of the merge candidate indicated by the skip index, the reference screen index, and the reference direction.

[0621] 4) Current screen reference mode

[0622] The current frame reference mode can represent a prediction mode that uses the previously reconstructed area in the target frame to which the target block belongs.

[0623] Motion vectors can be used to specify previously reconstructed areas. The reference frame index of the target block can be used to determine whether the target block has been encoded in the current frame reference mode.

[0624] A flag or index indicating whether a target block is encoded in the current screen reference mode can be sent by the encoding device 100 to the decoding device 200. Optionally, whether a target block is encoded in the current screen reference mode can be inferred from the reference screen index of the target block.

[0625] When a target block is encoded in the current frame reference mode, the current frame can exist in a fixed position or any position in the reference frame list for the target block.

[0626] For example, the fixed position could be the position where the reference screen index value is 0 or the last position.

[0627] When the target screen exists at any position in the list of reference screens, an additional reference screen index indicating such an arbitrary position can be sent by the encoding device 100 to the decoding device 200 by a signal.

[0628] 5) Sub-block merging mode

[0629] Sub-block merging mode can be a mode that derives motion information from sub-blocks of the CU.

[0630] When applying the sub-block merging mode, the motion information of col-sub-blocks of the target sub-block in the reference image (i.e., based on the temporal merging candidate of the sub-block) and / or affine control point motion vector merging candidates can be used to generate a list of sub-block merging candidates.

[0631] 6) Triangular partitioning mode

[0632] In the triangular partitioning mode, the target block can be partitioned diagonally, and sub-target blocks generated through partitioning can be produced. For each sub-target block, motion information of the corresponding sub-target block can be exported, and the exported motion information can be used to derive the prediction samples of each sub-target block. The prediction samples of the target block can be derived by weighted summing of the prediction samples of the sub-target blocks generated through partitioning.

[0633] 7) Combined inter-frame and intra-frame prediction modes

[0634] The combined inter-frame-intra-frame prediction mode can be a mode that uses a weighted sum of prediction samples generated via inter-frame prediction and prediction samples generated via intra-frame prediction to derive prediction samples for the target block.

[0635] In the above mode, the decoding device 200 can autonomously correct the derived motion information. For example, the decoding device 200 can search for motion information with the minimum sum of absolute differences (SAD) in a specific region based on a reference block indicated by the derived motion information, and can derive the found motion information as corrected motion information.

[0636] In the above mode, the decoding device 200 can use optical flow to compensate for the prediction samples derived via inter-frame prediction.

[0637] In the AMVP mode, merge mode, skip mode, etc. described above, the index information of the list can be used to specify the motion information among multiple motion information in the list that will be used to predict the target block.

[0638] To improve coding efficiency, the coding device 100 can generate an index of the element with the minimum cost in inter-frame prediction of the target block using only the elements in the signal transmission list. The coding device 100 can encode this index and can transmit the encoded index via signal transmission.

[0639] Therefore, the encoding device 100 and the decoding device 200 must be able to derive the lists described above (i.e., the candidate list for predicted motion vectors and the candidate list for merging) using the same scheme and based on the same data. Here, the same data may include reconstructed frames and reconstructed blocks. Furthermore, in order to specify elements using indices, the order of elements in the list must be fixed.

[0640] Figure 10 Spatial candidates according to an embodiment are shown.

[0641] exist Figure 10 The image shows the locations of the spatial candidates.

[0642] The large block in the center of the graph represents the target block. The five smaller blocks represent spatial candidates.

[0643] The coordinates of the target block can be (xP, yP), and the size of the target block can be represented by (nPSW, nPSH).

[0644] Spatial candidate A0 can be a block adjacent to the lower left corner of the target block. A0 can be a block that occupies the pixel located at coordinates (xP-1, yP+nPSH).

[0645] Spatial candidate A1 can be the block that is adjacent to the left side of the target block. A1 can be the bottommost block among the blocks that are adjacent to the left side of the target block. Alternatively, A1 can be the block that is adjacent to the top side of A0. A1 can be the block that occupies the pixel located at coordinates (xP-1, yP+nPSH-1).

[0646] Spatial candidate B0 can be the block adjacent to the top right corner of the target block. B0 can be a block that occupies the pixel located at coordinates (xP+nPSW, yP-1).

[0647] Spatial candidate B1 can be the block that is top-adjacent to the target block. B1 can be the rightmost block among the blocks that are top-adjacent to the target block. Optionally, B1 can be the block that is left-adjacent to B0. B1 can be the block that occupies the pixel located at coordinates (xP+nPSW-1, yP-1).

[0648] Spatial candidate B2 can be a block adjacent to the top-left corner of the target block. B2 can be a block that occupies the pixel located at coordinates (xP-1, yP-1).

[0649] Determining the availability of spatial and temporal candidates

[0650] In order to include spatial or temporal motion information in the list, it must be determined whether the spatial or temporal motion information is available.

[0651] In the following text, candidate blocks may include spatial candidates and temporal candidates.

[0652] For example, the determination can be performed by sequentially applying steps 1) through 4).

[0653] Step 1) When the PU including the candidate block is outside the boundary of the screen, the availability of the candidate block can be set to "false". The expression "availability is set to false" can have the same meaning as "set to unavailable".

[0654] Step 2) When the PU including the candidate block is outside the boundary of the stripe, the availability of the candidate block can be set to "false". When the target block and the candidate block are in different stripes, the availability of the candidate block can be set to "false".

[0655] Step 3) When the PU including the candidate block is outside the boundary of the parallel block, the availability of the candidate block can be set to "false". When the target block and the candidate block are in different parallel blocks, the availability of the candidate block can be set to "false".

[0656] Step 4) When the prediction mode of the PU including the candidate block is intra-frame prediction mode, the availability of the candidate block can be set to "false". When the PU including the candidate block does not use inter-frame prediction, the availability of the candidate block can be set to "false".

[0657] Figure 11 The order in which motion information of spatial candidates is added to the merging list is shown according to an embodiment.

[0658] like Figure 11As shown, when multiple motion information entries from spatial candidates are added to the merge list, the order A1, B1, B0, A0, and B2 can be used. In other words, multiple motion information entries from available spatial candidates can be added to the merge list in the order A1, B1, B0, A0, and B2.

[0659] Methods for exporting merge lists in merge mode and skip mode

[0660] As described above, the maximum number of merge candidates in the merge list can be set. The maximum number can be indicated by "N". The set number can be sent from the encoding device 100 to the decoding device 200. The stripe header can include N. In other words, the maximum number of merge candidates in the merge list for the target block of the stripe can be set via the stripe header. For example, the value of N can essentially be 5.

[0661] Multiple motion information (i.e., merge candidates) can be added to the merge list in the order of steps 1) to 4).

[0662] Step 1) Within the space candidates, available space candidates can be added to the merge list. This can be done by... Figure 11 The order shown in the diagram is used to add multiple motion information entries from available space candidates to the merge list. Here, if the motion information from an available space candidate overlaps with other motion information already existing in the merge list, the motion information from the available space candidate may not be added to the merge list. The operation of checking whether the corresponding motion information overlaps with other motion information existing in the list can be simply referred to as "overlap check".

[0663] The maximum number of motion information entries that can be added is N.

[0664] Step 2) When the number of motion information entries in the merge list is less than N and time candidates are available, the motion information of the time candidates can be added to the merge list. However, if the available motion information of time candidates overlaps with other motion information already existing in the merge list, the available motion information of time candidates may not be added to the merge list.

[0665] Step 3) When the number of motion information entries in the merge list is less than N and the target strip type is "B", the combined motion information generated by combined bidirectional prediction (bidirectional prediction) can be added to the merge list.

[0666] The target strip can be a strip that includes the target block.

[0667] Combined motion information can be a combination of L0 motion information and L1 motion information. L0 motion information can be motion information that only refers to the L0 reference frame list. L1 motion information can be motion information that only refers to the L1 reference frame list.

[0668] The merged list may contain one or more L0 motion entries. Additionally, the merged list may contain one or more L1 motion entries.

[0669] Combined motion information may include one or more pieces of combined motion information. When generating combined motion information, the L0 motion information and L1 motion information that will be used in the step of generating combined motion information can be predefined from the one or more L0 motion information and the one or more L1 motion information. One or more pieces of combined motion information can be generated in a predefined order via bidirectional prediction using a pair of different motion information from a merge list. One of the different motion information in the pair can be L0 motion information, and the other of the different motion information in the pair can be L1 motion information.

[0670] For example, the combined motion information with the highest priority can be a combination of L0 motion information with a merge index of 0 and L1 motion information with a merge index of 1. When the motion information with a merge index of 0 is not L0 motion information, or when the motion information with a merge index of 1 is not L1 motion information, neither combined motion information is generated nor added. Next, the combined motion information with the next higher priority can be a combination of L0 motion information with a merge index of 1 and L1 motion information with a merge index of 0. Subsequent detailed combinations can conform to other combinations in the field of video encoding / decoding.

[0671] Here, when the combined motion information overlaps with other motion information already existing in the merge list, the combined motion information may not be added to the merge list.

[0672] Step 4) When the number of motion information entries in the merge list is less than N, the motion information of the zero vector can be added to the merge list.

[0673] Zero-vector motion information can be motion information where the motion vector is zero.

[0674] The number of zero-vector motion information entries can be one or more. The reference frame indices for one or more zero-vector motion information entries can be different from each other. For example, the reference frame index value for the first zero-vector motion information entry can be 0. The reference frame index value for the second zero-vector motion information entry can be 1.

[0675] The number of zero-vector motion information entries can be the same as the number of reference frames in the reference frame list.

[0676] The reference direction for zero-vector motion information can be bidirectional. Both motion vectors can be zero vectors. The number of zero-vector motion information entries can be the smaller of the number of reference frames in reference frame list L0 and the number of reference frames in reference frame list L1. Optionally, when the number of reference frames in reference frame list L0 and the number of reference frames in reference frame list L1 are different from each other, a unidirectional reference direction can be used for reference frame indexing that can be applied to only a single reference frame list.

[0677] The encoding device 100 and / or the decoding device 200 may subsequently add zero-vector motion information to the merge list while changing the reference frame index.

[0678] When zero-vector motion information overlaps with other motion information already existing in the merge list, the zero-vector motion information may not be added to the merge list.

[0679] The order of steps 1) to 4) above is merely exemplary and can be changed. Furthermore, some steps in the above steps may be omitted based on predefined conditions.

[0680] Method for exporting a candidate list of predicted motion vectors in AMVP mode

[0681] The maximum number of predicted motion vector candidates in the candidate list can be predefined. N can be used to indicate the predefined maximum number. For example, the predefined maximum number could be 2.

[0682] Multiple pieces of motion information (i.e., predicted motion vector candidates) can be added to the predicted motion vector candidate list in the order of steps 1) to 3).

[0683] Step 1) Available spatial candidates can be added to the list of predicted motion vector candidates. Spatial candidates may include a first spatial candidate and a second spatial candidate.

[0684] The first spatial candidate can be one of A0, A1, scaled A0, and scaled A1. The second spatial candidate can be one of B0, B1, B2, scaled B0, scaled B1, and scaled B2.

[0685] Multiple motion information entries from available spatial candidates can be added to the predicted motion vector candidate list in the order of first spatial candidates and second spatial candidates. In this case, if the motion information of an available spatial candidate overlaps with other motion information already existing in the predicted motion vector candidate list, the motion information of the available spatial candidate may not be added to the predicted motion vector candidate list. In other words, when the value of N is 2, if the motion information of the second spatial candidate is the same as the motion information of the first spatial candidate, then the motion information of the second spatial candidate may not be added to the predicted motion vector candidate list.

[0686] The maximum number of motion information entries that can be added is N.

[0687] Step 2) When the number of motion information entries in the predicted motion vector candidate list is less than N and a time candidate is available, the motion information of the time candidate can be added to the predicted motion vector candidate list. In this case, if the motion information of the available time candidate overlaps with other motion information already existing in the predicted motion vector candidate list, the available time candidate may not be added to the predicted motion vector candidate list.

[0688] Step 3) When the number of motion information entries in the candidate list of predicted motion vectors is less than N, zero-vector motion information can be added to the candidate list of predicted motion vectors.

[0689] Zero-vector motion information may include one or more zero-vector motion information pieces. The reference frame indices of the one or more zero-vector motion information pieces may be different from each other.

[0690] The encoding device 100 and / or the decoding device 200 can sequentially add multiple zero-vector motion information to the candidate list of predicted motion vectors while changing the reference frame index.

[0691] When zero-vector motion information overlaps with other motion information already existing in the candidate list of predicted motion vectors, the zero-vector motion information may not be added to the candidate list of predicted motion vectors.

[0692] The description of zero-vector motion information presented above, combined with the merged list, can also be applied to zero-vector motion information. Repeated descriptions will be omitted.

[0693] The order of steps 1) to 3) described above is merely exemplary and can be changed. Furthermore, some steps may be omitted based on predefined conditions.

[0694] Figure 12 The transformation and quantization processes are shown based on the example.

[0695] like Figure 12As shown, quantization levels can be generated by performing transformation and / or quantization processing on the residual signal.

[0696] The residual signal can be generated as the difference between the original block and the predicted block. Here, the predicted block can be a block generated via intra-frame prediction or inter-frame prediction.

[0697] The residual signal can be transformed into a signal in the frequency domain through a transformation process that is part of the quantization process.

[0698] The transform kernel used for the transform can include various DCT kernels, such as Discrete Cosine Transform (DCT) Type 2 (DCT-II) kernel and Discrete Sine Transform (DST) kernel.

[0699] These transform kernels can perform separable or two-dimensional (2D) non-separable transforms on the residual signal. A separable transform can be a transform indicating that a one-dimensional (1D) transform is performed on the residual signal in each of the horizontal and vertical directions.

[0700] In addition to DCT-II, the DCT and DST types adaptively used for 1D transformations may also include DCT-V, DCT-VIII, DST-I, and DST-VII, as shown in each of Tables 3 and 4 below.

[0701] Table 3

[0702]

[0703]

[0704] Table 4

[0705] Transform set Transform candidate 0 DST-VII, DCT-VIII, DST-I 1 DST-VII, DST-I, DCT-VIII 2 DST-VII, DCT-V, DST-I

[0706] As shown in Tables 3 and 4, transform sets can be used when deriving the DCT or DST type to be used for the transform. Each transform set can include multiple transform candidates. Each transform candidate can be a DCT type or a DST type.

[0707] Table 5 below shows examples of the transform sets that will be applied in the horizontal direction and the transform sets that will be applied in the vertical direction according to the intra-frame prediction mode.

[0708] Table 5

[0709] Intra-prediction mode 0 1 2 3 4 5 6 7 8 9 Vertical Transformation Set 2 1 0 1 0 1 0 1 0 1 Horizontal transformation set 2 1 0 1 0 1 0 1 0 1 Intra-prediction mode 10 11 12 13 14 15 16 17 18 19 Vertical Transformation Set 0 1 0 1 0 0 0 0 0 0 Horizontal transformation set 0 1 0 1 2 2 2 2 2 2 Intra-prediction mode 20 21 22 23 24 25 26 27 28 29 Vertical Transformation Set 0 0 0 1 0 1 0 1 0 1 Horizontal transformation set 2 2 2 1 0 1 0 1 0 1 Intra-prediction mode 30 31 32 33 34 35 36 37 38 39 Vertical Transformation Set 0 1 0 1 0 1 0 1 0 1 Horizontal transformation set 0 1 0 1 0 1 0 1 0 1 Intra-prediction mode 40 41 42 43 44 45 46 47 48 49 Vertical Transformation Set 0 1 0 1 0 1 2 2 2 2 Horizontal transformation set 0 1 0 1 0 1 0 0 0 0 Intra-prediction mode 50 51 52 53 54 55 56 57 58 59 Vertical Transformation Set 2 2 2 2 2 1 0 1 0 1 Horizontal transformation set 0 0 0 0 0 1 0 1 0 1 Intra-prediction mode 60 61 62 63 64 65 66 Vertical Transformation Set 0 1 0 1 0 1 0 Horizontal transformation set 0 1 0 1 0 1 0

[0710] Table 5 shows the numbers of the vertical transform set and horizontal transform set that will be applied to the horizontal direction of the residual signal according to the intra-frame prediction mode of the target block.

[0711] As illustrated in Table 5, the transform sets to be applied in the horizontal and vertical directions can be predefined based on the intra-prediction mode of the target block. Encoding device 100 can use transforms included in the transform set corresponding to the intra-prediction mode of the target block to perform transforms and inverse transforms on the residual signal. Furthermore, decoding device 200 can use transforms included in the transform set corresponding to the intra-prediction mode of the target block to perform an inverse transform on the residual signal.

[0712] In the transform and inverse transform, as illustrated in Tables 3, 4, and 5, the set of transforms to be applied to the residual signal can be determined and may not be transmitted as a signal. Transform indication information can be transmitted from the encoding device 100 to the decoding device 200 as a signal. The transform indication information may be information indicating which of the multiple transform candidates included in the set of transforms to be applied to the residual signal is used.

[0713] For example, when the target block size is 64×64 or smaller, a transform set with three transforms can be configured separately according to the intra-frame prediction mode. The optimal transform method can be selected from a total of nine multi-transform methods generated by combinations of three transforms in the horizontal direction and three transforms in the vertical direction. With such an optimal transform method, the residual signal can be encoded and / or decoded, thus improving coding efficiency.

[0714] Here, information indicating which of the multiple transformations belonging to each transform set has been used for at least one of the vertical and horizontal transformations can be entropy encoded and / or entropy decoded. Here, truncated univariate binarization can be used to encode and / or decode such information.

[0715] As mentioned above, various transformation methods can be applied to residual signals generated via intra-frame prediction or inter-frame prediction.

[0716] The transformation may include at least one of a first transformation and a second transformation. Transform coefficients can be generated by performing a first transformation on the residual signal, and second transformation coefficients can be generated by performing a second transformation on the transform coefficients.

[0717] The first transformation can be referred to as the "primary transformation". Furthermore, the first transformation can also be referred to as the "Adaptive Multitransformation (AMT) scheme". As mentioned above, AMT can represent applying different transformations to various 1D directions (i.e., the vertical and horizontal directions).

[0718] A secondary transformation can be a transformation used to increase the energy concentration of the transformation coefficients generated by the first transformation. Similar to the first transformation, a secondary transformation can be a separable transformation or a non-separable transformation. Such a non-separable transformation can be a non-separable secondary transformation (NSST).

[0719] The first transformation can be performed using at least one of a predefined plurality of transformation methods. For example, the predefined plurality of transformation methods may include the Discrete Cosine Transform (DCT), the Discrete Sine Transform (DST), the Karhunen-Loeve Transform (KLT), etc.

[0720] Furthermore, depending on the kernel function defined for the Discrete Cosine Transform (DCT) or Discrete Sine Transform (DST), the first transform can be of various types.

[0721] For example, the transform type can be determined based on at least one of the following: 1) the prediction mode of the target block (e.g., one of intra-frame prediction and inter-frame prediction), 2) the size of the target block, 3) the shape of the target block, 4) the intra-frame prediction mode of the target block, 5) the components of the target block (e.g., one of luma component and chroma component), and 6) the partition type applied to the target block (e.g., one of quadtree, binary tree and ternary tree).

[0722] For example, based on the transform kernels presented in Table 6 below, the first transform may include transforms such as DCT-2, DCT-5, DCT-7, DST-7, DST-1, DST-8, and DCT-8. Table 6 below illustrates various transform types and transform kernel functions used for Multiple Transform Selection (MTS).

[0723] MTS can refer to the selection of a combination of one or more DCT and / or DST cores to transform the residual signal in the horizontal and / or vertical directions.

[0724] Table 6

[0725]

[0726] In Table 6, i and j can be integer values ​​that are equal to or greater than 0 and less than or equal to N-1.

[0727] A secondary transformation can be performed on the transformation coefficients generated by performing the first transformation.

[0728] For example, in the first transformation, a transformation set can also be defined in the secondary transformation. The methods used to derive and / or determine the above transformation set can be applied not only to the first transformation but also to the secondary transformation.

[0729] The first and second transformations can be determined for a specific target.

[0730] For example, the first and second transformations can be applied to signal components corresponding to one or more of the luma and chroma components. Whether to apply the first and / or second transformations can be determined based on at least one of the coding parameters for the target block and / or neighboring blocks. For example, whether to apply the first and / or second transformations can be determined based on the size and / or shape of the target block.

[0731] In the encoding device 100 and the decoding device 200, transformation information indicating the transformation method to be used for the target can be derived by using specified information.

[0732] For example, the transformation information may include transformation indices that will be used for primary and / or secondary transformations. Optionally, the transformation information may indicate that primary and / or secondary transformations are not used.

[0733] For example, when the target of the primary and secondary transforms is a target block, the transform method to be applied to the primary and / or secondary transforms, as indicated by the transform information, can be determined based on at least one of the encoding parameters for the target block and / or blocks adjacent to the target block.

[0734] Optionally, the encoding device 100 may send transformation information indicating the transformation method for a specific target to the decoding device 200 via a signal.

[0735] For example, for a single CU, the decoding device 200 can derive transformation information such as whether a primary transformation is used, the index indicating the primary transformation, whether a secondary transformation is used, and the index indicating the secondary transformation. Alternatively, for a single CU, transformation information indicating the following can be transmitted via signals: whether a primary transformation is used, the index indicating the primary transformation, whether a secondary transformation is used, and the index indicating the secondary transformation.

[0736] Quantized transform coefficients (i.e., quantization levels) can be generated by quantizing the result produced by performing a first transform and / or a secondary transform, or by quantizing the residual signal.

[0737] Figure 13 This shows a diagonal scan based on an example.

[0738] Figure 14 The horizontal scan is shown based on the example.

[0739] Figure 15 The vertical scan is shown according to the example.

[0740] The quantized transform coefficients can be scanned via at least one of (top right) diagonal scan, vertical scan, and horizontal scan, based on at least one of intra-frame prediction mode, block size, and block shape. The block can be a transform unit (TU).

[0741] Each scan can be started at a specific start point and terminated at a specific end point.

[0742] For example, by using Figure 13 A diagonal scan is used to scan the coefficients of the block to transform the quantized transform coefficients into a 1D vector form. Optionally, this can be used depending on the block size and / or intra-frame prediction mode. Figure 14 Horizontal scan or Figure 15 It uses vertical scanning instead of diagonal scanning.

[0743] A vertical scan can be an operation that scans 2D block-type coefficients in the column direction. A horizontal scan can be an operation that scans 2D block-type coefficients in the row direction.

[0744] In other words, the choice between diagonal, vertical, and horizontal scans can be determined based on the block size and / or inter-frame prediction mode.

[0745] like Figure 13 , Figure 14 and Figure 15 As shown, the quantized transform coefficients can be scanned along the diagonal, horizontal, or vertical direction.

[0746] The quantized transformation coefficients can be represented by block shapes. Each block can include multiple sub-blocks. Each sub-block can be defined based on either the minimum block size or the minimum block shape.

[0747] During scanning, the scanning order, based on the type or direction of the scan, can be applied first to the sub-blocks. Furthermore, the scanning order, based on the direction of the scan, can be applied to the quantized transform coefficients within each sub-block.

[0748] For example, such as Figure 13 , Figure 14 and Figure 15 As shown, when the target block size is 8×8, quantized transform coefficients can be generated by a first transform, a second transform, and quantization of the residual signal of the target block. Therefore, one of the three types of scan sequences can be applied to four 4×4 sub-blocks, and the quantized transform coefficients can be scanned for each 4×4 sub-block according to the scan sequence.

[0749] The encoding device 100 can generate entropy-coded quantized transform coefficients by performing entropy coding on scanned quantized transform coefficients, and can generate a bit stream including the entropy-coded quantized transform coefficients.

[0750] The decoding device 200 can extract entropy-encoded quantized transform coefficients from the bitstream, and can generate quantized transform coefficients by performing entropy decoding on the entropy-encoded quantized transform coefficients. The quantized transform coefficients can be arranged in a 2D block format via inverse scanning. Here, as a method of inverse scanning, at least one of upper right diagonal scanning, vertical scanning, and horizontal scanning can be performed.

[0751] In the decoding device 200, inverse quantization can be performed on the quantized transform coefficients. A secondary inverse transform can be performed on the result generated by inverse quantization, depending on whether a secondary inverse transform is performed. Furthermore, a first inverse transform can be performed on the result generated by the secondary inverse transform, depending on whether a first inverse transform will be performed. The reconstructed residual signal can be generated by performing a first inverse transform on the result generated by the secondary inverse transform.

[0752] For the luminance component reconstructed via intra-frame prediction or inter-frame prediction, an inverse mapping with dynamic range can be performed before loop filtering.

[0753] The dynamic range can be divided into 16 equal segments, and the mapping functions for the corresponding segments can be signaled. These mapping functions can be signaled at the stripe level or the parallel block group level.

[0754] An inverse mapping function can be derived from the mapping function to perform the inverse mapping.

[0755] Loop filtering, reference frame storage, and motion compensation can be performed in the inverse mapping region.

[0756] Predicted blocks generated via inter-frame prediction can be transformed to a mapped region using a mapping function, and the transformed predicted blocks can be used to generate reconstructed blocks. However, since intra-frame prediction is performed in the mapped region, predicted blocks generated via intra-frame prediction can be used to generate reconstructed blocks without requiring mapping and / or inverse mapping.

[0757] For example, when the target block is a residual block of the chrominance component, the residual block can be transformed into the inverse mapping region by scaling the chrominance component of the mapping region.

[0758] Scaling availability can be signaled at the stripe level or the parallel block group level.

[0759] For example, scaling can be applied only when the mapping is available for the luminance component and the partitions for the luminance and chrominance components follow the same tree structure.

[0760] Scaling can be performed based on the average value of the samples in the luminance prediction block corresponding to the chrominance prediction block. Here, when the target block uses inter-frame prediction, the luminance prediction block can represent the mapped luminance prediction block.

[0761] The scaling values ​​can be derived by using an index-referenced lookup table of the segment to which the average of the sample values ​​of the brightness prediction block belongs.

[0762] The residual block can be transformed into the inverse mapping region by scaling the residual block using the final derived values. Subsequently, for blocks of the chroma components, reconstruction, intra-frame prediction, inter-frame prediction, loop filtering, and storage of reference frames can be performed in the inverse mapping region.

[0763] For example, information indicating whether the mapping and / or inverse mapping of the luminance and chrominance components is available can be sent via a sequence parameter set using signals.

[0764] A predicted block for the target block can be generated based on a block vector. The block vector indicates the displacement between the target block and a reference block. The reference block can be a block in the target image.

[0765] In this way, the prediction mode that generates prediction blocks by referencing the target image can be called the "intra-block copy (IBC) mode".

[0766] The IBC mode can be applied to CUs with specific dimensions. For example, the IBC mode can be applied to an M×N CU. Here, M and N can be less than or equal to 64.

[0767] IBC modes can include skip mode, merge mode, AMVP mode, etc. In skip mode or merge mode, a merge candidate list can be configured, and the merge index is signaled, allowing a single merge candidate to be specified from among the existing merge candidates in the merge candidate list. The block vector of the specified merge candidate can be used as the block vector of the target block.

[0768] In AMVP mode, the differential block vector can be signaled. Additionally, the prediction block vector can be derived from the target block's left and top neighboring blocks. Furthermore, the index of which neighboring block will be used can be signaled.

[0769] In IBC mode, the predicted block can be included in the target CTU or the left CTU, and can be limited to blocks within the previously reconstructed region. For example, the value of the block vector can be restricted such that the predicted block of the target block is located in a specific region. The specific region can be defined by three 64×64 blocks that are encoded and / or decoded before the 64×64 block including the target block. Restricting the value of the block vector in this way reduces memory consumption and device complexity caused by the implementation of IBC mode.

[0770] Figure 16 This is a configuration diagram of the encoding apparatus according to an embodiment.

[0771] Encoding device 1600 may correspond to encoding device 100 described above.

[0772] Encoding device 1600 may include a processing unit 1610, a memory 1630, a user interface (UI) input device 1650, a UI output device 1660, and a storage 1640 that communicate with each other via a bus 1690. Encoding device 1600 may also include a communication unit 1620 connected to a network 1699.

[0773] The processing unit 1610 may be a central processing unit (CPU) or a semiconductor device for executing processing instructions stored in memory 1630 or storage 1640. The processing unit 1610 may be at least one hardware processor.

[0774] The processing unit 1610 can generate and process signals, data, or information input to, output from, or used in the encoding device 1600, and can perform checks, comparisons, determinations, etc., related to the signals, data, or information. In other words, in this embodiment, the processing unit 1610 can perform the generation and processing of data or information, as well as the checks, comparisons, and determinations related to the data or information.

[0775] The processing unit 1610 may include an inter-frame prediction unit 110, an intra-frame prediction unit 120, a switcher 115, a subtractor 125, a transform unit 130, a quantization unit 140, an entropy coding unit 150, an inverse quantization unit 160, an inverse transform unit 170, an adder 175, a filter unit 180, and a reference frame buffer 190.

[0776] At least some of the following components—inter-frame prediction unit 110, intra-frame prediction unit 120, switcher 115, subtractor 125, transform unit 130, quantization unit 140, entropy coding unit 150, inverse quantization unit 160, inverse transform unit 170, adder 175, filter unit 180, and reference frame buffer 190—may be program modules and capable of communicating with external devices or systems. These program modules may be included in the encoding device 1600 in the form of an operating system, application module, or other program modules.

[0777] The program modules can be physically stored in various types of known storage devices. Furthermore, at least some of the program modules can also be stored in a remote storage device capable of communicating with the encoding device 1600.

[0778] Program modules may include, but are not limited to, routines, subroutines, programs, objects, components, and data structures for performing functions or operations according to an embodiment or for implementing abstract data types according to an embodiment.

[0779] The program module can be implemented using instructions or code that are executed by at least one processor of the encoding device 1600.

[0780] The processing unit 1610 can execute instructions or codes in the inter-frame prediction unit 110, intra-frame prediction unit 120, switcher 115, subtractor 125, transform unit 130, quantization unit 140, entropy coding unit 150, dequantization unit 160, inverse transform unit 170, adder 175, filter unit 180 and reference frame buffer 190.

[0781] The storage unit may represent memory 1630 and / or storage 1640. Each of memory 1630 and storage 1640 may be any of a variety of volatile or non-volatile storage media. For example, memory 1630 may include at least one of read-only memory (ROM) 1631 and random access memory (RAM) 1632.

[0782] The storage unit can store data or information used for the operation of the encoding device 1600. In an embodiment, the data or information of the encoding device 1600 can be stored in the storage unit.

[0783] For example, storage units can store images, blocks, lists, motion information, inter-frame prediction information, bitstreams, etc.

[0784] The encoding device 1600 can be implemented in a computer system that includes a computer-readable storage medium.

[0785] The storage medium may store at least one module required for the operation of the encoding device 1600. The memory 1630 may store at least one module and may be configured such that the at least one module is operated by the processing unit 1610.

[0786] The communication unit 1620 can be used to perform functions related to communication of data or information with the encoding device 1600.

[0787] For example, communication unit 1620 can send a bit stream to decoding device 1700, which will be described later.

[0788] Figure 17 This is a configuration diagram of a decoding apparatus according to an embodiment.

[0789] Decoding device 1700 can correspond to the decoding device 200 described above.

[0790] The decoding device 1700 may include a processing unit 1710, a memory 1730, a user interface (UI) input device 1750, a UI output device 1760, and a storage 1740 that communicate with each other via a bus 1790. The decoding device 1700 may also include a communication unit 1720 connected to a network 1799.

[0791] Processing unit 1710 may be a central processing unit (CPU) or semiconductor device for executing processing instructions stored in memory 1730 or storage 1740. Processing unit 1710 may be at least one hardware processor.

[0792] The processing unit 1710 can generate and process signals, data, or information input to, output from, or used in the decoding device 1700, and can perform checks, comparisons, determinations, etc., related to the signals, data, or information. In other words, in this embodiment, the processing unit 1710 can perform the generation and processing of data or information, as well as the checks, comparisons, and determinations related to the data or information.

[0793] The processing unit 1710 may include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an intra-frame prediction unit 240, an inter-frame prediction unit 250, a switcher 245, an adder 255, a filter unit 260, and a reference frame buffer 270.

[0794] At least some of the entropy decoding unit 210, inverse quantization unit 220, inverse transform unit 230, intra-frame prediction unit 240, inter-frame prediction unit 250, adder 255, switcher 245, filter unit 260, and reference frame buffer 270 of the decoding device 200 may be program modules and are capable of communicating with external devices or systems. These program modules may be included in the decoding device 1700 in the form of an operating system, application module, or other program modules.

[0795] The program modules can be physically stored in various types of known storage devices. Furthermore, at least some of the program modules can also be stored in a remote storage device capable of communicating with the decoding device 1700.

[0796] Program modules may include, but are not limited to, routines, subroutines, programs, objects, components, and data structures for performing functions or operations according to an embodiment or for implementing abstract data types according to an embodiment.

[0797] The program module can be implemented using instructions or code executed by at least one processor of the decoding device 1700.

[0798] The processing unit 1710 can run instructions or codes in the entropy decoding unit 210, the dequantization unit 220, the inverse transform unit 230, the intra-frame prediction unit 240, the inter-frame prediction unit 250, the switcher 245, the adder 255, the filter unit 260, and the reference frame buffer 270.

[0799] The storage unit may represent memory 1730 and / or storage 1740. Each of memory 1730 and storage 1740 may be any of a variety of volatile or non-volatile storage media. For example, memory 1730 may include at least one of ROM 1731 and RAM 1732.

[0800] The storage unit can store data or information used for the operation of the decoding device 1700. In an embodiment, the data or information of the decoding device 1700 can be stored in the storage unit.

[0801] For example, storage units can store images, blocks, lists, motion information, inter-frame prediction information, bitstreams, etc.

[0802] The decoding device 1700 can be implemented in a computer system that includes a computer-readable storage medium.

[0803] The storage medium may store at least one module required for the operation of the decoding device 1700. The memory 1730 may store at least one module and may be configured such that the at least one module is operated by the processing unit 1710.

[0804] The communication unit 1720 can be used to perform functions related to communication of data or information with the decoding device 1700.

[0805] For example, communication unit 1720 can receive bit streams from encoding device 1600.

[0806] In the following text, "processing unit" may refer to processing unit 1610 of encoding device 1600 and / or processing unit 1710 of decoding device 1700. For example, regarding prediction-related functions, the processing unit may represent switch 115 and / or switch 245. Regarding inter-frame prediction-related functions, the processing unit may represent inter-frame prediction unit 110, subtractor 125, and adder 175, and may also represent inter-frame prediction unit 250 and adder 255. Regarding intra-frame prediction-related functions, the processing unit may represent intra-frame prediction unit 120, subtractor 125, and adder 175, and may also represent intra-frame prediction unit 240 and adder 255. Regarding transform-related functions, the processing unit may represent transform unit 130 and inverse transform unit 170, and may also represent inverse transform unit 230. Regarding quantization-related functions, the processing unit may represent quantization unit 140 and inverse quantization unit 160, and may also indicate inverse quantization unit 220. Regarding functions related to entropy encoding and / or entropy decoding, the processing unit may represent entropy encoding unit 150 and / or entropy decoding unit 210. Regarding functions related to filtering, the processing unit may represent filter unit 180 and / or filter unit 260. Regarding functions related to reference frames, the processing unit may instruct reference frame buffer 190 and / or reference frame buffer 270.

[0807] Based on the foregoing description, a prediction method based on bilateral matching (BM) according to this disclosure will be described in detail. The terms used in the embodiments included in this disclosure are defined as follows.

[0808] Target block: The block that will be encoded / decoded. It can be referred to as the current block.

[0809] Target image: The image to which the block to be encoded / decoded belongs. It can be referred to as the current image.

[0810] BM Reference Image List: A list of reference images used in BM-based encoding / decoding.

[0811] BM Reference Image List 0 (or First BM Reference Image List): A list of reference images used in predictions of BM using P-stripes or the first of two lists of reference images used in predictions of BM using B-stripes.

[0812] BM Reference Image List 1 (or Second BM Reference Image List): The second of two lists of reference images used in BM prediction using B-strips.

[0813] BM search area: The area being searched to find the BM reference block.

[0814] BM reference block: The block in the BM reference image that has the highest correlation with the target block. It can refer to the block that is most similar to the first BM candidate reference block and the second candidate reference block.

[0815] First BM reference block: The BM reference block in the first BM reference image.

[0816] Second BM reference block: BM reference block in the second BM reference image.

[0817] BM motion vector: the motion difference between the BM reference block and the target block.

[0818] BM Optimal Block: The value calculated by weighted combination or weighted summation of the first BM reference block and the second BM reference block. It can refer to the BM prediction signal for the target block.

[0819] Information on the encoding method based on BM: It includes at least one of the syntax elements proposed in this invention.

[0820] In the above definition, terms prefixed with "first" can be information about either direction L0 or L1, and terms prefixed with "second" can be information about the other direction. Optionally, at least one of the terms prefixed with "first" or "second" can refer to the target image to which the target block belongs or information within the target image.

[0821] According to this disclosure, BM can be one of the prediction methods for obtaining the predicted block of the current block. The BM-based prediction method will be described in detail below.

[0822] Figure 18 This is a flowchart of a method for obtaining a prediction block based on the BM method according to an embodiment of the present disclosure.

[0823] It can be omitted Figure 18 The BM method is executed by at least one of steps [D1] to [D3] shown. Optionally, the decision to execute at least one of steps [D1] to [D3] may be based on at least one of the following: encoding parameters, frame information, stripe information, parallel block information, quantization parameters (QP), code block flag (CBF), block size, block depth, block shape, entropy coding method, block size, block shape (square, rectangle), or time level. In this case, the block may be at least one of a coding tree block, a coding block, a prediction block, a transform block, or a block of a predetermined size.

[0824] When encoding / decoding the current block, you can use Figure 18 The prediction method based on the BM method shown in the figure obtains the prediction block of the current block.

[0825] [D1] Steps for selecting a BM reference image

[0826] To perform predictions based on the BM method, a BM reference image can be selected from a list of BM reference images. In this case, "BM reference image" indicates the reference image used for predictions using BM. BM reference images can be selected implicitly or explicitly.

[0827] In addition, the "BM reference image list" represents a list of reference images used for predictions using BM. The BM reference image list can be indicated as "bm_ref_pic_list".

[0828] For each of the L0 and L1 directions, a BM reference image list can be configured. In this case, one of the L0 BM reference image list and the L1 reference image list can be referred to as BM reference image list 0. In other words, the first of the two reference image lists used in the prediction of BM using P-stripes or in the prediction of BM using B-stripes can be referred to as "BM reference image list 0 (bm_ref_pic_lists[0])".

[0829] The BM reference image list 0 may include at least one of the target image, a decoded past image, or a decoded future image.

[0830] In addition, another one in the L0 BM reference image list and the L1 reference image list can be called BM reference image list 1. In other words, the second list of the two reference image lists used in the prediction of BM using B-strips can be called "BM reference image list 1 (bm_ref_pic_lists[1])".

[0831] The BM reference image list 1 may include at least one of the target image, a decoded past image, or a decoded future image.

[0832] Furthermore, after determining whether to perform BM-based prediction on the target block, if it is determined that BM-based prediction is not performed, it can be determined whether to perform general inter-frame prediction on the target block.

[0833] In this case, the BM reference image list can be generated separately from the reference image list used for general inter-frame prediction. As an example, when performing prediction based on the BM method, the BM reference image list can be used to select the BM reference image, while when performing general inter-frame prediction, the general reference image list can be used to select the reference image.

[0834] Optionally, when inter-frame prediction is applied to a target block, it can be determined whether to perform BM-based prediction on the target block. In this case, the BM reference image list can be set to be the same as the reference image list used to perform inter-frame prediction.

[0835] As another example, the BM prediction image for the target block in the images belonging to BM reference image list 0 can be referred to as the "first BM reference image". Additionally, the BM prediction image for the target block in the images belonging to BM reference image list 1 can be referred to as the "second BM reference image".

[0836] To perform predictions based on the BM method, at least one BM reference image can be selected. As an example, the predicted block of the target block can be derived using one BM reference image, or the predicted block of the target block can be derived using a first BM reference image and a second BM reference image.

[0837] Furthermore, the index that identifies each reference image belonging to the BM reference image list can be called the BM reference index (bm_ref_idx). In this case, the BM reference index of the first BM reference image can be called bm_ref_idx_l0, and the BM reference index of the second BM reference image can be called bm_ref_idx_l1.

[0838] The BM reference index used to identify the BM reference image for predicting the target block within the BM reference image list can be explicitly encoded and transmitted using signals. In other words, for the target block, at least one of bm_ref_idx_l0 and bm_ref_idx_l1 can be encoded and transmitted using signals.

[0839] The first BM reference index (i.e., bm_ref_idx_l0) used to identify the first BM reference image and the second BM reference index (i.e., bm_ref_idx_l1) used to identify the second BM reference image can be encoded and transmitted with signals respectively.

[0840] Optionally, only the first BM reference index for the first BM reference image can be signaled, and the signaling of the second BM reference index for selecting the second BM reference image can be omitted. In this case, a BM reference image whose time direction is opposite to that of the first BM reference image and whose distance from the target image is the same as that of the first BM reference image can be set as the second BM reference image.

[0841] As another example, a BM reference image can be selected based on a reference index used for general inter-frame prediction (general reference index). For example, the BM reference index can be derived to the same value as the general reference index that identifies one of the reference images. In other words, bm_ref_idx_l0 can be derived to the same value as ref_idx_l0 that identifies one of the reference images included in the L0 reference image list, and bm_ref_idx_l1 can be derived to the same value as ref_idx_l1 that identifies one of the reference images included in the L1 reference image list.

[0842] When the BM reference image list is distinguished from the general reference image list, the BM reference index can be encoded / decoded. On the other hand, when the general reference image list is used as the BM reference image list, the general reference index can be encoded / decoded, and the BM reference index can be derived to the same value as the general reference index.

[0843] Therefore, the reference image used for general inter-frame prediction can be set as the BM reference image. As an example, when a first reference image from a first reference image list and a second reference image from a second reference image list are selected for performing inter-frame prediction of the target block, the first reference image and the second reference image can be set as the first BM reference image and the second BM reference image, respectively.

[0844] As another example, a first BM reference image can be derived as an image in the decoded images that is temporally earlier than the target image, and a second BM reference image can be derived as an image in the decoded images that is temporally later than the target image. Here, a past image represents a reference image with a frame order count (POC) less than the target image, and a future image represents a reference image with a POC greater than the target image. As an example, the first BM reference image can be derived as a reference image that is temporally immediately preceding the target image, or a reference image among reference images with a POC less than the target image that has the smallest POC difference with the target image. The second BM reference image can be derived as a reference image that is temporally immediately following the target image, or a reference image among reference images with a POC greater than the target image that has the smallest POC difference with the target image.

[0845] As another example, the target image can be set as the BM reference image. In other words, at least one of the first BM reference image or the second BM reference image can be the target image.

[0846] Furthermore, the BM reference image can be determined by the prediction mode of the target block. Here, the prediction mode can be at least one of intra-frame prediction, IBC prediction, or inter-frame prediction. As an example, when the prediction mode of the target block is intra-frame prediction, the target image can be selected as the BM reference image.

[0847] Optionally, when the prediction mode of the target block is IBC prediction, the target image can be selected as the BM reference image.

[0848] Optionally, when the prediction mode of the target block is inter-frame prediction, at least one of a first BM reference image or a second BM reference image can be selected from the reference images decoded before the target image.

[0849] As another example, the BM reference image can be determined by the type (slice_type) of the slice to which the target image belongs. For example, when the slice to which the target block belongs is of type I, the target image can be set as the BM reference image.

[0850] Optionally, when the strip to which the target block belongs is of type B, at least one of a first BM reference image or a second BM reference image can be selected from the reference images decoded before the target image.

[0851] As another example, the encoder / decoder can adaptively select / derive a BM reference image based on at least one of the encoding parameters.

[0852] Let us assume there exist a target image P(t) at time t, decoded past reference images (e.g., P(t-n1), P(t-n2), P(t-n3)), and decoded future reference images (e.g., P(t+n4), P(t+n5), P(t+n6)). In this case, n x This represents the difference in Proof of Conformity (POC) with the target image, and can be a predetermined positive integer. For example, n... x It can be as follows.

[0853] n1=1, n2=2, n3=3, n4=1, n5=2, n6=3

[0854] In this case, the BM reference image list bm_ref_pic_list can be configured as follows.

[0855] bm_ref_pic_list={P(t),P(t-n1),P(t+n4),P(t+n6)}

[0856] Table 7 shows the configuration of the BM reference image list.

[0857] Table 7

[0858]

[0859]

[0860] Optionally, when there exists a target image P(t) at time t, decoded past images (P(t-n1), P(t-n2), P(t-n3)) and decoded future images (P(t+n4), P(t+n5), P(t+n6)), the BM reference image list bm_ref_pic_list can be configured as follows.

[0861] bm_ref_pic_list={P(t-n1),P(t+n4),P(t-n2),P(t+n6)}

[0862] Table 8 shows the configuration of the BM reference image list.

[0863] Table 8

[0864] bm_ref_idx BM reference image 0 <![CDATA[P(t-n1)]]> 1 <![CDATA[P(t+n4)]]> 2 <![CDATA[P(t-n2)]]> 3 <![CDATA[P(t+n6)]]>

[0865] In the examples in Tables 7 and 8, bm_ref_idx represents the index that identifies the BM reference image included in the BM reference image list. Furthermore, the bm_ref_idx of the BM reference image indicating the target block can be encoded and transmitted using a signal.

[0866] Optionally, when there exists a target image P(t) at time t, the decoded past images (P(t-n1), P(t-n2), P(t-n3)) and the decoded future images (P(t+n4), P(t+n5), P(t+n6)), BM reference image list 0 (i.e., bm_ref_pic_lists[0]) and BM reference image list 1 (i.e., bm_ref_pic_lists[1]) can be configured as follows.

[0867] bm_ref_l0={P(t),P(t-n1),P(t-n2),P(t-n3)}

[0868] bm_ref_l1={P(t+n4),P(t+n5),P(t+n6),P(t-n1)}

[0869] Tables 9 and 10 show the configurations of BM reference image list 0 and BM reference image list 1.

[0870] Table 9

[0871]

[0872]

[0873] Table 10

[0874] bm_ref_idx_l1 BM reference image 0 <![CDATA[P(t+n4)]]> 1 <![CDATA[P(t+n5)]]> 2 <![CDATA[P(t+n6)]]> 3 <![CDATA[P(t-n1)]]>

[0875] Optionally, when there exists a target image P(t) at time t, the decoded past images (P(t-n1), P(t-n2), P(t-n3)) and the decoded future images (P(t+n4), P(t+n5), P(t+n6)), BM reference image list 0 (i.e., bm_ref_pic_lists[0]) and BM reference image list 1 (i.e., bm_ref_pic_lists[1]) can be configured as follows.

[0876] bm_ref_l0={P(t-n1),P(t-n2),P(t-n3),P(t+n4)}

[0877] bm_ref_l1={P(t+n4),P(t+n5),P(t+n6),P(t-n1)}

[0878] Tables 11 and 12 show the configurations of BM reference image list 0 and BM reference image list 1.

[0879] Table 11

[0880] bm_ref_idx_l0 BM reference image 0 <![CDATA[P(t-n1)]]> 1 <![CDATA[P(t-n2)]]> 2 <![CDATA[P(t-n3)]]> 3 <![CDATA[P(t+n4)]]>

[0881] Table 12

[0882]

[0883]

[0884] In the examples in Tables 9 to 12, bm_ref_idx_l0 represents the index identifying a BM reference frame included in BM reference frame list 0, and bm_ref_idx_l1 represents the index identifying a BM reference frame included in BM reference frame list 1. Furthermore, at least one of bm_ref_idx_l0 or bm_ref_idx_l1 of the BM reference image indicating the target block can be encoded and transmitted using a signal.

[0885] [D2] Steps to search BM

[0886] It is possible to search for BM reference blocks within a BM reference image. Specifically, it is possible to search for the reference block in the BM reference image that has the highest correlation with the target block, i.e., the BM reference block.

[0887] The first BM reference block and the second BM reference block can be derived from the pair with the highest similarity between the first BM candidate reference block contained in the first BM reference image and the second BM candidate reference block contained in the second BM reference image. Here, the pair with the highest similarity can refer to the pair with the lowest matching criterion.

[0888] In addition, in the search processing of the BM reference block, at least one of the following can be used: target block encoding information, BM-based encoding method information, matching criteria, initial motion vector for BM search, initial motion vector acquisition method, BM search region information, first BM reference image, second BM reference image, early termination condition for BM search, or BM search method.

[0889] The BM search area can include the location indicated by the initial motion vector used for the BM search and the neighboring area at that location.

[0890] Reference blocks can be searched by symmetrically applying the MV offset to the first initial motion vector MV0 of the first BM reference image and the second initial motion vector of the second BM reference image. In other words, the correspondence in Equation 1 can be established between MV0' (the position of the first BM candidate reference block in the first BM reference image) and MV1' (the position of the second BM candidate reference block in the second BM reference image).

[0891] Equation 1

[0892] MV0′=MV0+MV_offset

[0893] MV1′ = MV1 - MV_offset

[0894] In Equation 1, MV_offset can represent the MV offset. As an example, the range of MV_offset values ​​can be {-2, -1, 0, 1, 2}.

[0895] The maximum value of MV_offset can be predefined in the encoder and decoder. Alternatively, information indicating the maximum value of MV_offset can be encoded and transmitted as a signal, or the maximum value of MV_offset can be adaptively set based on the POC of the first BM reference image and the second reference image. Here, the maximum value of MV_offset can represent the size of the BM search region.

[0896] The search process for the BM reference block in the BM search area may include searching for integer sample offsets and adjusting decimal samples.

[0897] In the step of searching for integer sample offsets, a matching criterion is calculated between a first reference block in the first BM reference frame indicated by a first initial motion vector MV0 and a second reference block in the second BM reference frame indicated by a second initial motion vector MV1. Here, the matching criterion is a value calculated by a cost function, and as an example, it can be the sum of absolute differences (SAD) between the first and second reference blocks.

[0898] When the matching standard value between the first reference block and the second reference block is less than the threshold, the step of searching for integer sample offsets can be terminated early. When the step of searching for integer sample offsets is terminated early, the first initial motion vector MV0 and the second initial motion vector MV1 can be set as the first final motion vector and the second final motion vector, respectively.

[0899] When the matching criterion value between the first reference block and the second reference block is equal to or greater than the threshold, the first integer motion vector MV0' and the second integer motion vector MV1' indicating the reference block pair with the minimum matching criterion are searched while the MV_offset value is changed. The first integer motion vector MV0' and the second integer motion vector MV1' indicating the reference block pair with the minimum matching criterion can be referred to as an integer motion vector pair. The integer motion vector pair can be indicated as (MV0_int, MV1_int).

[0900] Next, we can proceed with the steps to adjust the decimal sample points.

[0901] In the step of adjusting decimal samples, the correlation in the decimal sample cells is searched based on the integer motion vector pairs derived from the step of searching for integer sample offsets. In other words, the integer motion vector used to derive the minimum matching criterion in the step of searching for integer sample offsets is set as the center point, and the matching criterion at the decimal sample positions around the center point is calculated. As an example, when the integer motion vector pair used to derive the minimum matching criterion is (MV0_int, MV1_int), MV0_int is the center point in the first BM reference image, and MV1_int is the center point in the second BM reference image.

[0902] Furthermore, the decimal sample point positions can be calculated using the relevant positions at the center point location in the BM reference image and the matching standard values ​​at integer positions around the center point. As an example, the decimal sample point positions can be calculated as shown in Equation 2 below.

[0903] Equation 2

[0904]

[0905] In equation 2, E i (0,0) represents the matching standard value at the center point of the (i+1)th BM reference image. i can have a value of 0 or 1. In other words, for each of the first and second BM reference images, the decimal sample point position can be determined.

[0906] E i (-1,0) and E i (1,0) represent the matching criteria values ​​at the left and right integer positions of the center point, respectively. Furthermore, E i (0,-1) and E i (0,1) represent the matching standard values ​​at the top and bottom integer positions of the center point, respectively.

[0907] As shown in Equation 2, the horizontal decimal position x can be calculated by using the matching criteria of the center point and the matching criteria of the integer positions adjacent to the left and right of the center point. i min Furthermore, the vertical decimal position y can be calculated using the matching criteria of the center point and the matching criteria at integer positions adjacent to the bottom and top of the center point. i min .

[0908] When exporting the decimal sample point position (x i min ,y i min When calculating the motion vector, the decimal sample point position can be added to the integer sample point position. As an example, the final motion vector can be calculated by adding the first decimal motion vector (x... 0 min ,y 0 min The first final motion vector is derived by adding the second decimal motion vector (x) to the first integer motion vector MV0_int, and can be derived by adding the second decimal motion vector (x) to the first integer motion vector MV0_int. 1 min ,y 1 min The second final motion vector is derived from the second integer motion vector MV1_int. Furthermore, the derived final motion vector may be referred to as the BM motion vector. The block indicated by the BM motion vector is set as the BM reference block.

[0909] Furthermore, the step of adjusting decimal samples can be selectively performed. As an example, the step of adjusting decimal samples may not be performed when the minimum matching criterion derived from the step of searching for integer sample offsets is less than a threshold. In this case, the integer motion vector used to derive the minimum matching criterion in the step of searching for integer sample offsets can be set as the BM motion vector.

[0910] In addition, the first BM reference image and the second BM reference image can be selected from the first BM reference image list and the second BM reference image list, respectively.

[0911] In this case, the first BM reference image and the second BM reference image can be identical to each other. In other words, a reference image can be used as both the first BM reference image and the second BM reference image.

[0912] Optionally, at least one of the first BM reference image or the second BM reference image may be the target image.

[0913] Figure 19 and Figure 20 An example is shown where the BM reference block is determined using a BM reference image.

[0914] Figure 19 An example is shown where the first BM reference image and the second BM reference image are the same as the target image.

[0915] When the target image is set as the BM reference image, such as in Figure 19 In the example shown, when the BM reference image is the target image, a first BM reference block and a second BM reference block with the best correlation to the target block can be searched within the target image.

[0916] Figure 20 Examples are shown where the first BM reference image and the second BM reference image each have a different POC from the target image.

[0917] When each of the first BM reference image and the second BM reference image is different from the target image, such as in Figure 20 In the example shown, a first BM reference block can be derived from a first BM reference image, and a second BM reference block can be derived from a second BM reference image.

[0918] As described above, the first BM reference image can be the target image, a decoded past image, or a decoded future image. Similarly, the second BM reference image can be the target image, a decoded past image, or a decoded future image.

[0919] The BM motion vector represents the positional difference between the target block and the BM reference block.

[0920] Figure 21 and Figure 22 The BM motion vector derived through BM search is shown.

[0921] Figure 21 This illustrates the case where the BM reference image is the target image. When the top-left pixel position of the target block is (x... t y t And the top-left pixel position of the BM reference block is (x bm y bm When ), BM motion vector MV bm It can be (x) bm -x t y bm -y t ).

[0922] Figure 22 This is an example of a situation where the first BM reference image and the second BM reference image are different from the target image.

[0923] When the top left pixel position of the target block is (x t y t And the top-left pixel position of the first BM reference block is (xbm_1 y bm_1 At that time, the first BM motion vector MV bm_1 It can be (x) bm_1 -x t y bm_1 -y t ).

[0924] Additionally, when the top-left pixel position of the second BM reference block is (x bm_2 y bm_2 At that time, the second BM motion vector MV bm_2 It can be (x) bm_2 -x t y bm_2 -y t ).

[0925] The BM-optimal block of the target block can be obtained by using at least one of a first BM reference block indicated by a first BM motion vector or a second BM reference block indicated by a second BM motion vector.

[0926] As an example, the optimal BM block can be obtained by weighted combination or weighted summation of the first BM reference block and the second BM reference block.

[0927] As an example, as shown in Equation 3 below, the pixel P at position (x,y) within the target block can be obtained by averaging or weighted summing the pixels within the first BM reference block indicated by the first BM motion vector in the first BM reference image and the pixels within the second BM reference block indicated by the second BM motion vector in the second BM reference image. BM(x,y) .

[0928] Equation 3

[0929]

[0930] In equation 3 above, P BM(x,y) This represents the pixel value at position (x, y) within the optimal block of the BM. P0(a, b) represents the pixel value at position (a, b) within the first BM reference block, and P1(a, b) represents the value at position (a, b) within the second BM reference block. i BM y represents the horizontal component of the (i+1)th BM motion vector. i BM This represents the vertical component of the (i+1)th BM motion vector. Here, i can be 0 or 1.

[0931] `shift` represents the shift parameter used to calculate the average value.

[0932] The shift pointer can have predefined values ​​in the encoder and decoder. As an example, shift can be 1.

[0933] Alternatively, shift can be determined based on at least one of internal bit depth (IBD) or sample bit depth (bitDepth).

[0934] Optionally, shift can be set to Max(3, (1+IBD)-bitDepth). Here, the function MAX(a, b) is a function that returns the larger of a and b. When IBD is 14 and bitDepth is 10, shift can be set to 5. Optionally, when IBD is 14 and bitDepth is 8, shift can be set to 7.

[0935] In Equation 3, offset represents the offset parameter used for rounding. Offset can be determined based on at least one of the Internal Bit Depth (IBD) or the Sample Bit Depth (bitDepth).

[0936] As an example, the offset can be set to the median of an N-bit image. In other words, for a 10-bit image, the offset can be set to 512, and for an 8-bit image, the offset can be set to 128.

[0937] As another example, the offset can be exported as (1<<(shift-1)).

[0938] Pixels within the best block of the BM can also be exported without performing rounding. When rounding is set to not be performed, the offset parameter in Equation 3 above can have a value of 0.

[0939] Furthermore, the BM optimal block (or BM optimal signal) can refer to the prediction block (or prediction signal) of the target block. In other words, a sample point at position (x,y) within the BM optimal block can represent a prediction sample point at position (x,y) within the target block.

[0940] Furthermore, a BM reference block can be searched only within a BM search region of the BM reference image. Here, the BM search region can represent the entire or a portion of the BM reference image. As an example, the BM search region can include the location indicated by the initial motion vector and its neighboring region.

[0941] The starting position of the search (i.e., the starting position of the BM search) can be the position indicated by the initial motion vector. In other words, after calculating the matching criteria for the position indicated by the initial motion vector, the matching criteria for the modified position can be calculated while changing MV_offset.

[0942] Optionally, the location of the target block within the BM reference image can be set as the BM search start location.

[0943] The size of the BM search region can be adaptively determined based on the size of the target block. As an example, the BM search region could be an area corresponding to N times the size of the target block. In this case, the maximum size of the BM search region can be predefined in the encoder and decoder. As an example, the maximum value of the distance from the position indicated by the initial motion vector to one boundary of the BM search region could be 2, 4, 8, or 16.

[0944] Optionally, the BM search area can be set to a region corresponding to a predetermined number of blocks. In other words, the remaining region within the BM reference area, excluding the region corresponding to the predetermined number of blocks, may not be included in the BM search area. The predetermined number can be a natural number greater than or equal to 1.

[0945] As an example, the BM search area can be a region corresponding to a predetermined number of blocks existing around the target block. The blocks existing around the target block can include at least one of the blocks adjacent to the target block or blocks that are not adjacent to the target block but are at a distance from the target block less than or equal to a threshold.

[0946] Here, the block unit configured for the BM search region can be at least one of a coding tree block, coding block, prediction block, transform block, virtual pipeline data unit (VPDU), or block of a predetermined size.

[0947] As an example, the BM search region can consist of multiple spatially contiguous blocks.

[0948] Optionally, the BM search region can consist of multiple blocks existing at spatially separated locations.

[0949] Furthermore, multiple blocks configured in the BM search area can have the same size and / or shape. Alternatively, multiple blocks configured in the BM search area can have different sizes and / or shapes.

[0950] As another example, the smallest quadrilateral region comprising multiple blocks can be set as the BM search region.

[0951] Optionally, the BM search region can be set by taking into account the size of the VPDU. As an example, the BM search region can consist of N VPDU regions surrounding the target block within the BM reference image. Optionally, the BM search region can consist of N VPDU regions surrounding the position indicated by the initial motion vector within the BM reference image. Optionally, the BM search region can consist of the target VPDU region to which the target block belongs and (N-1) coding tree regions surrounding the target VPDU region.

[0952] In other words, the remaining region within the BM reference image, excluding the N VPDU regions, may not be the BM search region. Therefore, the BM reference block can be searched outside the remaining region excluding the N VPDU regions. Furthermore, N can be a natural number such as 1, 2, 3, or 4.

[0953] Optionally, the BM search region can be set by taking into account the size of the coding tree blocks. As an example, the BM search region can consist of N coding tree block regions surrounding the target block within the BM reference image. Optionally, the BM search region can consist of N coding tree block regions surrounding the position indicated by the initial motion vector within the BM reference image. Optionally, the BM search region can consist of the target coding tree block to which the target block belongs and (N-1) coding tree regions surrounding the target coding tree block.

[0954] In other words, the remaining region within the BM reference image, excluding the N coded tree block regions, may not be the BM study region. Therefore, the BM reference block can be searched outside the remaining region excluding the N coded tree block regions. Furthermore, N can be a natural number such as 1, 2, 3, or 4.

[0955] The BM search region can be a quadrilateral region, such as a rectangle or a rhombus. In this case, the dimensions of a rectangular BM search region can be defined by its horizontal and vertical dimensions. Conversely, the dimensions of a rhombus-shaped BM search region can be defined by the dimensions of its two diagonals.

[0956] The size of the BM search region can also be defined by the horizontal and vertical distances from the starting position of the BM search. Here, the starting position of the BM search can be the initial motion vector of the BM search.

[0957] Furthermore, non-decoded regions in the BM reference image can be set to be excluded from the BM search region. As an example, when at least a portion of the target block is included in a minimal quadrilateral region comprising multiple blocks, the region within the quadrilateral region excluding the area occupied by the target block can be set as the BM search region.

[0958] When at least a portion of the region to be set as the BM search region extends beyond the image boundary, the remaining region other than the region extending beyond the image boundary may be set as the BM search region. Optionally, when at least a portion of the region to be set as the BM search region extends beyond the image boundary, the region extending beyond the image boundary may be replaced with a region corresponding to the region extending beyond the image boundary on the relative boundary of the image boundary.

[0959] At least one of the shape or size of the BM search region can be encoded and signaled using a high-level syntax such as encoding parameters, sequence parameters, image parameters, or stripe headers, or it can be encoded and signaled using a coding tree unit or a predetermined block unit. Here, the predetermined block unit can be a coding block, a transform block, or a prediction block.

[0960] Optionally, the size and / or shape of the BM search area can be adaptively determined based on at least one of the following: whether the BM reference image is the target image, the distance between the target block and the position indicated by the initial motion vector, or the size / shape of the target block.

[0961] Figures 23 to 34 An example related to the BM search area is shown.

[0962] Figures 23 to 28 An example indicating that the BM search area is rectangular.

[0963] exist Figure 23 The image shown is set as the BM reference image. In this case, the BM search area can be rectangular.

[0964] exist Figure 23 In this context, the BM search region can be all or part of the decoding region in the target image.

[0965] exist Figure 24 In the example, a reference image different from the target image (i.e., a reference image with a different POC than the target image) is shown as the BM reference image. (See example in...) Figure 24 In the example shown, when the BM reference image is a decoded past image or a decoded future image, the BM search area can be rectangular.

[0966] exist Figure 24 In this context, the BM search region can be all or part of the image in the BM reference image.

[0967] exist Figure 25 In the example, the target image is shown as the BM reference image. (As shown in...) Figure 25 In the example shown, when the region set as the BM search region includes undecoded regions such as the target block, the remaining regions other than the undecoded regions can be set as the BM search region.

[0968] exist Figure 26 In the example, a reference image different from the target image (i.e., a reference image with a different POC than the target image) is shown as the BM reference image. (See example in...) Figure 26 In the example shown, when the region set as the BM search region exists outside the image boundary, the remaining region other than the region outside the image boundary can be set as the BM search region.

[0969] Figure 27 and Figure 28 Information elements used to determine the rectangular BM search area are shown.

[0970] As in Figure 27 In the example shown, the size of the rectangular BM search region can be defined by its width (rec_search_width) and height (rec_search_height). At least one of the variables rec_search_width (representing the horizontal value) and rec_search_height (representing the vertical value) can be encoded and signaled using high-level syntax such as encoding parameters, sequence parameters, image parameters, or strip headers. Alternatively, at least one of the variables rec_search_width (representing the horizontal value) and rec_search_height (representing the vertical value) can be encoded and signaled according to at least one of coding tree units or predetermined block units. Here, the predetermined block unit can be a coding block, a transform block, or a prediction block.

[0971] Furthermore, the BM search region can be a square shape with the same width and height. In this case, either the width `rec_search_width` or the height `rec_search_height` can be encoded and signaled.

[0972] Optionally, the ratio of the width to the height of the BM search region can be fixed in the encoder and decoder. As an example, the width could be twice the height. In this case, only the height `rec_search_height` can be encoded and signaled, and the width can be derived based on the height. Alternatively, only the width `rec_search_width` can be encoded and signaled, and the height can be derived based on the width.

[0973] As in Figure 28In the example shown, the size of the rectangular BM search region can be defined by its horizontal dimension (del_width) and vertical dimension (del_height) from the start position of the BM search. Here, the horizontal dimension del_width can represent the distance from the start position of the BM search to the left or right boundary of the BM search region, and the vertical dimension del_height can represent the distance from the start position of the BM search to the top or bottom boundary of the BM search region. The size of the BM search region can be (1+2*del_width)×(1+2*del_height). At least one of the horizontal dimension del_width or the vertical dimension del_height can be encoded and signaled using a high-level syntax such as encoding parameters, sequence parameters, image parameters, or stripe headers, or it can be encoded and signaled according to at least one of the coding tree units or predetermined block units.

[0974] The BM search region can be a square shape with the same width and height. In this case, only the horizontal length del_width or the vertical length del_height can be encoded and transmitted as a signal.

[0975] Optionally, the ratio of the width to the height of the BM search region can be fixed in the encoder and decoder. As an example, the width could be twice the height. In this case, only the vertical length `del_height` can be encoded and signaled, and the horizontal length can be derived from the vertical length. Optionally, only the horizontal length `del_width` can be encoded and signaled, and the vertical length can be derived from the horizontal length.

[0976] Figures 29 to 34 This shows an example of a diamond-shaped search region for BM.

[0977] exist Figure 29 The image shown is set as the BM reference image. In this case, the BM search area can be a rhombus.

[0978] exist Figure 29 In this context, the BM search region can be all or part of the decoding region in the target image.

[0979] exist Figure 30 In the example, a reference image different from the target image (i.e., a reference image with a different POC than the target image) is shown as the BM reference image. (See example in...) Figure 30 In the example shown, when the BM reference image is a decoded past image or a decoded future image, the BM study area can be a rhombus.

[0980] exist Figure 30In this context, the study area of ​​BM can be a portion of the BM reference image.

[0981] exist Figure 31 In the example, the target image is shown as the BM reference image. (As shown in...) Figure 31 In the example shown, when the region set as the BM search region includes undecoded regions such as the target block, the remaining regions other than the undecoded regions can be set as the BM search region.

[0982] exist Figure 32 In the example, a reference image different from the target image (i.e., a reference image with a different POC than the target image) is shown as the BM reference image. (See example in...) Figure 32 In the example shown, when the region set as the BM search region exists outside the image boundary, the remaining region other than the region outside the image boundary can be set as the BM search region.

[0983] Figure 33 and Figure 34 Information elements used to determine the diamond-shaped BM search area are shown.

[0984] As in Figure 33 In the example shown, the size of the diamond-shaped BM search region can be defined by its width (dia_search_width) and height (dia_search_height). At least one of the variables dia_search_width (indicating horizontal values) and dia_search_height (indicating vertical values) can be encoded and signaled using high-level syntax such as encoding parameters, sequence parameters, image parameters, or strip headers. Alternatively, at least one of the variables dia_search_width (indicating horizontal values) and dia_search_height (indicating vertical values) can be encoded and signaled according to at least one of encoding tree units or predetermined block units. Here, the predetermined block unit can be an encoding block, a transform block, or a prediction block.

[0985] As in Figure 34In the example shown, the size of the rhombus-shaped BM search region can be defined by its horizontal dimension (del_width) and vertical dimension (del_height) from the start position of the BM search. Here, the horizontal dimension del_width represents the diagonal length from the start position of the BM search to the vertices of the rhombus that lie on the horizontal line starting from the start position. The vertical dimension del_height represents the diagonal length from the start position of the BM search to the vertices of the rhombus that lie on the vertical line starting from the start position. The BM search region is a rhombus with diagonal lengths of (2*del_width) and (2*del_height), and the size of the BM search region can be 1 / 2*(1+2*del_width)×(1+2*del_height). At least one of the horizontal dimension del_width or the vertical dimension del_height can be encoded and signaled using a high-level syntax such as encoding parameters, sequence parameters, image parameters, or stripe headers, or can be encoded and signaled according to at least one of the coding tree unit or predetermined block unit. As in the example above, the BM search region can be rectangular or rhombus-shaped. In this case, information about the shape of the specified BM search region can be encoded and transmitted using signals. As an example, this information could be an index indicating one of several shape candidates (e.g., bm_shape_idx). For example, when bm_shape_idx has a first value, it could indicate that the BM search region is a rectangle, and when bm_shape_idx has a second value, it could indicate that the BM search region is a rhombus.

[0986] The information can be encoded and signaled using high-level syntax such as encoding parameters, sequence parameters, image parameters, or stripe headers, or it can be encoded and signaled according to at least one of coding tree units or predetermined block units. Here, the predetermined block unit can be a coding block, a transform block, or a prediction block.

[0987] Optionally, the shape of the BM search area can be adaptively determined based on at least one of the following: whether the BM reference image is the target image, the distance between the target block and the BM search start position, whether the area set as the BM search area includes the target block or an undecoded area, or whether the area set as the BM search area extends beyond the screen boundary.

[0988] Furthermore, the search for a BM reference block within the BM reference image can begin from the BM search start position. In this case, the BM search start position can be specified by the initial motion vector of the target block.

[0989] As an example, the starting position of a BM search can be set to the position indicated by the initial motion vector, the integer position closest to the position indicated by the initial motion vector, or the position obtained by adding the position to the offset / subtracting the offset from the position.

[0990] In addition, the initial motion vector can be determined based on the coordinate information of the target block.

[0991] Figure 35 This shows an example where the location of the target block is configured as the start location of the BM search.

[0992] As in Figure 35 In the example shown, when the top-left coordinate of the target block is (x t y t When ), the corresponding position (x) of the target block in the BM reference image can be determined. t y t ) is set as the starting position for the BM search. In other words, when the position of the target block in the BM reference image is the BM search position, it can be the same as when the initial motion vector is (0, 0).

[0993] When there is a block that has been encoded / decoded using BM in a block adjacent to the target block, the initial motion vector for BM search can be obtained based on the BM motion vector of the adjacent block.

[0994] Figure 36 An example is shown of deriving the initial motion vector from a block adjacent to the target block.

[0995] Let's assume that there exists a block adjacent to the target block that is encoded / decoded using BM, and the BM motion vector of the block encoded / decoded using BM is (MV bm_x MV bm_y In this case, the initial motion vector of the target block can be set to (MV). bm_x MV bm_y ), which is the same as the BM motion vector of the adjacent block.

[0996] When the position of the target block is (x t y t And the initial motion vector is (MV bm_x MV bm_y When ), the BM search start position in the BM reference image can be set to (x t +MV bm_x y t +MV bm_y ).

[0997] When multiple neighboring blocks exist via BM encoding / decoding, the initial motion vector of the target block can be derived from the block first found when searching the multiple neighboring blocks in a predetermined order. Optionally, information from one of the specified multiple neighboring blocks can be encoded and transmitted using a signal.

[0998] As another example, the initial motion vector of the target block can be obtained based on at least one of the merge encoding information. Here, the merge encoding information may include at least one of merge_idx, merge_subblock_flag, merge_subblock_idx, regular_merge_flag, mmvd_merge_flag, mmvd_cand_flag, mmvd_distance_idx, mmvd_direction_idx, merge_gpm_partition_idx, merge_gpm_idx0, or merge_gpm_idx1.

[0999] As an example, after configuring the merge candidate list for the target block, the motion vector of one of the merge candidates included in the merge candidate list can be set as the initial motion vector of the target block.

[1000] In this case, the motion vector of the first merge candidate (i.e., the merge candidate with index 0) in the merge candidate list of the target block can be set as the initial motion vector of the target block.

[1001] Alternatively, the motion vector of a merge candidate selected by the merge index (merge_idx) decoded from the bitstream can be set as the initial motion vector of the target block from the merge candidate list.

[1002] As an example, let's assume the location of the target block is (x t y t ), and the motion vector of the first merge candidate in the merge candidate list or the merge candidate selected by the merge index (merge_idx) is (x m y m In this case, the starting position of the BM search in the BM reference image can be set to (x). t +x m yt+y m ).

[1003] Optionally, MMVD can be used to modify the motion vector of a merge candidate by using an offset vector, and the modified motion vector can be set as the initial motion vector of the target block. In this case, the merge candidate can be selected via merge_cand_flag, and the magnitude and direction of the offset vector can be determined via mmvd_distance_idx and mmvd_direction_idx, respectively.

[1004] Furthermore, two motion vectors are required to perform BM prediction. Therefore, when the merge candidate from which the initial motion vector is to be derived (e.g., the first merge candidate or a merge candidate selected by the merge index) only has unidirectional motion information, the motion vector in the other direction can be derived based on the unidirectional information of the corresponding merge candidate. Then, the first motion vector from the unidirectional information of the corresponding merge candidate and the second motion vector in the other direction derived from the unidirectional information can be set as the initial motion vectors in the first and second directions, respectively.

[1005] As an example, when the first merge candidate in the merge candidate list only has L0 motion information, a reference frame in the L1 reference frame whose direction is opposite to that of the first merge candidate's L0 reference frame and whose distance is the same as the distance between the L0 reference frame and the target frame is set as the L1 reference frame. Then, a motion vector with the same size as the first merge candidate's L0 motion vector but opposite in direction can be set as the L1 motion vector. The first merge candidate's L0 motion vector can then be set as the initial motion vector for the L0 direction, and the L1 motion vector can be set as the initial motion vector for the L1 direction.

[1006] As another example, the initial motion vector of the target block can also be derived from the merge candidates with bidirectional motion information in the merge candidate list. For example, the motion vector of the first merge candidate with available bidirectional motion information in the merge candidate list of the target block can be set as the initial motion vector of the target block. In other words, the initial motion vector of the target block can be derived from the merge candidate with the lowest index among the merge candidates with bidirectional motion information in the merge candidate list.

[1007] Alternatively, the merge candidate list can be reconfigured using only merge candidates with bidirectional motion information, and one of the merge candidates to be included in the reconstructed merge candidate list can be selected via a merge index (merge_idx). The motion vector of the merge candidate selected via the merge index can be set as the initial motion vector of the target block.

[1008] As another example, the initial motion vector of the target block can be derived based on the AMVP mode of inter-frame prediction. Specifically, the initial motion vector of the target block can be obtained based on AMVP encoding information. Here, the AMVP encoding information may include at least one of mvp_l0_flag, mvp_l1_flag, abs_mvd_greater0_flag, abs_mvd_greater1_flag, abs_mvd_minus2, or mvd_sign_flag.

[1009] The motion vector (x) obtained based on AMVP encoding information can be used to... AMVP y AMVP ) is set as the initial motion vector of the target block. Here, the motion vector (x) obtained based on AMVP encoding information can be obtained by adding the motion vector prediction value or motion vector difference to the motion vector prediction value. AMVP y AMVP Here, motion vector prediction values ​​can be obtained through motion vector prediction factors (mvp_l0_flag / mvp_l1_flag), and motion vector differences can be obtained through information related to motion vector differences (at least one of abs_mvd_greater0_flag, abs_mvd_greater1_flag, abs_mvd_minus2, or mvd_sign_flag).

[1010] It can be derived by adding the motion vector difference to the motion vector prediction.

[1011] As an example, when the position of the target block is (x t y t And the motion vector obtained based on AMVP encoding information is (x AMVP y AMVP When ), the position (x) indicated by the initial motion vector at the location of the target block in the BM reference image. t +x AMVP y t + yAMVP ) can be set as the starting position for a BM search.

[1012] Information used to determine the initial motion vector for BM search (e.g., merged encoding information or AMVP encoding information) can be encoded and signaled by high syntax elements such as sequence parameters, image parameters, or strip headers, or it can be encoded and signaled according to encoding tree units or predetermined block units.

[1013] As in the example above, the initial motion vector of the target block can be set to be the same as the target block's coordinates, or it can be derived based on at least one of the merge mode or AMVP mode. In this case, information indicating which of the target block's coordinates, merge mode, or AMVP mode was used to derive the target block's motion vector can be encoded and signaled. This information can be encoded and signaled using high-level syntax elements such as sequence parameters, image parameters, or stripe headers, or it can be encoded and signaled according to coding tree units or predetermined block units.

[1014] Table 13 shows examples of each method in which different indices are assigned to the initial motion vectors used to derive the target block. To specify the method used to derive the initial motion vectors for the target block, bm_init_idx shown in Table 13 can be encoded and sent with a signal.

[1015] Table 13

[1016] bm_init_idx BM searches for the initial motion vector 0 (0,0) 1 Use merged encoding information 2 Use AMVP to encode information

[1017] In the example shown in Table 13, when bm_init_idx is 0, it indicates that the initial motion vector of the target block is (0,0). In this case, the position of the target block in the BM reference image can be set as the BM search start position.

[1018] When bm_init_idx is 1, it indicates the initial motion vector of the target block obtained based on the merged encoding information. In this case, the merged encoding information may exist in the bitstream.

[1019] When bm_init_idx is 2, it indicates the initial motion vector of the target block obtained based on AMVP encoding information. In this case, the AMVP encoding information can exist in the bitstream.

[1020] Optionally, a method for determining the initial motion vector of the target block can be adaptively determined based on at least one of the following: the size / shape of the target block, the prediction mode of the target block, whether the neighboring blocks adjacent to the target block are encoded by BM, and the position of the target block in the target image.

[1021] As an example, when the prediction mode of the target block is intra-frame prediction or IBC prediction, the initial motion vector can be set to (0, 0). On the other hand, when the prediction mode of the target block is inter-frame prediction, the initial motion vector can be derived based on either the merge mode or the AMVP mode.

[1022] Methods for searching a BM reference block within a BM search region may include at least one of full search, 2D logarithmic search, N-step search, rhombus search, hexagonal search, or test region search. Here, full search refers to searching all pixels within the BM search region. Furthermore, detailed descriptions of 2D logarithmic search, N-step search, rhombus search, hexagonal search, and test region search are omitted in this disclosure.

[1023] The search method within the BM search region can be predefined in the encoder and decoder.

[1024] Optionally, a search method within the BM search area can be adaptively determined based on at least one of the following: the size / shape of the target block, the accuracy of the motion vector, whether the BM reference image is the target image, the position of the current block within the target image, the position of the BM search region within the BM reference image, the size / shape of the BM search region, or the POC of the L0 BM reference image and the L1 BM reference image.

[1025] Optionally, syntax elements specifying the search method within the BM search area (e.g., bm_search_idx) can be encoded and sent using signals.

[1026] As an example, based on the value of bm_search_idx, the search method within the BM search area can be determined as follows.

[1027] - First value: Full search

[1028] -Second value: 2D logarithmic search

[1029] -Third value: N-step search

[1030] -Fourth value: Diamond search

[1031] - Fifth value: Hexagonal search

[1032] -Sixth value: Test area search

[1033] Furthermore, the relevance used to search for the BM reference block can be calculated based on a cost function. Here, the cost function can be one of the following: Sum of Absolute Differences (SAD), Sum of Absolute Transformed Differences (SATD), Sum of Absolute Differences with Mean Removed (MR-SAD), Mean Squared Error (MSE), or Sum of Squared Errors (SSE). However, the type of cost function is not limited to the examples listed.

[1034] Furthermore, the value calculated through the cost function can be referred to as the matching criterion.

[1035] According to the cost function, the block pair with the best relevance (i.e., the pair of the first BM candidate reference block and the second BM candidate reference block) can be the block pair with the highest matching criterion value (i.e., the block pair with the highest cost function) or the block pair with the lowest matching criterion value (i.e., the block pair with the lowest cost function).

[1036] However, in this disclosure, it is assumed that the block pair with the minimum cost function value (i.e., the block pair with the lowest matching criterion value) has the best relevance. In other words, in the BM search, it can be determined that the candidate block pair with the lowest matching criterion value has the best relevance to the target block.

[1037] Alternatively, instead of selecting the block pair with the lowest matching criterion value, when the matching criterion values ​​are sorted in ascending order, all the top N candidates or one of these N candidates can be selected to derive the BM reference block. When multiple candidates are selected, the BM reference block of the target block can be derived by a weighted sum or average of the multiple candidate blocks. Here, N is a natural number and can have values ​​such as 2 or 3.

[1038] In a BM search, information specifying the type of cost function (or matching criterion) used to calculate relevance can be encoded and signaled. As an example, this information could be an index (e.g., bm_criteria_idx) indicating one of several cost function candidates, and the cost function candidate indicated by the index could include at least one of SAD, SSD, SATD, MR_SAD, or SSE.

[1039] As an example, the cost function candidate indicated by the value of bm_criteria_idx can be set as follows.

[1040] -First value: SAD

[1041] -Second value: SSD

[1042] -Third value: SATD

[1043] - Fourth value: MR_SAD

[1044] - Fifth value: SSE

[1045] When calculating at least one of the aforementioned correlation or cost functions (i.e., matching criteria), instead of pixels, coding parameters such as motion vectors or intra-frame prediction modes (directions) can be used. For example, it can be determined that the smaller the motion vector difference between the first BM candidate reference block and the second BM candidate reference block, the higher the correlation; or the smaller the intra-frame prediction mode difference between the first candidate BM reference block and the second BM candidate reference block, the higher the correlation. As described above, the correlation or cost function (i.e., matching criteria) can be calculated using coding parameters such as the motion vectors of candidate blocks or the intra-frame prediction modes (directions).

[1046] Optionally, the sum of a first cost function calculated based on pixel values ​​and a second cost function calculated based on encoding parameters can be set as the matching criterion. As an example, the first cost function can represent the SAD between a first BM candidate block and a second BM candidate block, and the second cost function can be the sum of the distance between the first BM candidate blocks at a first initial motion vector position and the distance between the second BM candidate blocks at a second initial motion vector position.

[1047] Furthermore, the method for deriving the matching criteria can be adaptively determined based on the size of the target block. For example, when the target block size is greater than a threshold, the sum of the first cost function and the second cost function can be derived as the matching criterion value. Conversely, when the target block size is less than the threshold, the first cost function can be derived as the matching criterion value.

[1048] When determining the coding parameters of a target block, the coding parameter with the highest relevance (i.e., the coding parameter with the lowest matching criterion (cost function)) can be used as a candidate coding parameter for the target block. As an example, the intra-prediction mode with the highest relevance (i.e., the intra-prediction mode with the lowest matching criterion (i.e., cost function)) can be set as an MPM candidate for the list of most probable modes (MPMs) used for entropy coding / decoding of the intra-prediction modes of the target block. Here, the intra-prediction mode with the highest relevance can represent the intra-prediction mode of the candidate block pair with the highest relevance (i.e., at least one of the first BM candidate reference block or the second BM candidate reference block).

[1049] Optionally, the motion information with the highest relevance (i.e., the motion information with the lowest matching criterion (i.e., the cost function)) can be set as a candidate for the merge candidate list or AMVP list used for entropy encoding / decoding of the motion information of the target block. Here, the motion information with the highest relevance can represent the motion information of the candidate block pair with the highest relevance (i.e., at least one of the first BM candidate reference block or the second BM candidate reference block).

[1050] To reduce the computational complexity of the matching criteria, the matching criteria can be calculated using only pixels at sub-sampling locations within the BM candidate reference block.

[1051] Figure 37 An example is shown where matching criteria are calculated using pixels at subsampling locations.

[1052] As in Figure 37 In the example shown, when calculating the matching criteria between a first BM candidate reference block in a first BM reference image and a first BM candidate reference block in a second BM reference image, instead of calculating the matching criteria using all pixels belonging to the first / second BM candidate reference block, the matching criteria can be calculated using only the pixels at the subsampled locations. In other words, the matching criteria can be calculated using only some pixels in the BM candidate reference block instead of all pixels, and here, some pixels can represent at least one pixel in the BM candidate reference block.

[1053] In addition, such as in Figure 37 In the example shown in (a), subsampling can be performed only in the vertical direction. As an example, when assuming the top-left position of the block is (0, 0), the matching criteria can be calculated by using only the pixels at position (n1, 2×n2) in a block with width W and height H, respectively. Here, n1 can be an integer of {0, 1, 2, ..., W-1}, and n2 can be an integer of {0, 1, 2, ..., H / 2-1}.

[1054] Optionally, as in Figure 37 In the example shown in (b), subsampling can be performed only in the horizontal direction. As an example, when assuming the top-left position of the block is (0, 0), the matching criteria can be calculated by using only the pixels at positions (2×n1, n2) within the block with widths and heights of W and H, respectively.

[1055] Optionally, as in Figure 37 In the example shown in (c), subsampling can be performed for both the horizontal and vertical directions. As an example, when assuming the top-left position of the block is (0, 0), the matching criteria can be calculated by using only the pixels at positions (2×n1, 2×n2) within the block with widths of W and heights of H.

[1056] When performing subsampling, a bilinear filter can be applied to each pair of adjacent pixels. As an example, such as in... Figure 37In the examples shown, when performing 1 / 2 subsampling in the horizontal or vertical direction, the pixel value at the subsampling location can be set to the value obtained by applying a bilinear filter to two pixel pairs. As an example, when performing subsampling in the horizontal direction, the pixel value at the subsampling location can be derived by applying a bilinear filter to the target pixel and its right or left adjacent pixel. Furthermore, when performing subsampling in the vertical direction, the pixel value at the subsampling location can be derived by applying a bilinear filter to the target pixel and its top or bottom adjacent pixel.

[1057] As an example, in Figure 37 In the example shown in (a), the pixel at position (n1, 2×n2) can be obtained by applying a bilinear filter to the pair of pixels at position (n1, 2×n2) and at position (n1, 2×n2+1).

[1058] As an example, in Figure 37 In the example shown in (a), the pixel at position (2×n1, n2) can be obtained by applying a bilinear filter to the pair of pixels at position (2×n1, n2) and at position (2×n1+1, n2).

[1059] The subsampling rate or subsampling method can be predefined in the encoder and decoder. Alternatively, the subsampling method can be adaptively determined based on at least one of the target block size / shape, motion vector accuracy, or the size of the BM search region.

[1060] As an example, when the target block size is 4×4 or smaller, subsampling may not be performed. In this case, the matching criteria can be calculated using all pixels in the BM reference block.

[1061] Optionally, subsampling can be performed when at least one of the width or height of the target block is 8 or greater. In this case, the matching criteria can be calculated using only pixels at some locations within the block.

[1062] Furthermore, whether to perform subsampling in the vertical direction, whether to perform subsampling in the horizontal direction, or whether to perform subsampling in both the vertical and horizontal directions can be variably determined by the size / shape of the target block.

[1063] As an example, when the target block is a square, subsampling can be performed in both the vertical and horizontal directions.

[1064] Optionally, when the target block has a non-square shape with a width greater th...

Claims

1. A method for decoding an image, the method comprising: Derive the first and second initial motion vectors of the target block; A BM search is performed on the first bilaterally matched BM reference image and the second BM reference image using the first initial motion vector and the second initial motion vector; as well as The target block is reconstructed using the first BM reference block and the second BM reference block obtained through the BM search.

2. The method as described in claim 1, wherein, The optimal BM block is derived based on the average operation or weighted summation operation of the first BM reference block and the second BM reference block.

3. The method as described in claim 2, wherein, The BM optimal block is configured as the prediction block of the target block.

4. The method according to claim 2, wherein, The final predicted block of the target block is obtained by performing a weighted summation operation between the BM block and the intra-predicted block obtained by performing intra-prediction on the target block.

5. The method of claim 4, wherein, The intra-prediction block is obtained by using a predefined intra-prediction mode in the decoder.

6. The method of claim 4, wherein, The intra-prediction mode used for the intra-prediction is derived from the first BM reference block or the second BM reference block encoded by the intra-prediction.

7. The method according to claim 4, wherein, The first weight applied to the intra-prediction block and the second weight applied to the best BM block are determined by comparing a threshold with a matching criterion between the first BM reference block and the second BM reference block.

8. The method according to claim 2, wherein, The final prediction block of the target block is obtained by performing a weighted summation operation between the BM block and the inter-prediction block obtained by performing inter-frame prediction on the target block.

9. The method of claim 8, wherein, The first BM reference image or the second BM reference image is configured as a reference image for the inter-frame prediction.

10. The method of claim 2, wherein, The residual block of the target block is derived as the sum of the prediction residual block and the BM residual block. The BM residual block is obtained based on the residual coefficients decoded from the bitstream, and The residual block of the BM optimal block is configured as the prediction residual block.

11. The method of claim 10, wherein, The residual block of the BM optimal block is derived by subtracting the prediction block of the target block from the BM optimal block.

12. The method according to claim 10, wherein, The residual samples in the residual block are reconstructed by using only the predicted residual samples in the predicted residual block in a portion of the target block.

13. The method of claim 1, wherein, The first initial motion vector is derived based on the Advanced Motion Vector Prediction (AMVP) encoding information. The second initial motion vector is derived based on the merged encoding information.

14. The method according to claim 1, wherein, The first initial motion vector is configured to be the same as the motion vector of the merged candidate included in the merged candidate list, and The vector with the same magnitude but opposite direction as the first initial motion vector is configured as the second initial motion vector.

15. The method according to claim 1, wherein, The BM search is performed only in the BM search regions within the first BM reference image and the second BM reference image.

16. The method according to claim 15, wherein, The shape of the BM search region is determined based on index information indicating one of the shape candidates. Among the shape candidates, the first shape candidate is a rectangle and the second shape candidate is a rhombus.

17. The method according to claim 15, wherein, Based on the position indicated by the initial motion vector, the BM search region is configured to be N times the size of the basic unit, where N is a natural number greater than or equal to 1. The basic unit is at least one of the target unit, the coding tree unit, or the Virtual Pipeline Data Unit (VPDU).

18. The method of claim 17, wherein, When at least a portion of the region to be configured as the BM search region extends beyond the image boundary, the residual region other than the region extending beyond the image boundary is configured as the BM search region.

19. A method for encoding an image, the method comprising: Derive the first and second initial motion vectors of the target block; A BM search is performed on the first bilaterally matched BM reference image and the second BM reference image using the first initial motion vector and the second initial motion vector; as well as The target block is encoded using a first BM reference block and a second BM reference block obtained through the BM search.

20. A recording medium for storing a bitstream generated by an image encoding method, the recording medium comprising: Derive the first and second initial motion vectors of the target block; A BM search is performed on the first bilaterally matched BM reference image and the second BM reference image using the first initial motion vector and the second initial motion vector; as well as The target block is encoded using a first BM reference block and a second BM reference block obtained through the BM search.