Method, apparatus, and recording medium for image encoding / decoding
By constructing a filter candidate list and determining the final filter based on the matching cost, the flexibility and efficiency problems of filter selection in the existing technology are solved, and efficient image encoding/decoding is achieved to meet the needs of high-resolution and high-definition images.
Patent Information
- Application Number
- CN202480016670.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-01-04
- Filing Date
- 2024-01-04
- Publication Date
- 2025-10-03
AI Technical Summary
Existing image encoding/decoding technologies are difficult to effectively adapt to the needs of high-resolution and high-definition images, especially the lack of flexibility and efficiency in filter selection.
By constructing a filter candidate list, the final filter is determined based on the matching cost, and by reconstructing and reordering the filter candidate list, a filter is adaptively selected from multiple filters for encoding/decoding.
It improves the efficiency and quality of image encoding/decoding, adapts to the encoding requirements of high-resolution and high-definition images, and achieves more efficient filter selection.
Smart Images

Figure CN120752915A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to methods, devices, and storage media for image encoding / decoding.
[0002] This application claims the benefit of Korean Patent Application No. 10-2023-0001441, filed on January 4, 2023, and Korean Patent Application No. 10-2024-0001763, filed on January 4, 2024, which are hereby incorporated by reference in their entirety into this application. Background Art
[0003] With the continuous development of the information and communications industry, broadcast services supporting high-definition (HD) resolution have become popular around the world. As a result of this popularity, a large number of users have become accustomed to high-resolution and high-definition images and / or videos.
[0004] To meet user demands for higher definition, numerous organizations have accelerated the development of next-generation imaging devices. In addition to high-definition television (HDTV) and full-high-definition (FHD) television, interest in ultra-high-definition television (UHDTV), which boasts a resolution over four times that of full-high-definition (FHD), has also grown. This growing interest is driving a demand for image encoding and decoding technologies that deliver higher resolution and definition.
[0005] As image compression technology, there are various technologies such as inter-frame prediction technology, intra-frame prediction technology, transform, quantization technology, and entropy coding technology.
[0006] Inter-frame prediction technology is a technology for predicting the values of pixels included in the current picture using pictures before and / or after the current picture. Intra-frame prediction technology is a technology for predicting the values of pixels included in the current picture using information about the pixels in the current picture. Transformation and quantization technology can be a technology for compressing the energy of the residual signal. Entropy coding technology is a technology for assigning short codewords to frequently occurring values and long codewords to less frequently occurring values.
[0007] By utilizing these image compression technologies, data regarding images can be efficiently compressed, transmitted, and stored. Summary of the Invention
[0008] Technical issues Embodiments may provide an apparatus, method, and storage medium for performing encoding / decoding on a target block using filtering.
[0009] The embodiments may provide an apparatus, method, and storage medium for performing encoding / decoding on a target block by adaptively selecting a filter from a plurality of filters.
[0010] Technical Solution According to one aspect, there is provided a decoding method, including: constructing a filter candidate list including a plurality of filter candidates; and determining a final filter among the plurality of filter candidates.
[0011] A final filter may be determined based on matching costs of the plurality of filter candidates.
[0012] The filter candidate list can be reconstructed based on the matching costs.
[0013] Reordering of the plurality of filter candidates in the filter candidate list may be performed based on the matching costs.
[0014] The matching cost may be a calculation result using a cost function for samples existing in a template generated using a plurality of filter candidates.
[0015] Multiple filter candidates may be applied to a template region of a template.
[0016] A filter candidate used in a neighboring block of a target block may be used as one of a plurality of filter candidates of the target block.
[0017] According to an aspect, there is provided an encoding method including: constructing a filter candidate list including a plurality of filter candidates; and determining a final filter among the plurality of filter candidates.
[0018] A final filter may be determined based on matching costs of the plurality of filter candidates.
[0019] The filter candidate list can be reconstructed based on the matching costs.
[0020] Reordering of the plurality of filter candidates in the filter candidate list may be performed based on the matching costs.
[0021] The matching cost may be a calculation result using a cost function for samples existing in a template generated using a plurality of filter candidates.
[0022] Multiple filter candidates may be applied to a template region of a template.
[0023] A filter candidate used in a neighboring block of a target block may be used as one of a plurality of filter candidates of the target block.
[0024] According to another aspect, a computer-readable storage medium for storing a bitstream for image decoding is provided, wherein the bitstream includes filter information, a filter candidate list including multiple filter candidates is constructed, and based on the filter information, a final filter is determined among the multiple filter candidates.
[0025] The final filter may be determined based on the matching costs of the plurality of filter candidates.
[0026] The filter candidate list can be reconstructed based on the matching cost.
[0027] Reordering of the plurality of filter candidates in the filter candidate list may be performed based on the matching costs.
[0028] The matching cost may be a calculation result using a cost function for samples existing in a template generated using a plurality of filter candidates.
[0029] Multiple filter candidates may be applied to a template region of a template.
[0030] Beneficial effects Provided are an apparatus, method, and storage medium for performing encoding / decoding on a target block using filtering.
[0031] Provided are an apparatus, method, and storage medium for performing encoding / decoding on a target block by adaptively selecting a filter from a plurality of filters. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 is a block diagram showing a configuration of an embodiment of an encoding device to which the present disclosure is applied; Figure 2 is a block diagram showing a configuration of an embodiment of a decoding device to which the present disclosure is applied; Figure 3 is a diagram schematically illustrating a partition structure of an image when the image is encoded and decoded; Figure 4 is a diagram illustrating a form of prediction units (PUs) that a coding unit (CU) can include; Figure 5 is a diagram illustrating a form of a transform unit (TU) that can be included in a CU; Figure 6 shows the division of blocks according to an example; Figure 7 is a diagram for explaining an embodiment of an intra prediction process; Figure 8 is a diagram showing reference samples used in the intra prediction process; Figure 9 is a diagram for explaining an embodiment of an inter-frame prediction process; Figure 10 shows spatial candidates according to an embodiment; Figure 11 shows the order in which motion information of spatial candidates is added to a merge list according to an embodiment; Figure 12 shows a transform and quantization process according to an example; Figure 13 shows a diagonal scan according to an example; Figure 14 shows a horizontal scan according to an example; Figure 15 shows vertical scanning according to an example; Figure 16 is a configuration diagram of an encoding device according to an embodiment; Figure 17 is a configuration diagram of a decoding device according to an embodiment; Figure 18 illustrating partition boundaries in a geometric partitioning scheme according to an example; Figure 19 shows partition boundaries, partition offsets, and partition angles in a geometric partitioning pattern according to an example; Figure 20 shows a weight map used in various prediction blocks according to specific partition boundaries according to an example; Figure 21 illustrates template matching according to an example; Figures 22a to 22t Various examples of subsampling methods in template matching are shown; Figures 23a to 23n Various other examples of subsampling methods in template matching are shown; Figures 24a to 24n Various other examples of subsampling methods in template matching are shown; Figures 25a to 25j shows subsampling using sub-block units for a region according to an example; Figures 26a to 26l shows subsampling using sub-block units for an area adjacent to a block according to an example; Figures 27 to 31 shows a search pattern in template matching according to an example; Figure 27 showing a first relationship between a search pattern and a resolution according to an example; Figure 28 showing a second relationship between search pattern and resolution according to an example; Figure 29 showing a third relationship between search pattern and resolution according to an example; Figure 30 showing a fourth relationship between search pattern and resolution according to an example; Figure 31 showing a fifth relationship between search pattern and resolution according to an example; Figure 32 showing a sixth relationship between search pattern and resolution according to an example; Figure 33A first template configuration method in an affine mode according to an example is shown; Figure 34 A second template configuration method in an affine mode according to an example is shown; Figure 35a showing a first location from which motion information is derived according to an example; Figure 35b showing a second location from which motion information is derived according to an example; Figure 36 shows two-sided matching according to an example; Figure 37 shows a sample point at which current filtering is performed and samples adjacent to the sample point according to an example; Figure 38 is a flowchart of an encoding method according to an embodiment; Figure 39 is a flowchart of a decoding method according to an embodiment; Figure 40 shows determination of a final filter not including encoding / decoding according to an example; Figure 41 shows the determination of the final filter according to an example without the construction of a list in encoding; Figure 42 shows determination of a final filter according to an example without construction of a list in decoding; Figure 43 shows the determination of the final filter according to an example without encoding / decoding and construction of a list; Figure 44 is a flowchart of an encoding method using template matching cost according to an embodiment; Figure 45 is a flowchart of a decoding method using template matching cost according to an embodiment; Figure 46 shows determination of a final filter not including encoding / decoding according to an example; Figure 47 shows vertical sample boundaries and horizontal sample boundaries for a target block according to an example; Figure 48 shows vertical sample boundaries and horizontal sample boundaries for a target block according to another example; Figure 49 shows values required for determining filter candidates according to an example; Figure 50 showing blocks that are not adjacent to a target block according to an example; Figure 51 shows luma component positions corresponding to positions of chroma component samples according to an example; Figure 52shows intra prediction of a template region according to an example; Figure 53 showing the locations of neighboring spatial candidates according to an example; Figure 54 showing the positions of temporal candidates according to an example; Figure 55 shows the positions of shift time candidates according to an example; Figure 56 shows positions for selecting neighboring motion vectors; and Figure 57 The derivation of shifted temporal candidates based on the current motion vector according to an example is shown. DETAILED DESCRIPTION
[0033] The present invention can be variously modified and can have various embodiments, and specific embodiments will be described in detail below with reference to the accompanying drawings. However, it should be understood that these embodiments are not intended to limit the present invention to the specific disclosed forms, and they include all changes, equivalent forms or modified forms included in the spirit and scope of the present invention.
[0034] The following exemplary embodiments will be described in detail with reference to the accompanying drawings showing specific embodiments. These embodiments are described so that those of ordinary skill in the art to which the present disclosure pertains can easily implement them. It should be noted that the various embodiments are different from one another, but do not need to be mutually exclusive. For example, the specific shapes, structures, and characteristics described herein may be implemented as the other embodiments without departing from the spirit and scope of the other embodiments associated with one embodiment. In addition, it should be understood that the position or arrangement of the various components in each disclosed embodiment can be changed without departing from the spirit and scope of the embodiments. Therefore, the attached detailed description is not intended to limit the scope of the present disclosure, and the scope of the exemplary embodiments is limited only by the attached claims and their equivalents (as long as they are appropriately described).
[0035] In the accompanying drawings, like reference numerals are used to designate the same or similar functions in various aspects. The shapes, sizes, etc. of components in the accompanying drawings may be exaggerated to make the description clear.
[0036] Terms such as "first" and "second" may be used to describe various components, but the components are not limited by the terms. The terms are only used to distinguish one component from another. For example, a first component may be referred to as a second component without departing from the scope of this specification. Similarly, a second component may be referred to as a first component. The term "and / or" may include a combination of multiple related description items or any one of the multiple related description items.
[0037] It will be understood that when a component is referred to as being “connected” or “coupled” to another component, the two components may be directly connected or coupled to each other, or intervening components may be present between the two components. On the other hand, it will be understood that when a component is referred to as being “directly connected or coupled,” there are no intervening components between the two components.
[0038] The components described in the embodiments are shown separately to indicate different feature functions, but this does not mean that each component is formed by a separate piece of hardware or software. That is, for the convenience of description, multiple components are arranged and included separately. For example, at least two components of the multiple components can be integrated into a single component. Conversely, a component can be divided into multiple components. As long as it does not depart from the essence of this specification, embodiments in which multiple components are integrated or embodiments in which some components are separated are included in the scope of this specification.
[0039] The terms used in the embodiments are only used to describe specific embodiments and are not intended to limit the present invention. Singular expressions include plural expressions, unless the contrary description is specifically indicated in the context. In the embodiments, it should be understood that terms such as "including" or "having" are only intended to indicate the presence of features, numbers, steps, operations, components, parts or a combination thereof, and are not intended to exclude the possibility of one or more other features, numbers, steps, operations, components, parts or a combination thereof being present or added. That is, in the embodiments, the expression describing a component "including" a specific component means that other components may be included in the practice of the present invention or the scope of the technical spirit of the present invention, but does not exclude the presence of components other than the specific components.
[0040] In an embodiment, the term "at least one" may mean one of one or more quantities (such as 1, 2, 3, and 4). In an embodiment, the term "plurality" may mean one of two or more quantities (such as 2, 3, and 4).
[0041] Some components of the embodiments are not essential components for performing essential functions, but may be optional components used only to improve performance. The embodiments may be implemented using only the components necessary to achieve the essence of the embodiments. For example, a structure including only the essential components (excluding optional components used only to improve performance) is also included in the scope of the embodiments.
[0042] The embodiments will be described in detail below with reference to the accompanying drawings so that those skilled in the art can easily implement the embodiments. In the following description of the embodiments, detailed descriptions of well-known functions or configurations that are deemed to obscure the main points of this specification will be omitted. In addition, the same reference numerals are used to designate the same components throughout the drawings, and repeated descriptions of the same components will be omitted.
[0043] Hereinafter, "image" may refer to a single frame constituting a video, or may refer to the video itself. For example, "encoding and / or decoding an image" may refer to "encoding and / or decoding a video" or "encoding and / or decoding any one of a plurality of images constituting a video."
[0044] Hereinafter, the terms "video" and "moving picture" may be used to have the same meaning and may be used interchangeably with each other.
[0045] Hereinafter, a target image may be an encoding target image to be encoded and / or a decoding target image to be decoded. Furthermore, a target image may be an input image input to an encoding device or an input image input to a decoding device. Furthermore, a target image may be a current image, that is, a target image currently to be encoded and / or decoded. For example, the terms "target image" and "current image" may be used interchangeably to have the same meaning.
[0046] Hereinafter, the terms "image," "picture," "frame," and "screen" may be used to have the same meaning and may be used interchangeably with each other.
[0047] Hereinafter, a target block may be an encoding target block (i.e., a target to be encoded) and / or a decoding target block (i.e., a target to be decoded). Furthermore, a target block may be a current block, i.e., a target to be currently encoded and / or decoded. Herein, the terms "target block" and "current block" may be used to have the same meaning and may be used interchangeably. The current block may refer to an encoding target block that is an encoding target during encoding and / or a decoding target block that is a decoding target during decoding. Furthermore, the current block may be at least one of a coding block, a prediction block, a residual block, and a transform block.
[0048] Hereinafter, the terms "block" and "unit" may be used to have the same meaning and may be used interchangeably with each other. Alternatively, a "block" may refer to a specific unit.
[0049] Hereinafter, the terms "region" and "segment" are used interchangeably with each other.
[0050] In the following embodiments, specific information, data, flags, indexes, elements, and attributes may have their own values. The value "0" corresponding to each of the information, data, flags, indexes, elements, and attributes may indicate false, logically false, or a first predefined value. In other words, the values "0", false, logically false, and the first predefined value may be used interchangeably. The value "1" corresponding to each of the information, data, flags, indexes, elements, and attributes may indicate true, logically true, or a second predefined value. In other words, the values "1", true, logically true, and the second predefined value may be used interchangeably.
[0051] When a variable such as i or j is used to indicate a row, column, or index, the value i may be an integer 0 or greater than 0, or may be an integer 1 or greater than 1. In other words, in an embodiment, each of the rows, columns, and indexes may be counted starting from 0, or may be counted starting from 1.
[0052] In an embodiment, the term “one or more” or the term “at least one” may mean the term “plurality.” The term “one or more” or the term “at least one” may be used interchangeably with “plurality.”
[0053] Hereinafter, terms to be used in the embodiments will be described.
[0054] Encoder: An encoder refers to a device for performing encoding. In other words, an encoder can refer to an encoding device.
[0055] Decoder: A decoder refers to a device for performing decoding. In other words, a decoder can refer to a decoding device.
[0056] Unit: A unit may refer to a unit of image encoding and decoding. The terms "unit" and "block" may be used to have the same meaning and may be used interchangeably with each other.
[0057] – A cell may be an M×N array of samples. Each of M and N may be a positive integer. A cell may generally represent a two-dimensional array of samples.
[0058] During image encoding and decoding, a "unit" may be a region generated by partitioning an image. In other words, a "unit" may be a specified region within an image. A single image may be partitioned into multiple units. Alternatively, an image may be partitioned into sub-parts, and a unit may represent each sub-part when encoding or decoding is performed on the partitioned sub-parts.
[0059] – During the encoding and decoding of an image, predefined processing can be performed on each unit according to the type of the unit.
[0060] – According to functions, unit types may be classified into macro units, coding units (CUs), prediction units (PUs), residual units, transform units (TUs), and the like. Alternatively, according to functions, a unit may represent a block, a macro block, a coding tree unit, a coding tree block, a coding unit, a coding block, a prediction unit, a prediction block, a residual unit, a residual block, a transform unit, a transform block, and the like. For example, a target unit as a target of encoding and / or decoding may be at least one of a CU, a PU, a residual unit, and a TU.
[0061] The term “unit” may mean information including a luma component block, a chroma component block corresponding to the luma component block, and syntax elements for each block, such that the unit is designated to be distinguished from the block.
[0062] The size and shape of the unit can be implemented differently. Furthermore, the unit can have any of a variety of sizes and shapes. Specifically, the shape of the unit can include not only a square but also geometric shapes that can be represented in two dimensions (2D), such as a rectangle, trapezoid, triangle, and pentagon.
[0063] In addition, the unit information may include one or more of a unit type, a unit size, a unit depth, a unit encoding order, a unit decoding order, etc. For example, the unit type may indicate one of CU, PU, residual unit, and TU.
[0064] - A cell can be partitioned into sub-cells, each sub-cell having a size smaller than that of the associated cell.
[0065] Depth: Depth may indicate the degree to which a cell is partitioned. In addition, the depth of a cell may indicate the level at which the corresponding cell exists when the cell is represented by a tree structure.
[0066] - The cell partition information may include a depth indicating the depth of the cell. The depth may indicate the number of times the cell is partitioned and / or the extent to which the cell is partitioned.
[0067] – In a tree structure, the root node can be considered to have the smallest depth and the leaf nodes can be considered to have the largest depth. The root node can be the highest (top) node. The leaf nodes can be the lowest nodes.
[0068] A single cell can be hierarchically partitioned into multiple sub-cells, with the sub-cell having depth information based on a tree structure. In other words, a cell and the sub-cells generated by partitioning the cell may correspond to a node and its child nodes, respectively. Each partitioned sub-cell may have a cell depth. Since the depth indicates the number of times a cell has been partitioned and / or the extent to which a cell has been partitioned, the sub-cell partition information may include information about the size of the sub-cell.
[0069] In the tree structure, the top node may correspond to the initial node before partitioning. The top node may be referred to as the "root node." Furthermore, the root node may have the smallest depth value. Here, the depth of the top node may be level "0."
[0070] - A node at a depth level of "1" may represent a cell generated when the original cell is partitioned once. A node at a depth level of "2" may represent a cell generated when the original cell is partitioned twice.
[0071] - Leaf nodes at depth level "n" may represent cells generated when the initial cell is partitioned n times.
[0072] A leaf node may be a bottom node that cannot be partitioned further. The depth of a leaf node may be a maximum level. For example, a predefined value for the maximum level may be 3.
[0073] – QT depth may indicate the depth for a four-partition zone. BT depth may indicate the depth for a two-partition zone. TT depth may indicate the depth for a three-part zone.
[0074] – Sample: Sample can be the basic unit of building blocks. Available from 0 to 2 according to the bit depth (Bd). Bd A value of -1 is used to represent a sample point.
[0075] – A sample can be a pixel or a pixel value.
[0076] - Hereinafter, the terms "pixel" and "sample" may be used to have the same meaning and may be used interchangeably with each other.
[0077] Coding Tree Unit (CTU): A CTU can be composed of a single luma component (Y) coding tree block and two chroma components (i.e., Cb and Cr) coding tree blocks related to the luma component coding tree block. In addition, a CTU can represent information including the above blocks and syntax elements for each block.
[0078] – Each coding tree unit (CTU) may be partitioned using one or more partitioning methods, such as quadtree (QT), binary tree (BT), and ternary tree (TT), to configure subunits such as coding units, prediction units, and transform units. Quadtree may represent a quadtree. In addition, each coding tree unit may be partitioned using a multi-type tree (MTT) using one or more partitioning methods.
[0079] – “CTU” may be used as a term to designate a pixel block as a processing unit in image decoding and encoding processes (such as in the case of partitioning an input image).
[0080] Coding Tree Block (CTB): “CTB” may be used as a term designating any one of a Y coding tree block, a Cb coding tree block, and a Cr coding tree block.
[0081] Neighboring block: A neighboring block (or neighboring block) may refer to a block adjacent to the target block. A neighboring block may refer to a reconstructed neighboring block.
[0082] Hereinafter, the terms “neighboring block” and “adjacent block” may be used to have the same meaning and may be used interchangeably with each other.
[0083] The neighboring block may represent a reconstructed neighboring block.
[0084] Spatially neighboring blocks: Spatially neighboring blocks may be blocks that are spatially adjacent to the target block. Neighboring blocks may include spatially neighboring blocks.
[0085] – A target block and spatially neighboring blocks may be included in a target picture.
[0086] - The spatially neighboring block may mean a block whose boundary contacts the target block or a block located within a predetermined distance from the target block.
[0087] - A spatially adjacent block may refer to a block adjacent to a vertex of the target block. Here, a block adjacent to a vertex of the target block may refer to a block vertically adjacent to a neighboring block horizontally adjacent to the target block or a block horizontally adjacent to a neighboring block vertically adjacent to the target block.
[0088] Temporally neighboring blocks: Temporally neighboring blocks may be blocks that are temporally adjacent to the target block. Neighboring blocks may include temporally neighboring blocks.
[0089] – Temporally neighboring blocks may include co-located blocks (col blocks).
[0090] The col block may be a block in a previously reconstructed co-located picture (col picture). The position of the col block in the col picture may correspond to the position of the target block in the target picture. Alternatively, the position of the col block in the col picture may be equal to the position of the target block in the target picture. The col picture may be a picture included in the reference picture list.
[0091] - The temporally neighboring block may be a block temporally adjacent to the spatially neighboring block of the target block.
[0092] Prediction mode: The prediction mode may be information indicating a mode for intra prediction or a mode for inter prediction.
[0093] Prediction unit: A prediction unit may be a basic unit for prediction such as inter prediction, intra prediction, inter compensation, intra compensation, and motion compensation.
[0094] A single prediction unit can be divided into multiple partitions or sub-prediction units of smaller size. The multiple partitions can also be the basic units when performing prediction or compensation. Partitions generated by dividing the prediction unit can also be prediction units.
[0095] Prediction unit partition: A prediction unit partition may be a shape into which a prediction unit is divided.
[0096] Reconstructed neighboring unit: The reconstructed neighboring unit may be a unit that is adjacent to the target unit and has been decoded and reconstructed.
[0097] The reconstructed neighboring cell may be a cell that is spatially adjacent to the target cell or temporally adjacent to the target cell.
[0098] - The reconstructed spatially neighboring unit may be a unit included in the target picture that has been reconstructed through encoding and / or decoding.
[0099] –The reconstructed temporally adjacent unit may be a unit included in the reference image and already reconstructed by encoding and / or decoding. The position of the reconstructed temporally adjacent unit in the reference image may be the same as the position of the target unit in the target picture, or may correspond to the position of the target unit in the target picture. In addition, the reconstructed temporally adjacent unit may be a block adjacent to the corresponding block in the reference image. Here, the position of the corresponding block in the reference image may correspond to the position of the target block in the target image. Here, the fact that the positions of the blocks correspond to each other may mean that the positions of the blocks are the same as each other, may mean that one block is included in another block, or may mean that one block occupies a specific position in another block.
[0100] Sub-picture: A picture can be divided into one or more sub-pictures. A sub-picture can be composed of one or more tile rows and one or more tile columns.
[0101] A sub-picture may be a region in a picture having a square shape or a rectangular (ie, non-square rectangular) shape. In addition, a sub-picture may include one or more CTUs.
[0102] - A sub-picture can be a rectangular area of one or more slices in a picture.
[0103] - A sub-picture may include one or more tiles, one or more bricks and / or one or more slices.
[0104] Tile: A tile may be an area in a picture having a square shape or a rectangular (ie, non-square rectangular) shape.
[0105] – A tile may include one or more CTUs.
[0106] – A parallel block can be partitioned into one or more partitions.
[0107] Block: A block may represent one or more CTU rows in a tile.
[0108] A tile can be partitioned into one or more partitions, each of which can include one or more CTU rows.
[0109] – Parallel blocks that are not partitioned into two parts can also represent partitions.
[0110] Slice: A slice may include one or more tiles in a picture. Alternatively, a slice may include one or more partitions in a tile.
[0111] - A sub-picture may comprise one or more slices that together cover a rectangular area of the picture. Thus, every sub-picture boundary is also always a slice boundary, and every vertical sub-picture boundary is also always a vertical tile boundary.
[0112] Parameter set: The parameter set may correspond to header information in the internal structure of the bitstream.
[0113] The parameter set may include at least one of a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptation parameter set (APS), a decoding parameter set (DPS), and the like.
[0114] – Information signaled by each parameter set can be applied to pictures that reference the corresponding parameter set. For example, information in a VPS can be applied to pictures that reference the VPS. Information in an SPS can be applied to pictures that reference the SPS. Information in a PPS can be applied to pictures that reference the PPS.
[0115] – Each parameter set can refer to a higher parameter set. For example, PPS can refer to SPS. SPS can refer to VPS.
[0116] – In addition, the parameter set may include a parallel block group, slice header information, and parallel block header information. A parallel block group may be a group including multiple parallel blocks. In addition, the meaning of "parallel block group" may be the same as the meaning of "slice".
[0117] Rate-distortion optimization: The encoding device may use rate-distortion optimization to provide high encoding efficiency by utilizing a combination of the following: the size of a coding unit (CU), prediction mode, the size of a prediction unit (PU), motion information, and the size of a transform unit (TU).
[0118] – The rate-distortion optimization scheme can calculate the rate-distortion cost of each combination to select the best combination from these combinations. The equation “ " is used to calculate the rate-distortion cost. Generally, the combination that minimizes the rate-distortion cost can be selected as the optimal combination under the rate-distortion optimization scheme.
[0119] – D may represent distortion. D may be the average of the squares of the differences between the original transform coefficients and the reconstructed transform coefficients in the transform unit (i.e., the mean square error).
[0120] – R may represent the rate, which may use relevant context information to represent the bit rate.
[0121] – R may include not only encoding parameter information (such as prediction mode, motion information, and coding block flag), but also bits generated by encoding transform coefficients.
[0122] – The encoding device may perform processes such as inter-frame prediction and / or intra-frame prediction, transformation, quantization, entropy coding, inverse quantization (dequantization) and / or inverse transformation in order to calculate accurate D and R. These processes significantly increase the complexity of the encoding device.
[0123] – Bitstream: A bitstream may refer to a stream of bits including encoded image information.
[0124] Parsing: Parsing can be the decision of the value of the syntax element by performing entropy decoding on the bitstream. Alternatively, the term "parsing" can refer to such entropy decoding itself.
[0125] Symbol: A symbol may be at least one of a syntax element, a coding parameter, and a transform coefficient of a coding target unit and / or a decoding target unit. In addition, a symbol may be a target of entropy coding or a result of entropy decoding.
[0126] Reference picture: A reference picture may be an image referenced by a unit in order to perform inter-frame prediction or motion compensation. Alternatively, a reference picture may be an image including a reference unit referenced by a target unit in order to perform inter-frame prediction or motion compensation.
[0127] Hereinafter, the terms “reference picture” and “reference image” may be used to have the same meaning and may be used interchangeably with each other.
[0128] Reference picture list: A reference picture list may be a list including one or more reference images used for inter prediction or motion compensation.
[0129] – The types of reference picture lists may include combined list (LC), list 0 (L0), list 1 (L1), list 2 (L2), list 3 (L3), etc.
[0130] – For inter prediction, one or more reference picture lists may be used.
[0131] Inter-frame prediction indicator: The inter-frame prediction indicator may indicate the inter-frame prediction direction for the target unit. Inter-frame prediction can be one of unidirectional prediction and bidirectional prediction. Optionally, the inter-frame prediction indicator may indicate the number of reference pictures used to generate the prediction unit for the target unit. Optionally, the inter-frame prediction indicator may indicate the number of prediction blocks used for inter-frame prediction or motion compensation for the target unit.
[0132] Prediction list utilization flag: The prediction list utilization flag may indicate whether to use at least one reference picture in a specific reference picture list to generate a prediction unit.
[0133] – The prediction list utilization flag can be used to derive the inter-frame prediction indicator. Conversely, the prediction list utilization flag can be used to derive the inter-frame prediction indicator. For example, if the prediction list utilization flag indicates "0" (as a first value), it can indicate that the prediction block is not generated using the reference pictures in the reference picture list for the target unit. If the prediction list utilization flag indicates "1" (as a second value), it can indicate that the prediction unit is generated using the reference picture list for the target unit.
[0134] Reference picture index: A reference picture index may be an index indicating a specific reference picture in a reference picture list.
[0135] Picture Order Count (POC): The POC value of a picture may indicate the order in which the corresponding picture is displayed.
[0136] Motion Vector (MV): A motion vector is a 2D vector used for inter-frame prediction or motion compensation. A motion vector represents the offset between a target image and a reference image.
[0137] – For example, you can use a command such as (mv x , mv y ) to express MV. mv x Can indicate the horizontal component, mv y May indicate the vertical component.
[0138] -Search range: The search range may be a 2D area where a search for an MV is performed during inter prediction. For example, the size of the search range may be M×N. M and N may be positive integers, respectively.
[0139] Motion vector candidate: A motion vector candidate may be a block that is a prediction candidate when a motion vector is predicted or a motion vector of a block that is a prediction candidate.
[0140] – A motion vector candidate may be included in a motion vector candidate list.
[0141] Motion vector candidate list: A motion vector candidate list may be a list configured using one or more motion vector candidates.
[0142] Motion vector candidate index: The motion vector candidate index may be an indicator for indicating a motion vector candidate in the motion vector candidate list. Alternatively, the motion vector candidate index may be an index of a motion vector predictor.
[0143] Motion information: The motion information may be information including at least one of a reference picture list, a reference image, a motion vector candidate, a motion vector candidate index, a merge candidate, and a merge index, as well as a motion vector, a reference picture index, and an inter prediction indicator.
[0144] Merge candidate list: A merge candidate list may be a list configured using one or more merge candidates.
[0145] Merge candidate: A merge candidate may be a spatial merge candidate, a temporal merge candidate, a combined merge candidate, a combined bi-predictive merge candidate, a history-based candidate, an average merge candidate based on the average of two candidates, a zero merge candidate, etc. The merge candidate may include an inter prediction indicator and may include motion information such as prediction type information, reference picture index for each list, motion vector, prediction list utilization flag, and an inter prediction indicator.
[0146] Merge index: The merge index may be an indicator for indicating a merge candidate in a merge candidate list.
[0147] – The merge index may indicate a reconstructed unit for deriving a merge candidate among reconstructed units spatially adjacent to the target unit and reconstructed units temporally adjacent to the target unit.
[0148] – The merge index may indicate at least one of a plurality of pieces of motion information of the merge candidate.
[0149] Transform unit: A transform unit can be a basic unit for residual signal encoding and / or residual signal decoding (such as transformation, inverse transformation, quantization, inverse quantization, transform coefficient encoding, and transform coefficient decoding). A single transform unit can be partitioned into multiple sub-transform units of smaller size. Here, the transform may include one or more of a primary transform and a secondary transform, and the inverse transform may include one or more of a primary inverse transform and a secondary inverse transform.
[0150] Scaling: Scaling can be referred to as the process of multiplying the levels of the transform coefficients by a factor.
[0151] – As a result of scaling the transform coefficient levels, transform coefficients may be generated. Scaling may also be referred to as “inverse quantization”.
[0152] Quantization Parameter (QP): A quantization parameter can be a value used to generate transform coefficient levels for a transform coefficient during quantization. Alternatively, a quantization parameter can also be a value used to generate a transform coefficient by scaling the transform coefficient levels during inverse quantization. Alternatively, a quantization parameter can be a value mapped to a quantization step size.
[0153] Delta quantization parameter: Delta quantization parameter represents the difference between the target unit's quantization parameter and the predicted quantization parameter.
[0154] Scan: Scanning can refer to a method of arranging the order of coefficients in a cell, block, or matrix. For example, a method for arranging a 2D array as a one-dimensional (1D) array can be called "scanning." Alternatively, a method for arranging a 1D array as a 2D array can also be called "scanning" or "inverse scanning."
[0155] Transform coefficient: The transform coefficient may be a coefficient value generated when the encoding device performs transformation. Alternatively, the transform coefficient may be a coefficient value generated when the decoding device performs at least one of entropy decoding and inverse quantization.
[0156] - A quantized level or a quantized transform coefficient level generated by applying quantization to a transform coefficient or a residual signal may also be included in the meaning of the term "transform coefficient".
[0157] Quantization level: The quantization level may be a value generated when the encoding device performs quantization on the transform coefficient or residual signal. Alternatively, the quantization level may be a value that is a target of inverse quantization when the decoding device performs inverse quantization.
[0158] - A quantized transform coefficient level as a result of transformation and quantization may also be included in the meaning of the quantization level.
[0159] Non-zero transform coefficient: A non-zero transform coefficient may be a transform coefficient having a value other than 0, or may be a transform coefficient level having a value other than 0. Alternatively, a non-zero transform coefficient may be a transform coefficient having a value with a magnitude other than 0, or may be a transform coefficient level having a value with a magnitude other than 0.
[0160] Quantization Matrix: A quantization matrix may be a matrix used in a quantization process or an inverse quantization process to improve the subjective or objective image quality of an image. A quantization matrix may also be referred to as a "scaling list."
[0161] Quantization matrix coefficient: A quantization matrix coefficient can be each element in the quantization matrix. A quantization matrix coefficient can also be called a "matrix coefficient."
[0162] Default matrix: The default matrix may be a quantization matrix predefined by the encoding device and the decoding device.
[0163] Non-default matrix: A non-default matrix may be a quantization matrix that is not pre-defined by the encoding device and the decoding device. The non-default matrix may refer to a quantization matrix that is signaled by a user from the encoding device to the decoding device.
[0164] Most Probable Mode (MPM): The MPM may represent an intra prediction mode that is highly likely to be used for intra prediction for a target block.
[0165] The encoding apparatus and the decoding apparatus may determine one or more MPMs based on encoding parameters related to the target block and properties of an entity related to the target block.
[0166] The encoding device and the decoding device may determine one or more MPMs based on the intra-frame prediction mode of the reference block. The reference block may include multiple reference blocks. The multiple reference blocks may include a spatially neighboring block adjacent to the left of the target block and a spatially neighboring block adjacent to the top of the target block. In other words, depending on which intra-frame prediction mode has been used for the reference block, one or more different MPMs may be determined.
[0167] - One or more MPMs may be determined in the same manner in both the encoding device and the decoding device. That is, the encoding device and the decoding device may share the same MPM list including one or more MPMs.
[0168] MPM list: The MPM list may be a list including one or more MPMs. The number of the one or more MPMs in the MPM list may be predefined.
[0169] MPM indicator: The MPM indicator may indicate an MPM to be used for intra prediction for a target block among one or more MPMs in an MPM list. For example, the MPM indicator may be an index for the MPM list.
[0170] Since the MPM list is determined in the same manner in both the encoding device and the decoding device, there may be no need to transmit the MPM list itself from the encoding device to the decoding device.
[0171] – The MPM indicator may be signaled from the encoding apparatus to the decoding apparatus. Since the MPM indicator is signaled, the decoding apparatus may determine an MPM to be used for intra prediction for a target block among the MPMs in the MPM list.
[0172] MPM usage indicator: The MPM usage indicator may indicate whether an MPM usage mode is to be used for prediction for a target block. The MPM usage mode may be a mode of determining an MPM to be used for intra prediction for a target block using an MPM list.
[0173] – An MPM usage indicator may be signaled from an encoding device to a decoding device.
[0174] Signaling: "Signaling" may mean that information is sent from an encoding device to a decoding device. Alternatively, "signaling" may mean that information is included in a bitstream or recording medium by the encoding device. Information signaled by the encoding device can be used by the decoding device.
[0175] The encoding device may generate coded information by encoding information to be signaled. The coded information may be transmitted from the encoding device to the decoding device. The decoding device may obtain the information by decoding the transmitted coded information. Here, the encoding may be entropy encoding, and the decoding may be entropy decoding.
[0176] Selective signaling: Information can be selectively signaled. Selective signaling of information can mean that an encoding device selectively includes information in a bitstream or recording medium (depending on specific conditions). Selective signaling of information can mean that a decoding device selectively extracts information from a bitstream (depending on specific conditions).
[0177] Omission of signaling: Signaling for information may be omitted. Regarding information, omission of signaling for information may mean that the encoding device (under certain conditions) does not include the information in the bitstream or recording medium. Omission of signaling for information may also mean that the decoding device (under certain conditions) does not extract the information from the bitstream.
[0178] Statistics: Variables, encoding parameters, constants, and the like may have computable values. Statistics are values generated by performing calculations (operations) on the values of a specified target. For example, a statistic may indicate one or more of the following: the average, weighted average, weighted sum, minimum, maximum, mode, median, and interpolated values of the values of a specific variable, encoding parameter, constant, and the like.
[0179] Figure 1 is a block diagram showing a configuration of an embodiment of an encoding device to which the present disclosure is applied.
[0180] The encoding device 100 may be an encoder, a video encoding device, or an image encoding device. A video may include one or more images (pictures). The encoding device 100 may sequentially encode one or more images of a video.
[0181] Reference Figure 1 , the encoding device 100 includes an inter-frame prediction unit 110, an intra-frame prediction unit 120, a switch 115, a subtractor 125, a transform unit 130, a quantization unit 140, an entropy encoding unit 150, an inverse quantization (dequantization) unit 160, an inverse transform unit 170, an adder 175, a filter unit 180 and a reference picture buffer 190.
[0182] The encoding apparatus 100 may perform encoding on a target image using an intra mode and / or an inter mode. In other words, the prediction mode of the target block may be one of the intra mode and the inter mode.
[0183] Hereinafter, the terms “intra mode”, “intra prediction mode”, “intra-screen mode” and “intra prediction mode” may be used to have the same meaning and may be used interchangeably with each other.
[0184] Hereinafter, the terms “inter mode”, “inter prediction mode”, “inter-picture mode” and “inter-picture prediction mode” may be used to have the same meaning and may be used interchangeably with each other.
[0185] Hereinafter, the term "image" may refer to only a portion of an image, or may refer to a block. Furthermore, processing of an "image" may refer to sequential processing of a plurality of blocks.
[0186] In addition, the encoding device 100 can generate a bitstream including the encoded information by encoding the target image, and can output and store the generated bitstream. The generated bitstream can be stored in a computer-readable storage medium and can be streamed via a wired and / or wireless transmission medium.
[0187] When the intra mode is used as the prediction mode, the switch 115 may switch to the intra mode. When the inter mode is used as the prediction mode, the switch 115 may switch to the inter mode.
[0188] The encoding apparatus 100 may generate a prediction block of the target block. In addition, after having generated the prediction block, the encoding apparatus 100 may encode a residual block for the target block using a residual between the target block and the prediction block.
[0189] When the prediction mode is intra mode, the intra prediction unit 120 may use pixels of a previously encoded / decoded neighboring block adjacent to the target block as reference samples. The intra prediction unit 120 may use the reference samples to perform spatial prediction on the target block and may generate prediction samples for the target block via spatial prediction. The prediction samples may represent samples in the prediction block.
[0190] The inter prediction unit 110 may include a motion prediction unit and a motion compensation unit.
[0191] When the prediction mode is inter mode, the motion prediction unit may search for an area in the reference image that best matches the target block during the motion prediction process, and may derive a motion vector for the target block and the area based on the area found. Here, the motion prediction unit may use the search range as the target area for the search.
[0192] The reference image may be stored in the reference picture buffer 190. More specifically, when encoding and / or decoding of a reference image has been processed, the encoded and / or decoded reference image may be stored in the reference picture buffer 190.
[0193] Since decoded pictures are stored, the reference picture buffer 190 may be a decoded picture buffer (DPB).
[0194] The motion compensation unit may generate a prediction block for the target block by performing motion compensation using a motion vector. Here, the motion vector may be a two-dimensional (2D) vector used for inter-frame prediction. In addition, the motion vector may indicate the offset between the target image and the reference image.
[0195] When the motion vector has a value other than an integer, the motion prediction unit and the motion compensation unit may generate a prediction block by applying an interpolation filter to a partial area of the reference image. In order to perform inter-frame prediction or motion compensation, it may be determined which mode among the skip mode, merge mode, advanced motion vector prediction (AMVP) mode, and current picture reference mode corresponds to a method for predicting and compensating the motion of the PU included in the CU based on the CU, and inter-frame prediction or motion compensation may be performed according to the mode.
[0196] The subtractor 125 may generate a residual block, which is the difference between the target block and the prediction block. The residual block may also be referred to as a "residual signal."
[0197] The residual signal may be the difference between the original signal and the predicted signal. Alternatively, the residual signal may be a signal generated by transforming or quantizing the difference between the original signal and the predicted signal, or a signal generated by transforming and quantizing the difference. The residual block may be a residual signal for a block unit.
[0198] The transform unit 130 may generate a transform coefficient by transforming the residual block, and may output the generated transform coefficient. Here, the transform coefficient may be a coefficient value generated by transforming the residual block.
[0199] The transform unit 130 may use one of a plurality of predefined transform methods when performing the transform.
[0200] The plurality of predefined transform methods may include discrete cosine transform (DCT), discrete sine transform (DST), Karhunen-Loeve transform (KLT), and the like.
[0201] The transform method used to transform the residual block may be determined based on at least one of the encoding parameters for the target block and / or the neighboring blocks. For example, the transform method may be determined based on at least one of the inter-prediction mode for the PU, the intra-prediction mode for the PU, the size of the TU, and the shape of the TU. Alternatively, transform information indicating the transform method may be signaled from the encoding device 100 to the decoding device 200.
[0202] When the transform skip mode is used, the transform unit 130 may omit the operation of transforming the residual block.
[0203] By performing quantization on the transform coefficients, quantized transform coefficient levels or quantized levels may be generated. Hereinafter, in an embodiment, each of the quantized transform coefficient levels and the quantized levels may also be referred to as a 'transform coefficient'.
[0204] The quantization unit 140 may generate quantized transform coefficient levels (i.e., quantized levels or quantized coefficients) by quantizing the transform coefficients according to the quantization parameters. The quantization unit 140 may output the generated quantized transform coefficient levels. In this case, the quantization unit 140 may quantize the transform coefficients using a quantization matrix.
[0205] The entropy encoding unit 150 may generate a bitstream by performing entropy encoding based on a probability distribution based on a value calculated by the quantization unit 140 and / or an encoding parameter value calculated during encoding. The entropy encoding unit 150 may output the generated bitstream.
[0206] The entropy encoding unit 150 may perform entropy encoding on information about pixels of an image and information required for decoding the image. For example, the information required for decoding the image may include syntax elements and the like.
[0207] When entropy coding is applied, fewer bits are allocated to symbols that appear more frequently, and more bits are allocated to symbols that appear less frequently. Since symbols are represented by this allocation, the size of the bit string used to encode the target symbol can be reduced. Therefore, entropy coding can improve the compression performance of video encoding.
[0208] Furthermore, to perform entropy coding, the entropy coding unit 150 may use a coding method such as Exponential Golomb, Context-Adaptive Variable Length Coding (CAVLC), or Context-Adaptive Binary Arithmetic Coding (CABAC). For example, the entropy coding unit 150 may use a variable length coding / code (VLC) table to perform entropy coding. For example, the entropy coding unit 150 may derive a binarization method for a target symbol. Furthermore, the entropy coding unit 150 may derive a probability model for a target symbol / bin. The entropy coding unit 150 may use the derived binarization method, probability model, and context model to perform arithmetic coding.
[0209] The entropy encoding unit 150 may transform coefficients in a 2D block form into a 1D vector form through a transform coefficient scanning method in order to encode quantized transform coefficient levels.
[0210] Coding parameters may be information required for encoding and / or decoding. Coding parameters may include information encoded by encoding device 100 and transmitted from encoding device 100 to a decoding device, and may also include information that can be derived during the encoding or decoding process. For example, the information transmitted to the decoding device may include syntax elements.
[0211] The coding parameters may include not only information (or flags or indices) such as syntax elements that are encoded by the encoding device and sent by the encoding device to the decoding device using a signal, but also information derived in the encoding or decoding process. In addition, the coding parameters may include information required for encoding or decoding an image. For example, the coding parameters may include at least one value of the following items, a combination of the following items, or statistics: the size of the unit / block, the shape / form of the unit / block, the depth of the unit / block, the partition information of the unit / block, the partition structure of the unit / block, information indicating whether the unit / block is partitioned in a quadtree structure, information indicating whether the unit / block is partitioned in a binary tree structure, the partition direction of the binary tree structure (horizontal or vertical), the partition form of the binary tree structure (symmetric partitioning or asymmetric partitioning), information indicating whether the unit / block is partitioned in a ternary tree structure, the partition direction of the ternary tree structure (horizontal or vertical), the partition form of the ternary tree structure (symmetric partitioning or asymmetric partitioning), symmetrical partitioning, etc.), information indicating whether the unit / block is partitioned in a multi-type tree structure, a combination and direction of partitions of the multi-type tree structure (horizontal direction or vertical direction, etc.), a partition form of the multi-type tree structure (symmetrical partitioning or asymmetrical partitioning, etc.), a partition tree of the multi-type tree form (binary tree or ternary tree), a prediction type (intra-frame prediction or inter-frame prediction), an intra-frame prediction mode / direction, an intra-frame luminance prediction mode / direction, an intra-frame chrominance prediction mode / direction, an intra-frame partition information, an inter-frame partition information, a coding block partition flag, a prediction block partition flag, a transform block partition flag, a reference sample filtering method, a reference sample filter tap, a reference sample filter coefficient, a prediction block filtering method method, prediction block filter taps, prediction block filter coefficients, prediction block boundary filtering method, prediction block boundary filter taps, prediction block boundary filter coefficients, inter prediction mode, motion information, motion vector, motion vector difference, reference picture index, inter prediction direction, inter prediction indicator, prediction list utilization flag, reference picture list, reference image, POC, motion vector predictor, motion vector prediction index, motion vector prediction candidate, motion vector candidate list, information indicating whether merge mode is used, merge index, merge candidate, merge candidate list, information indicating whether skip mode is used, type of interpolation filter, taps of interpolation filter, interpolation filter coefficients of a filter, size of a motion vector, precision of motion vector representation, transform type, transform size, information indicating whether a first transform is used, information indicating whether an additional (second) transform is used, first transform selection information (or first transform index), second transform selection information (or second transform index), information indicating the presence or absence of a residual signal, coding block pattern, coding block flag, quantization parameter, residual quantization parameter, quantization matrix, information about a loop filter, information indicating whether a loop filter is applied, coefficients of a loop filter, taps of a loop filter, shape / form of a loop filter, information indicating whether a deblocking filter is applied,Coefficients of a deblocking filter, taps of a deblocking filter, strength of a deblocking filter, shape / form of a deblocking filter, information indicating whether an adaptive sample offset is applied, value of an adaptive sample offset, category of an adaptive sample offset, type of an adaptive sample offset, information indicating whether an adaptive loop filter is applied, coefficients of an adaptive loop filter, taps of an adaptive loop filter, shape / form of an adaptive loop filter, binarization / debinarization method, context model, context model determination method, context model update method, information indicating whether a normal mode is performed, information indicating whether a bypass mode is performed, a significant coefficient flag, a last significant coefficient flag, a coding flag of a coefficient group, position of a last significant coefficient, information indicating whether a value of a coefficient is greater than 1, information indicating whether a value of a coefficient is greater than 2, information indicating whether a value of a coefficient is greater than 3, residual coefficient value information, sign information, reconstructed luminance samples, reconstructed chrominance samples, context binary bits, Bypass binary bits, residual luminance samples, residual chrominance samples, transform coefficients, luminance transform coefficients, chrominance transform coefficients, quantization levels, luminance quantization levels, chrominance quantization levels, transform coefficient levels, transform coefficient level scanning methods, size of the motion vector search area on the decoding device side, shape / form of the motion vector search area on the decoding device side, number of motion vector searches on the decoding device side, size of CTU, minimum block size, maximum block size, maximum block depth, minimum block depth, image display / output order, stripe identification information, stripe type, stripe partition information, parallel block group identification information, parallel block group type, parallel block group partition information, parallel block identification information, parallel block type, parallel block partition information, picture type, bit depth, input sample bit depth, reconstructed sample bit depth, residual sample bit depth, transform coefficient bit depth, quantization level bit depth, information about the luminance signal, information about the chrominance signal, color space of the target block, and color space of the residual block. In addition, the above-mentioned coding parameter related information may also be included in the coding parameters. Information used to calculate and / or derive the above-mentioned coding parameters may also be included in the coding parameters. Information calculated or derived using the above encoding parameters may also be included in the encoding parameters.
[0212] The first transform selection information may indicate a first transform applied to the target block.
[0213] The second transform selection information may indicate a second transform applied to the target block.
[0214] The residual signal may represent the difference between the original signal and the predicted signal. Alternatively, the residual signal may be a signal generated by transforming the difference between the original signal and the predicted signal. Alternatively, the residual signal may be a signal generated by transforming and quantizing the difference between the original signal and the predicted signal. The residual block may be a residual signal for a block.
[0215] Here, signaling information may mean that the encoding device 100 includes entropy-coded information generated by performing entropy encoding on a flag or index in the bitstream, and may mean that the decoding device 200 obtains information by performing entropy decoding on the entropy-coded information extracted from the bitstream. Here, the information may include a flag, an index, etc.
[0216] A signal may refer to information to be transmitted using a signal. Hereinafter, information for an image and a block may be referred to as a "signal." Furthermore, hereinafter, the terms "information" and "signal" may be used interchangeably to have the same meaning. For example, a specific signal may be a signal representing a specific block. An original signal may be a signal representing a target block. A predicted signal may be a signal representing a predicted block. A residual signal may be a signal representing a residual block.
[0217] The bitstream may include information based on a specific syntax. The encoding apparatus 100 may generate a bitstream including information according to the specific syntax. The decoding apparatus 200 may acquire information from the bitstream according to the specific syntax.
[0218] Since the encoding apparatus 100 performs encoding via inter-frame prediction, the encoded target image can be used as a reference image for another image to be subsequently processed. Therefore, the encoding apparatus 100 can reconstruct or decode the encoded target image and store the reconstructed or decoded image as a reference image in the reference picture buffer 190. For decoding, the encoded target image can be dequantized and inversely transformed.
[0219] The quantized levels may be dequantized by the dequantization unit 160 and inversely transformed by the inverse transform unit 170. The dequantization unit 160 may generate dequantized coefficients by performing inverse transform on the quantized levels. The inverse transform unit 170 may generate dequantized and inversely transformed coefficients by performing inverse transform on the dequantized coefficients.
[0220] The dequantized and inverse-transformed coefficients may be added to the prediction block by the adder 175. The dequantized and inverse-transformed coefficients are added to the prediction block, and then a reconstructed block may be generated. Here, the dequantized and / or inverse-transformed coefficients may refer to coefficients on which one or more of dequantization and inverse transformation have been performed, and may also refer to a reconstructed residual block. Here, the reconstructed block may refer to a restored block or a decoded block.
[0221] The reconstructed block may be filtered by the filter unit 180. The filter unit 180 may apply one or more filters selected from the group consisting of a deblocking filter, a sample adaptive offset (SAO) filter, an adaptive loop filter (ALF), and a non-local filter (NLF) to the reconstructed samples, the reconstructed block, or the reconstructed picture. The filter unit 180 may also be referred to as a "loop filter."
[0222] The deblocking filter can eliminate block distortion that occurs at boundaries between blocks in a reconstructed picture. To determine whether to apply the deblocking filter, the number of columns or rows of pixels included in the block and including the basis for determining whether to apply the deblocking filter to the target block can be determined.
[0223] When a deblocking filter is applied to a target block, the applied filter may be different depending on the required strength of the deblocking filter. In other words, among different filters, a filter determined in consideration of the strength of the deblocking filter may be applied to the target block. When a deblocking filter is applied to a target block, one or more filters selected from a long tap filter, a strong filter, a weak filter, and a Gaussian filter may be applied to the target block depending on the required strength of the deblocking filter.
[0224] Also, when vertical filtering and horizontal filtering are performed on a target block, the horizontal filtering and the vertical filtering may be performed in parallel.
[0225] SAO can add an appropriate offset to pixel values to compensate for coding errors. SAO can perform pixel-by-pixel correction on an image to which deblocking has been applied, wherein the correction uses an offset that is the difference between the original image and the image to which deblocking has been applied. To perform offset correction on an image, a method can be used for dividing the pixels included in the image into a specific number of regions, determining the regions to which the offset is applied within the divided regions, and applying the offset to the determined regions. A method can also be used for applying the offset while taking into account edge information for each pixel.
[0226] The ALF can perform filtering based on values obtained by comparing a reconstructed image with an original image. After the pixels included in the image have been divided into a predetermined number of groups, the filter to be applied to each group can be determined, and filtering can be performed differently for each group. Information related to whether an adaptive loop filter is applied can be signaled for each CU. This information can be signaled for the luma signal. The shape and filter coefficients of the ALF to be applied to each block can be different for each block. Alternatively, an ALF having a fixed form can be applied to the block regardless of its characteristics.
[0227] A non-local filter can perform filtering based on a reconstructed block similar to a target block. A region similar to the target block can be selected from the reconstructed image, and statistical properties of the selected similar region can be used to filter the target block. Information on whether to apply a non-local filter can be signaled for each coding unit (CU). Furthermore, the shape and filter coefficients of the non-local filter applied to a block can vary depending on the block.
[0228] The reconstructed block or reconstructed image filtered by the filter unit 180 may be stored as a reference picture in the reference picture buffer 190. The reconstructed block filtered by the filter unit 180 may be part of a reference picture. In other words, the reference picture may be a reconstructed picture composed of the reconstructed blocks filtered by the filter unit 180. The stored reference picture may then be used for inter-frame prediction or motion compensation.
[0229] Figure 2 is a block diagram showing a configuration of an embodiment of a decoding device to which the present disclosure is applied.
[0230] The decoding device 200 may be a decoder, a video decoding device, or an image decoding device.
[0231] Reference Figure 2 , the decoding apparatus 200 may include an entropy decoding unit 210 , an inverse quantization (dequantization) unit 220 , an inverse transform unit 230 , an intra prediction unit 240 , an inter prediction unit 250 , a switch 245 , an adder 255 , a filter unit 260 , and a reference picture buffer 270 .
[0232] The decoding apparatus 200 may receive a bitstream output from the encoding apparatus 100. The decoding apparatus 200 may receive a bitstream stored in a computer-readable storage medium, and may receive a bitstream streamed through a wired / wireless transmission medium.
[0233] The decoding apparatus 200 may perform decoding on a bitstream in an intra mode and / or an inter mode. In addition, the decoding apparatus 200 may generate a reconstructed image or a decoded image through decoding, and may output the reconstructed image or the decoded image.
[0234] For example, an operation of switching to intra mode or inter mode based on the prediction mode used for decoding may be performed by the switch 245. When the prediction mode used for decoding is intra mode, the switch 245 may be operated to switch to intra mode. When the prediction mode used for decoding is inter mode, the switch 245 may be operated to switch to inter mode.
[0235] The decoding device 200 can obtain a reconstructed residual block by decoding the input bit stream and can generate a prediction block. When the reconstructed residual block and the prediction block are obtained, the decoding device 200 can generate a reconstructed block as a target to be decoded by adding the reconstructed residual block to the prediction block.
[0236] The entropy decoding unit 210 may generate symbols by performing entropy decoding on the bitstream based on the probability distribution of the bitstream. The generated symbols may include symbols in the form of quantized transform coefficient levels (i.e., quantized levels or quantized coefficients). Here, the entropy decoding method may be similar to the entropy encoding method described above. That is, the entropy decoding method may be an inverse process of the entropy encoding method described above.
[0237] The entropy decoding unit 210 may change a coefficient having a one-dimensional (1D) vector form into a 2D block shape through a transform coefficient scanning method in order to decode quantized transform coefficient levels.
[0238] For example, the coefficients of the block can be changed to a 2D block shape by scanning the block coefficients using an upper right diagonal scan. Optionally, which of the upper right diagonal scan, vertical scan, and horizontal scan to use can be determined based on the size of the corresponding block and / or the intra prediction mode.
[0239] The quantized coefficients may be dequantized by the dequantization unit 220. The dequantization unit 220 may generate dequantized coefficients by performing dequantization on the quantized coefficients. Furthermore, the dequantized coefficients may be inversely transformed by the inverse transform unit 230. The inverse transform unit 230 may generate a reconstructed residual block by performing inverse transform on the dequantized coefficients. As a result of performing dequantization and inverse transform on the quantized coefficients, a reconstructed residual block may be generated. Here, when generating the reconstructed residual block, the dequantization unit 220 may apply a quantization matrix to the quantized coefficients.
[0240] When the intra mode is used, the intra prediction unit 240 may generate a prediction block by performing spatial prediction on a target block using pixel values of a previously decoded neighboring block adjacent to the target block.
[0241] The inter-frame prediction unit 250 may include a motion compensation unit. Alternatively, the inter-frame prediction unit 250 may be designated as a "motion compensation unit."
[0242] When the inter mode is used, the motion compensation unit may generate a prediction block by performing motion compensation on the target block using a motion vector and a reference image stored in the reference picture buffer 270 .
[0243] The motion compensation unit may apply an interpolation filter to a partial area of a reference image when a motion vector has a value other than an integer, and may generate a prediction block using the reference image to which the interpolation filter is applied. To perform motion compensation, the motion compensation unit may determine which mode of skip mode, merge mode, advanced motion vector prediction (AMVP) mode, and current picture reference mode corresponds to a motion compensation method for a PU included in the CU based on the CU, and may perform motion compensation according to the determined mode.
[0244] The reconstructed residual block and the prediction block may be added to each other by the adder 255. The adder 255 may generate a reconstructed block by adding the reconstructed residual block and the prediction block.
[0245] The reconstructed block may be filtered by the filter unit 260. The filter unit 260 may apply at least one of a deblocking filter, an SAO filter, an ALF, and an NLF to the reconstructed block or the reconstructed image. The reconstructed image may be a picture including the reconstructed block.
[0246] The filter unit may output a reconstructed image.
[0247] The reconstructed image and / or reconstructed block filtered by the filter unit 260 may be stored as a reference picture in the reference picture buffer 270. The reconstructed block filtered by the filter unit 260 may be part of a reference picture. In other words, the reference picture may be an image composed of the reconstructed blocks filtered by the filter unit 260. The stored reference picture may then be used for inter-frame prediction or motion compensation.
[0248] Figure 3 is a diagram schematically showing a partition structure of an image when the image is encoded and decoded.
[0249] Figure 3 An example in which a single unit is partitioned into a plurality of subunits may be schematically shown.
[0250] To efficiently partition an image, coding units (CUs) are used in encoding and decoding. The term "unit" can be used to collectively designate 1) a block containing image samples and 2) a syntax element. For example, "a partition of a unit" can mean "a partition of blocks corresponding to the unit."
[0251] A CU may be used as a basic unit for image encoding / decoding. A CU may be used as a unit to which one mode selected from intra mode and inter mode is applied in image encoding / decoding. In other words, in image encoding / decoding, it may be determined which mode of intra mode and inter mode is to be applied to each CU.
[0252] Also, a CU may be a basic unit for predicting, transforming, quantizing, inversely transforming, dequantizing, and encoding / decoding a transform coefficient.
[0253] Reference Figure 3 , the image 300 may be sequentially partitioned into units corresponding to a largest coding unit (LCU), and a partition structure may be determined for each LCU. Here, LCU may be used to have the same meaning as a coding tree unit (CTU).
[0254] Partitioning a cell may mean partitioning a block corresponding to the cell. Block partition information may include depth information regarding the depth of the cell. The depth information may indicate the number of times the cell is partitioned and / or the extent to which the cell is partitioned. A single cell may be hierarchically partitioned into multiple sub-cells, with the single cell having depth information based on a tree structure.
[0255] Each partitioned sub-unit may have depth information. The depth information may be information indicating the size of the CU. The depth information may be stored for each CU.
[0256] Each CU may have depth information. When a CU is partitioned, the depth of the CU generated from the partition may increase by 1 from the depth of the partitioned CU.
[0257] The partition structure may indicate the distribution of coding units (CUs) within the LCU 310 for efficient image encoding. This distribution may be determined based on whether a single CU is partitioned into multiple CUs. The number of CUs generated by partitioning may be a positive integer of 2 or greater, including 2, 3, 4, 8, 16, and the like.
[0258] Depending on the number of CUs generated by partitioning, the horizontal size and vertical size of each CU generated by partitioning may be smaller than the horizontal size and vertical size of the CU before partitioning. For example, the horizontal size and vertical size of each CU generated by partitioning may be half the horizontal size and vertical size of the CU before partitioning.
[0259] Each partitioned CU may be recursively partitioned into four CUs in the same manner. Compared to at least one of the horizontal size and the vertical size of the CU before partitioning, at least one of the horizontal size and the vertical size of each partitioned CU may be reduced through recursive partitioning.
[0260] Partitioning of the CU may be performed recursively up to a predefined depth or a predefined size.
[0261] For example, the depth of the CU may have a value ranging from 0 to 3. The size of the CU may range from a size of 64×64 to a size of 8×8 according to the depth of the CU.
[0262] For example, the depth of the LCU 310 may be 0, and the depth of the minimum coding unit (SCU) may be a predefined maximum depth. Here, as described above, the LCU may be a CU having a maximum coding unit size, and the SCU may be a CU having a minimum coding unit size.
[0263] Partitioning may begin at the LCU 310, and each time the horizontal and / or vertical dimensions of the CU are reduced by partitioning, the depth of the CU may increase by one.
[0264] For example, for each depth, a non-partitioned CU may have a size of 2N×2N. In addition, when a CU is partitioned, a CU of size 2N×2N may be partitioned into four CUs of size N×N. Whenever the depth increases by 1, the value of N may be halved.
[0265] Reference Figure 3 , an LCU with a depth of 0 may have 64×64 pixels or a 64×64 block. 0 may be the minimum depth. An SCU with a depth of 3 may have 8×8 pixels or an 8×8 block. 3 may be the maximum depth. Here, a CU with a 64×64 block as an LCU may be represented by a depth of 0. A CU with a 32×32 block may be represented by a depth of 1. A CU with a 16×16 block may be represented by a depth of 2. A CU with an 8×8 block as an SCU may be represented by a depth of 3.
[0266] Information about whether the corresponding CU is partitioned can be represented by the partition information of the CU. The partition information can be 1-bit information. All CUs except the SCU can include partition information. For example, the value of the partition information of a non-partitioned CU can be a first value. The value of the partition information of a partitioned CU can be a second value. When the partition information indicates whether the CU is partitioned, the first value can be "0" and the second value can be "1".
[0267] For example, when a single CU is partitioned into four CUs, the horizontal size and vertical size of each of the four CUs generated by partitioning can be half the horizontal size and vertical size of the CU before partitioning. When a CU of size 32×32 is partitioned into four CUs, the size of each of the four partitioned CUs can be 16×16. When a single CU is partitioned into four CUs, it can be considered that the CU has been partitioned in a quadtree structure. In other words, it can be considered that quadtree partitioning has been applied to the CU.
[0268] For example, when a single CU is partitioned into two CUs, the horizontal size or vertical size of each of the two CUs generated by partitioning may be half the horizontal size or vertical size of the CU before partitioning. When a CU of size 32×32 is partitioned vertically into two CUs, the size of each of the two partitioned CUs may be 16×32. When a CU of size 32×32 is partitioned horizontally into two CUs, the size of each of the two partitioned CUs may be 32×16. When a single CU is partitioned into two CUs, it can be considered that the CU has been partitioned in a binary tree structure. In other words, it can be considered that binary tree partitioning has been applied to the CU.
[0269] For example, when a single CU is partitioned (or split) into three CUs, the original CU before partitioning is partitioned so that its horizontal size or vertical size is divided at a ratio of 1:2:1, thereby enabling the generation of three sub-CUs. For example, when a CU of size 16×32 is partitioned horizontally into three sub-CUs, the three sub-CUs generated by the partitioning may have sizes of 16×8, 16×16, and 16×8, respectively, from top to bottom. For example, when a CU of size 32×32 is partitioned vertically into three sub-CUs, the three sub-CUs generated by the partitioning may have sizes of 8×32, 16×32, and 8×32, respectively, from left to right. When a single CU is partitioned into three CUs, the CU may be considered to be partitioned in a ternary tree form. In other words, ternary tree partitioning may be considered to have been applied to the CU.
[0270] Both quadtree partitioning and binary tree partitioning are applied to Figure 3 LCU 310.
[0271] In the encoding apparatus 100 , a coding tree unit (CTU) of size 64×64 may be partitioned into multiple smaller CUs using a recursive quadtree structure. A single CU may be partitioned into four CUs of the same size. Each CU may be recursively partitioned and may have a quadtree structure.
[0272] By recursive partitioning of CUs, the optimal partitioning method that incurs the minimum rate-distortion cost can be selected.
[0273] Figure 3 A coding tree unit (CTU) 320 in is an example of a CTU to which quadtree partitioning, binarytree partitioning, and ternarytree partitioning are all applied.
[0274] As described above, in order to partition a CTU, at least one of quadtree partitioning, binary tree partitioning, and ternary tree partitioning may be applied to the CTU. Partitioning may be applied based on a specific priority.
[0275] For example, quadtree partitioning may be preferentially applied to CTUs. CUs that cannot be further partitioned in quadtree form may correspond to leaf nodes of the quadtree. CUs corresponding to leaf nodes of the quadtree may be root nodes of a binary tree and / or ternary tree. That is, CUs corresponding to leaf nodes of the quadtree may be partitioned in a binary tree or ternary tree form, or may not be further partitioned. In this case, each CU generated by applying binary tree partitioning or ternary tree partitioning to a CU corresponding to a leaf node of the quadtree is prevented from being partitioned again by the quadtree, thereby efficiently performing block partitioning and / or signaling block partition information.
[0276] The quad partition information may be used to signal the partition of the CU corresponding to each node of the quadtree. The quad partition information having a first value (e.g., "1") may indicate that the corresponding CU is partitioned in a quadtree manner. The quad partition information having a second value (e.g., "0") may indicate that the corresponding CU is not partitioned in a quadtree manner. The quad partition information may be a flag having a specific length (e.g., 1 bit).
[0277] There may be no priority between binary tree partitioning and ternary tree partitioning. That is, the CU corresponding to the leaf node of the quadtree may be partitioned in a binary tree form or a ternary tree form. In addition, the CU generated by binary tree partitioning or ternary tree partitioning may be further partitioned in a binary tree form or a ternary tree form, or may not be further partitioned.
[0278] Partitioning performed when there is no priority between binary tree partitioning and ternary tree partitioning may be referred to as "multi-type tree partitioning." That is, the CU corresponding to the leaf node of the quadtree may be the root node of the multi-type tree. The partitioning of the CU corresponding to each node of the multi-type tree may be signaled using at least one of information indicating whether the CU is partitioned according to the multi-type tree, partition direction information, and partition tree information. For the partitioning of the CU corresponding to each node of the multi-type tree, information indicating whether partitioning according to the multi-type tree is performed, partition direction information, and partition tree information may be sequentially signaled.
[0279] For example, information indicating whether a CU is partitioned in a multi-type tree and having a first value (e.g., "1") may indicate that the corresponding CU is partitioned in a multi-type tree form. Information indicating whether a CU is partitioned in a multi-type tree and having a second value (e.g., "0") may indicate that the corresponding CU is not partitioned in a multi-type tree form.
[0280] When a CU corresponding to each node of the multi-type tree is partitioned in the multi-type tree form, the corresponding CU may further include partition direction information.
[0281] The partition direction information may indicate the partition direction of the multi-type tree partition. Partition direction information having a first value (e.g., "1") may indicate that the corresponding CU is partitioned in the vertical direction. Partition direction information having a second value (e.g., "0") may indicate that the corresponding CU is partitioned in the horizontal direction.
[0282] When a CU corresponding to each node of a multi-type tree is partitioned in a multi-type tree form, the corresponding CU may further include partition tree information. The partition tree information may indicate a tree used for multi-type tree partitioning.
[0283] For example, partition tree information having a first value (eg, "1") may indicate that the corresponding CU is partitioned in a binary tree format. Partition tree information having a second value (eg, "0") may indicate that the corresponding CU is partitioned in a ternary tree format.
[0284] Here, each of the above-mentioned information indicating whether partitioning by the multi-type tree is performed, the partition tree information, and the partition direction information may be a flag having a specific length (for example, 1 bit).
[0285] At least one of the four partition information, the information indicating whether partitioning by a multi-type tree is performed, the partition direction information, and the partition tree information may be entropy encoded and / or decoded. To perform entropy encoding / decoding of such information, information of neighboring CUs adjacent to the target CU may be used.
[0286] For example, it can be considered that the partition form (i.e., partitioned / non-partitioned, partition tree, and / or partition direction) of the left CU and / or the upper CU is likely to be similar to the partition form of the target CU. Therefore, based on the information of the neighboring CU, context information for entropy encoding and / or entropy decoding of the information of the target CU can be derived. Here, the information of the neighboring CU may include at least one of the following: 1) four partition information of the neighboring CU, 2) information indicating whether the neighboring CU is partitioned according to a multi-type tree, 3) partition direction information of the neighboring CU, and 4) partition tree information of the neighboring CU.
[0287] In another embodiment of binary tree partitioning and ternary tree partitioning, binary tree partitioning may be performed first. That is, binary tree partitioning may be applied first, and then the CU corresponding to the leaf node of the binary tree may be set as the root node of the ternary tree. In this case, quadtree partitioning or binary tree partitioning may not be performed on the CU corresponding to the node of the ternary tree.
[0288] A CU that is not further partitioned by quadtree partitioning, binary tree partitioning, and / or ternary tree partitioning may be a unit for coding, prediction, and / or transformation. That is, the CU may not be further partitioned for prediction and / or transformation. Therefore, the partition structure for partitioning the CU into prediction units (PUs) and / or transform units (TUs), its partition information, etc. may not be present in the bitstream.
[0289] However, when the size of the CU as the unit of partition is larger than the size of the maximum transform block, the CU may be recursively partitioned until the size of the CU becomes smaller than or equal to the size of the maximum transform block. For example, when the size of the CU is 64×64 and the size of the maximum transform block is 32×32, the CU may be partitioned into four 32×32 blocks to perform transform. For example, when the size of the CU is 32×64 and the size of the maximum transform block is 32×32, the CU may be partitioned into two 32×32 blocks.
[0290] In this case, information indicating whether the CU is partitioned for transform may not be separately signaled. Without signaling, whether the CU is partitioned may be determined by comparing the horizontal size (and / or vertical size) of the CU with the horizontal size (and / or vertical size) of the largest transform block. For example, when the horizontal size of the CU is larger than the horizontal size of the largest transform block, the CU may be vertically bisected. Furthermore, when the vertical size of the CU is larger than the vertical size of the largest transform block, the CU may be horizontally bisected.
[0291] Information about the maximum size and / or minimum size of a CU and information about the maximum size and / or minimum size of a transform block may be signaled or determined at a level higher than the level of the CU. For example, the higher level may be a sequence level, a picture level, a tile level, a tile group level, or a slice level. For example, the minimum size of a CU may be set to 4×4. For example, the maximum size of a transform block may be set to 64×64. For example, the maximum size of a transform block may be set to 4×4.
[0292] Information about the minimum size of the CU corresponding to the leaf node of the quadtree (i.e., the minimum size of the quadtree) and / or information about the maximum depth of the path from the root node to the leaf node of the multi-type tree (i.e., the maximum depth of the multi-type tree) may be signaled or determined at a level higher than the CU level. For example, the higher level may be the sequence level, the picture level, the slice level, the tile group level, or the tile level. Information about the minimum size of the quadtree and / or information about the maximum depth of the multi-type tree may be signaled or determined separately at each of the intra-slice level and the inter-slice level.
[0293] Information about the difference between the size of a CTU and the maximum size of a transform block may be signaled or determined at a level higher than the CU level. For example, the higher level may be the sequence level, picture level, slice level, tile group level, or tile block level. Information about the maximum size of the CU corresponding to each node of the binary tree (i.e., the maximum size of the binary tree) may be determined based on the size of the CTU and the difference information. The maximum size of the CU corresponding to each node of the ternary tree (i.e., the maximum size of the ternary tree) may have different values depending on the slice type. For example, the maximum size of the ternary tree at the intra-slice level may be 32×32. For example, the maximum size of the ternary tree at the inter-slice level may be 128×128. For example, the minimum size of the CU corresponding to each node of the binary tree (i.e., the minimum size of the binary tree) and / or the minimum size of the CU corresponding to each node of the ternary tree (i.e., the minimum size of the ternary tree) may be set to the minimum size of the CU.
[0294] In another example, the maximum size of the binary tree and / or the maximum size of the ternary tree may be signaled or determined at the slice level. In addition, the minimum size of the binary tree and / or the minimum size of the ternary tree may be signaled or determined at the slice level.
[0295] Based on the various block sizes and depths described above, quad partition information, information indicating whether partitioning by a multi-type tree is performed, partition tree information, and / or partition direction information may or may not exist in a bitstream.
[0296] For example, when the size of the CU is not greater than the minimum size of the quadtree, the CU may not include quad partition information, and the quad partition information of the CU may be inferred as the second value.
[0297] For example, when the size (horizontal and vertical) of the CU corresponding to each node of the multi-type tree is larger than the maximum size (horizontal and vertical) of the binary tree and / or the maximum size (horizontal and vertical) of the ternary tree, the CU may not be partitioned in a binary tree and / or ternary tree format. In this determination method, information indicating whether partitioning by the multi-type tree is performed may not be signaled, but may be inferred as the second value.
[0298] Alternatively, when the size (horizontal and vertical dimensions) of the CU corresponding to each node of the multi-type tree is equal to the minimum size (horizontal and vertical dimensions) of the binary tree, or when the size (horizontal and vertical dimensions) of the CU is equal to twice the minimum size (horizontal and vertical dimensions) of the ternary tree, the CU may not be partitioned in a binary tree and / or ternary tree format. With this determination, information indicating whether partitioning according to the multi-type tree is performed may not be signaled but may be inferred as the second value. This is because when the CU is partitioned in a binary tree and / or ternary tree format, a CU smaller than the minimum size of the binary tree and / or the minimum size of the ternary tree is generated.
[0299] Optionally, binary or ternary tree partitioning may be limited based on the size of a virtual pipeline data unit (i.e., the size of a pipeline buffer). For example, when a CU is partitioned into sub-CUs that do not fit within the size of a pipeline buffer using binary or ternary tree partitioning, binary or ternary tree partitioning may be limited. The size of a pipeline buffer may be equal to the maximum size of a transform block (e.g., 64×64).
[0300] For example, when the size of the pipeline buffer is 64×64, the following partitions may be limited.
[0301] - Ternary tree partitioning for N×M CUs (where N and / or M is 128) - Horizontal binary tree partitioning for 128×N CUs (where N<=64) - Vertical binary tree partitioning for N×128 CUs (where N<=64) Alternatively, when the depth of the CU corresponding to each node of the multi-type tree is equal to the maximum depth of the multi-type tree, the CU may not be partitioned in a binary tree form and / or a ternary tree form. In this determination method, information indicating whether partitioning by the multi-type tree is performed may not be signaled but may be inferred as the second value.
[0302] Alternatively, information indicating whether partitioning according to a multi-type tree is performed may be signaled only when at least one of vertical binary tree partitioning, horizontal binary tree partitioning, vertical ternary tree partitioning, and horizontal ternary tree partitioning is possible for the CU corresponding to each node of the multi-type tree. Otherwise, the CU may not be partitioned in a binary tree form and / or a ternary tree form. With this determination, information indicating whether partitioning according to a multi-type tree is performed may not be signaled but may be inferred as a second value.
[0303] Alternatively, for the CU corresponding to each node of the multi-type tree, partition direction information may be signaled only when both vertical binary tree partitioning and horizontal binary tree partitioning are possible, or only when both vertical ternary tree partitioning and horizontal ternary tree partitioning are possible. Otherwise, the partition direction information may not be signaled but may be inferred as a value indicating the direction in which the CU may be partitioned.
[0304] Alternatively, for a CU corresponding to each node of a multi-type tree, partition tree information may be signaled only when both vertical binary tree partitioning and vertical ternary tree partitioning are possible, or only when both horizontal binary tree partitioning and horizontal ternary tree partitioning are possible. Otherwise, the partition tree information may not be signaled but may be inferred as a value indicating a tree of partitions applicable to the CU.
[0305] Figure 4 is a diagram illustrating a form of prediction units that a coding unit can include.
[0306] In a CU partitioned from an LCU, the CU that is no longer partitioned can be divided into one or more prediction units (PUs). This division is also called "partitioning."
[0307] PU can be a basic unit for prediction. PU can be encoded and decoded in any one of skip mode, inter mode and intra mode. PU can be partitioned into various shapes according to each mode. For example, Figure 1 The target block described above refers to Figure 2 The target blocks described may all be PUs.
[0308] A CU may not be split into PUs. When a CU is not split into PUs, the size of the CU and the size of the PU may be equal to each other.
[0309] In skip mode, partitioning may not exist in a CU.In skip mode, a 2Nx2N mode 410 may be supported without partitioning, wherein in the 2Nx2N mode 410, the size of the PU and the size of the CU are the same as each other.
[0310] In inter mode, eight types of partition shapes may exist in a CU. For example, in inter mode, 2N×2N mode 410, 2N×N mode 415, N×2N mode 420, N×N mode 425, 2N×nU mode 430, 2N×nD mode 435, nL×2N mode 440, and nR×2N mode 445 may be supported.
[0311] In intra mode, 2N×2N mode 410 and N×N mode 425 may be supported.
[0312] In 2N×2N mode 410, a PU of size 2N×2N may be encoded. A PU of size 2N×2N may represent a PU of the same size as a CU. For example, a PU of size 2N×2N may have a size of 64×64, 32×32, 16×16, or 8×8.
[0313] In NxN mode 425, PUs of size NxN may be encoded.
[0314] For example, in intra prediction, when the size of a PU is 8×8, four partitioned PUs may be encoded, and the size of each partitioned PU may be 4×4.
[0315] When encoding a PU in intra mode, any one of multiple intra prediction modes may be used to encode the PU. For example, HEVC technology may provide 35 intra prediction modes, and a PU may be encoded in any of the 35 intra prediction modes.
[0316] Which mode of the 2Nx2N mode 410 and the NxN mode 425 is to be used to encode the PU may be determined based on the rate-distortion penalty.
[0317] The encoding device 100 may perform an encoding operation on a PU of size 2N×2N. Here, the encoding operation may be an operation of encoding the PU in each of a plurality of intra-prediction modes that can be used by the encoding device 100. Through the encoding operation, an optimal intra-prediction mode for the PU of size 2N×2N may be derived. The optimal intra-prediction mode may be an intra-prediction mode that, among the plurality of intra-prediction modes that can be used by the encoding device 100, results in a minimum rate-distortion cost when encoding the PU of size 2N×2N.
[0318] In addition, the encoding device 100 may sequentially perform encoding operations on each PU obtained by performing N×N partitioning. Here, the encoding operation may be an operation of encoding the PU in each of a plurality of intra-frame prediction modes that can be used by the encoding device 100. Through the encoding operation, an optimal intra-frame prediction mode for a PU of size N×N may be derived. The optimal intra-frame prediction mode may be an intra-frame prediction mode that results in the minimum rate-distortion cost when encoding a PU of size N×N among a plurality of intra-frame prediction modes that can be used by the encoding device 100.
[0319] The encoding apparatus 100 may determine which of the PU of size 2N×2N and the PU of size N×N to be encoded based on a comparison between the rate-distortion cost of the PU of size 2N×2N and the rate-distortion cost of the PU of size N×N.
[0320] A single CU may be partitioned into one or more PUs, and a PU may be partitioned into multiple PUs.
[0321] For example, when a single PU is partitioned into four PUs, the horizontal and vertical sizes of each of the four PUs generated by partitioning can be half the horizontal and vertical sizes of the PU before partitioning. When a PU of size 32×32 is partitioned into four PUs, the size of each of the four partitioned PUs can be 16×16. When a single PU is partitioned into four PUs, it can be considered that the PU has been partitioned in a quadtree structure.
[0322] For example, when a single PU is partitioned into two PUs, the horizontal size or vertical size of each of the two PUs generated by partitioning may be half the horizontal size or vertical size of the PU before partitioning. When a PU of size 32×32 is partitioned vertically into two PUs, the size of each of the two partitioned PUs may be 16×32. When a PU of size 32×32 is partitioned horizontally into two PUs, the size of each of the two partitioned PUs may be 32×16. When a single PU is partitioned into two PUs, the PU may be considered to have been partitioned in a binary tree structure.
[0323] Figure 5 is a diagram illustrating forms of transformation units that can be included in a coding unit.
[0324] A transform unit (TU) may be a basic unit used for processes such as transformation, quantization, inverse transformation, inverse quantization, entropy encoding, and entropy decoding in a CU.
[0325] A TU may have a square shape or a rectangular shape. The shape of a TU may be determined based on the size and / or shape of a CU.
[0326] Among the CUs partitioned from the LCU, the CUs that are no longer partitioned into CUs may be partitioned into one or more TUs. Here, the partition structure of the TU may be a quadtree structure. For example, Figure 5 As shown in , a single CU 510 can be partitioned one or more times according to a quadtree structure. Through this partitioning, a single CU 510 can be composed of TUs of various sizes.
[0327] It can be considered that a single CU is recursively split when it is split two or more times. Through splitting, a single CU can be composed of transform units (TUs) of various sizes.
[0328] Alternatively, a single CU may be split into one or more TUs based on the number of vertical lines and / or horizontal lines that partition the CU.
[0329] A CU may be divided into symmetric TUs or asymmetric TUs. To divide into asymmetric TUs, information about the size and / or shape of each TU may be signaled from the encoding apparatus 100 to the decoding apparatus 200. Alternatively, the size and / or shape of each TU may be derived from the information about the size and / or shape of the CU.
[0330] A CU may not be divided into TUs. When a CU is not divided into TUs, the size of the CU and the size of the TU may be equal to each other.
[0331] A single CU may be partitioned into one or more TUs, and a TU may be partitioned into multiple TUs.
[0332] For example, when a single TU is partitioned into four TUs, the horizontal size and vertical size of each of the four TUs generated by partitioning can be half the horizontal size and vertical size of the TU before partitioning. When a 32×32 TU is partitioned into four TUs, the size of each of the four partitioned TUs can be 16×16. When a single TU is partitioned into four TUs, the TU can be considered to have been partitioned in a quadtree structure.
[0333] For example, when a single TU is partitioned into two TUs, the horizontal size or vertical size of each of the two TUs generated by partitioning may be half the horizontal size or vertical size of the TU before partitioning. When a 32×32 TU is partitioned vertically into two TUs, the size of each of the two partitioned TUs may be 16×32. When a 32×32 TU is partitioned horizontally into two TUs, the size of each of the two partitioned TUs may be 32×16. When a single TU is partitioned into two TUs, the TU may be considered to have been partitioned in a binary tree structure.
[0334] Can be used with Figure 5 The CU is divided in different ways as shown in FIG.
[0335] For example, a single CU may be split into three CUs, and the horizontal sizes or vertical sizes of the three CUs generated by the splitting may be 1 / 4, 1 / 2, and 1 / 4 of the horizontal size or vertical size of the original CU before the splitting, respectively.
[0336] For example, when a CU of size 32×32 is vertically split into three CUs, the sizes of the three CUs generated by the splitting may be 8×32, 16×32, and 8×32, respectively. In this way, when a single CU is split into three CUs, it can be considered that the CU is split in the form of a ternary tree.
[0337] One of the exemplary partitioning forms (i.e., quadtree partitioning, binary tree partitioning, and ternary tree partitioning) may be applied to the partitioning of a CU, and multiple partitioning schemes may be combined and used together for the partitioning of a CU. Here, the case where multiple partitioning schemes are combined and used together may be referred to as "composite tree form partitioning."
[0338] Figure 6 The division of blocks according to an example is shown.
[0339] In the video encoding and / or decoding process, such as Figure 6 As shown in , the target block can be divided. For example, the target block can be a CU.
[0340] For splitting of the target block, an indicator indicating splitting information may be signaled from the encoding apparatus 100 to the decoding apparatus 200. The splitting information may be information indicating how the target block is split.
[0341] The split information may be one or more of a split flag (hereinafter referred to as "split_flag"), a quad-binary flag (hereinafter referred to as "QB_flag"), a quadtree flag (hereinafter referred to as "quadtree_flag"), a binary tree flag (hereinafter referred to as "binarytree_flag"), and a binary type flag (hereinafter referred to as "Btype_flag").
[0342] "split_flag" may be a flag indicating whether a block is split. For example, a split_flag value of 1 may indicate that the corresponding block is split. A split_flag value of 0 may indicate that the corresponding block is not split.
[0343] "QB_flag" may be a flag indicating which of the quadtree and binary tree forms corresponds to the shape of the block partition. For example, a QB_flag value of 0 may indicate that the block is partitioned in a quadtree form. A QB_flag value of 1 may indicate that the block is partitioned in a binary tree form. Alternatively, a QB_flag value of 0 may indicate that the block is partitioned in a binary tree form. A QB_flag value of 1 may indicate that the block is partitioned in a quadtree form.
[0344] "quadtree_flag" may be a flag indicating whether the block is partitioned in a quadtree form. For example, a quadtree_flag value of 1 may indicate that the block is partitioned in a quadtree form. A quadtree_flag value of 0 may indicate that the block is not partitioned in a quadtree form.
[0345] "binarytree_flag" may be a flag indicating whether the block is partitioned in a binary tree form. For example, a binarytree_flag value of 1 may indicate that the block is partitioned in a binary tree form. A binarytree_flag value of 0 may indicate that the block is not partitioned in a binary tree form.
[0346] "Btype_flag" may be a flag indicating which of vertical and horizontal divisions corresponds to the division direction when the block is divided in a binary tree form. For example, a Btype_flag value of 0 may indicate that the block is divided in the horizontal direction. A Btype_flag value of 1 may indicate that the block is divided in the vertical direction. Alternatively, a Btype_flag value of 0 may indicate that the block is divided in the vertical direction. A Btype_flag value of 1 may indicate that the block is divided in the horizontal direction.
[0347] For example, it may be derived by signaling at least one of quadtree_flag, binarytree_flag, and Btype_flag. Figure 6 The partitioning information of the blocks in is shown in Table 1 below.
[0348] Table 1
[0349] For example, the QB_flag may be derived by signaling at least one of the split_flag, QB_flag, and Btype_flag. Figure 6 The partitioning information of the blocks in is shown in Table 2 below.
[0350] Table 2
[0351] The partitioning method may be limited to a quadtree or a binary tree depending on the size and / or shape of the block. When this restriction is applied, split_flag may be a flag indicating whether the block is partitioned in a quadtree or a flag indicating whether the block is partitioned in a binary tree. The size and shape of the block may be derived based on the depth information of the block, and the depth information may be signaled from the encoding device 100 to the decoding device 200.
[0352] When the block size falls within a specific range, partitioning only in quadtree form is possible. For example, the specific range may be defined by at least one of a maximum block size and a minimum block size that can be partitioned only in quadtree form.
[0353] Information indicating the maximum block size and the minimum block size that can be divided only in the quadtree form may be signaled from the encoding apparatus 100 to the decoding apparatus 200 through a bitstream. In addition, this information may be signaled for at least one of units such as a video, a sequence, a picture, a parameter, a tile group, and a slice (or segment).
[0354] Alternatively, the maximum block size and / or the minimum block size may be a fixed size predefined by the encoding device 100 and the decoding device 200. For example, when the size of the block is larger than 64×64 and smaller than 256×256, only quadtree division is possible. In this case, split_flag may be a flag indicating whether quadtree division is performed.
[0355] When the size of the block is larger than the maximum size of the transform block, only partitioning in a quadtree form is possible. Here, the sub-block generated by the partitioning may be at least one of a CU and a TU.
[0356] In this case, split_flag may be a flag indicating whether the CU is partitioned in a quadtree form.
[0357] When the size of the block falls within a specific range, it is possible to divide it only in the binary tree form or the ternary tree form. For example, the specific range can be defined by at least one of the maximum block size and the minimum block size that can be divided only in the binary tree form or the ternary tree form.
[0358] Information indicating the maximum block size and / or the minimum block size that can be partitioned only in a binary tree form or in a ternary tree form may be signaled from the encoding apparatus 100 to the decoding apparatus 200 through a bitstream. In addition, this information may be signaled for at least one of units such as a sequence, a picture, and a slice (or segment).
[0359] Alternatively, the maximum block size and / or the minimum block size may be a fixed size predefined by the encoding device 100 and the decoding device 200. For example, when the size of the block is greater than 8×8 and less than 16×16, only splitting in a binary tree form is possible. In this case, split_flag may be a flag indicating whether to perform splitting in a binary tree form or a ternary tree form.
[0360] The above description of partitioning in a quadtree form may be equally applied to a binary tree form and / or a ternary tree form.
[0361] The partitioning of a block may be limited by the previous partitioning. For example, when a block is partitioned in a specific binary tree format and multiple sub-blocks are generated from the partitioning, each sub-block can be partitioned only in a specific tree format. Here, the specific tree format can be at least one of a binary tree format, a ternary tree format, and a quadtree format.
[0362] When the horizontal size or the vertical size of the partition block is a size that cannot be further divided, the above-mentioned indicator may not be signaled.
[0363] Figure 7 is a diagram for explaining an embodiment of an intra prediction process.
[0364] from Figure 7 The arrow extending radially from the center of the diagram in indicates the prediction direction of the intra prediction mode. Furthermore, numbers appearing near the arrows indicate examples of mode values assigned to the intra prediction mode or the prediction direction of the intra prediction mode.
[0365] exist Figure 7 In the example, number 0 may indicate planar mode, which is a non-directional intra prediction mode, and number 1 may indicate DC mode, which is a non-directional intra prediction mode.
[0366] Intra-frame encoding and / or decoding may be performed using reference samples of neighboring blocks of a target block. The neighboring blocks may be reconstructed neighboring blocks. The reference samples may represent neighboring samples.
[0367] For example, intra encoding and / or decoding may be performed using values of reference samples included in the reconstructed neighboring block or encoding parameters of the reconstructed neighboring block.
[0368] The encoding device 100 and / or the decoding device 200 may generate a prediction block by performing intra-frame prediction on the target block based on information about samples in the target image. When intra-frame prediction is performed, the encoding device 100 and / or the decoding device 200 may generate a prediction block for the target block by performing intra-frame prediction based on information about samples in the target image. When intra-frame prediction is performed, the encoding device 100 and / or the decoding device 200 may perform directional prediction and / or non-directional prediction based on at least one reconstructed reference sample.
[0369] The prediction block may be a block generated as a result of performing intra prediction. The prediction block may correspond to at least one of a CU, a PU, and a TU.
[0370] The prediction block may have a size corresponding to at least one of a CU, a PU, and a TU. The prediction block may have a square shape of 2N×2N or N×N. The N×N size may include 4×4, 8×8, 16×16, 32×32, 64×64, etc.
[0371] Alternatively, the prediction block may be a square block of size 2×2, 4×4, 8×8, 16×16, 32×32, 64×64, etc. or a rectangular block of size 2×8, 4×8, 2×16, 4×16, 8×16, etc.
[0372] Intra-frame prediction may be performed considering the intra-frame prediction mode used for the target block. The number of intra-frame prediction modes that the target block can have may be a predefined fixed value, or may be a value determined differently according to the properties of the prediction block. For example, the properties of the prediction block may include the size of the prediction block, the type of the prediction block, etc. In addition, the properties of the prediction block may indicate encoding parameters used for the prediction block.
[0373] For example, regardless of the size of the prediction block, the number of intra prediction modes may be fixed to N. Alternatively, the number of intra prediction modes may be 3, 5, 9, 17, 34, 35, 36, 65, 67, or 95, for example.
[0374] The intra prediction mode can be a non-directional mode or a directional mode.
[0375] For example, intra prediction modes may include Figure 7 The numbers 0 to 66 shown in the figure correspond to two non-directional modes and 65 directional modes.
[0376] For example, in the case of using a specific intra prediction method, the intra prediction mode may include Figure 7 The numbers -14 to 80 shown correspond to two non-directional modes and 93 directional modes.
[0377] The two non-directional modes may include a DC mode and a planar mode.
[0378] A directional mode may be a prediction mode with a specific direction or a specific angle. A directional mode may also be referred to as an "angle mode."
[0379] The intra-frame prediction mode may be represented by at least one of a mode number, a mode value, a mode angle, and a mode direction. In other words, the terms "intra-frame prediction mode (mode) number," "intra-frame prediction mode (mode) value," "intra-frame prediction mode (mode) angle," and "intra-frame prediction mode (mode) direction" may be used to have the same meaning and may be used interchangeably.
[0380] The number of intra prediction modes may be M. The value of M may be 1 or greater. In other words, the number of intra prediction modes may be M, where M includes the number of non-directional modes and the number of directional modes.
[0381] Regardless of the size of the block and / or the color components, the number of intra prediction modes may be fixed to M. For example, the number of intra prediction modes may be fixed to any one of 35 and 67 regardless of the size of the block.
[0382] Alternatively, the number of intra prediction modes may differ according to the shape, size, and / or type of color component of a block.
[0383] For example, in Figure 7 , the directional prediction mode shown by the dotted line can be applied only to prediction for non-square blocks.
[0384] For example, the larger the block size, the greater the number of intra-frame prediction modes. Alternatively, the larger the block size, the fewer the number of intra-frame prediction modes. When the block size is 4×4 or 8×8, the number of intra-frame prediction modes may be 67. When the block size is 16×16, the number of intra-frame prediction modes may be 35. When the block size is 32×32, the number of intra-frame prediction modes may be 19. When the block size is 64×64, the number of intra-frame prediction modes may be 7.
[0385] For example, the number of intra prediction modes may differ depending on whether the color component is a luma signal or a chroma signal. Alternatively, the number of intra prediction modes corresponding to a luma component block may be greater than the number of intra prediction modes corresponding to a chroma component block.
[0386] For example, in vertical mode with a mode value of 50, prediction may be performed in the vertical direction based on the pixel values of the reference samples. For example, in horizontal mode with a mode value of 18, prediction may be performed in the horizontal direction based on the pixel values of the reference samples.
[0387] Even in a directional mode other than the above-described modes, the encoding apparatus 100 and the decoding apparatus 200 may perform intra prediction on a target unit using reference samples according to an angle corresponding to the directional mode.
[0388] An intra-frame prediction mode located to the right relative to the vertical mode may be referred to as a "vertical-right mode". An intra-frame prediction mode located below the horizontal mode may be referred to as a "horizontal-below mode". For example, in Figure 7 , the intra prediction mode whose mode value is one of 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, and 66 may be the vertical-right mode. The intra prediction mode whose mode value is one of 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, and 17 may be the horizontal-down mode.
[0389] The non-directional mode may include a DC mode and a planar mode. For example, the value of the DC mode may be 1. The value of the planar mode may be 0.
[0390] The directional mode may include an angular mode. Among the plurality of intra prediction modes, the remaining modes except the DC mode and the planar mode may be directional modes.
[0391] When the intra prediction mode is the DC mode, the prediction block may be generated based on the average value of the pixel values of the plurality of reference pixels. For example, the pixel values of the prediction block may be determined based on the average value of the pixel values of the plurality of reference pixels.
[0392] The number of intra prediction modes and the mode values of each intra prediction mode described above are merely exemplary and may be defined differently depending on embodiments, implementations, and / or requirements.
[0393] In order to perform intra prediction on a target block, a step of checking whether samples included in a reconstructed neighboring block can be used as reference samples for the target block may be performed. When a sample that cannot be used as a reference sample for the target block exists among the samples in the neighboring block, a value generated by interpolation and / or replication using at least one sample value among the samples included in the reconstructed neighboring block may replace the sample value of the sample that cannot be used as a reference sample. When the value generated by replication and / or interpolation replaces the sample value of the existing sample, the sample may be used as a reference sample for the target block.
[0394] When intra prediction is used, a filter may be applied to at least one of a reference sample and a prediction sample based on at least one of a size of a target block and an intra prediction mode.
[0395] The type of filter to be applied to at least one of the reference sample and the prediction sample may differ according to at least one of the intra prediction mode of the target block, the size of the target block, and the shape of the target block. The type of filter may be classified according to one or more of the length of the filter tap, the value of the filter coefficient, and the filter strength. The length of the filter tap may indicate the number of filter taps. In addition, the number of filter taps may indicate the length of the filter.
[0396] When the intra prediction mode is the planar mode, the sample value of the predicted target block may be generated by using a weighted sum of the upper reference sample of the target block, the left reference sample of the target block, the upper right reference sample of the target block, and the lower left reference sample of the target block according to the position of the predicted target sample in the prediction block when generating the prediction block of the target block.
[0397] When the intra prediction mode is the DC mode, an average value of reference samples above the target block and reference samples to the left of the target block may be used when generating a prediction block for the target block. Furthermore, filtering using the values of the reference samples may be performed on specific rows or columns in the target block. The specific rows may be one or more rows above the reference sample. The specific columns may be one or more columns to the left of the reference sample.
[0398] When the intra prediction mode is a directional mode, the prediction block may be generated using the upper reference sample, the left reference sample, the upper right reference sample, and / or the lower left reference sample of the target block.
[0399] In order to generate the above-mentioned prediction samples, real number-based interpolation may be performed.
[0400] The intra prediction mode of the target block may be predicted from intra prediction modes of neighboring blocks adjacent to the target block, and information used for the prediction may be entropy encoded / decoded.
[0401] For example, when the intra prediction modes of the target block and the neighboring block are identical to each other, a predefined flag may be used to signal that the intra prediction modes of the target block and the neighboring block are the same.
[0402] For example, an indicator indicating the same intra prediction mode as that of the target block among the intra prediction modes of a plurality of neighboring blocks may be signaled.
[0403] When the intra prediction modes of the target block and the neighboring blocks are different from each other, information about the intra prediction mode of the target block may be encoded and / or decoded using entropy encoding and / or entropy decoding.
[0404] Figure 8 is a diagram showing reference samples used in the intra prediction process.
[0405] The reconstructed reference samples used for intra prediction of the target block may include a lower left reference sample, a left reference sample, an upper left corner reference sample, an upper reference sample, and an upper right reference sample.
[0406] For example, a left reference sample may represent a reconstructed reference pixel adjacent to the left side of the target block. An upper reference sample may represent a reconstructed reference pixel adjacent to the top of the target block. An upper-left corner reference sample may represent a reconstructed reference pixel located at the upper-left corner of the target block. A lower-left corner reference sample may represent a reference sample located below a left sample line among samples located on the same line as a left sample line composed of left reference samples. An upper-right corner reference sample may represent a reference sample located to the right of an upper sample line among samples located on the same line as an upper sample line composed of upper reference samples.
[0407] When the size of the target block is N×N, the numbers of the lower left reference sample, the left reference sample, the upper reference sample, and the upper right reference sample may all be N.
[0408] By performing intra prediction on the target block, a prediction block can be generated. The process of generating the prediction block may include determining the values of the pixels in the prediction block. The target block and the prediction block may be the same size.
[0409] The reference samples used for intra-frame prediction of the target block may change according to the intra-frame prediction mode of the target block. The direction of the intra-frame prediction mode may indicate the dependency between the reference samples and the pixels of the prediction block. For example, the value of a specified reference sample may be used as the value of one or more specified pixels in the prediction block. In this case, the specified reference sample and the one or more specified pixels in the prediction block may be samples and pixels located on a straight line along the direction of the intra-frame prediction mode. In other words, the value of the specified reference sample may be copied as the value of a pixel located in a direction opposite to the direction of the intra-frame prediction mode. Alternatively, the value of a pixel in the prediction block may be the value of a reference sample located in the direction of the intra-frame prediction mode relative to the position of the pixel.
[0410] In this example, when the intra prediction mode of the target block is vertical mode, the upper reference sample can be used for intra prediction. When the intra prediction mode is vertical mode, the value of a pixel in the prediction block can be the value of a reference sample located vertically above the position of the pixel. Therefore, the upper reference sample adjacent to the top of the target block can be used for intra prediction. In addition, the value of the pixels in a row of the prediction block can be the same as the value of the pixels of the upper reference sample.
[0411] In this example, when the intra prediction mode of the target block is horizontal, the left reference sample can be used for intra prediction. When the intra prediction mode is horizontal, the value of a pixel in the prediction block can be the value of a reference sample located horizontally to the left of the pixel. Therefore, the left reference sample adjacent to the left side of the target block can be used for intra prediction. In addition, the value of a pixel in a column of the prediction block can be the same as the value of the pixel of the left reference sample.
[0412] In an example, when the mode value of the intra prediction mode of the current block is 34, at least some of the left reference samples, the upper left reference samples, and at least some of the upper reference samples may be used for intra prediction. When the mode value of the intra prediction mode is 34, the value of a pixel in the prediction block may be the value of a reference sample diagonally located at the upper left corner of the pixel.
[0413] Also, in the case of an intra prediction mode whose mode value is a value ranging from 52 to 66, at least a portion of the upper right reference samples may be used for intra prediction.
[0414] Also, in the case of an intra prediction mode whose mode value is a value ranging from 2 to 17, at least a portion of the lower left reference samples may be used for intra prediction.
[0415] Also, in the case of an intra prediction mode whose mode value is a value ranging from 19 to 49, the top left corner reference sample may be used for intra prediction.
[0416] The number of reference samples used to determine the pixel value of one pixel in the prediction block may be 1 or 2 or more.
[0417] As described above, the pixel value of a pixel in the prediction block may be determined based on the position of the pixel and the position of the reference sample indicated by the direction of the intra prediction mode. When the position of the pixel and the position of the reference sample indicated by the direction of the intra prediction mode are integer positions, the value of one reference sample indicated by the integer position may be used to determine the pixel value of the pixel in the prediction block.
[0418] When the position of a pixel and the position of a reference sample indicated by the direction of the intra-frame prediction mode are not integer positions, an interpolated reference sample based on the two reference samples closest to the position of the reference sample may be generated. The value of the interpolated reference sample may be used to determine the pixel value of a pixel in the prediction block. In other words, when the position of a pixel in the prediction block and the position of a reference sample indicated by the direction of the intra-frame prediction mode indicate a position between two reference samples, an interpolated value based on the values of these two samples may be generated.
[0419] The prediction block generated via prediction may be different from the original target block. In other words, there may be a prediction error, which is the difference between the target block and the prediction block, and there may also be a prediction error between the pixels of the target block and the pixels of the prediction block.
[0420] Hereinafter, the terms “difference”, “error” and “residual” may be used to have the same meaning and may be used interchangeably with each other.
[0421] For example, in the case of directional intra prediction, the longer the distance between the pixels of the prediction block and the reference samples, the greater the prediction error that may occur. This prediction error can lead to discontinuities between the generated prediction block and neighboring blocks.
[0422] To reduce prediction error, a filtering operation may be applied to the prediction block. The filtering operation may be configured to adaptively apply a filter to areas of the prediction block that are considered to have a large prediction error. For example, the areas considered to have a large prediction error may be boundaries of the prediction block. Furthermore, the areas of the prediction block considered to have a large prediction error may vary depending on the intra-frame prediction mode, and the characteristics of the filter may also vary depending on the intra-frame prediction mode.
[0423] like Figure 8 As shown, for intra prediction of a target block, at least one of reference lines 0 to 3 may be used.
[0424] exist Figure 8 Each reference line in may indicate a reference sample line including one or more reference samples. When the number of a reference line is smaller, a reference sample line closer to the target block may be indicated.
[0425] The samples in segments A and F may be obtained by padding using the samples in segments B and E that are closest to the target block, rather than from the reconstructed neighboring blocks.
[0426] Index information indicating a reference sample line to be used for intra prediction of a target block may be signaled. The index information may indicate a reference sample line to be used for intra prediction of the target block among a plurality of reference sample lines. For example, the index information may have a value corresponding to any one of 0 to 3.
[0427] When the upper boundary of the target block is the boundary of the CTU, only reference sample line 0 may be available. Therefore, in this case, index information may not be signaled. When additional reference sample lines other than reference sample line 0 are used, filtering of the prediction block, which will be described later, may not be performed.
[0428] In case of inter-color intra prediction, a prediction block of a target block of a second color component may be generated based on a corresponding reconstructed block of a first color component.
[0429] For example, the first color component may be a luminance component, and the second color component may be a chrominance component.
[0430] To perform inter-color intra prediction, parameters of a linear model between the first color component and the second color component may be derived based on the template.
[0431] The template may include a reference sample above the target block (upper reference sample) and / or a reference sample to the left of the target block (left reference sample), and may include an upper reference sample and / or a left reference sample of a reconstructed block of a first color component corresponding to the reference samples.
[0432] For example, the parameters of the linear model may be derived using the following values: 1) the value of the sample of the first color component having the maximum value among the samples in the template, 2) the value of the sample of the second color component corresponding to the sample of the first color component, 3) the value of the sample of the first color component having the minimum value among the samples in the template, and 4) the value of the sample of the second color component corresponding to the sample of the first color component.
[0433] When the parameters of the linear model are derived, a prediction block of the target block may be generated by applying the corresponding reconstructed block to the linear model.
[0434] Depending on the image format, subsampling may be performed on samples adjacent to a reconstructed block of a first color component and the corresponding reconstructed block of the first color component. For example, when one sample of a second color component corresponds to four samples of a first color component, one corresponding sample may be calculated by subsampling the four samples of the first color component. When subsampling is performed, linear model parameter derivation and inter-color intra prediction may be performed based on the subsampled corresponding samples.
[0435] Information on whether to perform inter-color intra prediction and / or the range of the template may be signaled in the intra prediction mode.
[0436] The target block may be partitioned into two or four sub-blocks in a horizontal direction and / or a vertical direction.
[0437] The subblocks generated by partitioning can be reconstructed sequentially. That is, when intra prediction is performed on each subblock, a subprediction block for the subblock can be generated. Furthermore, when inverse quantization (dequantization) and / or inverse transformation is performed on each subblock, a subresidual block for the corresponding subblock can be generated. A reconstructed subblock can be generated by adding the subprediction block to the subresidual block. The reconstructed subblock can be used as a reference sample for intra prediction of the subblock with the next higher priority.
[0438] A subblock may be a block including a specific number (e.g., 16) or more samples. For example, when the target block is an 8×4 block or a 4×8 block, the target block may be partitioned into two subblocks. Furthermore, when the target block is a 4×4 block, the target block cannot be partitioned into subblocks. When the target block has another size, the target block may be partitioned into four subblocks.
[0439] Information on whether to perform intra prediction based on these subblocks and / or information on a partition direction (horizontal direction or vertical direction) may be signaled.
[0440] This subblock-based intra prediction may be restricted such that it is performed only when reference sample line 0 is used. When subblock-based intra prediction is performed, filtering of the prediction block, which will be described below, may not be performed.
[0441] A final prediction block may be generated by performing filtering on a prediction block generated through intra prediction.
[0442] Filtering may be performed by applying specific weights to filtering target samples, left reference samples, upper reference samples, and / or upper-left reference samples that are targets to be filtered.
[0443] The weight and / or reference samples (eg, range of reference samples, positions of reference samples, etc.) used for filtering may be determined based on at least one of block size, intra prediction mode, and positions of filtering target samples in a prediction block.
[0444] For example, filtering may be performed only in specific intra prediction modes (eg, DC mode, planar mode, vertical mode, horizontal mode, diagonal mode, and / or adjacent diagonal mode).
[0445] The adjacent diagonal pattern may be a pattern with a number obtained by adding k to the diagonal pattern number, or may be a pattern with a number obtained by subtracting k from the diagonal pattern number. In other words, the number of the adjacent diagonal pattern may be the sum of the diagonal pattern number and k, or may be the difference between the diagonal pattern number and k. For example, k may be a positive integer of 8 or less.
[0446] The intra prediction mode of the target block may be derived using intra prediction modes of neighboring blocks occurring around the target block, and such derived intra prediction mode may be entropy encoded and / or entropy decoded.
[0447] For example, when the intra prediction mode of the target block is the same as that of the neighboring block, specific flag information may be used to signal information indicating that the intra prediction mode of the target block is the same as that of the neighboring block.
[0448] Also, for example, indicator information of a neighboring block having the same intra prediction mode as the intra prediction mode of the target block among intra prediction modes of a plurality of neighboring blocks may be signaled.
[0449] For example, when the intra-frame prediction mode of the target block is different from the intra-frame prediction mode of the neighboring block, entropy encoding and / or entropy decoding can be performed on information about the intra-frame prediction mode of the target block by performing entropy encoding and / or entropy decoding based on the intra-frame prediction mode of the neighboring block.
[0450] Figure 9 is a diagram for explaining an embodiment of an inter-frame prediction process.
[0451] Figure 9 The rectangle shown in may represent an image (or picture). Figure 9 In the example, an arrow may indicate a prediction direction. An arrow pointing from a first picture to a second picture indicates that the second picture refers to the first picture. That is, each picture may be encoded and / or decoded according to the prediction direction.
[0452] Images can be classified into intra-frame pictures (I pictures), uni-predictive pictures or predictive coded pictures (P pictures), and bi-predictive pictures or bi-predictive coded pictures (B pictures) according to their coding types. Each picture can be encoded and / or decoded according to its coding type.
[0453] When the target image to be encoded is an I picture, the target image can be encoded using data included in the image itself without performing inter-frame prediction with reference to other images. For example, the I picture can be encoded only via intra-frame prediction.
[0454] When the target image is a P picture, the target image may be encoded through inter-frame prediction using a reference picture existing in one direction. Here, the one direction may be a forward direction or a backward direction.
[0455] When the target image is a B picture, the image may be encoded by inter-frame prediction using reference pictures existing in both directions, or may be encoded by inter-frame prediction using reference pictures existing in one of the forward direction and the backward direction. Here, the two directions may be the forward direction and the backward direction.
[0456] P-pictures and B-pictures encoded and / or decoded using reference pictures may be regarded as images using inter-frame prediction.
[0457] Hereinafter, inter prediction in inter mode according to an embodiment will be described in detail.
[0458] Inter-frame prediction or motion compensation may be performed using a reference image and motion information.
[0459] In the inter mode, the encoding apparatus 100 may perform inter prediction and / or motion compensation on the target block, and the decoding apparatus 200 may perform inter prediction and / or motion compensation corresponding to the inter prediction and / or motion compensation performed by the encoding apparatus 100 on the target block.
[0460] The motion information of the target block may be derived separately during inter prediction by the encoding apparatus 100 and the decoding apparatus 200. The motion information may be derived using the motion information of the reconstructed neighboring block, the motion information of the col block, and / or the motion information of the blocks adjacent to the col block.
[0461] For example, the encoding apparatus 100 or the decoding apparatus 200 may perform prediction and / or motion compensation by using the motion information of the spatial candidate and / or the temporal candidate as the motion information of the target block. The target block may represent a PU and / or a PU partition.
[0462] The spatial candidate may be a reconstructed block that is spatially adjacent to the target block.
[0463] The temporal candidate may be a reconstructed block corresponding to the target block in a previously reconstructed co-located picture (col picture).
[0464] In inter-frame prediction, the encoding device 100 and the decoding device 200 can improve encoding efficiency and decoding efficiency by utilizing motion information of spatial candidates and / or temporal candidates. The motion information of the spatial candidate may be referred to as "spatial motion information." The motion information of the temporal candidate may be referred to as "temporal motion information."
[0465] Hereinafter, the motion information of a spatial candidate may be the motion information of a PU including the spatial candidate. The motion information of a temporal candidate may be the motion information of a PU including the temporal candidate. The motion information of a candidate block may be the motion information of a PU including the candidate block.
[0466] Inter prediction may be performed using reference pictures.
[0467] The reference picture may be at least one of a picture before the target picture and a picture after the target picture.The reference picture may be an image used for prediction of the target block.
[0468] In inter prediction, a region in a reference picture may be specified using a reference picture index (or refIdx) indicating a reference picture, a motion vector to be described later, etc. Here, the region specified in the reference picture may indicate a reference block.
[0469] Inter prediction can select a reference picture and further select a reference block corresponding to the target block from the reference picture. In addition, inter prediction can use the selected reference block to generate a prediction block for the target block.
[0470] Motion information may be derived by each of the encoding apparatus 100 and the decoding apparatus 200 during inter prediction.
[0471] A spatial candidate may be a block that 1) exists in the target picture 2) has been previously reconstructed via encoding and / or decoding and 3) is adjacent to the target block or located at a corner of the target block. Here, a "block located at a corner of the target block" may be a block vertically adjacent to a neighboring block horizontally adjacent to the target block, or a block horizontally adjacent to a neighboring block vertically adjacent to the target block. In addition, a "block located at a corner of the target block" may have the same meaning as a "block adjacent to a corner of the target block." The meaning of "block located at a corner of the target block" may be included in the meaning of "block adjacent to the target block."
[0472] For example, the spatial candidate can be a reconstructed block located to the left of the target block, a reconstructed block located above the target block, a reconstructed block located at the lower left corner of the target block, a reconstructed block located at the upper right corner of the target block, or a reconstructed block located at the upper left corner of the target block.
[0473] Each of the encoding apparatus 100 and the decoding apparatus 200 may identify a block existing in the col picture at a position spatially corresponding to the target block. The position of the target block in the target picture and the position of the identified block in the col picture may correspond to each other.
[0474] Each of the encoding apparatus 100 and the decoding apparatus 200 may determine a col block existing at a predefined relative position with respect to the identified block as a temporal candidate. The predefined relative position may be a position existing inside and / or outside the identified block.
[0475] For example, the col block may include a first col block and a second col block. When the coordinates of the identified block are (xP, yP) and the size of the identified block is represented by (nPSW, nPSH), the first col block may be a block located at the coordinates (xP+nPSW, yP+nPSH). The second col block may be a block located at the coordinates (xP+(nPSW>>1), yP+(nPSH>>1)). When the first col block is unavailable, the second col block may be selectively used.
[0476] The motion vector of the target block may be determined based on the motion vector of the col block. Each of the encoding apparatus 100 and the decoding apparatus 200 may scale the motion vector of the col block. The scaled motion vector of the col block may be used as the motion vector of the target block. In addition, the motion vector of the motion information of the temporal candidate stored in the list may be the scaled motion vector.
[0477] The ratio of the motion vector of the target block to the motion vector of the col block may be the same as the ratio of the first temporal distance to the second temporal distance. The first temporal distance may be the distance between the reference picture and the target picture of the target block. The second temporal distance may be the distance between the reference picture and the col picture of the col block.
[0478] The scheme used to derive motion information may vary depending on the inter-frame prediction mode of the target block. For example, inter-frame prediction modes used for inter-frame prediction include Advanced Motion Vector Predictor (AMVP) mode, Merge mode, Skip mode, Merge mode with motion vector difference, Sub-block Merge mode, Triangle Partitioning mode, Inter-Intra combined prediction mode, Affine Inter mode, and Current Picture Reference mode. Merge mode may also be referred to as "motion merge mode." Each mode will be described in detail below.
[0479] 1) AMVP model When using the AMVP mode, the encoding device 100 may search for similar blocks in the vicinity of the target block. The encoding device 100 may obtain a prediction block by performing prediction on the target block using the motion information of the found similar blocks. The encoding device 100 may encode a residual block that is the difference between the target block and the prediction block.
[0480] 1-1) Create a list of predicted motion vector candidates When the AMVP mode is used as the prediction mode, each of the encoding device 100 and the decoding device 200 can create a list of prediction motion vector candidates using the motion vector of the spatial candidate, the motion vector of the temporal candidate, and the zero vector. The prediction motion vector candidate list may include one or more prediction motion vector candidates. At least one of the motion vector of the spatial candidate, the motion vector of the temporal candidate, and the zero vector may be determined and used as the prediction motion vector candidate.
[0481] Hereinafter, the terms “predicted motion vector (candidate)” and “motion vector (candidate)” may be used to have the same meaning and may be used interchangeably with each other.
[0482] Hereinafter, the terms “predicted motion vector candidate” and “AMVP candidate” may be used to have the same meaning and may be used interchangeably with each other.
[0483] Hereinafter, the terms “motion vector prediction candidate list” and “AMVP candidate list” may be used to have the same meaning and may be used interchangeably with each other.
[0484] The spatial candidate may include a reconstructed spatial neighboring block. In other words, the motion vector of the reconstructed neighboring block may be referred to as a "spatial prediction motion vector candidate."
[0485] The temporal candidate may include the col block and blocks adjacent to the col block. In other words, the motion vector of the col block or the motion vector of the block adjacent to the col block may be referred to as a "temporal prediction motion vector candidate".
[0486] The zero vector may be the (0,0) motion vector.
[0487] The predicted motion vector candidate may be a motion vector predictor for predicting a motion vector. In addition, in the encoding apparatus 100, each predicted motion vector candidate may be an initial search position for a motion vector.
[0488] 1-2) Searching for a motion vector using a list of motion vector prediction candidates The encoding device 100 may determine a motion vector to be used for encoding the target block within the search range using the list of predicted motion vector candidates. In addition, the encoding device 100 may determine a predicted motion vector candidate to be used as the predicted motion vector of the target block from among the predicted motion vector candidates in the predicted motion vector candidate list.
[0489] The motion vector to be used for encoding the target block may be a motion vector that can be encoded at a minimum cost.
[0490] Also, the encoding apparatus 100 may determine whether to encode the target block using the AMVP mode.
[0491] 1-3) Transmission of inter-frame prediction information The encoding apparatus 100 may generate a bitstream including inter prediction information required for inter prediction, and the decoding apparatus 200 may perform inter prediction on a target block using the inter prediction information of the bitstream.
[0492] The inter prediction information may include 1) mode information indicating whether the AMVP mode is used, 2) a prediction motion vector index, 3) a motion vector difference (MVD), 4) a reference direction, and 5) a reference picture index.
[0493] Hereinafter, the terms “predicted motion vector index” and “AMVP index” may be used to have the same meaning and may be used interchangeably with each other.
[0494] Furthermore, the inter prediction information may include a residual signal.
[0495] When the mode information indicates that the AMVP mode is used, the decoding apparatus 200 may acquire a predicted motion vector index, an MVD, a reference direction, and a reference picture index from a bitstream through entropy decoding.
[0496] The predicted motion vector index may indicate a predicted motion vector candidate to be used for prediction of the target block among the predicted motion vector candidates included in the predicted motion vector candidate list.
[0497] 1-4) Inter-frame prediction in AMVP mode using inter-frame prediction information The decoding apparatus 200 may derive a prediction motion vector candidate using the prediction motion vector candidate list, and may determine motion information of the target block based on the derived prediction motion vector candidate.
[0498] The decoding apparatus 200 may determine a motion vector candidate for the target block from among the predicted motion vector candidates included in the predicted motion vector candidate list using the predicted motion vector index. The decoding apparatus 200 may select the predicted motion vector candidate indicated by the predicted motion vector index from among the predicted motion vector candidates included in the predicted motion vector candidate list as the predicted motion vector of the target block.
[0499] The encoding apparatus 100 may generate an entropy-coded predicted motion vector index by applying entropy coding to the predicted motion vector index, and may generate a bitstream including the entropy-coded predicted motion vector index. The entropy-coded predicted motion vector index may be signaled from the encoding apparatus 100 to the decoding apparatus 200 through the bitstream. The decoding apparatus 200 may extract the entropy-coded predicted motion vector index from the bitstream, and may obtain the predicted motion vector index by applying entropy decoding to the entropy-coded predicted motion vector index.
[0500] The motion vector actually used for inter-frame prediction of the target block may not match the predicted motion vector. MVD may be used to indicate the difference between the motion vector actually used for inter-frame prediction of the target block and the predicted motion vector. The encoding apparatus 100 may derive a predicted motion vector similar to the motion vector actually used for inter-frame prediction of the target block in order to use the smallest possible MVD.
[0501] The motion vector difference (MVD) may be the difference between the motion vector of the target block and the predicted motion vector. The encoding apparatus 100 may calculate the MVD and may generate an entropy-encoded MVD by applying entropy encoding to the MVD. The encoding apparatus 100 may generate a bitstream including the entropy-encoded MVD.
[0502] The MVD may be transmitted from the encoding apparatus 100 to the decoding apparatus 200 through a bitstream. The decoding apparatus 200 may extract the entropy-encoded MVD from the bitstream, and may obtain the MVD by applying entropy decoding to the entropy-encoded MVD.
[0503] The decoding apparatus 200 may derive a motion vector of the target block by summing the MVD and the predicted motion vector. In other words, the motion vector of the target block derived by the decoding apparatus 200 may be the sum of the MVD and the motion vector candidate.
[0504] In addition, the encoding device 100 may generate entropy-encoded MVD resolution information by applying entropy encoding to the calculated MVD resolution information, and may generate a bitstream including the entropy-encoded MVD resolution information. The decoding device 200 may extract the entropy-encoded MVD resolution information from the bitstream, and may obtain the MVD resolution information by applying entropy decoding to the entropy-encoded MVD resolution information. The decoding device 200 may use the MVD resolution information to adjust the resolution of the MVD.
[0505] In addition, the encoding apparatus 100 may calculate MVD based on an affine model. The decoding apparatus 200 may derive an affine controlled motion vector of a target block by summing MVD and an affine controlled motion vector candidate, and may derive a motion vector of a subblock using the affine controlled motion vector.
[0506] The reference direction may indicate a list of reference pictures to be used for prediction of the target block. For example, the reference direction may indicate one of reference picture list L0 and reference picture list L1.
[0507] The reference direction only indicates the reference picture list to be used for prediction of the target block and may not mean that the direction of the reference picture is limited to the forward direction or the backward direction. In other words, each of the reference picture lists L0 and L1 may include pictures in the forward direction and / or the backward direction.
[0508] A unidirectional reference direction may mean that a single reference picture list is used. A bidirectional reference direction may mean that two reference picture lists are used. In other words, the reference direction may indicate one of the following cases: a case where only reference picture list L0 is used, a case where only reference picture list L1 is used, or a case where two reference picture lists are used.
[0509] The reference picture index may indicate a reference picture used to predict the target block among the reference pictures in the reference picture list. The encoding apparatus 100 may generate an entropy-coded reference picture index by applying entropy coding to the reference picture index, and may generate a bitstream including the entropy-coded reference picture index. The entropy-coded reference picture index may be signaled from the encoding apparatus 100 to the decoding apparatus 200 via the bitstream. The decoding apparatus 200 may extract the entropy-coded reference picture index from the bitstream, and may obtain the reference picture index by applying entropy decoding to the entropy-coded reference picture index.
[0510] When two reference picture lists are used to predict a target block, a single reference picture index and a single motion vector can be used for each reference picture list. Furthermore, when two reference picture lists are used to predict a target block, two prediction blocks can be specified for the target block. For example, the (final) prediction block for the target block can be generated using the average or weighted sum of the two prediction blocks for the target block.
[0511] The motion vector of the target block may be derived through the predicted motion vector index, MVD, reference direction, and reference picture index.
[0512] The decoding apparatus 200 may generate a prediction block for the target block based on the derived motion vector and the reference picture index. For example, the prediction block may be a reference block indicated by the derived motion vector in the reference picture indicated by the reference picture index.
[0513] Since the predicted motion vector index and the MVD are encoded but the motion vector itself of the target block is not encoded, the number of bits transmitted from the encoding apparatus 100 to the decoding apparatus 200 may be reduced and encoding efficiency may be improved.
[0514] For the target block, the motion information of reconstructed neighboring blocks can be used. In certain inter-frame prediction modes, the encoding device 100 may not separately encode the actual motion information of the target block. Instead of encoding the motion information of the target block, additional information may be encoded, where the additional information enables the motion information of the target block to be derived using the motion information of the reconstructed neighboring blocks. Since this additional information is encoded, the number of bits transmitted to the decoding device 200 can be reduced, and encoding efficiency can be improved.
[0515] For example, as an inter-frame prediction mode in which the motion information of the target block is not directly encoded, a skip mode and / or a merge mode may exist. Here, each of the encoding device 100 and the decoding device 200 may use an identifier and / or an index indicating a unit among the reconstructed neighboring units whose motion information is to be used as the motion information of the target unit.
[0516] 2) Merge mode Merge is a scheme for deriving motion information of a target block. The term "merge" may mean combining the motions of multiple blocks. "Merge" may also mean that the motion information of one block is also applied to other blocks. In other words, merge mode may be a mode in which the motion information of a target block is derived from the motion information of neighboring blocks.
[0517] When using merge mode, the encoding device 100 may use the motion information of the spatial candidate and / or the motion information of the temporal candidate to predict the motion information of the target block. The spatial candidate may include a reconstructed spatial neighboring block that is spatially adjacent to the target block. The spatial neighboring block may include a left neighboring block and an upper neighboring block. The temporal candidate may include a col block.
[0518] The terms “spatial candidate” and “spatial merging candidate” may be used to have the same meaning and may be used interchangeably with each other. The terms “temporal candidate” and “temporal merging candidate” may be used to have the same meaning and may be used interchangeably with each other.
[0519] The encoding apparatus 100 may obtain a prediction block through prediction. The encoding apparatus 100 may encode a residual block that is a difference between the target block and the prediction block.
[0520] 2-1) Create a merge candidate list When using merge mode, each of the encoding device 100 and the decoding device 200 can use the motion information of the spatial candidate and / or the motion information of the temporal candidate to create a merge candidate list. The motion information may include 1) a motion vector, 2) a reference picture index, and 3) a reference direction. The reference direction can be unidirectional or bidirectional. The reference direction may indicate an inter-frame prediction indicator.
[0521] The merge candidate list may include a merge candidate. The merge candidate may be motion information. In other words, the merge candidate list may be a list storing multiple pieces of motion information.
[0522] The merge candidate may be a plurality of pieces of motion information of temporal candidates and / or spatial candidates. In other words, the merge candidate list may include motion information of temporal candidates and / or spatial candidates, etc.
[0523] In addition, the merge candidate list may include new merge candidates generated by combining merge candidates already in the merge candidate list. In other words, the merge candidate list may include new motion information generated by combining multiple pieces of motion information previously in the merge candidate list.
[0524] In addition, the merge candidate list may include a history-based merge candidate. The history-based merge candidate may be motion information of a block that was encoded and / or decoded before the target block.
[0525] Furthermore, the merge candidate list may include a merge candidate based on an average of two merge candidates.
[0526] A merge candidate may be a specific mode for deriving inter-frame prediction information. The merge candidate may be information indicating a specific mode for deriving inter-frame prediction information. Inter-frame prediction information for a target block may be derived according to the specific mode indicated by the merge candidate. Furthermore, the specific mode may include a process for deriving a series of inter-frame prediction information. This specific mode may be an inter-frame prediction information deriving mode or a motion information deriving mode.
[0527] Inter prediction information of the target block may be derived according to a mode indicated by a merge candidate selected from among merge candidates in the merge candidate list by a merge index.
[0528] For example, the motion information derivation mode in the merge candidate list may be at least one of the following modes: 1) a motion information derivation mode for a sub-block unit and 2) an affine motion information derivation mode.
[0529] In addition, the merge candidate list may include motion information of a zero vector. A zero vector may also be referred to as a "zero merge candidate."
[0530] In other words, the multiple motion information in the merge candidate list can be at least one of the following information: 1) motion information of spatial candidates, 2) motion information of temporal candidates, 3) motion information generated by combining multiple motion information previously existing in the merge candidate list, and 4) zero vector.
[0531] Motion information may include 1) a motion vector, 2) a reference picture index, and 3) a reference direction. The reference direction may also be referred to as an "inter prediction indicator." The reference direction may be unidirectional or bidirectional. A unidirectional reference direction may indicate L0 prediction or L1 prediction.
[0532] A merge candidate list may be created before performing prediction in merge mode.
[0533] The number of merge candidates in the merge candidate list may be predefined. Each of the encoding device 100 and the decoding device 200 may add merge candidates to the merge candidate list according to a predefined scheme and a predefined priority, so that the merge candidate list has a predefined number of merge candidates. The merge candidate list of the encoding device 100 and the merge candidate list of the decoding device 200 may be made identical to each other using the predefined scheme and the predefined priority.
[0534] Merging may be applied on a CU or PU basis. When performing merging on a CU or PU basis, the encoding device 100 may transmit a bitstream including predefined information to the decoding device 200. For example, the predefined information may include 1) information indicating whether merging is performed for each block partition, and 2) information about blocks on which merging is to be performed among blocks that are spatial candidates and / or temporal candidates for the target block.
[0535] 2-2) Searching for motion vectors using the merge candidate list The encoding device 100 may determine a merge candidate to be used for encoding the target block. For example, the encoding device 100 may perform prediction on the target block using a merge candidate in the merge candidate list and may generate a residual block for the merge candidate. The encoding device 100 may encode the target block using the merge candidate that generates the minimum cost in encoding the prediction and residual block.
[0536] Also, the encoding apparatus 100 may determine whether to encode the target block using the merge mode.
[0537] 2-3) Transmission of inter-frame prediction information The encoding device 100 may generate a bitstream including inter-frame prediction information required for inter-frame prediction. The encoding device 100 may generate entropy-coded inter-frame prediction information by performing entropy coding on the inter-frame prediction information, and may transmit the bitstream including the entropy-coded inter-frame prediction information to the decoding device 200. The entropy-coded inter-frame prediction information may be signaled by the encoding device 100 to the decoding device 200 through the bitstream. The decoding device 200 may extract the entropy-coded inter-frame prediction information from the bitstream, and may obtain the inter-frame prediction information by applying entropy decoding to the entropy-coded inter-frame prediction information.
[0538] The decoding apparatus 200 may perform inter prediction on a target block using inter prediction information of a bitstream.
[0539] The inter prediction information may include 1) mode information indicating whether a merge mode is used, 2) a merge index, and 3) correction information.
[0540] Furthermore, the inter prediction information may include a residual signal.
[0541] The decoding apparatus 200 may acquire a merge index from a bitstream only when the mode information indicates that the merge mode is used.
[0542] The mode information may be a merge flag. The unit of the mode information may be a block. The information about the block may include the mode information, and the mode information may indicate whether the merge mode is applied to the block.
[0543] The merge index may indicate a merge candidate to be used for predicting the target block among the merge candidates included in the merge candidate list. Alternatively, the merge index may indicate a block to be merged with the target block among neighboring blocks that are spatially or temporally adjacent to the target block.
[0544] The encoding apparatus 100 may select a merge candidate having the highest encoding performance among the merge candidates included in the merge candidate list, and may set a value of the merge index to indicate the selected merge candidate.
[0545] The correction information may be information for correcting a motion vector. The encoding apparatus 100 may generate the correction information. The decoding apparatus 200 may correct a motion vector of a merge candidate selected by a merge index based on the correction information.
[0546] The correction information may include at least one of information indicating whether correction is to be performed, correction direction information, and correction size information. A prediction mode that corrects a motion vector based on signaled correction information may be referred to as a 'merge mode with motion vector difference'.
[0547] 2-4) Inter-frame prediction using merge mode of inter-frame prediction information The decoding apparatus 200 may perform prediction on the target block using a merge candidate indicated by a merge index among merge candidates included in the merge candidate list.
[0548] The motion vector of the target block may be specified by the motion vector of the merge candidate indicated by the merge index, the reference picture index, and the reference direction.
[0549] 3) Skip mode Skip mode can be a mode in which the motion information of a spatial candidate or a temporal candidate is applied to the target block without change. In addition, skip mode can be a mode in which a residual signal is not used. In other words, when skip mode is used, the reconstructed block can be the same as the predicted block.
[0550] The difference between the merge mode and the skip mode is whether the residual signal is transmitted or used. That is, the skip mode is similar to the merge mode except that the residual signal is not transmitted or used.
[0551] When the skip mode is used, the encoding device 100 may transmit information about a block whose motion information is to be used as the motion information of the target block among blocks that are spatial candidates or temporal candidates to the decoding device 200 through a bitstream. The encoding device 100 may generate entropy-coded information by performing entropy coding on the information, and may signal the entropy-coded information to the decoding device 200 through a bitstream. The decoding device 200 may extract the entropy-coded information from the bitstream, and may obtain information by applying entropy decoding to the entropy-coded information.
[0552] In addition, when the skip mode is used, the encoding device 100 may not transmit other syntax information (such as MVD) to the decoding device 200. For example, when the skip mode is used, the encoding device 100 may not signal syntax elements related to at least one of the MVD, the coded block flag, and the transform coefficient level to the decoding device 200.
[0553] 3-1) Create a merge candidate list The merge candidate list can also be used in skip mode. In other words, the merge candidate list can be used in both merge mode and skip mode. In this regard, the merge candidate list can also be referred to as a "skip candidate list" or a "merge / skip candidate list."
[0554] Alternatively, the skip mode may use an additional candidate list different from the candidate list of the merge mode. In this case, in the following description, the merge candidate list and the merge candidate may be replaced by the skip candidate list and the skip candidate, respectively.
[0555] A merge candidate list may be created before performing prediction in skip mode.
[0556] 3-2) Searching for motion vectors using the merge candidate list The encoding device 100 may determine a merge candidate to be used for encoding the target block. For example, the encoding device 100 may perform prediction on the target block using the merge candidate in the merge candidate list. The encoding device 100 may encode the target block using the merge candidate that generates the minimum cost in prediction.
[0557] Also, the encoding apparatus 100 may determine whether to encode the target block using the skip mode.
[0558] 3-3) Transmission of inter-frame prediction information The encoding apparatus 100 may generate a bitstream including inter prediction information required for inter prediction, and the decoding apparatus 200 may perform inter prediction on a target block using the inter prediction information of the bitstream.
[0559] The inter prediction information may include 1) mode information indicating whether the skip mode is used and 2) a skip index.
[0560] The skip index may be the same as the merge index described above.
[0561] When skip mode is used, the target block may be encoded without using a residual signal. The inter prediction information may not include a residual signal. Alternatively, the bitstream may not include a residual signal.
[0562] The decoding device 200 may obtain the skip index from the bitstream only when the mode information indicates that the skip mode is used. As described above, the merge index and the skip index may be the same as each other. The decoding device 200 may obtain the skip index from the bitstream only when the mode information indicates that the merge mode or the skip mode is used.
[0563] The skip index may indicate a merge candidate to be used for prediction of a target block among merge candidates included in the merge candidate list.
[0564] 3-4) Inter-frame prediction in skip mode using inter-frame prediction information The decoding apparatus 200 may perform prediction on the target block using the merge candidate indicated by the skip index among the merge candidates included in the merge candidate list.
[0565] The motion vector of the target block may be specified by the motion vector of the merge candidate indicated by the skip index, the reference picture index, and the reference direction.
[0566] 4) Current picture reference mode The current picture reference mode may denote a prediction mode that uses a previously reconstructed area in a target picture to which the target block belongs.
[0567] A motion vector for specifying a previously reconstructed area may be used.A reference picture index of a target block may be used to determine whether the target block has been encoded in a current picture reference mode.
[0568] A flag or index indicating whether the target block is a block encoded in the current picture reference mode may be signaled by the encoding apparatus 100 to the decoding apparatus 200. Alternatively, whether the target block is a block encoded in the current picture reference mode may be inferred through the reference picture index of the target block.
[0569] When the target block is encoded in the current picture reference mode, the current picture may exist at a fixed position or an arbitrary position in the reference picture list for the target block.
[0570] For example, the fixed position may be a position where the value of the reference picture index is 0 or a last position.
[0571] When a target picture exists at an arbitrary position in a reference picture list, an additional reference picture index indicating such an arbitrary position may be signaled by the encoding apparatus 100 to the decoding apparatus 200 .
[0572] 5) Sub-block merging mode The subblock merge mode may be a mode in which motion information is derived from subblocks of a CU.
[0573] When the subblock merging mode is applied, a subblock merging candidate list may be generated using the motion information of the col-sub-block of the target subblock (i.e., the subblock-based temporal merging candidate) in the reference image and / or the affine control point motion vector merging candidate.
[0574] 6) Triangular partitioning mode In the triangular partitioning mode, the target block can be partitioned in a diagonal direction, and sub-target blocks generated by the partitioning can be generated. For each sub-target block, motion information of the corresponding sub-target block can be derived, and the derived motion information can be used to derive prediction samples for each sub-target block. The prediction samples of the target block can be derived by taking the weighted sum of the prediction samples of the sub-target blocks generated by the partitioning.
[0575] 7) Combined inter-intra prediction mode The combined inter-intra prediction mode may be a mode in which a prediction sample of a target block is derived using a weighted sum of a prediction sample generated via inter prediction and a prediction sample generated via intra prediction.
[0576] In the above-described mode, the decoding device 200 may autonomously correct the derived motion information. For example, the decoding device 200 may search for motion information having the minimum sum of absolute differences (SAD) in a specific area based on the reference block indicated by the derived motion information, and may derive the found motion information as corrected motion information.
[0577] In the above-described mode, the decoding apparatus 200 may use optical flow to compensate for prediction samples derived through inter-frame prediction.
[0578] In the above-described AMVP mode, merge mode, skip mode, etc., index information of a list may be used to specify motion information to be used for prediction of a target block among a plurality of pieces of motion information in the list.
[0579] To improve encoding efficiency, the encoding apparatus 100 may signal only the index of the element generating the minimum cost in inter-frame prediction of the target block among the elements in the list. The encoding apparatus 100 may encode the index and signal the encoded index.
[0580] Therefore, the encoding device 100 and the decoding device 200 must be able to derive the above-described lists (i.e., the predicted motion vector candidate list and the merge candidate list) using the same scheme based on the same data. Here, the same data may include reconstructed pictures and reconstructed blocks. In addition, in order to specify an element using an index, the order of the elements in the list must be fixed.
[0581] Figure 10 Illustrated are spatial candidates according to an embodiment.
[0582] exist Figure 10 , the positions of the spatial candidates are shown.
[0583] The large block in the center of the figure may represent the target block, and the five small blocks may represent spatial candidates.
[0584] The coordinates of the target block may be (xP, yP), and the size of the target block may be represented by (nPSW, nPSH).
[0585] The spatial candidate A0 may be a block adjacent to the lower left corner of the target block. A0 may be a block occupying a pixel located at coordinates (xP-1, yP+nPSH).
[0586] Spatial candidate A1 may be a block adjacent to the left side of the target block. A1 may be the lowest block among the blocks adjacent to the left side of the target block. Alternatively, A1 may be a block adjacent to the top of A0. A1 may be a block occupying the pixel at coordinates (xP-1, yP+nPSH-1).
[0587] The spatial candidate B0 may be a block adjacent to the upper right corner of the target block. B0 may be a block occupying a pixel located at coordinates (xP+nPSW, yP-1).
[0588] Spatial candidate B1 may be a block adjacent to the top of the target block. B1 may be the rightmost block among the blocks adjacent to the top of the target block. Alternatively, B1 may be a block adjacent to the left of B0. B1 may be a block occupying the pixel located at coordinates (xP+nPSW-1, yP-1).
[0589] The spatial candidate B2 may be a block adjacent to the upper left corner of the target block. B2 may be a block occupying a pixel located at coordinates (xP-1, yP-1).
[0590] Determination of availability of spatial and temporal candidates In order to include the motion information of the spatial candidate or the motion information of the temporal candidate in the list, it is necessary to determine whether the motion information of the spatial candidate or the motion information of the temporal candidate is available.
[0591] Hereinafter, candidate blocks may include spatial candidates and temporal candidates.
[0592] For example, the determination may be performed by sequentially applying the following steps 1) to 4) below.
[0593] Step 1) When the PU including the candidate block is located outside the boundary of the picture, the availability of the candidate block may be set to “false.” The expression “availability is set to false” may have the same meaning as “set to unavailable.”
[0594] Step 2) When the PU including the candidate block is located outside the boundary of the slice, the availability of the candidate block may be set to “false.” When the target block and the candidate block are located in different slices, the availability of the candidate block may be set to “false.”
[0595] Step 3) When the PU including the candidate block is located outside the boundary of the tile, the availability of the candidate block may be set to “false.” When the target block and the candidate block are located in different tiles, the availability of the candidate block may be set to “false.”
[0596] Step 4) When the prediction mode of the PU including the candidate block is intra prediction mode, the availability of the candidate block may be set to “false.” When the PU including the candidate block does not use inter prediction, the availability of the candidate block may be set to “false.”
[0597] Figure 11 An order in which motion information of spatial candidates is added to a merge list according to an embodiment is shown.
[0598] like Figure 11 As shown in , when multiple pieces of motion information of spatial candidates are added to the merge list, the order of A1, B1, B0, A0, and B2 may be used. That is, multiple pieces of motion information of available spatial candidates may be added to the merge list in the order of A1, B1, B0, A0, and B2.
[0599] Methods for exporting merge lists in merge mode and skip mode As described above, the maximum number of merge candidates in a merge list can be set. The set maximum number can be indicated by "N". The set number can be transmitted from the encoding device 100 to the decoding device 200. The slice header of the slice may include N. In other words, the maximum number of merge candidates in the merge list for the target block of the slice can be set by the slice header. For example, the value of N can basically be 5.
[0600] A plurality of pieces of motion information (ie, merge candidates) may be added to the merge list in the order of the following steps 1) to 4) below.
[0601] Step 1) Among the space candidates, available space candidates can be added to the merge list. Figure 11 The plurality of pieces of motion information of the available spatial candidates are added to the merge list in the order shown in FIG. Here, when the motion information of the available spatial candidate overlaps with other motion information already in the merge list, the motion information of the available spatial candidate may not be added to the merge list. The operation of checking whether the corresponding motion information overlaps with other motion information in the list may be simply referred to as "overlap check."
[0602] The maximum number of pieces of motion information added may be N.
[0603] Step 2) When the number of motion information items in the merge list is less than N and a temporal candidate is available, the motion information of the temporal candidate may be added to the merge list. Here, when the motion information of the available temporal candidate overlaps with other motion information already in the merge list, the motion information of the available temporal candidate may not be added to the merge list.
[0604] Step 3) When the number of pieces of motion information in the merge list is less than N and the type of the target slice is 'B', combined motion information generated by combining bidirectional predictions (bi-prediction) may be added to the merge list.
[0605] The target slice may be a slice including the target block.
[0606] The combined motion information may be a combination of L0 motion information and L1 motion information. The L0 motion information may be motion information that refers only to the L0 reference picture list. The L1 motion information may be motion information that refers only to the L1 reference picture list.
[0607] In the merge list, there may be one or more pieces of L0 motion information. In addition, in the merge list, there may be one or more pieces of L1 motion information.
[0608] The combined motion information may include one or more pieces of combined motion information. When generating the combined motion information, the L0 motion information and the L1 motion information to be used in the step of generating the combined motion information may be predefined among the one or more pieces of L0 motion information and the one or more pieces of L1 motion information. The one or more pieces of combined motion information may be generated in a predefined order by combining bidirectional prediction using a pair of different motion information in a merge list. One piece of the pair of different motion information may be L0 motion information, and the other piece of the pair of different motion information may be L1 motion information.
[0609] For example, the combined motion information added with the highest priority may be a combination of L0 motion information with a merge index of 0 and L1 motion information with a merge index of 1. When the motion information with a merge index of 0 is not L0 motion information or when the motion information with a merge index of 1 is not L1 motion information, the combined motion information may be neither generated nor added. Next, the combined motion information added with the next highest priority may be a combination of L0 motion information with a merge index of 1 and L1 motion information with a merge index of 0. The subsequent detailed combinations may conform to other combinations in the field of video encoding / decoding.
[0610] Here, when the combined motion information overlaps with other motion information already present in the merge list, the combined motion information may not be added to the merge list.
[0611] Step 4) When the number of pieces of motion information in the merge list is less than N, motion information of a zero vector may be added to the merge list.
[0612] The zero-vector motion information may be motion information in which a motion vector is a zero vector.
[0613] The number of pieces of zero-vector motion information may be one or more. The reference picture indexes of one or more pieces of zero-vector motion information may be different from each other. For example, the reference picture index value of the first zero-vector motion information may be 0. The reference picture index value of the second zero-vector motion information may be 1.
[0614] The number of pieces of zero-vector motion information may be the same as the number of reference pictures in the reference picture list.
[0615] The reference direction of the zero-vector motion information can be bidirectional. Both motion vectors can be zero vectors. The number of pieces of zero-vector motion information can be the smaller of the number of reference pictures in reference picture list L0 and the number of reference pictures in reference picture list L1. Alternatively, when the number of reference pictures in reference picture list L0 and the number of reference pictures in reference picture list L1 are different, a unidirectional reference direction can be used for a reference picture index applicable only to a single reference picture list.
[0616] The encoding apparatus 100 and / or the decoding apparatus 200 may then add zero-vector motion information to the merge list while changing the reference picture index.
[0617] When the zero-vector motion information overlaps with other motion information already in the merge list, the zero-vector motion information may not be added to the merge list.
[0618] The order of steps 1) to 4) is merely exemplary and may be changed. In addition, some of the steps above may be omitted according to predefined conditions.
[0619] Method for deriving a candidate list of predicted motion vectors in AMVP mode The maximum number of motion vector predictor candidates in the motion vector predictor candidate list may be predefined. The predefined maximum number may be indicated by N. For example, the predefined maximum number may be 2.
[0620] A plurality of pieces of motion information (ie, motion vector predictor candidates) may be added to the motion vector predictor candidate list in the order of steps 1) to 3) below.
[0621] Step 1) Available spatial candidates among the spatial candidates may be added to the motion vector prediction candidate list.The spatial candidates may include a first spatial candidate and a second spatial candidate.
[0622] The first spatial candidate may be one of A0, A1, scaled A0, and scaled A1. The second spatial candidate may be one of B0, B1, B2, scaled B0, scaled B1, and scaled B2.
[0623] Multiple pieces of motion information of available spatial candidates may be added to the motion vector prediction candidate list in the order of the first spatial candidate and the second spatial candidate. In this case, if the motion information of an available spatial candidate overlaps with other motion information already in the motion vector prediction candidate list, the motion information of the available spatial candidate may not be added to the motion vector prediction candidate list. In other words, when the value of N is 2, if the motion information of the second spatial candidate is the same as the motion information of the first spatial candidate, the motion information of the second spatial candidate may not be added to the motion vector prediction candidate list.
[0624] The maximum number of pieces of motion information added may be N.
[0625] Step 2) When the number of pieces of motion information in the motion vector predictor candidate list is less than N and a temporal candidate is available, the motion information of the temporal candidate may be added to the motion vector predictor candidate list. In this case, if the motion information of the available temporal candidate overlaps with other motion information already in the motion vector predictor candidate list, the motion information of the available temporal candidate may not be added to the motion vector predictor candidate list.
[0626] Step 3) When the number of pieces of motion information in the motion vector predictor candidate list is less than N, zero vector motion information may be added to the motion vector predictor candidate list.
[0627] The zero-vector motion information may include one or more pieces of zero-vector motion information. Reference picture indices of the one or more pieces of zero-vector motion information may be different from each other.
[0628] The encoding apparatus 100 and / or the decoding apparatus 200 may sequentially add a plurality of pieces of zero-vector motion information to the prediction motion vector candidate list while changing the reference picture index.
[0629] When the zero-vector motion information overlaps with other motion information already existing in the motion vector predictor candidate list, the zero-vector motion information may not be added to the motion vector predictor candidate list.
[0630] The above description of the zero vector motion information in conjunction with the merge list is also applicable to the zero vector motion information, and its repeated description will be omitted.
[0631] The order of steps 1) to 3) described above is merely exemplary and may be changed. In addition, some of the steps may be omitted according to predefined conditions.
[0632] Figure 12 Transformation and quantization processes according to examples are shown.
[0633] like Figure 12 As shown in , the quantized levels may be generated by performing a transform and / or quantization process on the residual signal.
[0634] The residual signal may be generated as a difference between the original block and the predicted block.Here, the predicted block may be a block generated via intra prediction or inter prediction.
[0635] The residual signal may be transformed into a signal in the frequency domain through a transform process as part of the quantization process.
[0636] The transform kernel used for the transform may include various DCT kernels such as discrete cosine transform (DCT) type 2 (DCT-II) and discrete sine transform (DST) kernels.
[0637] These transform kernels may perform separable transform or two-dimensional (2D) non-separable transform on the residual signal. The separable transform may be a transform indicating that a one-dimensional (1D) transform is performed on the residual signal in each of horizontal and vertical directions.
[0638] The DCT type and DST type adaptively used for 1D transform may include DCT-V, DCT-VIII, DST-I, and DST-VII in addition to DCT-II, as shown in each of Table 3 and Table 4 below.
[0639] Table 3
[0640] Table 4
[0641] As shown in Tables 3 and 4, when deriving the DCT type or DST type to be used for transformation, a transform set can be used. Each transform set may include multiple transform candidates. Each transform candidate may be a DCT type or a DST type.
[0642] Table 5 below shows an example of a transform set to be applied to a horizontal direction and a transform set to be applied to a vertical direction according to an intra prediction mode.
[0643] Table 5
[0644] In Table 5, the numbers of a vertical transform set and a horizontal transform set to be applied to a horizontal direction of a residual signal according to an intra prediction mode of a target block are shown.
[0645] like Figure 4 and Figure 5 As illustrated in , a transform set to be applied in the horizontal and vertical directions may be predefined according to the intra prediction mode of the target block. The encoding apparatus 100 may perform transform and inverse transform on the residual signal using the transform included in the transform set corresponding to the intra prediction mode of the target block. In addition, the decoding apparatus 200 may perform inverse transform on the residual signal using the transform included in the transform set corresponding to the intra prediction mode of the target block.
[0646] In the transformation and inverse transformation, the transform set to be applied to the residual signal may be determined and may not be signaled as illustrated in Tables 3, 4, and 5. Transform indication information may be signaled from the encoding apparatus 100 to the decoding apparatus 200. The transform indication information may be information indicating which of a plurality of transform candidates included in the transform set to be applied to the residual signal is to be used.
[0647] For example, when the target block size is 64×64, a transform set each having three transforms can be configured according to the intra prediction mode. An optimal transform method can be selected from a total of nine multi-transform methods generated by combining three transforms in the horizontal direction and three transforms in the vertical direction. Using such an optimal transform method, the residual signal can be encoded and / or decoded, thereby improving coding efficiency.
[0648] Here, information indicating which of the multiple transforms belonging to each transform set has been used for at least one of the vertical transform and the horizontal transform may be entropy encoded and / or entropy decoded. Here, truncated unary binarization may be used to encode and / or decode such information.
[0649] As described above, methods using various transforms may be applied to a residual signal generated via intra prediction or inter prediction.
[0650] The transform may include at least one of a first transform and a secondary transform. A transform coefficient may be generated by performing a first transform on the residual signal, and a secondary transform coefficient may be generated by performing a secondary transform on the transform coefficient.
[0651] The first transform may be referred to as a “primary transform.” Furthermore, the first transform may also be referred to as an “adaptive multi-transform (AMT) scheme.” As described above, AMT may indicate that different transforms are applied to each 1D direction (ie, vertical and horizontal directions).
[0652] The secondary transform may be a transform for increasing the energy concentration of the transform coefficients generated by the first transform. Similar to the first transform, the secondary transform may be a separable transform or a non-separable transform. Such a non-separable transform may be a non-separable secondary transform (NSST).
[0653] The first transform may be performed using at least one of a plurality of predefined transform methods, such as discrete cosine transform (DCT), discrete sine transform (DST), Karhunen-Loeve transform (KLT), and the like.
[0654] Furthermore, the first transform may be a transform having various types according to a kernel function defining discrete cosine transform (DCT) or discrete sine transform (DST).
[0655] For example, the transform type may be determined based on at least one of the following: 1) a prediction mode of the target block (e.g., one of intra-frame prediction and inter-frame prediction), 2) a size of the target block, 3) a shape of the target block, 4) an intra-frame prediction mode of the target block, 5) a component of the target block (e.g., one of a luminance component and a chrominance component), and 6) a partition type applied to the target block (e.g., one of a quadtree, a binary tree, and a ternary tree).
[0656] For example, according to the transform kernels presented in Table 6 below, the first transform may include transforms such as DCT-2, DCT-5, DCT-7, DST-7, DST-1, DST-8, and DCT-8. In Table 6 below, various transform types and transform kernel functions for multi-transform selection (MTS) are exemplified.
[0657] MTS may refer to the selection of a combination of one or more DCT and / or DST kernels to transform the residual signal in the horizontal and / or vertical directions.
[0658] Table 6
[0659] In Table 6, i and j may be integer values equal to or greater than 0 and less than or equal to N-1.
[0660] A secondary transform may be performed on transform coefficients generated by performing the first transform.
[0661] As in the first transform, a transform set may also be defined in a secondary transform.The method for deriving and / or determining the above-mentioned transform set may be applied not only to the first transform but also to the secondary transform.
[0662] The primary transform and the secondary transform may be determined for a specific goal.
[0663] For example, the first transform and the secondary transform may be applied to signal components corresponding to one or more of a luma component and a chroma component. Whether to apply the first transform and / or the secondary transform may be determined based on at least one of the coding parameters for the target block and / or the neighboring blocks. For example, whether to apply the first transform and / or the secondary transform may be determined based on the size and / or shape of the target block.
[0664] In the encoding apparatus 100 and the decoding apparatus 200 , transform information indicating a transform method to be used for a target may be derived by using the designation information.
[0665] For example, the transform information may include a transform index to be used for the primary transform and / or the secondary transform. Alternatively, the transform information may indicate that the primary transform and / or the secondary transform is not to be used.
[0666] For example, when the target of the primary transform and the secondary transform is a target block, the transform method to be applied to the primary transform and / or the secondary transform indicated by the transform information can be determined based on at least one of the encoding parameters for the target block and / or blocks adjacent to the target block.
[0667] Alternatively, transformation information indicating a transformation method for a specific target may be signaled from the encoding apparatus 100 to the decoding apparatus 200 .
[0668] For example, for a single CU, whether a primary transform is used, an index indicating the primary transform, whether a secondary transform is used, and an index indicating the secondary transform may be derived as transform information by the decoding apparatus 200. Alternatively, for a single CU, transform information indicating whether a primary transform is used, an index indicating the primary transform, whether a secondary transform is used, and an index indicating the secondary transform may be signaled.
[0669] A quantized transform coefficient (ie, a quantization level) may be generated by performing quantization on a result generated by performing the first transform and / or the sub-transform or performing quantization on a residual signal.
[0670] Figure 13 A diagonal scan according to an example is shown.
[0671] Figure 14 A horizontal scan according to an example is shown.
[0672] Figure 15 A vertical scan according to an example is shown.
[0673] The quantized transform coefficients may be scanned via at least one of (upper right) diagonal scanning, vertical scanning, and horizontal scanning according to at least one of an intra prediction mode, a block size, and a block shape. The block may be a transform unit (TU).
[0674] Each scan may be initiated at a specific start point and may be terminated at a specific end point.
[0675] For example, by using Figure 13 The coefficients of the block are scanned by diagonal scanning to change the quantized transform coefficients into 1D vector form. Optionally, the quantized transform coefficients can be used according to the size of the block and / or the intra prediction mode. Figure 14 Horizontal scan or Figure 15 vertical scanning instead of diagonal scanning.
[0676] Vertical scanning may be an operation of scanning 2D block type coefficients in a column direction, and horizontal scanning may be an operation of scanning 2D block type coefficients in a row direction.
[0677] In other words, which of diagonal scanning, vertical scanning, and horizontal scanning is to be used may be determined according to the size of a block and / or an inter prediction mode.
[0678] like Figure 13 、 Figure 14 and Figure 15 As shown in , the quantized transform coefficients may be scanned along a diagonal direction, a horizontal direction, or a vertical direction.
[0679] The quantized transform coefficients can be represented by a block shape. Each block can include multiple sub-blocks. Each sub-block can be defined according to a minimum block size or a minimum block shape.
[0680] In the scanning, a scanning order according to a type or direction of scanning may be first applied to a subblock. In addition, a scanning order according to a direction of scanning may be applied to quantized transform coefficients in each subblock.
[0681] For example, Figure 13 、 Figure 14 and Figure 15As shown in , when the size of the target block is 8×8, quantized transform coefficients can be generated through the first transform, the secondary transform, and the quantization of the residual signal of the target block. Therefore, one of three types of scanning orders can be applied to the four 4×4 sub-blocks, and the quantized transform coefficients can also be scanned for each 4×4 sub-block according to the scanning order.
[0682] The encoding apparatus 100 may generate entropy-encoded quantized transform coefficients by performing entropy encoding on the scanned quantized transform coefficients, and may generate a bitstream including the entropy-encoded quantized transform coefficients.
[0683] The decoding apparatus 200 may extract entropy-encoded quantized transform coefficients from a bitstream and may generate quantized transform coefficients by performing entropy decoding on the entropy-encoded quantized transform coefficients. The quantized transform coefficients may be arranged in the form of 2D blocks via inverse scanning. Here, as the inverse scanning method, at least one of upper right diagonal scanning, vertical scanning, and horizontal scanning may be performed.
[0684] In the decoding apparatus 200, inverse quantization may be performed on the quantized transform coefficients. Depending on whether the inverse secondary transform is to be performed, a secondary inverse transform may be performed on the result generated by performing the inverse quantization. In addition, depending on whether the first inverse transform is to be performed, a first inverse transform may be performed on the result generated by performing the secondary inverse transform. A reconstructed residual signal may be generated by performing the first inverse transform on the result generated by performing the secondary inverse transform.
[0685] For luma components reconstructed via intra prediction or inter prediction, inverse mapping with a dynamic range may be performed before loop filtering.
[0686] The dynamic range can be divided into 16 equal segments and the mapping function of the corresponding segments can be signaled.Such mapping function can be signaled at the slice level or tile group level.
[0687] An inverse mapping function for performing inverse mapping may be derived based on the mapping function.
[0688] Loop filtering, storage of reference pictures, and motion compensation may be performed in the inverse mapping area.
[0689] The prediction block generated by inter prediction can be transformed to the mapping area by mapping using a mapping function, and the transformed prediction block can be used to generate a reconstructed block. However, since intra prediction is performed in the mapping area, the prediction block generated by intra prediction can be used to generate a reconstructed block without the need for mapping and / or inverse mapping.
[0690] For example, when the target block is a residual block of a chroma component, the residual block may be transformed to the inverse mapping area by scaling the chroma component of the mapping area.
[0691] Whether scaling is available can be signaled at the slice level or tile group level.
[0692] For example, scaling may be applied only if mapping is available for luma components and the partitions of the luma components and the partitions of the chroma components follow the same tree structure.
[0693] Scaling may be performed based on an average value of the values of samples in the luma prediction block corresponding to the chroma prediction block.Here, when the target block uses inter prediction, the luma prediction block may refer to a mapped luma prediction block.
[0694] The value required for scaling may be derived by referring to a lookup table using an index of a segment to which the average value of the sample values of the luma prediction block belongs.
[0695] The residual block may be transformed to the inverse mapping region by scaling the residual block using the final derived value. Thereafter, for the block of the chroma component, reconstruction, intra prediction, inter prediction, loop filtering, and storage of reference pictures may be performed in the inverse mapping region.
[0696] For example, information indicating whether mapping and / or inverse mapping of luma components and chroma components is available may be signaled through a sequence parameter set.
[0697] A prediction block for a target block may be generated based on a block vector. The block vector may indicate a displacement between the target block and a reference block. The reference block may be a block in the target image.
[0698] In this manner, a prediction mode that generates a prediction block by referring to a target image may be referred to as an 'intra block copy (IBC) mode'.
[0699] The IBC mode can be applied to a CU of a specific size. For example, the IBC mode can be applied to a CU of M×N size. Here, M and N can be less than or equal to 64.
[0700] IBC modes may include skip mode, merge mode, AMVP mode, and the like. In the case of skip mode or merge mode, a merge candidate list may be configured, and a merge index may be signaled, thereby specifying a single merge candidate among the merge candidates present in the merge candidate list. The block vector of the specified merge candidate may be used as the block vector of the target block.
[0701] In the case of AMVP mode, a differential block vector can be signaled. In addition, a prediction block vector can be derived from the left neighboring block and the upper neighboring block of the target block. In addition, an index indicating which neighboring block to use can be signaled.
[0702] The prediction block in IBC mode can be included in the target CTU or the left CTU and can be limited to blocks within the previously reconstructed area. For example, the value of the block vector can be limited so that the prediction block of the target block is located in a specific area. The specific area can be an area defined by three 64×64 blocks that are encoded and / or decoded before the 64×64 block including the target block. By limiting the value of the block vector in this way, the memory consumption and device complexity caused by the implementation of the IBC mode can be reduced.
[0703] Figure 16 is a configuration diagram of an encoding device according to an embodiment.
[0704] The encoding apparatus 1600 may correspond to the encoding apparatus 100 described above.
[0705] The encoding apparatus 1600 may include a processing unit 1610, a memory 1630, a user interface (UI) input device 1650, a UI output device 1660, and a storage 1640 communicating with each other through a bus 1690. The encoding apparatus 1600 may further include a communication unit 1620 connected to a network 1699.
[0706] The processing unit 1610 may be a central processing unit (CPU) or a semiconductor device for executing processing instructions stored in the memory 1630 or the storage 1640. The processing unit 1610 may be at least one hardware processor.
[0707] The processing unit 1610 may generate and process signals, data, or information input to, output from, or used in the encoding device 1600, and may perform checks, comparisons, determinations, etc. related to the signals, data, or information. In other words, in an embodiment, the generation and processing of data or information, and checks, comparisons, and determinations related to the data or information may be performed by the processing unit 1610.
[0708] The processing unit 1610 may include an inter-frame prediction unit 110, an intra-frame prediction unit 120, a switch 115, a subtractor 125, a transform unit 130, a quantization unit 140, an entropy encoding unit 150, an inverse quantization unit 160, an inverse transform unit 170, an adder 175, a filter unit 180 and a reference picture buffer 190.
[0709] At least some of the inter-frame prediction unit 110, the intra-frame prediction unit 120, the switch 115, the subtractor 125, the transform unit 130, the quantization unit 140, the entropy encoding unit 150, the inverse quantization unit 160, the inverse transform unit 170, the adder 175, the filter unit 180, and the reference picture buffer 190 may be program modules and may communicate with an external device or system. The program modules may be included in the encoding apparatus 1600 in the form of an operating system, an application module, or other program modules.
[0710] The program modules may be physically stored in various types of well-known storage devices. In addition, at least some of the program modules may also be stored in a remote storage device that can communicate with the encoding device 1600.
[0711] Program modules may include, but are not limited to, routines, subroutines, programs, objects, components, and data structures for performing functions or operations according to embodiments or for implementing abstract data types according to embodiments.
[0712] The program modules may be implemented using instructions or codes executed by at least one processor of the encoding device 1600 .
[0713] The processing unit 1610 can execute instructions or codes in the inter-frame prediction unit 110, the intra-frame prediction unit 120, the switch 115, the subtractor 125, the transform unit 130, the quantization unit 140, the entropy coding unit 150, the inverse quantization unit 160, the inverse transform unit 170, the adder 175, the filter unit 180 and the reference picture buffer 190.
[0714] The storage unit may represent the memory 1630 and / or the storage 1640. Each of the memory 1630 and the storage 1640 may be any of various types of volatile or non-volatile storage media. For example, the memory 1630 may include at least one of a read-only memory (ROM) 1631 and a random access memory (RAM) 1632.
[0715] The storage unit may store data or information used for the operation of the encoding apparatus 1600. In an embodiment, data or information of the encoding apparatus 1600 may be stored in the storage unit.
[0716] For example, the storage unit may store pictures, blocks, lists, motion information, inter-frame prediction information, bitstreams, and the like.
[0717] The encoding device 1600 may be implemented in a computer system including a computer-readable storage medium.
[0718] The storage medium may store at least one module required for the operation of the encoding apparatus 1600. The memory 1630 may store at least one module and may be configured such that the at least one module is executed by the processing unit 1610.
[0719] Functions related to communication of data or information of the encoding apparatus 1600 may be performed through the communication unit 1620 .
[0720] For example, the communication unit 1620 may transmit the bitstream to the decoding apparatus 1700 which will be described later.
[0721] Figure 17 is a configuration diagram of a decoding device according to an embodiment.
[0722] The decoding device 1700 may correspond to the decoding device 200 described above.
[0723] The decoding apparatus 1700 may include a processing unit 1710, a memory 1730, a user interface (UI) input device 1750, a UI output device 1760, and a storage 1740 communicating with each other through a bus 1790. The decoding apparatus 1700 may further include a communication unit 1720 connected to a network 1799.
[0724] The processing unit 1710 may be a central processing unit (CPU) or a semiconductor device for executing processing instructions stored in the memory 1730 or the storage 1740. The processing unit 1710 may be at least one hardware processor.
[0725] The processing unit 1710 may generate and process signals, data, or information input to, output from, or used in the decoding device 1700, and may perform checks, comparisons, determinations, etc. related to the signals, data, or information. In other words, in embodiments, the generation and processing of data or information, as well as checks, comparisons, and determinations related to the data or information, may be performed by the processing unit 1710.
[0726] The processing unit 1710 may include an entropy decoding unit 210 , an inverse quantization unit 220 , an inverse transform unit 230 , an intra prediction unit 240 , an inter prediction unit 250 , a switch 245 , an adder 255 , a filter unit 260 , and a reference picture buffer 270 .
[0727] At least some of the entropy decoding unit 210, the inverse quantization unit 220, the inverse transform unit 230, the intra-frame prediction unit 240, the inter-frame prediction unit 250, the adder 255, the switch 245, the filter unit 260, and the reference picture buffer 270 of the decoding device 200 may be program modules and may communicate with an external device or system. The program modules may be included in the decoding device 1700 in the form of an operating system, an application module, or other program modules.
[0728] The program modules may be physically stored in various types of well-known storage devices. In addition, at least some of the program modules may also be stored in a remote storage device that can communicate with the decoding device 1700.
[0729] Program modules may include, but are not limited to, routines, subroutines, programs, objects, components, and data structures for performing functions or operations according to embodiments or for implementing abstract data types according to embodiments.
[0730] The program modules may be implemented using instructions or codes executed by at least one processor of the decoding device 1700 .
[0731] The processing unit 1710 can execute instructions or codes in the entropy decoding unit 210, the inverse quantization unit 220, the inverse transform unit 230, the intra-frame prediction unit 240, the inter-frame prediction unit 250, the switch 245, the adder 255, the filter unit 260 and the reference picture buffer 270.
[0732] The storage unit may represent the memory 1730 and / or the storage 1740. Each of the memory 1730 and the storage 1740 may be any of various types of volatile or non-volatile storage media. For example, the memory 1730 may include at least one of the ROM 1731 and the RAM 1732.
[0733] The storage unit may store data or information used for the operation of the decoding apparatus 1700. In an embodiment, data or information of the decoding apparatus 1700 may be stored in the storage unit.
[0734] For example, the storage unit may store pictures, blocks, lists, motion information, inter-frame prediction information, bitstreams, and the like.
[0735] The decoding device 1700 may be implemented in a computer system including a computer-readable storage medium.
[0736] The storage medium may store at least one module required for the operation of the decoding apparatus 1700. The memory 1730 may store at least one module and may be configured such that the at least one module is executed by the processing unit 1710.
[0737] Functions related to communication of data or information of the decoding apparatus 1700 may be performed through the communication unit 1720 .
[0738] For example, the communication unit 1720 may receive a bitstream from the encoding apparatus 1700 .
[0739] Hereinafter, a processing unit may refer to the processing unit 1610 of the encoding apparatus 1600 and / or the processing unit 1710 of the decoding apparatus 1700. For example, with respect to functions related to prediction, a processing unit may refer to the switch 115 and / or the switch 245. With respect to functions related to inter-frame prediction, a processing unit may refer to the inter-frame prediction unit 110, the subtractor 125, and the adder 175, and may refer to the inter-frame prediction unit 250 and the adder 255. With respect to functions related to intra-frame prediction, a processing unit may refer to the intra-frame prediction unit 120, the subtractor 125, and the adder 175, and may refer to the intra-frame prediction unit 240 and the adder 255. With respect to functions related to transform, a processing unit may refer to the transform unit 130 and the inverse transform unit 170, and may refer to the inverse transform unit 230. With respect to functions related to quantization, a processing unit may refer to the quantization unit 140 and the inverse quantization unit 160, and may refer to the inverse quantization unit 220. Regarding functions related to entropy encoding and / or entropy decoding, the processing unit may indicate the entropy encoding unit 150 and / or the entropy decoding unit 210. Regarding functions related to filtering, the processing unit may indicate the filter unit 180 and / or the filter unit 260. Regarding functions related to reference pictures, the processing unit may indicate the reference picture buffer 190 and / or the reference picture buffer 270.
[0740] Adaptive filter selection In typical inter-frame prediction, a fixed motion compensation filter is used when generating a prediction block, and thus there may be a limitation in improving encoding efficiency.
[0741] To improve encoding efficiency, embodiments may provide an image encoding / decoding method, apparatus, and bitstream storage medium including an adaptive filter selection method.
[0742] Definitions of Terms Used in the Examples Neighboring blocks: Neighboring blocks can refer to blocks adjacent to the target block. Neighboring blocks can include spatially adjacent blocks and temporally adjacent blocks. Neighboring blocks can also refer to reconstructed adjacent blocks in a reference image. Neighboring blocks do not necessarily need to be in contact with the target block.
[0743] Spatially neighboring blocks: Spatially neighboring blocks may be blocks that are spatially adjacent to the target block.
[0744] The target block and spatially neighboring blocks may be included in a target image.
[0745] The spatially neighboring block may include a block at least a portion of a boundary of which contacts at least a portion of a boundary of the target block. Alternatively, the spatially neighboring block may include a block at a distance from the target block that is less than or equal to a reference value.
[0746] The spatially neighboring blocks may include blocks diagonally adjacent to vertices of the target block.
[0747] The spatially adjacent blocks may include an upper left block adjacent to the upper left of the target block, an upper block adjacent to the top of the target block, an upper right block adjacent to the upper right of the target block, a left block adjacent to the left of the target block, a right block adjacent to the right of the target block, a lower left block adjacent to the lower left of the target block, a lower block adjacent to the bottom of the target block, and a lower right block adjacent to the lower right of the target block.
[0748] Temporally neighboring blocks: Temporally neighboring blocks may be blocks that are temporally adjacent to the target block.
[0749] Temporally adjacent blocks may include colblocks. A colblock may be a block in a reconstructed image stored in a reference image buffer. A colpicture image (colpicture) may refer to a picture including a colblock. A colpicture image may be an image included in a reference image list.
[0750] The col block may be determined based on the position of the target block in the target image.The fact that two blocks are "temporally adjacent to each other" may mean that the positions of the two blocks satisfy a specific condition.
[0751] The position of the col block in the col image may be the same as the position of the target block in the target image. Alternatively, the position of the col block in the col image may correspond to the position of the target block in the target image. Here, the situation where the positions of the blocks correspond to each other may mean that the areas of the blocks are the same as each other, may mean that the area of one block is included in the area of another block, and may mean that one block occupies a specific position of another block.
[0752] For example, the position of the col block in the col image may be the same as the position of the target block in the target image. Alternatively, the col block may be a block including col pixels in the col image. The col pixels may be pixels having the same coordinates as a specific pixel in the target block.
[0753] The temporally neighboring block may be a block temporally adjacent to the spatially neighboring block of the target block.
[0754] Neighboring samples: Neighboring samples may refer to samples in neighboring blocks. Neighboring samples may include predicted samples, reconstructed samples, residual samples, and decoded samples.
[0755] Hereinafter, terms listed in one row may be used as the same meaning in the embodiments and may be used interchangeably with each other in the embodiments.
[0756] - "Motion information", "Motion vector", "Block vector" - "Bili prediction", "Bidirectional prediction", "Inter bi prediction" and "Bidirectional inter prediction" Predefined value: A predefined value may refer to a value commonly used by the encoding device 1600 and the decoding device 1700. For example, a predefined value may be interpreted as being limited to a fixed value. Alternatively, the predefined value may be a value shared between the encoding device 1600 and the decoding device 1700 through signaling. Alternatively, the predefined value may be a value derived through the same process in the encoding device 1600 and the decoding device 1700, so that the encoding device 1600 and the decoding device 1700 have a common value. Alternatively, the predefined value may be a value shared by the encoding device and the decoding device.
[0757] The values derived through the same process in the encoding device 1600 and the decoding device 1700 may include values derived through the same process for the same value and / or the same information in the encoding device 1600 and the decoding device 1700 .
[0758] The values derived by the encoding device 1600 and the decoding device 1700 through the same process may include values derived by using the same conditional statement with respect to the same value and / or the same information in the encoding device 1600 and the decoding device 1700 .
[0759] The description of predefined values can also be applied to predefined information. In the above description, "value" can be replaced by "information".
[0760] Motion information: Motion information may refer to information including at least one of the following items: reference picture list information, reference image, motion vector candidate, motion vector candidate index, merge candidate and merge index, block vector, block vector candidate and block vector candidate index, as well as motion vector, reference picture index and inter-frame prediction indicator.
[0761] In an embodiment, “a case where an indicator indicating whether a specific method is performed is true” may mean a case where it is true whether a specific method is performed in a prediction mode, motion information; encoding parameter; and / or position indicated by the indicator.
[0762] For example, the indicator indicating whether a specific mode is to be executed may have a value ranging from 0 to 3, and the specific mode may be executed only when the indicator has a value of 1 or 3. In this case, “a case where the indicator indicating whether a specific mode is to be executed is true” may mean a case where the indicator indicating whether a specific mode is to be executed has a value of 1 or 3.
[0763] In an embodiment, “a case where an indicator indicating whether a specific method is executed is false” may mean a case where an indicator indicating whether a specific method is executed is not true.
[0764] Extension of the embodiment of inter prediction mode The inter prediction mode, the intra block copy (IBC) mode, and the intra template matching prediction mode may have a common feature of referring to a specific reconstructed block for prediction for a target block.
[0765] In the following embodiments, “intra template matching prediction” and “intra template matching” may be used to have the same meaning and may be used interchangeably with each other in the embodiments.
[0766] Therefore, in an embodiment, the inter prediction mode can be replaced by the IBC mode or the intra template matching mode. The description of the case where the inter prediction mode is used for the target block can also be applied to the case where the IBC mode or the intra template matching mode is used for the target block. The description of the inter prediction mode can be applied to the IBC mode or the intra template matching mode, and the inter prediction mode can be replaced by the IBC mode or the intra template matching mode. In addition, information related to the inter prediction mode can be regarded as information related to the IBC mode or the intra template matching mode. The description of the information related to the inter prediction mode can be applied to information related to the IBC mode or the intra template matching mode. For example, when the IBC mode or the intra template matching mode is used for the target block, the value of the inter prediction mode indicator can be 0 (or false). However, in this case, the value of the IBC mode indicator or the intra template matching mode indicator can be 1 (or true).
[0767] Furthermore, in an embodiment, a motion vector (MV) used for inter-frame prediction may be replaced with a block vector (BV) used for IBC. The description of the case where an MV is used for a target block may also be applied to the case where a BV is used for a target block. The description of the MV may be applied to the BV, and the MV may be replaced with the BV. Furthermore, information related to the MV may be considered as information related to the BV. The description of information related to the MV may be applied to information related to the BV. However, the BV may be information indicating a specific reconstructed block in a target image (not a reference image) that includes the target block.
[0768] In an embodiment related to the inter-frame prediction mode, the reference block and reference template used for template matching are described as existing in the reference image. On the other hand, when the IBC mode or the intra-frame template matching mode is used for the target block, the reference block and reference template may only exist in the target image. Therefore, in an embodiment, the reference image described with respect to the inter-frame prediction mode can be regarded as the target image in the IBC mode and the intra-frame template matching mode. Alternatively, in an embodiment, the reference image described with respect to the inter-frame prediction mode can be limited to the target image in the IBC mode and the intra-frame template matching mode, and images other than the target image may not be referenced in the IBC mode and the intra-frame template matching mode.
[0769] Adaptive Motion Vector Resolution (AMVR) In an embodiment, the term "resolution" may refer to the term "motion vector resolution".
[0770] In adaptive motion vector resolution, the resolution of motion vector differences can be adjusted on a block basis.
[0771] The adaptive motion vector resolution information may indicate a resolution of a motion vector difference. The resolution of the motion vector difference for the target block may be determined by signaling / encoding / decoding regarding the adaptive motion vector resolution information.
[0772] The motion vector resolutions applicable to blocks may be the same as or different from each other.
[0773] For example, the resolution of the motion vector applicable to the target block may be determined based on at least one of an encoding parameter, motion information, and mode information of the target block.
[0774] Adaptive motion vector resolution can improve coding efficiency by adjusting the resolution of motion vector differences.
[0775] For example, the adjusted resolution may be one of 16-pixels (pel), 8-pixels, 4-pixels, full-pixels, half-pixels, and quarter-pixels, and is not limited to the pixel values listed above.
[0776] The pixel may refer to the number of pixels used as a unit. For example, when the adjusted resolution of the target block is 4-pixels, each component of the motion vector difference may represent a multiple of four pixels.
[0777] When the value of the component of the motion vector difference is changed by 1 when the adjusted resolution is n-pixels, the position indicated by the motion vector difference may change by n pixels. In other words, when the adjusted resolution of the target block is n-pixels, each component of the motion vector difference may indicate a reference block in units of n pixels.
[0778] The resolution of the motion vector differences may be predefined.
[0779] Subsampling Performing subsampling may mean selecting only some of the samples in a particular region.
[0780] Performing subsampling may mean, when selecting only some of the samples in a specific area, 1) selecting samples at the SUBSAMPLE_START_HOR-th position in the horizontal direction and the SUBSAMPLE_START_VER-th position in the vertical direction relative to the specific sample, and 2) selecting samples that have a sample interval that is a multiple of SUBSAMPLE_STEP_HOR in the horizontal direction and a multiple of SUBSAMPLE_STEP_VER in the vertical direction relative to the sample in 1). Alternatively, performing subsampling may mean selecting the samples in 1) and some of the samples in 2).
[0781] The specific position may indicate an upper left sampling point of the region where subsampling is performed. However, the specific position is not limited to the upper left sampling point of the region where subsampling is performed.
[0782] Each of SUBSAMPLE_START_HOR and SUBSAMPLE_START_VER may be 0 or a positive integer. Information about at least one of SUBSAMPLE_START_HOR and SUBSAMPLE_START_VER may be signaled / encoded / decoded, or alternatively, values of SUBSAMPLE_START_HOR and / or SUBSAMPLE_START_VER may be determined as predefined values without signaling / encoding / decoding of information.
[0783] Each of SUBSAMPLE_STEP_HOR and SUBSAMPLE_STEP_VER may be 0, 1, 2, or a positive integer. Information about at least one of SUBSAMPLE_STEP_HOR and SUBSAMPLE_STEP_VER may be signaled / encoded / decoded, or alternatively, values of SUBSAMPLE_STEP_HOR and / or SUBSAMPLE_STEP_VER may be determined as predefined values without signaling / encoding / decoding of information.
[0784] Subsampling methods can be classified according to the area where subsampling is performed, the position of the area where subsampling is performed, the size of the area where subsampling is performed, SUBSAMPLE_START_HOR, SUBSAMPLE_START_VER, SUBSAMPLE_STEP_HOR, and SUBSAMPLE_STEP_VER. However, the criteria based on which the subsampling methods are classified are not limited to the above values.
[0785] Information about each subsampling method may be signaled / encoded / decoded, or alternatively, a predefined subsampling method may be used without signaling / encoding / decoding.
[0786] The information about each subsampling method may be information for determining the corresponding subsampling method.
[0787] For example, the information about each subsampling method may be information about at least one of a region where subsampling is performed, a position of the region where subsampling is performed, a size of the region where subsampling is performed, SUBSAMPLE_START_HOR, SUBSAMPLE_START_VER, SUBSAMPLE_STEP_HOR, and SUBSAMPLE_STEP_VER.
[0788] For example, the subsampling method may be determined based on at least one of motion information, encoding parameters, size, and prediction mode of the target block.
[0789] The subsampling method may be determined in at least one unit of a sequence level, a picture level, a tile level, a tile group level, a slice level, a coding tree unit (CTU) level, a coding unit (CU) level, and a prediction unit (PU) level, but the unit for determining the subsampling method is not limited thereto, and the subsampling method may be determined for a specific unit described in the embodiments.
[0790] Geometric Partitioning Model (GPM) GPM can be a method for determining the partition boundaries of a target block and deriving a weighted sum of two prediction blocks (or reference blocks) as the final prediction block. Here, the partition boundaries can partition the target block along one of various directions. For the weighted sum, a weight map can be determined based on the partition boundaries.
[0791] For example, at least one prediction block (or at least one reference block) in the geometric partitioning mode may refer to a prediction block generated by unidirectional prediction and / or bidirectional prediction, or a reference block in at least one direction of unidirectional prediction and / or bidirectional prediction.
[0792] Alternatively, for example, at least one prediction block in the geometric partition mode may refer to a prediction block generated by intra prediction.
[0793] For example, one prediction block may be generated by inter prediction in the geometric partition mode, and another prediction block may be generated by intra prediction.
[0794] Figure 18 Partition boundaries in a geometric partitioning scheme according to an example are shown.
[0795] exist Figure 18 In the , rectangles can represent blocks. The solid or dashed lines in each rectangle can represent the partition boundaries of the geometric partitions.
[0796] exist Figure 18 , 20 square blocks for 20 predefined angles are shown. In other words, one block can indicate one angle. In each square block, four partition boundaries corresponding to four distances are shown as solid or dashed lines.
[0797] exist Figure 18 , the partition boundaries shown in one block may represent partition boundaries for a partition pattern that may be selected according to a particular angle θ and p to the angle θ.
[0798] Each partition mode may be a value indicating a partition boundary. Each partition mode may refer to a mode in which geometric partitioning will be performed.
[0799] The value of partition mode may indicate a combination of θ and ρ. A specific value of partition mode may indicate a combination of a specific θ and a specific ρ.
[0800] The partition mode may be an integer value. In other words, a combination of a specific θ and a specific ρ may be represented by a value of one partition mode and may be signaled between the encoding apparatus 1600 and the decoding apparatus 1700 .
[0801] A partition pattern of a predefined number of GPMs may be determined based on the constraints of θ and ρ. For example, a partition pattern of 80 GPMs may be defined and used based on the 20 predefined angles for θ and the four predefined distances for ρ described above.
[0802] A combination of a specific θ and a specific ρ may be excluded from the partitioning pattern of the geometric partitioning. For example, such an excluded combination may be a combination that overlaps with other combinations. Alternatively, such an excluded combination may refer to a partitioning method that is the same as another partitioning method for the target block.
[0803] exist Figure 18 In FIG, each dotted line indicates a combination of θ and ρ that is excluded from the partitioning pattern of the geometric partitioning. With this exclusion, a partitioning pattern of 64 GPMs can be defined and used.
[0804] The partition mode of a GPM may specify the shape of the GPM and the partition boundaries of the GPM. Based on this specification, hereinafter, the term "partition mode" of a GPM may have the same meaning as the "shape" of the GPM or the "partition boundary" of the GPM, and the terms "mode," "partition mode," "shape," and "partition boundary" may be used interchangeably.
[0805] Each partition mode of a GPM may be indicated by an integer value or an index. Hereinafter, a partition mode of a GPM may refer to an integer value and / or an index used to determine and / or identify a shape of the GPM and / or a partition boundary of the GPM.
[0806] Here, no value can be assigned to Figure 18 The partitioning pattern indicated by the dotted line in . That is, Figure 18 The partition modes indicated by dashed lines in are only used to determine the order of the partition modes and may not be actually used in the GPM.
[0807] The partition information candidate list for the geometric partition mode may include multiple partition information candidates. Each of the partition information candidates may include information for specifying the processing of the geometric partition mode. For example, each of the partition information candidates may include information for specifying a partition boundary / partition line. The partition information candidates may specify different processing.
[0808] Figure 19 Partition boundaries, partition offsets, and partition angles in a geometric partitioning pattern according to an example are shown.
[0809] Geometric partitions in GPM can be specified by partition angles and partition offsets.
[0810] Hereinafter, θ(Theta) may indicate the partition angle, and this may be specified by the partition offset ρ(Rho).
[0811] θ may be the angle of the partition boundary. For example, θ may be the angle between the bottom line of the target block and the partition boundary. Alternatively, θ may be the angle between the X axis and the partition boundary.
[0812] ρ can be the (shortest) distance between the specific location of the target block and the partition boundary. Alternatively, ρ can be the distance between two points on a line passing through the specific location of the target block and perpendicular to the partition boundary. These two points can be a point at the specific location of the target block and a point on the partition boundary.
[0813] For example, Figure 19 As shown, the specific point may be the lower right corner of the target block. The specific point may be the rightmost and lowermost pixel of the target block.
[0814] For example, the specific point may be the center of the target block. The specific position may be the center pixel of the target block.
[0815] For example, the specific point may be the lower left corner of the target block. The specific point may be the leftmost and lowermost pixel of the target block.
[0816] θ can be constrained to a predefined value. For example, θ can be one of 20 predefined angles.
[0817] ρ can be constrained to a predefined value. For example, ρ can be one of four predefined distances.
[0818] The predefined distance may be varied according to θ. Alternatively, the predefined distance may be determined based on θ.
[0819] The predefined distance may be changed according to the size of the target block. Alternatively, the predefined distance may be determined based on the size of the target block.
[0820] θ and ρ may be implemented as fixed point values and may be represented by integers.
[0821] The two partition areas of the target block can be specified based on the partition boundary of the geometric partition mode. The partition boundary can partition the target block into two partition areas. The first partition area can be the upper left area, the upper area, or the left area of the partition boundary. The second partition area can be the lower right area, the lower area, or the right area of the partition boundary. In other words, when the partition line is not a vertical line, the first partition area can be the area above the partition line, and the second partition area can be the area below the partition line. In other words, when the partition line is a vertical line, the first partition area can be the area to the left of the partition line, and the second partition area can be the area to the right of the partition line.
[0822] Figure 20 The weight maps used in various prediction blocks depending on specific partition boundaries according to an example are shown.
[0823] exist Figure 20 , a first weight map for a first prediction block and a second weight map for a second prediction block are shown.
[0824] The first prediction block may be one of the two prediction blocks generated in the geometric partitioning mode. The second prediction block may be the other of the two prediction blocks generated in the geometric partitioning mode. The first weight map for the first prediction block may indicate weights of pixels in the first prediction block. The second weight map for the second prediction block may indicate weights of pixels in the second prediction block.
[0825] By means of the first weight map and the second weight map, the weights of the pixels corresponding to the first prediction block and the second prediction block can be determined. Here, the corresponding pixels can be pixels with the same coordinates.
[0826] The first prediction block may be a prediction block for the first partitioned area. Here, "the first prediction block for the first partitioned area" may mean that the first prediction block is used to determine the value of each of all pixels in the first partitioned area. Among the weights in the first weight map for the first prediction block, the weights included in the first partitioned area may be 0 or greater. Among the weights in the first weight map for the first prediction block, at least some of the weights included in the second partitioned area may be 0. In other words, the first prediction block may not be used to determine the values of at least some of the pixels included in the second partitioned area. Alternatively, the values of at least some of the pixels included in the first partitioned area may be determined only by the first prediction block, regardless of the second prediction block. The values of at least some of the pixels included in the second partitioned area may be determined only by the second prediction block, regardless of the first prediction block.
[0827] The second prediction block may be a prediction block for the second partitioned area. Here, "the second prediction block for the second partitioned area" may mean that the second prediction block is used to determine the value of each of all pixels in the second partitioned area. Among the weights in the second weight map for the second prediction block, the weights included in the second partitioned area may be 0 or greater. Among the weights in the second weight map for the second prediction block, at least some of the weights included in the first partitioned area may be 0. In other words, the second prediction block may not be used to determine the values of at least some of the pixels included in the first partitioned area. Alternatively, the values of at least some of the pixels included in the second partitioned area may be determined only by the second prediction block, regardless of the first prediction block. The values of at least some of the pixels included in the first partitioned area may be determined only by the first prediction block, regardless of the second prediction block.
[0828] exist Figure 20 In the weight map for each prediction block, a white area may mean that the prediction block does not affect the configuration of the white area in the final prediction block. That is, the white area in the weight map for the prediction block may indicate that the weight of the pixels in the white area in the prediction block is 0.
[0829] The weight of a specific pixel in the first prediction block may be determined based on the distance between the specific pixel and the partition boundary.The weight of a specific pixel in the second prediction block may be determined based on the distance between the specific pixel and the partition boundary.
[0830] The weight of a specific pixel in the first prediction block may be determined according to the distance between the specific pixel and the partition boundary. The weight of a specific pixel in the second prediction block may be determined according to the distance between the specific pixel and the partition boundary.
[0831] For example, when the distance between a specific pixel in the target block and the partition boundary is less than a reference value, the value of the specific pixel may be a weighted sum of the value of the first pixel in the first prediction block and the value of the second pixel in the second prediction block. Here, the position of the specific pixel, the position of the first pixel, and the position of the second pixel may be the same as each other. In this case, when the first pixel is included in the first partition area, the first weight for the first pixel may be larger as the position of the first pixel is farther from the partition boundary. When the first pixel is included in the second partition area, the weight for the first pixel may be smaller as the position of the first pixel is farther from the partition boundary. When the second pixel is included in the second partition area, the weight for the second pixel may be larger as the position of the second pixel is farther from the partition boundary. When the second pixel is included in the first partition area, the weight for the second pixel may be smaller as the position of the second pixel is farther from the partition boundary.
[0832] For example, when the distance between a specific pixel in the target block and the partition boundary is greater than a reference value, the value of the specific pixel may be the value of the first pixel in the first prediction block or the value of the second pixel in the second prediction block. In this case, when the specific pixel is included in the first partition area, the value of the first pixel in the first prediction block may be used as the value of the specific pixel. When the specific pixel is included in the second partition area, the value of the second pixel in the second prediction block may be used as the value of the specific pixel. Here, the position of the specific pixel, the position of the first pixel, and the position of the second pixel may be the same as each other.
[0833] The following [Equation 1] indicates that Figure 20 The weight map shown in corresponds to the generation of the prediction signal of the GPM.
[0834] [Equation 1] P G =(W0 P0+ W1 P1+ 4)>>3 P0 may be a pixel value of a pixel at a specific position in the first prediction block.
[0835] W0 may be a weight for a specific position among the weights in the first weight map.
[0836] P1 may be a pixel value of a pixel at a specific position in the second prediction block.
[0837] W1 may be a weight for a specific position among the weights in the second weight map.
[0838] P G It can be the pixel value of the pixel at a specific position in the final prediction block.
[0839] Each partition information candidate in the partition information candidate list may include a plurality of pieces of information for specifying a weight map in a geometric partition mode.
[0840] The partition information candidates in the partition information candidate list may respectively specify different processes in the geometric partition mode. Here, the processes in the geometric partition mode may include partition boundaries and weight maps.
[0841] Template Matching(TM) Figure 21 Template matching according to an example is shown.
[0842] In template matching, the motion information of the target block may be determined and / or changed based on the result of calculating the cost function between the target template and the reference template. The determination and / or change of the motion information may refer to refinement of the motion information.
[0843] The target template may be a template of the target block, and the reference template may be a template of the reference block.
[0844] In an embodiment, the cost function may be at least one of the sum of absolute differences (SAD), the sum of absolute transformed differences (SATD), the mean removed sum of absolute differences (MR-SAD), the mean square error (MSE), and the sum of squared errors (SSE). However, the cost function is not limited to the items listed above.
[0845] The reference block may include at least one of the following: 1) a block indicated by initial motion information, 2) a block indicated by motion information derived in a search process of template matching, 3) a block indicated by motion information finally refined by template matching, 4) a block whose sample point (or position) within the search range of template matching is one of an upper left sample point (or position), a lower left sample point (or position), an upper right sample point (or position), a lower right sample point (or position), and a center sample point (or position), and 5) a block finally determined by template matching.
[0846] The size of the reference block may be equal to the size of the target block.
[0847] The motion information refined by template matching may be the motion information having the lowest matching cost derived in the search process of template matching.However, the method for deriving the motion information is not limited to the above-mentioned standard.
[0848] The template matching cost may refer to a calculation result using a cost function between a template of a target block and a template of a reference block used in template matching.
[0849] Each of the reference block, the reference template, and the reference region may include at least one of prediction samples, reconstructed samples, residual samples, and decoded samples for the reference image. Alternatively, each of the reference block, the reference template, and the reference region may include at least one of prediction samples, reconstructed samples, residual samples, and decoded samples for the target image.
[0850] Template configuration in template matching Target template configuration The target template may include surrounding samples of the target block.
[0851] The reference area for the target block may include surrounding samples of the target block.
[0852] For example, the reference area for the target block may include at least one of samples located in a lower left area, a left area, an upper left area, an upper area, and an upper right area around the target block.
[0853] For example, the target template in template matching may be the same as the reference region for the target block.
[0854] For example, samples in a target template specified based on a target block as a reference may be samples corresponding to samples in a reference template specified based on a reference block.
[0855] In an embodiment, specifying a specific sample point based on a specific block may mean determining the specific sample point according to a relative position with respect to the specific block.
[0856] For example, when configuring a target template in template matching, some of the samples in the reference area of the target block may be selected, and the target template may be configured using the selected samples.
[0857] For example, the samples selected to configure the target template based on the target block may be samples corresponding to the samples selected to configure the template of the reference block based on the reference block.
[0858] For example, the reference area of the target block specified based on the target block may be an area corresponding to the reference area of the reference block specified based on the reference block.
[0859] Reference template configuration The reference template may include surrounding samples of the reference block.
[0860] The reference area of the reference block may include surrounding samples of the reference block.
[0861] For example, the reference area of the reference block may include at least one of samples located in a lower left area, a left area, an upper left area, an upper area, and an upper right area around the reference block.
[0862] For example, the reference template in template matching can be the same as the reference area of the reference block.
[0863] For example, samples in a reference template specified based on a reference block as a reference may be samples corresponding to samples in a target template specified based on a target block.
[0864] In an embodiment, specifying a specific sample point based on a specific block may mean determining the specific sample point according to a relative position with respect to the specific block.
[0865] For example, when configuring a reference template in template matching, some sample points in a reference area of a reference block may be selected and the reference template may be configured using the selected sample points.
[0866] For example, the samples selected to configure the reference template based on the reference block may be samples corresponding to the samples selected to configure the template of the target block based on the target block.
[0867] For example, the reference area of the reference block specified based on the reference block may be an area corresponding to the reference area of the target block specified based on the target block.
[0868] The template matching method may include at least one of an intra-frame template matching mode and an inter-frame template matching mode.
[0869] The intra template matching mode may refer to a template matching method in which each of a reference block, a reference template, and a reference region includes at least one of a prediction sample, a reconstructed sample, a residual sample, and a decoded sample for a target image.
[0870] The inter template matching mode may refer to a template matching method in which each of a reference block, a reference template, and a reference region includes at least one of a prediction sample, a reconstructed sample, a residual sample, and a decoded sample for a reference image.
[0871] The target / reference template in template matching may include at least one of the following: 1) at least one sample point in the TMSIZE_LEFT line adjacent to the left side of the target / reference block; and 2) at least one sample point in the TMSIZE_ABOVE line adjacent to the top of the target / reference block.
[0872] However, the positional relationship between each sample point in the template and the target / reference block and / or the template configuration method are not limited to the above-mentioned relationship or method.
[0873] Each of TMSIZE_LEFT and TMSIZE_ABOVE may be 0, 1, 2, 3, 4, or a positive integer of 4 or greater.
[0874] TMSIZE_LEFT and TMSIZE_ABOVE may be the same as each other. Alternatively, TMSIZE_LEFT and TMSIZE_ABOVE may be different from each other.
[0875] Each of TMSIZE_LEFT and TMSIZE_ABOVE may be a predefined value or a value determined based on signaled / encoded / decoded information.
[0876] Each of TMSIZE_LEFT and TMSIZE_ABOVE may be determined based on at least one of motion information, encoding parameters, size, and prediction mode of the target block.
[0877] Subsampling in Template Matching Subsampling for template configuration When configuring a template for template matching, all samples in the reference region may be used, or alternatively, only some of the samples in the reference region may be used.
[0878] The template used for template matching may refer to at least one of a template of a target block and a template of a reference block.
[0879] The reference region may refer to at least one of a reference region of a target block for template matching and a reference region of a reference block.
[0880] When configuring a template for template matching using only some samples located in a reference region, subsampling may be performed on all or part of the reference region.
[0881] When configuring a template for template matching using only some samples located in a reference region, the reference region may be divided into two or more regions. Each of the divided regions may be one of the following: 1) a first region for which subsampling is performed, 2) a second region for which subsampling is not performed and which is used for template configuration, and 3) a third region that is not used for template configuration. The template for template matching may be configured using samples selected by subsampling in the first region and samples in the second region.
[0882] For example, the first region may be a region located on the left and / or upper left of the block among regions in the reference region.
[0883] For example, the first corresponding region may be a region located above and / or to the upper left of the block among regions in the reference region.
[0884] Alternatively, when configuring a template for template matching using only some samples located in a reference region, the reference region may be divided into two or more regions. Each of the divided regions may be one of: 1) a first region on which subsampling is performed and 2) a third region not used for configuring a template. The template for template matching may be configured using samples selected in the first region by subsampling.
[0885] Subsampling for cost function calculation When calculating the cost function between the target template and the reference template in template matching, all the samples in each template may be used, or alternatively, only some of the samples in each template may be used. In other words, the cost function calculation may be performed only for some of the samples.
[0886] When only some samples in the templates are used to calculate the cost function between templates, subsampling may be performed on all or part of the template region.
[0887] When only some samples in the templates are used to calculate the cost function between templates, the area of each template used for template matching can be divided into two or more areas. Each of the divided areas can be one of the following: 1) a first area for which subsampling is performed, 2) a second area for which subsampling is not performed and used to calculate the cost function, and 3) a third area not used to calculate the cost function. The cost function between templates in template matching can be calculated using 1) the samples selected by subsampling in the first area and 2) the samples in the second area.
[0888] Alternatively, when only some samples in the templates are used to calculate the cost function between templates, the region of each template used for template matching can be divided into two or more regions. Each of the divided regions can be one of the following: 1) a first region on which subsampling is performed and 2) a third region not used for cost function calculation. The cost function between templates in template matching can be calculated using the samples selected by subsampling in the first region.
[0889] Subsampling of the search area When performing a search process in template matching, all samples / positions in the search area may be used, or alternatively, only some of the samples / positions in the search area may be selected. Searching and / or matching cost calculation may be performed only on the selected samples / positions. Alternatively, searching and / or matching cost calculation may be performed only on a plurality of pieces of motion information indicating the selected samples / positions.
[0890] When a search process in template matching is performed using only some of the samples / positions in the search area, subsampling may be performed on the entirety or a portion of the search area.
[0891] When configuring the search process in template matching using only some of the samples / positions in the search region, each search region may be divided into two or more regions. Each divided region may be one of: 1) a first region for which subsampling is performed, 2) a second region for which the search process is performed without subsampling, and 3) a third region for which the search process is not performed. The search process in template matching may be performed on pixels and / or positions selected by subsampling in the first region and on samples / positions in the second region. Alternatively, the search process in template matching may be performed on multiple pieces of motion information indicating samples / positions selected by subsampling in the first region and samples / positions in the second region.
[0892] Alternatively, when configuring the search process in template matching using only some of the samples / positions in the search region, each search region may be divided into two or more regions. Each of the divided regions may be one of: 1) a first region for which subsampling is performed, and 2) a first region for which the search process is not performed. The search process in template matching may be performed using the samples / positions selected by subsampling in the first region. Alternatively, the search process in template matching may be performed on multiple pieces of motion information indicating the samples / positions selected by subsampling in the first region.
[0893] Example of subsampling Figures 22a to 22t Various examples of subsampling methods in template matching are shown.
[0894] Figures 23a to 23n Various other examples of subsampling methods in template matching are shown.
[0895] Figures 24a to 24n Various other examples of subsampling methods in template matching are shown.
[0896] Figures 22a to 22t 、 Figures 23a to 23n as well as Figures 24a to 24n Each of the figures may show an area where subsampling is applied. The small rectangles in each figure may represent sample points or positions. The area may be a reference area, an area of a reference block, and / or a template area.
[0897] For example, in Figures 22a to 22t In each of the figures, the leftmost uppermost rectangle in each region may be a point or position having coordinates (0,0) in the region.
[0898] exist Figures 22a to 22t Although the area size is shown as 8 8, but this size is only an example, and Figures 22a to 22t The subsampling method shown in can also be applied to regions of various sizes.
[0899] Figures 22a to 22t 、 Figures 23a to 23n as well as Figures 24a to 24n Each of the figures in the figure may show a block and region to which sampling is applied. The small rectangles in each figure may represent sample points or locations. The block may be a target block or a reference block. The region may be a reference region, a region of a reference block, and / or a template region.
[0900] For example, in Figures 23a to 23n and Figures 24a to 24n In each of , the top left coordinate of the block may be (0,0).
[0901] Figures 23a to 23n and Figures 24a to 24n The dimensions of each region are merely exemplary and Figures 23a to 23n and Figures 24a to 24n The sampling method shown in the accompanying drawings can also be applied to areas with different sizes.
[0902] Sample points (or positions) indicated by shading in the area of each drawing may represent sample points (or positions) selected by subsampling. Sample points (or positions) indicated in white in the area of each drawing may represent sample points (or positions) not selected by subsampling.
[0903] like Figure 22a to Figure 22a As shown in , subsampling may be applied to all or part of the reference region, and the template may be configured using only the sample points (or positions) selected by subsampling.
[0904] like Figures 23a to 23n and Figures 24a to 24n As shown in , subsampling may be performed on all or part of the template region in template matching, and calculation of the cost function may be performed only on sample points (or positions) selected by subsampling.
[0905] like Figure 22a to Figure 22a As shown in , subsampling may be performed on all or part of a search area in template matching, and searching and / or matching cost calculation may be performed only on pixels and / or positions selected by subsampling.
[0906] like Figure 22a to Figure 22a As shown in , subsampling may be performed on all or part of a search area in template matching, and searching and / or matching cost calculation may be performed only on pieces of motion information indicating pixels and / or positions selected by subsampling.
[0907] Subsampling using sub-block units Figures 25a to 25jSubsampling using sub-block units for a region according to an example is shown.
[0908] Figures 26a to 26l It is shown that sub-sampling using sub-block units for an area adjacent to a block according to an example.
[0909] In the embodiments, when subsampling is applied to a region, the selection unit is described as a pixel (or position). However, the selection unit in subsampling can be a subblock. In other words, in the description of subsampling, a pixel or position can be replaced by a subblock.
[0910] In an embodiment, each sub-block may be a partitioned portion of a block as described in the embodiment. A block may be partitioned into multiple sub-blocks. Furthermore, each sub-block may be a partitioned portion of a region as described in the embodiment. The region may be partitioned into multiple sub-blocks.
[0911] In other words, each sub-block may indicate a plurality of pixels (or positions) in a block or region. Here, the blocks or positions may be adjacent to each other.
[0912] In other words, the description of selection of pixels or positions in sub-sampling described in the embodiments can also be applied to sub-blocks.
[0913] Each of the left area, the upper left area, and the upper area described in the embodiments may also be regarded as a sub-block.
[0914] The sub-block may have a specific size. The description of the size of the block or region described in the embodiments may also be applied to the sub-block.
[0915] The shapes of blocks and regions described in the embodiments may also be applied to sub-blocks.
[0916] For example, Figures 25a to 25j As shown, the sub-blocks may be squares having sizes such as 16×16, 8×8, 4×4, and 2×2.
[0917] For example, Figures 26a to 26l As shown, the sub-blocks may be rectangular with sizes such as 8×4, 4×8, 4×4, 4×2, 2×4, and 2×2.
[0918] Search Methods in Template Matching Definition of Search For example, the search in an embodiment may be performed using calculation of a cost function for determining similarities between NUM_TEMPLATE_COMPARE templates.
[0919] In an embodiment, the search may include a process of determining at least one piece of motion information satisfying a specific condition within a specific search range. Based on the at least one piece of motion information determined by the search, the motion information of the target block may be determined and / or changed.
[0920] For example, the motion information that satisfies the specific condition may refer to the motion information having the lowest matching cost among the multiple pieces of motion information within the search range. However, the motion information that satisfies the specific condition is not limited thereto.
[0921] In an embodiment, the search may include a process of determining at least one block satisfying a specific condition within a specific search range. Motion information indicating the block determined by the search may be used as motion information of the target block.
[0922] For example, a block satisfying a specific condition may be one of the reference blocks within the search range.
[0923] In an embodiment, NUM_TEMPLATE_COMPARE may be 0, 1, 2, or a positive integer.
[0924] Cost function The cost function may refer to a function used to determine similarity between at least one sample point in a target template and at least one sample point in a reference template.
[0925] The similarity between the first value and the second value may be determined using at least one of: 1) a difference between the two values, 2) a ratio of the two values, and 3) an operation that compares the difference between the two values to a specific value.
[0926] The cost function may be a function for determining similarity between at least one sample point in the target template and a sample point in the corresponding reference template.
[0927] The cost function may be one or more of the sum of absolute differences (SAD), the sum of absolute transformed differences (SATD), the mean removed sum of absolute differences (MR-SAD), the mean square error (MSE), and the sum of squared errors (SSE). However, the cost function is not limited to the items listed above.
[0928] The cost function used in template matching may be predefined or may be determined based on signaled / encoded / decoded information.
[0929] For example, when the target block satisfies an enabling condition and / or a part of an enabling condition for bilateral matching, or when bilateral matching is performed on the target block, MR-SAD may be used as a cost function in template matching.
[0930] For example, when the target block does not satisfy the enabling conditions and / or part of the enabling conditions for bilateral matching, or when bilateral matching is not performed on the target block, SAD may be used as a cost function in template matching.
[0931] For example, the type of cost function in bilateral matching may be determined based on whether specific conditions are met in bilateral matching. In this case, the type of cost function in template matching may be determined based on whether 1) a condition for enabling bilateral matching and 2) specific conditions for determining the type of cost function in bilateral matching are met.
[0932] For example, when the target block meets 1) enabling conditions and 2) specific conditions for bilateral matching, MR-SAD can be used as the cost function in bilateral matching, while when the target block does not meet those conditions, SAD can be used as the cost function in bilateral matching.
[0933] For example, when the target block meets the enabling conditions of bilateral matching; and inter-frame weighted bi-prediction is performed or the number of samples in the target block is greater than a specific value, MR-SAD can be used as the cost function in template matching, otherwise SAD can be used as the cost function in template matching.
[0934] Search Scope In an embodiment, the search range may be a specific range centered at the position indicated by the initial motion information. In other words, the center of the search range may be the position indicated by the initial motion information. The specific range may be a range with a predefined area.
[0935] The search range may be a specific range centered at the position indicated by the initial motion information. In other words, the center position of the search range may be the position indicated by the initial motion information.
[0936] Alternatively, the search range may be a specific range in which the position indicated by the initial motion information is the upper left position. In other words, the upper left position of the search range may be the position indicated by the initial motion information.
[0937] Alternatively, the search range may include at least one of samples located in a lower left area, a left area, an upper left area, an upper area, and an upper right area around the target block.
[0938] Alternatively, the search range may consist of a previously reconstructed area around the target block. For example, the search range may include at least one of the samples (or positions of the samples) in the lower left area, the left area, the upper left area, the upper area, and the upper right area around the target block.
[0939] At least one of a size and a shape of the search range for the target block may be predefined by the encoding apparatus 1600 and the decoding apparatus 1700 .
[0940] Alternatively, at least one of the size and shape of the search range of the target block may be determined based on at least one of the size of the target block, encoding parameters of the target block, motion information of the target block, and a prediction mode of the target block.
[0941] Alternatively, information indicating one of a size and a shape of a search range for a target block may be encoded / decoded / signaled.
[0942] The search range may have a rectangular shape with a horizontal length of SR_X and a vertical length of SR_Y. Alternatively, the search range may have a diamond shape with a horizontal length of SR_X and a vertical length of SR_Y. However, the shape and size of the search range are not limited to the above embodiment.
[0943] Each of SR_X and SR_Y may be a positive integer. Each of SR_X and SR_Y may be a predefined value or a value determined based on signaled / encoded / decoded information.
[0944] The initial motion information may be determined based on at least one of the following: motion information of the target block, encoding parameters of the target block, a motion vector of the target block, a reference image of the target block, a block vector of the target block, a motion vector predictor of the target block, a block vector predictor of the target block, motion information of at least one neighboring block of the target block, a merge candidate of the target block, a motion vector difference of the target block, and a block vector difference of the target block.
[0945] Search Methods The types of search methods may be classified according to one or more of the following: 1) search pattern, 2) search resolution, 3) search range, 4) initial motion information, and 5) unit for deriving initial motion information. However, the criteria based on which the types of search methods are classified are not limited to such conditions.
[0946] Each search method may be determined based on at least one of: motion information of a target block, encoding parameters of the target block, a size of the target block, a prediction mode of the target block, a reference image of the target block, a sample value of at least one sample within the target block, a target template, a sample value of at least one sample within the target template, and an area of the target template.
[0947] Search Style The search pattern may be one of a diamond pattern, a cross pattern, and a full search pattern. However, the search pattern is not limited to the patterns listed above.
[0948] Searching using a diamond pattern may refer to an operation of searching for one or more of the positions of (0, 2×RR), (RR, RR), (2×RR, 0), (RR, -RR), (0, -RR), (-RR, -RR), (-RR, 0), (-RR, RR), and (0, 0) when (0, 0) represents the position indicated by the initial motion information.
[0949] The search using the cross pattern may be an operation of searching for one or more of the positions of (0, RR), (RR, 0), (0, -RR), (-RR, 0), and (0, 0) when (0, 0) represents the position indicated by the initial motion information.
[0950] RR may refer to search resolution and may be a predefined positive number.
[0951] Searching using the full search style can be an operation that searches all locations within a predefined search scope.
[0952] For example, assuming that FS_i has a value ranging from -FS_X to FS_X and FS_j has a value ranging from -FS_Y to FS_Y, searching using the full search pattern may refer to an operation of searching for a position corresponding to (FS_i × RR, FX_j × RR). Here, (0, 0) may be the position indicated by the initial motion information. However, the search range is not limited to the above positions. Each of FS_X and FS_Y may be a predefined positive number.
[0953] Search resolution The search resolution may be one of 4 pixels, full pixels, half pixels, and quarter pixels. However, the search resolution is not limited to the above pixels.
[0954] The search resolution may be predefined. Alternatively, the search resolution may be determined based on at least one piece of information about the resolution of the adaptive motion vector. Furthermore, the search resolution may be determined based on a signaled / encoded / decoded value.
[0955] Export of motion information In an embodiment, the unit for deriving motion information may be an entire block or a sub-block.
[0956] In other words, motion information may be derived for the entire block. A plurality of pieces of motion information may be derived for each sub-block.
[0957] Search Methods in Template Matching Figures 27 to 31 Shown are search patterns in template matching according to an example.
[0958] Figure 27A first relationship between search pattern and resolution according to an example is shown.
[0959] Figure 28 A second relationship between search pattern and resolution according to an example is shown.
[0960] Figure 29 A third relationship between search pattern and resolution according to an example is shown.
[0961] Figure 30 A fourth relationship between search pattern and resolution according to an example is shown.
[0962] Figure 31 A fifth relationship between search pattern and resolution according to an example is shown.
[0963] Figure 32 A sixth relationship between search pattern and resolution according to an example is shown.
[0964] In an embodiment, the search pattern and resolution can be configured based on the motion information, encoding parameters, prediction mode, and adaptive motion vector resolution of the target block. The search pattern and resolution can be changed according to the motion information, encoding parameters, prediction mode, and adaptive motion vector resolution of the target block.
[0965] The search pattern and resolution in the search step of template matching can be based on Figures 27 to 32 The table shown is configured.
[0966] Based on the motion information of the target block, coding parameters, prediction mode and adaptive motion vector resolution Figures 27 to 32 A specific column is selected in the table shown in . A search using a search pattern and a search resolution corresponding to rows marked with "v" in order from the top to the bottom of the selected column may be performed.
[0967] For example, in Figure 27 In the embodiment, when the AMVP mode is used for the target block and the resolution determined by the adaptive motion vector resolution is 4 pixels, a diamond pattern search using a 4-pixel search resolution is performed, after which a cross pattern search using a 4-pixel search resolution may be performed.
[0968] ALT_IF may represent the index of an adaptive interpolation filter. To calculate the pixel value at a sample point location of a specific resolution, an interpolation filter may be applied. The adaptive interpolation filter may be an interpolation filter selected from a plurality of interpolation filters by an index. In other words, when an adaptive interpolation filter is applied, different interpolation filters may be used depending on the index to calculate the pixel value at a sample point location of a specific resolution.
[0969] For example, the specific resolution may be half a pixel. However, the specific resolution is not limited to half a pixel.
[0970] For example, the interpolation filter determined by the index may be one of a 6-tap interpolation filter and an 8-tap interpolation filter. However, the method for determining the interpolation filter is not limited to the above-described determination method.
[0971] Template configuration method in affine mode Figure 33 A first template configuration method in an affine mode according to an example is shown.
[0972] Figure 34 A second template configuration method in the affine mode according to an example is shown.
[0973] CPMV may refer to affine control point motion vector. The motion vector of each sub-block in the target block may be derived using CPMV.
[0974] When the affine mode is used for the target block, the target block may be partitioned into sub-block units. Here, the width of each sub-block may be N, and the height thereof may be M.
[0975] The motion information of each subblock may be determined based on at least one of the motion information, encoding parameters, and size of the target block.
[0976] The template matching cost of the target block may be determined based on at least one of the template matching costs for the sub-blocks of the partition. For example, the template matching cost of the target block may be the sum of the template matching costs for the sub-blocks of the partition, or the average of the template matching costs for the sub-blocks.
[0977] Each of N and M can be 2, 4, 8, or a positive integer.
[0978] Each of N and M may be a predefined value, or may be a value determined based on signaled / encoded / decoded information.
[0979] Template matching in bidirectional prediction blocks In an embodiment, the value of a specific information of a specific target can be used as the value of the specific information of another target. Alternatively, the specific information of another target can be determined based on the specific information of the specific target. This use and determination can be represented by "inheritance".
[0980] For example, the motion information of the target block may be determined based on the motion information of the neighboring blocks. The determination based on the dependency relationship may be represented by "the target block has inherited the motion information from the neighboring blocks."
[0981] For example, when the merge mode is used for the target block, one merge candidate may be specified from the merge candidate list based on the merge index, and motion information of the specified merge candidate may be used as motion information of the target block.
[0982] For example, when the AMVP mode is used for the target block, one MV candidate may be specified from the MV candidate list based on the MV candidate index, and motion information of the specified MV candidate may be used as motion information of the target block.
[0983] When the motion information inherited by the target block from the neighboring blocks indicates bidirectional prediction, an embodiment of performing template matching on the target block may be described by the following steps: [Step 1] Template matching may be performed in each of the L0 direction and the L1 direction. Template matching costs C0 and C1 for the plurality of pieces of motion information determined in the L0 direction and the L1 direction may be calculated.
[0984] Here, when template matching for each direction is performed, template matching may be performed according to the same method as a method for performing template matching in unidirectional prediction for a corresponding direction without considering motion information in other directions.
[0985] Here, when the target block satisfies a predefined condition, MR-SAD may be used as a cost function, whereas when the target block does not satisfy the predefined condition, SAD may be used as a cost function.
[0986] The predefined condition may be a condition based on at least one of: whether a model-based prediction method is performed on the target block; an indicator indicating whether a model-based prediction method is performed on the target block; whether bilateral matching is performed on the target block; an indicator indicating whether bilateral matching is performed on the target block; motion information of the target block; the size of the target block; encoding parameters of the target block; motion information of neighboring blocks of the target block; encoding parameters of neighboring blocks of the target block; and the type of cost function in template matching in neighboring blocks of the target block.
[0987] For example, when a model-based prediction method is performed on a target block or when an indicator indicating whether a model-based prediction method is performed on a target block is true, MR-SAD may be used as a cost function, otherwise SAD may be used as a cost function.
[0988] For example, when bilateral matching is performed on a target block or when an indicator indicating whether bilateral matching is performed on a target block is true; and the number of samples in the target block is equal to or greater than a specific value, MR-SAD may be used as a cost function, otherwise SAD may be used as a cost function.
[0989] The cost function may refer to a cost function for search in template matching; and / or a cost function for calculating at least one of C0, C1, and C'. The cost function for search in template matching and the cost function for calculating at least one of C0, C1, and C' may be the same as or different from each other.
[0990] For example, during the search in template matching, MR-SAD can be used as the cost function, and C0, C1, and C' can be calculated using SAD.
[0991] [Step 2] When C0 < C1, a new target template T' can be generated using the target template and the template in the L0 direction.
[0992] For example, T can be the target template, T0 can be the reference template in the L0 direction, and T1 can be the reference template in the L1 direction.
[0993] Here, T' can be determined by the following [Equation 2].
[0994] [Equation 2] T’ = w r × T + w r0 × T0 w r and w r0 Each of them can be a predefined value.
[0995] w r and w r0 Each of them can be: a value determined based on whether to perform inter-frame weighted bi-prediction on the target block; and / or a weight in inter-frame weighted bi-prediction. [09...
Claims
1. A decoding method, comprising: constructing a filter candidate list comprising a plurality of filter candidates; as well as A final filter is determined among the plurality of filter candidates.
2. The decoding method according to claim 1, wherein: The final filter is determined based on matching costs of the plurality of filter candidates.
3. The decoding method according to claim 2, wherein: The filter candidate list is reconstructed based on the matching cost.
4. The decoding method according to claim 2, wherein: Reordering of the plurality of filter candidates in the filter candidate list is performed based on the matching costs.
5. The decoding method according to claim 2, wherein: The matching cost is a calculation result using a cost function for samples existing in a template generated using the plurality of filter candidates. The decoding method according to claim 1 , wherein: The plurality of filter candidates are applied to a template region of a template.
7. The decoding method according to claim 1, wherein: A filter candidate used in a neighboring block of a target block is used as one of the plurality of filter candidates of the target block.
8. A coding method comprising: constructing a filter candidate list comprising a plurality of filter candidates; as well as A final filter is determined among the plurality of filter candidates.
9. The encoding method according to claim 8, wherein: The final filter is determined based on matching costs of the plurality of filter candidates.
10. The encoding method according to claim 9, wherein: The filter candidate list is reconstructed based on the matching cost.
11. The encoding method according to claim 9, wherein: Reordering of the plurality of filter candidates in the filter candidate list is performed based on the matching costs.
12. The encoding method according to claim 9, wherein: The matching cost is a calculation result using a cost function for samples existing in a template generated using the plurality of filter candidates.
13. The encoding method according to claim 8, wherein: The plurality of filter candidates are applied to a template region of a template.
14. The encoding method according to claim 8, wherein: A filter candidate used in a neighboring block of a target block is used as one of the plurality of filter candidates of the target block.
15. A computer-readable storage medium for storing a bitstream for image decoding, wherein: The bitstream includes filter information, A filter candidate list including a plurality of filter candidates is constructed, and Based on the filter information, a final filter is determined among the plurality of filter candidates.
16. The computer-readable storage medium of claim 15, wherein: The final filter is determined based on matching costs of the plurality of filter candidates.
17. The computer-readable storage medium of claim 16, wherein: The filter candidate list is reconstructed based on the matching cost.
18. The computer-readable storage medium of claim 16, wherein: Reordering of the plurality of filter candidates in the filter candidate list is performed based on the matching costs.
19. The computer-readable storage medium of claim 16, wherein: The matching cost is a calculation result using a cost function for samples existing in a template generated using the plurality of filter candidates.
20. The computer-readable storage medium of claim 15, wherein: The plurality of filter candidates are applied to a template region of a template.
Citation Information
Patent Citations
Pick-up nail UV lamp
KR1020230001441A
Mother substrate including a plurality of display device and method for fabrcating display device
KR1020240001763A